A content recommendation method and apparatus
By estimating content evaluation indicator information based on the historical behavior data of user groups, the calculation process of advertising recommendation is simplified, the problems of high traffic pressure and poor user experience in the prior art are solved, and the recommendation effect and user experience are improved.
Patent Information
- Application Number
- CN202010401115.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-05-13
AI Technical Summary
The prior art requires a large amount of experimental data and complex computing processes when recommending advertisements, resulting in high traffic pressure and poor user experience.
By determining the user group to which the target user belongs, the content evaluation index information is estimated based on the user historical behavior data in the group, the content to be recommended is evaluated and the recommended content is selected.
The calculation process is simplified, the amount of data is reduced, online experiments are avoided, traffic pressure is alleviated, and recommendation results and user experience are improved.
Smart Images

Figure CN113672797B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a content recommendation method and apparatus. Background Art
[0002] With the development of the Internet, online advertising has become a mainstream advertising delivery method. While media owners insert ad spaces for online ad display and ad trading platforms convert user traffic into cash revenue, they often need to take into account the user experience to ensure good development. Since users have different sensitivities and demands for ads, the focus of ad recommendation to users varies. For different users, personalized ad recommendation is required to achieve a balance between ad revenue and user experience when the ad categories do not fluctuate much.
[0003] In the existing ad recommendation selection process, generally, online tests are conducted for each user experience metric, and corresponding ads are selected for different users for recommendation. This method requires a large amount of experimental data. Especially when there are many user experience metrics, the required amount of experimental data is extremely large, the calculation process is complex, and the traffic pressure is high. Summary of the Invention
[0004] Embodiments of this application provide a content recommendation method and apparatus for improving the efficiency and effectiveness of content recommendation.
[0005] According to the first aspect of the embodiments of this application, a content recommendation method is provided, including:
[0006] Determine the user group to which the target user belongs;
[0007] Determine the content evaluation metric information corresponding to the user group, where the content evaluation metric information is estimated based on the historical behavior data of the users in the user group for the recommended content;
[0008] Evaluate each piece of content to be recommended for the target user according to the content evaluation metric information;
[0009] Determine the recommended content to be recommended to the target user from each piece of content to be recommended according to the evaluation result.
[0010] According to the second aspect of the embodiments of this application, a content recommendation apparatus is provided, where the apparatus includes:
[0011] A grouping unit for determining the user group to which the target user belongs;
[0012] An index unit for determining the content evaluation metric information corresponding to the user group, where the content evaluation metric information is calculated based on the historical behavior data of the users in the user group for the recommended content;
[0013] An evaluation unit, configured to evaluate each content to be recommended for the target user according to the content evaluation index information;
[0014] A determination unit, configured to determine the recommended content to be recommended to the target user from all the content to be recommended according to the evaluation result.
[0015] In an optional embodiment, the content evaluation index information includes at least two content evaluation indexes and the weight of each content evaluation index;
[0016] Wherein, the weight corresponding to each content evaluation index is calculated based on the sample content obtained from the historical behavior data of the users in the user group for the recommended content.
[0017] In an optional embodiment, the evaluation unit is specifically configured to:
[0018] For each content to be recommended, calculate the recommendation score of the content to be recommended by using the estimated value of the content evaluation index of the content to be recommended and the weight corresponding to each content evaluation index;
[0019] The estimated value of the content evaluation index of the content to be recommended is estimated based on the content feature value of the content to be recommended and the user feature value of the target user.
[0020] In an optional embodiment, the evaluation unit is specifically configured to determine the estimated value of the content evaluation index of the content to be recommended according to the following method:
[0021] Input the content feature value of the content to be recommended and the user feature value of the target user into the trained deep neural network model to obtain the estimated value of the content evaluation index of the content to be recommended;
[0022] The deep neural network model is trained according to the content feature value, user feature value of the sample content and the interaction behavior data of the user for the sample content to obtain the corresponding model parameters.
[0023] In an optional embodiment, the evaluation unit is specifically configured to determine the weight corresponding to the content evaluation index according to the following method:
[0024] At least obtain the estimated value of the content evaluation index of the recommended content and the final objective function corresponding to the user group, and the weight parameter of the content evaluation index of the recommended content is included in the final objective function;
[0025] Determine the gradient of the weight parameter of the content evaluation index with respect to the final objective function;
[0026] Using the estimated value of the content evaluation index of the recommended content, iterative calculation is performed on the gradient according to the gradient descent method. When the difference between two adjacent iterations is less than a preset threshold or the number of iterations is reached, the weight of the corresponding content evaluation index is determined.
[0027] In an alternative embodiment, the evaluation unit is specifically configured to determine the estimated value of the content evaluation index of the recommended content according to the following method:
[0028] Input the content feature value of the recommended content and the user feature value of the recommended content into the trained deep neural network model, and calculate the estimated value of the content evaluation index of the recommended content;
[0029] The deep neural network model is trained according to the content feature value, user feature value of the sample content, and the interaction behavior data of the user for the sample content to obtain the corresponding model parameters.
[0030] In an alternative embodiment, the evaluation unit is specifically configured to determine the final objective function corresponding to the user group according to the following method:
[0031] Determine the constraint conditions and the initial objective function corresponding to the user group;
[0032] Determine the final objective function according to the constraint conditions and the initial objective function.
[0033] In an alternative embodiment, the evaluation unit is specifically configured to:
[0034] Combine the constraint conditions and the initial objective function into a transitional objective function;
[0035] Convert the non-differentiable terms in the transitional objective function into differentiable terms to obtain the final objective function.
[0036] In an alternative embodiment, a filtering unit is further included, which is configured to:
[0037] Determine all relevant contents corresponding to the user group according to the content evaluation index information corresponding to the user group;
[0038] Determine the content to be recommended from all relevant contents according to the filtering rules.
[0039] According to the third aspect of the embodiments of the present application, a computing device is provided, including at least one processor and at least one memory. Wherein, the memory stores a computer program, and when the program is executed by the processor, the processor is caused to execute the steps of the content recommendation method provided by the embodiments of the present application.
[0040] According to a fourth aspect of the embodiments of the present application, a storage medium is provided. The storage medium stores computer instructions, and when the computer instructions run on a computer, the computer is caused to execute the steps of the content recommendation method provided by the embodiments of the present application.
[0041] In the embodiments of the present application, multiple user groups are set based on set rules. For each user group, according to the historical behavior data of the users in the user group for the recommended content, the content evaluation index information of the user group is estimated. During the process of content recommendation to a target user online, multiple candidate content items will be matched for the determined target user, and it is necessary to select recommended content items from these candidate content items to push to the user. The specific recommendation method is to determine the user group to which the target user belongs, and obtain the content evaluation index information corresponding to the user group. According to the content evaluation index information, each candidate content item for the target user is evaluated, and the recommended content item to be recommended to the target user is determined from all candidate content items according to the evaluation results. In the embodiments of the present application, the content evaluation index information corresponds to the user group and is calculated based on the historical behavior data of the users in the user group. In this way, compared with calculating for each individual user separately, the calculation process is simplified and the amount of calculation data is reduced. Moreover, when determining the recommended content item for the target user, only by determining the user group of the target user, the content evaluation index information of the user group can be directly used to evaluate multiple candidate content items respectively, so as to select the recommended content item from multiple candidate content items. The method is simple and easy to implement, avoiding the online experiment process, alleviating the pressure on online traffic, and recommending content according to the characteristics of the user group, so that the recommended content is not limited to the individual user, but according to the group characteristics related to the user, making the recommended content expand the scope of recommended content on the basis of ensuring relevance to the user individual, enhancing the user experience, and thus improving the recommendation effect. Therefore, the technical solution of the present application performs well in improving the efficiency and effect of content recommendation. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application.
[0043] Figure 1 Shows a system architecture diagram of a content recommendation system in the embodiments of the present application;
[0044] Figure 2 Shows a flowchart of a content recommendation method in the embodiments of the present application;
[0045] Figure 3 Shows a structural schematic diagram of an optimization algorithm model in the embodiments of the present application;
[0046] Figure 4 shows a schematic flowchart of an advertisement recommendation method in a specific embodiment of the present application;
[0047] Figure 5 shows a structural block diagram of a content recommendation device in an embodiment of the present application;
[0048] Figure 6 shows a structural block diagram of a server provided in an embodiment of the present application. Detailed implementation manners
[0049] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, rather than all, of the embodiments of the technical solutions of the present application. Based on the embodiments described in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the technical solutions of the present application.
[0050] The terms "first" and "second" in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0051] Some concepts involved in the embodiments of the present application are introduced below.
[0052] 1. Artificial intelligence
[0053] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, a theory, method, technology and application system that perceives the environment, acquires knowledge and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology mainly includes several major directions such as computer vision technology, speech processing technology, and machine learning / deep learning.
[0054] 2. Machine learning
[0055] Machine learning is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.
[0056] 3. Cloud Technology
[0057] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.
[0058] Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires a powerful system back-end support, which can only be achieved through cloud computing.
[0059] 4. Online Advertising, Media Owners, Advertisers, Advertising Exchange Platforms
[0060] Online advertising, also known as Internet advertising, refers to the advertisements placed on advertising spaces on Internet platforms (such as WeChat Moments, official accounts, news apps, etc.).
[0061] Media owners refer to entities that own Internet platforms (such as WeChat Moments, official accounts, news apps, etc.). Generally, they already have a large number of user visits (also known as user traffic) and hope to convert user traffic into cash income. Therefore, they insert advertising spaces in the platforms.
[0062] Advertisers refer to entities that display their advertisements through advertising spaces on Internet platforms.
[0063] An advertising exchange platform (Advertising Exchange, ADX) refers to an entity that connects media owners and advertisers and places advertisers' advertisements on the advertising spaces provided by media owners.
[0064] 5. CPM, eCPM, pCTR
[0065] CPM (Cost Per Mille) refers to the cost that an advertiser needs to pay after an advertisement is shown to one thousand visiting users on an Internet platform.
[0066] eCPM (effective Cost Per Mile) is the advertising revenue that can be obtained for every one thousand impressions, which is used to reflect the profitability of the platform and can be regarded as an effective pre - estimate of CPM.
[0067] pCTR (predict Click Through Rate) is the probability that an online advertising system estimates an advertisement will be clicked after it is placed in a certain situation.
[0068] 6. Multi - objective optimization
[0069] Multi - objective optimization, also known as multi - objective programming or Pareto optimization, is a field of multi - objective decision - making. When there are trade - offs between two or more conflicting objectives, it is necessary to find a balance point to make the optimal decision. When one of the main objectives is used as the objective function, the remaining objectives can be used as constraint conditions, and the variables representing the decision - making solutions are used to impose a restricted range on the decision - making solutions.
[0070] 7. Relaxation
[0071] In the field of optimization, for problems with high solution complexity caused by non - differentiable or discontinuous functions, other differentiable or continuous functions can be used to replace the original function. This process is called relaxation. After relaxation, the optimization problem is easier to solve, and the optimal solution can be regarded as an approximation of the optimal solution of the original problem.
[0072] 8. Content evaluation metrics, content evaluation metric information, weights of content evaluation metrics
[0073] Select the advertisement with the best comprehensive performance in terms of advertising revenue and user experience from multiple advertisements to be recommended, and score each advertisement to be recommended according to the set indicators, and determine the advertisement with the highest score from all advertisements to be recommended and recommend it to the user. The content evaluation indicators in the embodiments of the present application correspond to the set indicators and are used to characterize the characteristics of the advertisement and the user's experience of the advertisement. The content evaluation indicator information includes the content evaluation indicators of the advertisement and the weight of each content evaluation indicator. The content evaluation indicator information corresponds to the user grouping, that is, for the users in the same user grouping, the content evaluation indicators and the corresponding weights are the same. The content evaluation indicators can include two types of indicators. On the one hand, they are related to the characteristics of the advertisement itself, such as eCPM, etc. On the other hand, they correspond to the user's interaction behavior with the advertisement, such as the like rate, negative feedback rate, etc. For the former type of content evaluation indicator, its value is only related to the characteristics of the advertisement itself, such as eCPM, and can be directly estimated according to the characteristics of the advertisement, generally evaluated by the advertiser. For the latter type of content evaluation indicator, its value is related not only to the characteristics of the advertisement itself but also to the users interacting with the advertisement, such as the negative feedback rate. Therefore, a machine learning model can be used for prediction. That is to say, the values of the content evaluation indicators in the embodiments of the present application are all estimated values.
[0074] The weight of the content evaluation indicator reflects the influence of each content evaluation indicator on whether to recommend the advertisement. The greater the weight, the more important the content evaluation indicator. In the embodiments of the present application, a multi-objective optimization method is used to calculate the weight of each content evaluation indicator. The calculated weight corresponds to the user grouping, that is, the weight values of the content evaluation indicators of different advertisements corresponding to the same user grouping are the same. Among different advertisements, the estimated values of the content evaluation indicators are different. Therefore, the weight value of the content evaluation indicator and the estimated value of the content evaluation indicator can be used to score each advertisement, and it can be determined whether the advertisement is worthy of recommendation according to the obtained score.
[0075] The following introduces the basic concept of the present application.
[0076] In the related art, when recommending advertisements and other content to users, the balance between different interaction behaviors of users and recommended content is generally mainly considered. Taking the information flow scenario as an example, the optimization objectives are the click-through rate and like rate of the recommended content. Therefore, it is necessary to score the recommended content based on the click-through rate and like rate. In the model for predicting the score of the recommended content, both the click-through rate and like rate are positive samples, but different weights are set, thereby affecting the final score. However, the selection of the weight here is generally based on empirical values or adjusted by means of online A / B testing.
[0077] The concept of A / B testing originated from the double-blind testing in biomedicine. In double-blind testing, patients are randomly divided into two groups and given a placebo and a test drug respectively without their knowledge. After a period of experiment, the performances of the two groups of patients are compared to determine whether there are significant differences, so as to decide whether the test drug is effective. A / B testing also adopts a similar concept: multiple advertisements are recommended to multiple users for experiment in the same time dimension, and the interactive behavior data of each user for each advertisement are collected. Finally, the best advertisement is selected through analysis and evaluation.
[0078] A / B testing is a process of repeated iteration and optimization, and its basic steps can be divided into:
[0079] Step 1: Set the project goal, that is, the goal of A / B testing, to measure the advantages and disadvantages of each advertisement;
[0080] Step 2: Design an iterative development plan for optimization and determine multiple weight combinations;
[0081] Step 3: Determine multiple advertisements to be recommended and the traffic diversion ratio for each advertisement to be recommended;
[0082] Step 4: Open the online traffic for testing according to the traffic diversion ratio;
[0083] Step 5: Collect experimental data for validity and effect judgment;
[0084] Step 6: Determine the weight combination according to the test results, adjust the traffic diversion ratio to continue testing, or if the test effect is not achieved, continue the optimization and iteration plan in Step 2 to re-develop and go online for testing.
[0085] In the above-mentioned online A / B testing method, each group of weights needs to correspond to a group of online experiments. When there are many content evaluation indicators (such as click-through rate, like rate, comment rate, negative feedback rate, sharing rate, etc.), as the number of weight dimensions increases, the number of required experiments increases exponentially, resulting in a waste of traffic.
[0086] Based on this, in the embodiments of the present application, the content evaluation index information corresponding to each group is directly determined according to the user group to which the target user belongs, so that the content evaluation index information can be used to evaluate each content to be recommended for the target user in the group, and then the recommended content for the target user can be determined from all the content to be recommended according to the evaluation results. Thus, the trial-and-error process of multiple groups of online experiments is avoided, the amount of data is reduced, and the calculation difficulty is lowered.
[0087] Furthermore, the embodiments of the present application select recommended content from the content to be recommended according to the characteristics of user groups, not limited to personal characteristics of users, but selected according to the group characteristics related to users. Thus, while ensuring the relevance of the recommended content to individual users, the scope of the recommended content is expanded, the user experience is enhanced, and the recommendation effect is improved.
[0088] In the embodiments of the present application, the recommended content is determined by scoring the content to be recommended. The specific scoring method is to weight the weights and estimated values of all content evaluation indicators of each content to be recommended, so as to obtain a score. Then, all the content to be recommended is sorted according to the scores, and the content to be recommended with the highest score is recommended to the user as the recommended content.
[0089] On the one hand, the estimated value of the content evaluation indicator of the content to be recommended can be calculated, which is calculated according to the content feature value of the content to be recommended and the user feature value of the target user. For example, for the target user, the user feature value of the target user and the content feature value of a content to be recommended are input into the trained deep neural network model, so as to obtain the estimated value of the content evaluation indicator of the content to be recommended. The deep neural network model is trained according to the content feature value, user feature value of the sample content and the interaction behavior data of the user for the sample content. That is to say, in the embodiments of the present application, the estimated value of the content evaluation indicator of each content to be recommended is estimated based on the sample content.
[0090] On the other hand, the weight of the content evaluation indicator of the content to be recommended can be calculated by an optimization algorithm model. Each user group corresponds to a set of weights of the content evaluation indicators, and the weights can be calculated in advance and stored in the database. For a user group, the final objective function corresponding to the user group and the estimated values of the content evaluation indicators of the recommended content are obtained. Among them, the final objective function is a function containing weight parameters, and the estimated values of the content evaluation indicators of the recommended content are also estimated by the above trained deep neural network model. Using the final objective function and the estimated values of the content evaluation indicators of the recommended content, iterative calculation is performed according to the multi-objective optimization algorithm to determine the optimal weight combination as the weight of the user group.
[0091] Generally speaking, the selection of recommended content depends on the content evaluation indicators that need to be optimized in the actual scenario. Since there is often a restrictive relationship between different content evaluation indicators, for example, for users with a consistently low interaction frequency with advertisements, it is necessary to increase the user's interaction behavior frequency, such as click-through rate, like rate, etc. Therefore, the revenue of the advertising trading platform can be appropriately sacrificed to improve the pertinence of advertising recommendations and achieve a better advertising display effect. Therefore, in the embodiments of the present invention, when calculating the weight for the same user group, there are multiple optimization objectives, and the optimization order between multiple objectives is determined. One objective is used as the main objective to construct the objective function, and the remaining objectives are used as secondary objectives to construct the constraint conditions, so as to achieve the purpose of balancing various relationships.
[0092] In addition, for different user groups, since the optimization objectives are different, preferably, the objective functions between different user groups are set differently, and the constraint conditions are not exactly the same, so that the weights of the content evaluation indicators are different for different user groups, and further the content recommendation is more targeted.
[0093] In this way, for the target user, first determine the user group of the target user, and obtain the weight of the content evaluation indicator corresponding to this user group. And determine multiple content to be recommended that match the target user. For each content to be recommended, calculate the estimated value of the content evaluation indicator of the content to be recommended. Then, for each content to be recommended, calculate the score of the content to be recommended according to the estimated value of the content evaluation indicator of the content to be recommended and the weight of the content evaluation indicator. Finally, sort the multiple content to be recommended according to the scores, and select the content to be recommended with the highest score as the recommended content to send to the user.
[0094] In addition, after sending the recommended content to the user, the interaction and feedback data of the user on the recommended content can be collected and used as sample content to be input into the deep neural network model and the optimization algorithm model for the training and update of the two models.
[0095] In the embodiments of the present application, the multi-objective optimization algorithm can use optimization algorithms such as simulated annealing algorithm, genetic algorithm, and ant colony algorithm to calculate the weights of the content evaluation indicators. Preferably, the embodiments of the present application combine the slack variable and the gradient descent algorithm for iterative optimization and output the optimal solution of the weight.
[0096] The following specifically introduces the slack variable and the gradient descent algorithm.
[0097] If all the constraint conditions of the linear programming model are of the less-than type, then M non-negative slack variables can be introduced through the standardization process. The introduction of slack variables is often to facilitate the solution within a larger feasible region. If it is 0, it converges to the original state; if it is greater than zero, the constraint is relaxed.
[0098] Specifically, the research on linear programming problems is based on the standard form. Therefore, for a given mathematical model of a non-standard linear programming problem, it is necessary to convert it into the standard form. Generally, for linear programming models in different forms, some methods can be used to convert them into the standard form. Among them, when the constraint condition is a linear programming problem of the less-than type, a non-negative new variable can be added (or subtracted) to the left side of the inequality to convert it into an equation. This newly added non-negative variable is called a slack variable (or surplus variable), and can also be collectively referred to as a slack variable. In the objective function, it is generally considered that the coefficient of the newly added slack variable is zero.
[0099] For the embodiments of the present application, since in the process of calculating the weights of the content evaluation indicators, the types of some constraint conditions are of the less-than type, where the less-than type includes several cases such as the constraint condition being "≤" or "<" or "≥" or ">". In these cases, it is necessary to convert these constraint conditions into the standard form, so as to convert the non-differentiable terms into differentiable terms. Specifically, in the embodiments of the present application, the softmax function is used to relax the objective function.
[0100] The Softmax function is a function in the following form:
[0101]
[0102] where θ i and x are column vectors, can be replaced with a function f i (x) of x. Through the softmax function, the range of P(i) can be made between [0,1]. In regression and classification problems, usually θ is the parameter to be solved, and by finding the θ that makes P(i) the largest i as the optimal parameter.
[0103] Simply put, it is to map the output value to the interval of [0,1] through the softmax function, and the sum of these mapped output values is 1. Thus, when selecting the output value finally, the model parameter with the largest probability (that is, the largest corresponding output value) can be selected.
[0104] The Gradient Descent algorithm is a type of iterative method and can be used to solve the least squares problem. When solving the model parameters of machine learning algorithms, that is, the unconstrained optimization problem, the gradient descent is one of the most commonly used methods. When solving the minimum value of the loss function, the gradient descent method can be used to iteratively solve step by step to obtain the minimized loss function and the model parameter values. The calculation process of the gradient descent method is to solve the minimum value along the direction of the gradient descent (it can also solve the maximum value along the direction of the gradient ascent).
[0105] The iterative formula of the gradient descent algorithm is as follows:
[0106] a k+1 = a k + ρ k s -(k) ……Formula 2
[0107] Wherein, s -(k) represents the negative gradient direction, and ρ k represents the search step size in the gradient direction. The gradient direction can be obtained by taking the derivative of the function. Generally, the method for determining the step size is determined by a linear search algorithm, that is, regarding the coordinates of the next point as a function of a k+1 , and then finding the a k+1 that satisfies the minimum value of f(a k+1 ).
[0108] Generally, if the gradient vector is 0, it means reaching an extreme point, and at this time, the magnitude of the gradient is also 0. When using the gradient descent algorithm for optimization, the termination condition for algorithm iteration is that the magnitude of the gradient vector is close to 0, and a very small constant threshold can be set.
[0109] Since the gradient descent algorithm is generally used to solve the model parameters in unconstrained optimization problems, therefore, when constructing the final objective function in the embodiments of the present application, first determine the corresponding constraint conditions and the initial objective function according to user grouping, and combine the constraint conditions and the initial objective function to form a transitional objective function. In this way, the optimization problem with constraint conditions is transformed into an optimization problem without constraint conditions. Also, since some of the constraint conditions are in the form of non-differentiable terms, the softmax function is used to relax the transitional objective function, that is, converting the non-differentiable terms in the transitional objective function into differentiable terms to obtain the final objective function. Thus, the final objective function is a linear function of the weights, and the gradient of the final objective function with respect to the weights can be solved, and then the optimal solution of the weights can be output using the gradient descent algorithm.
[0110] In the embodiments of the present application, the calculation of the estimated value of the content evaluation index and the calculation of the weight of the content evaluation index can be performed offline; or the calculation of the estimated value of the content evaluation index is performed offline, and the calculation of the weight of the content evaluation index is performed online; or the calculation of the weight of the content evaluation index is performed offline, and the calculation of the estimated value of the content evaluation index is performed online; or both the calculation of the estimated value of the content evaluation index and the calculation of the weight of the content evaluation index are performed online. After calculating the weight of the content evaluation index for the user group according to the training samples corresponding to the user group, the weight of the content evaluation index is associated and stored with the user group, so that when calculating the score of the content to be recommended online, the corresponding weight of the content evaluation index can be directly obtained. In addition, the recommended content pushed to the user can also be used as a training sample, and the interaction data of the user for the recommended content is obtained and input into the optimization algorithm model for training, so as to further optimize the weight of the content evaluation index.
[0111] In the embodiments of the present application, the calculation process of the estimated value of the content evaluation index, the calculation process of the weight, and the sorting process of the content to be recommended are decoupled. After obtaining the estimated value of the content evaluation index (such as click-through rate, like rate, negative feedback rate, etc.) of the content to be recommended through the model fusion method, an optimization algorithm model is established for the multi-objective optimization problem, and the weights of the predicted values of each index are directly solved. Then, the score is calculated online using the estimated value and the weight, thus avoiding the process of trial and error of multiple online experiments to solve the weight, and reducing the calculation difficulty and calculation amount.
[0112] In the embodiments of the present application, based on artificial intelligence technology, the estimated value and weight of the content evaluation index are determined. Specifically, a machine learning algorithm model is used to calculate the estimated value of the content evaluation index, and a multi-objective optimization algorithm is used to calculate the weight of the content evaluation index. It should be noted that the model for calculating the estimated value of the content evaluation index in the present application is not limited to the deep neural network model, and it is not limited to using the machine learning algorithm to calculate the estimated value of the content evaluation index. For example, a statistical algorithm model, a logistic regression algorithm model, etc. can also be used to estimate the estimated value of the content evaluation index. On the other hand, the multi-objective optimization algorithm model in the embodiments is only an example and is not limited. In addition to the improved gradient descent algorithm, algorithms such as simulated annealing algorithm and genetic algorithm can also be used to solve the weight of the content evaluation index.
[0113] After introducing the design concept of the embodiments of the present application, the application scenarios set by the present application will be briefly described below. It should be noted that the following scenarios are only used to illustrate the embodiments of the present application rather than to limit them. In specific implementation, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.
[0114] Please refer to Figure 1, which is a schematic diagram of a content recommendation system provided by an embodiment of the present application. In this application scenario, there are a terminal device 101, a media main server 102, and an advertising trading platform server 103. The terminal device 101, the media main server 102, and the advertising trading platform server 103 can be directly or indirectly connected through wired or wireless communication methods, which are not limited in this application.
[0115] Among them, the terminal device 101 is used to send an advertisement request to the advertising trading platform server 103, and it can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. A client corresponding to the media main server 102 is installed in the terminal device 101. The client can be a web client, a client installed in the terminal device 101, or a light application embedded in a third-party application, etc. The type of the client is not limited in this application. An advertising space of the advertising trading platform is embedded in the client, and the advertising trading platform server 103 puts the advertisements of the advertisers on the advertising space of the client, so as to display them to users.
[0116] The advertising trading platform server 103 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms, etc., and is applied to advertising placement products to meet the processing requirements of personalized advertising placement big data.
[0117] When implemented based on cloud technology, the advertising trading platform server 103 can process user data and advertising data through cloud computing and cloud storage.
[0118] Cloud computing is a computing model that distributes computing tasks on a large number of resource pools, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0119] Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technology, and distributed file systems, and collaborates through application software or application interfaces to jointly provide data storage and business access functions to the outside world.
[0120] In a possible implementation, user feature data, content feature data, and user interaction data for recommended content are stored in a cloud storage manner. When training an optimization algorithm model, training samples are obtained from the storage system corresponding to the cloud storage, and the optimization algorithm model is trained using the training samples to obtain the weights of the content evaluation indicators. At this time, the computing tasks are distributed in a large number of resource pools through cloud computing, reducing the computing pressure and obtaining the training results at the same time.
[0121] For the training of the deep neural network model, training samples are obtained from the storage system corresponding to the cloud storage, and the deep neural network model is trained using the training samples. When training the deep neural network model, the computing tasks are distributed in a large number of resource pools through cloud computing, reducing the computing pressure and obtaining the training results at the same time. The trained deep neural network model can be used to estimate the estimated value of the content evaluation indicator of the content to be recommended. When determining the estimated value of the content evaluation indicator of the content to be recommended, the content feature value and user feature value of the content to be recommended are obtained from the storage system corresponding to the cloud storage, and the estimated value of the content evaluation indicator is estimated using the content feature value and user feature value. The estimation of the estimated value of the content evaluation indicator can be performed through a trained machine learning algorithm. At this time, the computing tasks are distributed in a large number of resource pools through cloud computing, reducing the computing pressure and obtaining the estimation results at the same time.
[0122] When it is necessary to perform a scoring and ranking on the content to be recommended, the weights and estimated values of the content evaluation indicators are obtained from the storage system corresponding to the cloud storage, and the score of the content to be recommended is calculated using the weights and estimated values of the content evaluation indicators.
[0123] The following describes the scenarios applicable to the content recommendation process.
[0124] Based on the user feature data, advertisement feature data, and user interaction behavior data for advertisement in the collected sample content, the advertisement trading platform server 103 trains the deep neural network model to obtain the corresponding model parameters. Using the trained deep neural network model, the estimated value of the content evaluation indicator of the advertisement can be calculated.
[0125] The advertisement trading platform server 103 calculates the weights of content evaluation metrics offline using an optimization algorithm model. The advertisement trading platform server 103 needs to determine the initial objective function and constraints for user grouping, as well as the estimated values of the content evaluation metrics for the recommended advertisements. Among them, the estimated values of the content evaluation metrics for the recommended advertisements are calculated by the above-trained deep neural network model. After forming the final objective function using the initial objective function and constraints, the optimal solution of the weight parameters in the final objective function is solved according to the gradient descent algorithm using the estimated values of the content evaluation metrics for the recommended advertisements, and used as the weights of the content evaluation metrics. After the advertisement trading platform server 103 calculates the weights of the content evaluation metrics corresponding to the user grouping, it stores the weights of the content evaluation metrics in association with the user grouping.
[0126] When the user operates on the terminal device, the client responds to the user's operation of triggering an advertisement request and sends an advertisement request to the advertisement trading platform server 103. Here, the terminal device can directly send an advertisement request to the advertisement trading platform server 103, or the terminal device can send an advertisement request to the media owner server 102, and then the media owner server 102 forwards the advertisement request to the advertisement trading platform server 103.
[0127] After receiving the advertisement request, the advertisement trading platform server 103 determines the user grouping of the target user and obtains the weights of the corresponding content evaluation metrics, determines multiple relevant advertisements that match the target user, filters the multiple relevant advertisements to obtain multiple advertisements to be recommended, and calculates the estimated values of the content evaluation metrics for each advertisement to be recommended. Here, the estimated values of the content evaluation metrics for the advertisements to be recommended are also calculated by the above-trained deep neural network model. For each advertisement to be recommended, the score of the advertisement to be recommended is calculated using the weights and estimated values of the content evaluation metrics. The advertisements to be recommended are sorted according to the scores, and the recommended advertisements are determined from them and sent to the terminal device 101.
[0128] After the terminal device 101 displays an advertisement to the user, the advertisement trading platform server 103 can use this advertisement as a training sample, collect the content feature values of this advertisement, the user feature values of the corresponding user, and the interaction behavior data of the user for this advertisement, for training and updating the above deep neural network model and optimization algorithm model.
[0129] It should be noted that the application scenarios mentioned above are only shown for the convenience of understanding the spirit and principle of this application, and the embodiments of this application are not limited in this regard. On the contrary, the embodiments of this application can be applied to any applicable scenario.
[0130] The following combines Figure 1 the application scenario shown, and describes the content recommendation method provided by the embodiments of this application.
[0131] Please refer to Figure 2 , an embodiment of the present application provides a content recommendation method, as Figure 2 shown, the method includes:
[0132] Step S200: The terminal device responds to a user operation and sends a content recommendation request to the server.
[0133] The content recommendation request here can be a request to push advertisements, or a request to push other data or information, such as pushing jokes, pushing weather forecasts, pushing service items, etc. The embodiment of the present application does not limit the specific pushed content, and only takes advertisements as an example for introduction and explanation, not limited to advertisement pushing only.
[0134] In the specific operation process, the user operation responded by the terminal device can be an operation of actively requesting advertisements by the user, such as clicking on the advertisement display area in the web page, or a special operation preset by the user, such as double-clicking the screen, etc. It can also be an operation bound when the user performs other operations. For example, when the user opens a certain web page, an advertisement request is triggered to be sent. Although the user does not actively initiate an advertisement request, the server will still send an advertisement to the terminal device so that the terminal device can display the advertisement to the user.
[0135] Step S201: The server determines the user group to which the target user belongs.
[0136] In the specific implementation process, users can be grouped according to their own characteristics, such as user activity, or grouped according to the interaction behavior between users and content, such as the interaction frequency between users and advertisements, or grouped according to other rules. Since the advertising trading platform converts user traffic into cash income and often needs to take into account the user experience to ensure the good development of the ecosystem, in the embodiment of the present application, the user grouping is mainly based on the historical interaction behavior data between users and content.
[0137] For example, count the number of advertisement exposures and the number of interactions with advertisements of each user in the past period (such as 90 days), and according to the exposure count threshold and the interaction count threshold, identify users with low, medium, and high interaction frequencies with advertisements. For users in different user groups, different objective functions and constraint conditions can be adopted to further obtain the weights of different content evaluation indicators. For example, set the objective function according to the principle of experience priority, or set the objective function according to the principle of income-experience balance, so that the objective functions of different user groups are different.
[0138] Step S202: Determine the content evaluation index information corresponding to the user group, and the content evaluation index information is calculated based on the historical behavior data of the users in the user group for the recommended content.
[0139] That is to say, the content evaluation index information is calculated based on the historical behavior data of users in the user group corresponding to the recommended content. Specifically, users will perform some operations on the recommended content, such as clicking and evaluating. These operations will be recorded in the logs or transaction data of the recommended content. In the embodiments of the present application, the historical behavior data is used to represent these operations, and then the content evaluation index information is calculated based on the historical behavior data.
[0140] Among them, the content evaluation index information includes at least two content evaluation indexes and the weight of each content evaluation index. Here, the weight corresponding to each content evaluation index is calculated based on the sample content obtained from the historical behavior data of users in the user group for the recommended content.
[0141] The content evaluation index is the basis for scoring and ranking the content to be recommended. Each content evaluation index corresponds to a term in the scoring formula when calculating the score of the content to be recommended, such as eCPM, pCTR, the like rate of users liking the advertisement, the negative feedback rate, etc.
[0142] In the specific implementation process, the content evaluation index information corresponding to different user groups can be the same or different. Among them, the content evaluation index information being different can be that the content evaluation indexes are different. For example, the content evaluation indexes of user group A include eCPM, pCTR, and the negative feedback rate, and the internal evaluation indexes of user group B include the click-through rate, the like rate, and the negative feedback rate. It can also be that the content evaluation indexes are the same but the weights of the content evaluation indexes are different. For example, the content evaluation index information of both user group A and user group B is eCPM, the click-through rate, and the negative feedback rate, but the weights of user group A are 20%, 40%, and 30% in sequence, and the weights of user group B are 50%, 10%, and 40% in sequence. Among them, the sum of all weights corresponding to the same content to be recommended is 1.
[0143] Step S203: Evaluate each content to be recommended for the target user according to the content evaluation index information.
[0144] In the specific implementation process, there is no limit to the way of evaluating the content to be recommended. For example, all the content to be recommended for the target user can be classified according to the content evaluation index information, and then the recommended content is determined from the classification results; or labels can be attached to each content to be recommended, and the recommended content is determined according to the label content; or each content to be recommended can be scored according to the content evaluation index information, and the recommended content is selected according to the scores.
[0145] Preferably, in the embodiments of the present application, the evaluation of each content to be recommended is presented in a scoring manner. Using the scoring method for evaluation and recommendation, the results are clear and straightforward, and it is simple and easy to operate.
[0146] For each content to be recommended, calculate the recommendation score of the content to be recommended by using the estimated value of the content evaluation index of the content to be recommended and the weight corresponding to each content evaluation index;
[0147] Among them, the estimated value of the content evaluation index of the content to be recommended is estimated based on the content feature value of the content to be recommended and the user feature value of the target user.
[0148] Specifically, the estimated value of the content evaluation index of each content to be recommended and the corresponding weight can be weighted. Since each user group corresponds to a set of weights, different contents to be recommended for the target user share a set of weights. For example, the content evaluation indexes of the content to be recommended 1, the content to be recommended 2, and the content to be recommended 3 are all eCPM, click-through rate, and negative feedback rate, which are represented by m 1 、m 2 、m 3 respectively, and the corresponding weights are 30%, 20%, and 50%. Then, the recommendation scores P of the content to be recommended 1, the content to be recommended 2, and the content to be recommended 3 all satisfy the following formula:
[0149] P = 30%m 1 + 20%m 2 + 30%m 3 ……Formula 3
[0150] Among them, m 1 is the eCPM of the content to be recommended, m 2 is the click-through rate of the content to be recommended, and m 3 is the negative feedback rate of the content to be recommended. The recommendation scores of the content to be recommended 1, the content to be recommended 2, and the content to be recommended 3 are determined by the corresponding values of the content evaluation indexes. In the embodiments of the present application, in order to simplify the process, reduce the amount of experiments and data, the estimated value is used as the value of the content evaluation index and substituted into the above Formula 3, so as to calculate the recommendation score of each content to be recommended. Here, the estimated value is estimated based on the content feature value of the corresponding content to be recommended and the user feature value of the target user.
[0151] Specifically, the estimated value of the content evaluation index of the content to be recommended is determined according to the following method:
[0152] Input the content feature value of the content to be recommended and the user feature value of the target user into the trained deep neural network model to obtain the estimated value of the content evaluation index of the content to be recommended;
[0153] The deep neural network model is trained according to the content feature value, user feature value of the sample content, and the interaction behavior data of the user for the sample content to obtain the corresponding model parameters.
[0154] In the specific implementation process, the content evaluation indicators include user experience indicators and electronic resource indicators related to the content to be recommended. User experience indicators such as pCTR, like rate, negative feedback rate, etc., and electronic resource indicators related to the content to be recommended such as CPM, eCPM, etc.
[0155] Among them, the estimated value of the electronic resource indicator can be directly calculated by conversion. For example, eCPM reflects the exposure bid of the advertiser for this advertisement. According to the different types of advertisements, it can be directly obtained by the advertiser's exposure bid, or after the advertiser bids for conversion, the platform converts it into an exposure bid.
[0156] The estimated value of the user experience indicator can be obtained through an online experiment. For example, for the click-through rate, a set number of experimental users can be selected, and each piece of content to be recommended is pushed to the experimental users, and the click-through rate of the experimental users is statistically calculated as the estimated value of the click-through rate. However, this method has a long cycle and is relatively difficult to implement. In the embodiments of the present application, the estimated value of the user experience indicator is calculated by an algorithm model, and the algorithm model can be a trained deep neural network model. Optionally, the algorithm model can also be a statistical model, or a logistic regression algorithm, etc.
[0157] The estimated values in the embodiments of the present application include the estimated value of the user experience indicator of the recommended content and the estimated value of the user experience indicator of the content to be recommended. These two estimated values can be calculated by different algorithm models. However, to ensure the unity of the estimation standard, the estimated value of the user experience indicator of the recommended content and the estimated value of the user experience indicator of the content to be recommended are calculated by the same algorithm model. In this way, not only can the accuracy be improved, but also the calculation steps can be reduced and the calculation difficulty can be lowered.
[0158] In addition, the weights of the content evaluation indicators can be stored after offline calculation. When calculating the score, they can be directly obtained from the storage area. Of course, the weights of the content evaluation indicators can also be directly calculated, and there is no limitation here.
[0159] Step S204: Determine the recommended content to be recommended to the target user from all the content to be recommended according to the evaluation result.
[0160] In the specific implementation process, all the content to be recommended can be sorted according to the scores, and the top N pieces of content to be recommended with the best ranking are used as the recommended content, or the piece of content to be recommended with the highest score can also be directly used as the recommended content.
[0161] Step S205: The server sends the recommended content to the terminal device.
[0162] In the embodiments of the present application, multiple user groups are set based on set rules. For each user group, information on content evaluation metrics of the user group is estimated according to the historical behavior data of the users in the user group regarding the recommended content. During the process of recommending content to a target user online, multiple content items to be recommended are matched for the determined target user, and it is necessary to select a recommended content item from these content items to be recommended and push it to the user. The specific recommendation method is to determine the user group to which the target user belongs and obtain the information on content evaluation metrics corresponding to the user group. According to the information on content evaluation metrics, each content item to be recommended for the target user is evaluated, and based on the evaluation results, the recommended content item to be recommended to the target user is determined from all the content items to be recommended. In the embodiments of the present application, the information on content evaluation metrics corresponds to the user group and is calculated based on the historical behavior data of the users in the user group. Compared with calculating separately for each individual user, this simplifies the calculation process and reduces the amount of calculation data. Moreover, when determining the recommended content for the target user, only by determining the user group of the target user, the information on content evaluation metrics of the user group can be directly used to evaluate each of the multiple content items to be recommended, so as to select the recommended content item from the multiple content items to be recommended. The method is simple and easy to implement, avoids the online experiment process, alleviates the pressure on online traffic, and recommends content based on the characteristics of the user group, so that the recommended content is not limited to the individual user, but is based on the group characteristics related to the user. On the basis of ensuring relevance to the individual user, the scope of the recommended content is expanded, the user experience is enhanced, and thus the recommendation effect is improved. Therefore, the technical solution of the present application has good performance in improving the efficiency and effect of content recommendation.
[0163] Further, before step 204 of determining the recommended content item to be recommended to the target user from all the content items to be recommended according to the evaluation results, it further includes:
[0164] Determine all relevant content items corresponding to the user group according to the information on content evaluation metrics corresponding to the user group;
[0165] Determine the content items to be recommended from all the relevant content items according to the filtering rules.
[0166] In the specific implementation process, after the server receives a content recommendation request from the target user, it will obtain multiple relevant content items matching the target user from the content storage area or the network. Then, through filtering in various dimensions, the relevant content items that do not meet the filtering rules are filtered out from all the relevant content items, and the remaining relevant content items are used as the content items to be recommended. The filtering rules include whether the advertising budget is greater than the budget threshold, whether the industry freshness is greater than the freshness threshold, whether the eCPM is greater than the revenue threshold, etc. Such a preliminary screening can save the amount of calculation and improve the accuracy of recommendation. The subsequent scoring and ranking processes are all executed among the content items to be recommended.
[0167] The following specifically introduces how to determine the weights corresponding to the content evaluation indicators.
[0168] In the embodiments of the present application, the weights corresponding to the content evaluation indicators correspond to user groups and can be determined in the following manner:
[0169] At least obtain the estimated values of the content evaluation indicators of the recommended content and the final objective function corresponding to the user group. The final objective function includes the weight parameters of the content evaluation indicators of the recommended content;
[0170] Determine the gradient of the weight parameters of the final objective function with respect to the content evaluation indicators;
[0171] Using the estimated values of the content evaluation indicators of the recommended content, perform iterative calculations on the gradient according to the gradient descent method. When the difference between two adjacent iterations is less than a preset threshold or the number of iterations is reached, determine the weights of the corresponding content evaluation indicators.
[0172] In the specific implementation process, users in the same user group correspond to the same content evaluation indicators and the weights of the content evaluation indicators. The specific content evaluation indicators and the determination of the weights of the content evaluation indicators depend on which content evaluation indicators of the to-be-recommended content to be optimized and won. For example, for users with a consistently low interaction frequency with advertisements, it is necessary to increase their interaction behavior frequencies, such as click-through rate, like rate, positive comments, etc. The platform revenue can be appropriately sacrificed to achieve better results. For users with a normal interaction frequency with advertisements, the platform revenue can be optimized on the basis of ensuring the experience indicators. Since the platform revenue and user experience indicators are often in a restrictive relationship, therefore, the embodiments of the present application need to clarify the optimization objectives and quantify the balance between these objectives through modeling.
[0173] Assume that under a certain user group, the total number of user requests for advertisements per unit time is K, that is, these users have a total of K to-be-recommended advertisement queues. For each to-be-recommended advertisement j in queue k, estimate values of content evaluation indicators such as eCPM, pCTR, like rate, negative feedback rate, etc. are determined, and the final recommended score is S k,j . After the advertisement scores in queue k are sorted, the recommended advertisement i k is the advertisement with the highest score, that is:
[0174] i k = argmax j {S k,j}... Formula 4
[0175] where S k,j is the recommended score of the to-be-recommended advertisement j in queue k.
[0176] Use the target users with normal advertising interaction frequency to illustrate the modeling process. For this part of target users, the optimization goal is to maximize the platform revenue on the basis of ensuring user experience indicators. Since the revenue generated by a single pull is measured by eCPM, therefore, the objective function is modeled as optimizing the average eCPM of recommended ads, which can be expressed by the formula as follows:
[0177]
[0178] where, ecpm ik is the eCPM of the ad i to be recommended in queue k.
[0179] In addition, it is necessary to set the baseline of user experience indicators and the tolerance range of ad category fluctuations for this user group. For the click-through rate pCTR, the baseline of the average pCTR of recommended ads is pctr 0 , that is, the average pCTR of all recommended ads needs to be greater than pctr 0 , then the corresponding constraint condition is:
[0180]
[0181] where, pctr ik is the click-through rate of the ad i to be recommended in queue k.
[0182] Similarly, other user experience indicators similar to formula 6 can also be added, such as like rate, social interaction rate, negative feedback rate, etc. In addition, let represent the categories that need to control fluctuations (such as follow-up ads, download ads, etc.), and represent whether the ad to be recommended belongs to this category (1 means the ad to be recommended belongs to this category, 0 means the ad to be recommended does not belong to this category). The upper and lower limits of the category proportion are r u and r l , then the constraint condition of category fluctuation is:
[0183]
[0184]
[0185] where, is the category that needs to control fluctuations, is that the ad to be recommended in queue k belongs to category.
[0186] Combining formulas 4-8, the optimization problem for this user group can be modeled. Among them, the recommendation score S k,j determines the score of the recommended ad i k , and then determines the values of the objective function and constraint conditions. According to this multi-objective optimization problem, the corresponding recommendation score Sk,j Decomposed into content evaluation metrics eCPM, pCTR, (representing whether ad j belongs to category If it is, it is 1; if not, it is 0) of linear weighting. The recommendation score S k,j The formula is as follows:
[0187]
[0188] Among them, ω 1 , ω 2 , ω 3 are the weights of the content evaluation metrics, ecpm k,j is the eCPM of the ad j to be recommended in queue k, pctr ik is the click-through rate of the ad j to be recommended in queue k, is the ad to be recommended in queue k belonging to category. When adding constraint conditions to the optimization problem, such as the like rate threshold, other category fluctuation thresholds, corresponding content evaluation metric terms and weights are added to Formula 9.
[0189] Assume that there is an objective function, m user experience metric constraints, and n category fluctuation constraints in the optimization problem. Then, there are m + n + 1 corresponding content evaluation metric terms in Formula 9, and also m + n + 1 weights. In the embodiments of the present application, through offline training, the optimal solutions of the m + n + 1 weights ω in Formula 9 are obtained, so that when the constraint conditions such as Formulas 6 - 8 are satisfied, the value of Formula 5 is maximized.
[0190] Similarly, for users with a consistently low ad interaction frequency, the objective function and constraint conditions of the multi-objective optimization problem can be changed. For example, the average number of likes of the user can be used as the objective function, and the average eCPM can be used as the constraint condition. Similarly, it is necessary to solve the optimal solutions of the m + n + 1 weights ω in Formula 9 so that the average pCTR of the recommended ads is maximized when the constraint conditions are satisfied.
[0191] In the embodiments of the present invention, after modeling the optimization problems for different user groups, a unified optimization algorithm model can be used for solving. Figure 3 Shows the structural schematic diagram of the optimization algorithm model in the embodiments of the present application.
[0192] As Figure 3 shown, the data for training the multi-objective optimization algorithm model includes the initial objective function and constraint conditions corresponding to the user groups. In the embodiments of the present application, to simultaneously achieve multiple optimization objectives through the multi-objective optimization algorithm, a most important optimization objective can be determined to form the corresponding initial objective function, and the remaining optimization objectives are used as constraint conditions. Figure 3In the example shown, the main optimization objectives are to optimize platform revenue or user experience. Therefore, based on optimizing platform revenue or user experience, an initial objective function is formed and input into the multi-objective optimization algorithm. The remaining optimization objectives may include optimizing user experience, optimizing platform revenue, and optimizing category fluctuations. Thus, one or more constraint conditions, such as user experience constraints, platform revenue constraints, and category fluctuation constraints, can be formed according to the above optimization objectives and input into the multi-objective optimization algorithm to achieve the purpose of balancing multiple optimization requirements. Further, after training by the multi-objective optimization algorithm, content evaluation indicators for rating the content to be recommended and their corresponding weights can be calculated. The process of obtaining the weights through training by the multi-objective optimization algorithm is specifically introduced below.
[0193] The offline data includes:
[0194] On the user side: user grouping, and the baselines of the objective function and constraint conditions corresponding to the user grouping. For example, for the user grouping with normal advertisement interaction frequency, the optimization objective is to maximize the average eCPM, and the baseline indicators of the constraint conditions are pctr 0 、r u and r l etc. For the user grouping with consistently low advertisement interaction frequency, the optimization objective is to maximize the average pCTR, and the baseline indicators of the constraint conditions are ecpm 0 、r u and r l etc.
[0195] On the advertisement side: the estimated values of the content evaluation indicators of each recommended advertisement in all the recommended advertisement queues. For example, the ecpm corresponding to each advertisement j in queue k k,j , pctr k,j , etc. Among them, the estimated values of the content evaluation indicators of the recommended content are determined in the following manner: input the content feature values and user feature values of the recommended content into the trained deep neural network model to calculate the estimated values of the content evaluation indicators of the recommended content.
[0196] In the embodiments of the present application, all the advertisement data in the advertisement queue corresponding to each user grouping is collected. Under the corresponding user grouping, according to the multi-objective optimization problem modeled thereby, the weights are solved based on the offline data, and finally the content evaluation indicators and their corresponding weights in Formula 9 are output.
[0197] In the specific implementation process, algorithms such as simulated annealing algorithm and genetic algorithm can be used to solve the weights. In the embodiments of the present application, the relaxation problem is combined with the gradient descent algorithm. Specifically, the objective function and the constraint conditions are combined in the form of a penalty term, so that the gradient descent algorithm can be used for calculation. Further, the non-differentiable terms in the calculation formula are converted into differentiable terms through the softmax function, thereby simplifying the calculation process on the basis of ensuring that the weights can be calculated.
[0198] Correspondingly, in the embodiments of the present application, the final objective function corresponding to the user grouping is determined according to the following method:
[0199] Determine the constraint conditions and the initial objective function corresponding to the user grouping;
[0200] Determine the final objective function according to the constraint conditions and the initial objective function.
[0201] Among them, determining the final objective function according to the constraint conditions and the initial objective function includes:
[0202] Combine the constraint conditions and the initial objective function into a transitional objective function;
[0203] Convert the non-differentiable terms in the transitional objective function into differentiable terms to obtain the final objective function.
[0204] In the specific implementation process, first, determine the initial objective function and the constraint conditions of the user grouping, and convert the constrained optimization problem into an unconstrained optimization problem. Add the relu function and the coefficient λ outside the constraint conditions, and combine them with the initial objective function (i.e., formula 4) to form a transitional objective function:
[0205]
[0206] Use the softmax function for relaxation to convert the non-differentiable terms in formula 7 and formula 8 into differentiable terms:
[0207]
[0208] Substitute formula 11 into formula 10, so that the objective function f(S k,j ) in formula 10 has a gradient with respect to S k,j . And, as can be seen from formula 9, S k,j is a linear function of the weights {w}. Therefore, through the chain rule, the gradient of the objective function f(S k,j ) with respect to {w} can be solved. After obtaining the numerical solution of this gradient, use the gradient descent method for iterative optimization to output the optimal solution {w *}, which is the weight.
[0209] From Figure 3For the multi-objective optimization algorithm model shown, the content evaluation indicators corresponding to Formulas 4-8 can be obtained as the eCPM item, the pCTR item, and the category factor The corresponding weights are in turn That is, the recommendation score calculation formula is:
[0210]
[0211] In this way, the weights {w *} of the content evaluation indicators output by the multi-objective optimization algorithm model can be utilized, and according to the estimated values of the content evaluation indicators, the recommendation score S k,j of each advertisement to be recommended can be calculated, and the advertisements to be recommended are sorted, and finally the recommended advertisements are determined.
[0212] The following uses a specific embodiment to introduce the above process in detail. The specific process of the specific embodiment is as Figure 4 shown, including:
[0213] The server receives an advertisement recommendation request sent by the terminal device, and the advertisement recommendation request contains the user identifier of the target user.
[0214] The server provides an online advertisement recommendation service to the target user based on the advertisement recommendation request. First, the server determines the user grouping of the target user, determines the sorting factors corresponding to the user grouping and the weight of each sorting factor. Here, the sorting factor is the content evaluation indicator in the above embodiment. The server determines all relevant advertisements matched by the target user and filters all relevant advertisements to obtain multiple advertisements to be recommended. The server uses the deep neural network model to estimate the estimated value of the sorting factor of each advertisement to be recommended, and obtains the sorting factor and weight corresponding to the user grouping. The server performs weighted calculation according to the estimated value of the sorting factor and the sorting factor weight to obtain the score of each advertisement to be recommended, and sorts all advertisements to be recommended according to the score. According to the sorting, the server determines the recommended advertisement to be sent to the target user and sends it to the user.
[0215] The estimated value of the sorting factor in the above specific embodiment is estimated through the deep neural network model, and the sorting factor and weight are calculated through the multi-objective optimization algorithm. Among them, the process of calculating the sorting factor weight is performed offline. After the server determines the sorting factor and the corresponding weight, the weight of the sorting factor can be associated and stored with the user grouping. When sorting the advertisements to be recommended, it can be directly queried and obtained.
[0216] In addition, the server can use the advertisements that have been recommended to the user as training samples, collect the user interaction behavior data of the user for the recommended advertisements, and train and update the deep neural network model and the multi-objective optimization algorithm.
[0217] The following is an embodiment of the device of the present application. For details not described in detail in the device embodiment, reference may be made to the corresponding method embodiments described above one by one.
[0218] Please refer to Figure 5 , which shows a structural block diagram of a content recommendation device provided by an embodiment of the present application. This cross-chain data processing device is implemented as all or part of the server 103 in Figure 1 The device includes: a grouping unit 501, an index unit 502, an evaluation unit 503, a determination unit 504, and a filtering unit 505.
[0219] The grouping unit is used to determine the user group to which the target user belongs;
[0220] The index unit is used to determine the content evaluation index information corresponding to the user group. The content evaluation index information is calculated based on the historical behavior data of the users in the user group for the recommended content;
[0221] The evaluation unit is used to evaluate each content to be recommended for the target user according to the content evaluation index information;
[0222] The determination unit is used to determine the recommended content to be recommended to the target user from all the content to be recommended according to the evaluation result.
[0223] In an optional embodiment, the content evaluation index information includes at least two content evaluation indexes and the weight of each content evaluation index;
[0224] Among them, the weight corresponding to each content evaluation index is calculated based on the sample content obtained from the historical behavior data of the users in the user group for the recommended content.
[0225] In an optional embodiment, the evaluation unit is specifically used for:
[0226] For each content to be recommended, use the estimated value of the content evaluation index of the content to be recommended and the weight corresponding to each content evaluation index to calculate the recommendation score of the content to be recommended;
[0227] The estimated value of the content evaluation index of the content to be recommended is estimated based on the content feature value of the content to be recommended and the user feature value of the target user.
[0228] In an optional embodiment, the evaluation unit is specifically used to determine the estimated value of the content evaluation index of the content to be recommended according to the following method:
[0229] Input the content feature value of the content to be recommended and the user feature value of the target user into the trained deep neural network model to obtain the estimated value of the content evaluation index of the content to be recommended;
[0230] The deep neural network model is trained based on the content feature values, user feature values of the sample content, and the interaction behavior data of the user for the sample content to obtain the corresponding model parameters.
[0231] In an alternative embodiment, the evaluation unit is specifically configured to determine the weight corresponding to the content evaluation index according to the following method:
[0232] At least obtain the estimated value of the content evaluation index of the recommended content and the final objective function corresponding to the user group. The final objective function includes the weight parameter of the content evaluation index of the recommended content;
[0233] Determine the gradient of the weight parameter of the final objective function with respect to the content evaluation index;
[0234] Using the estimated value of the content evaluation index of the recommended content, perform iterative calculation on the gradient according to the gradient descent method. When the difference between two adjacent iterations is less than the preset threshold or the number of iterations is reached, determine the weight of the corresponding content evaluation index.
[0235] In an alternative embodiment, the evaluation unit is specifically configured to determine the estimated value of the content evaluation index of the recommended content according to the following method:
[0236] Input the content feature value of the recommended content and the user feature value of the recommended content into the trained deep neural network model, and calculate the estimated value of the content evaluation index of the recommended content;
[0237] The deep neural network model is trained based on the content feature values, user feature values of the sample content, and the interaction behavior data of the user for the sample content to obtain the corresponding model parameters.
[0238] In an alternative embodiment, the evaluation unit is specifically configured to determine the final objective function corresponding to the user group according to the following method:
[0239] Determine the constraint conditions and the initial objective function corresponding to the user group;
[0240] Determine the final objective function according to the constraint conditions and the initial objective function.
[0241] In an alternative embodiment, the evaluation unit is specifically configured to:
[0242] Combine the constraint conditions and the initial objective function into a transitional objective function;
[0243] Convert the non-differentiable term in the transitional objective function into a differentiable term to obtain the final objective function.
[0244] In an alternative embodiment, it further includes a filtering unit for:
[0245] Determine all relevant content corresponding to the user group according to the content evaluation index information corresponding to the user group;
[0246] Determine the content to be recommended from all relevant content according to the filtering rules.
[0247] Please refer to Figure 6 , which shows a block diagram of the structure of a server provided by an embodiment of the present application. The server 1100 is implemented as Figure 1 the server 103 in
[0248] The server 1100 includes a central processing unit (CPU) 801, a system memory 1104 including a random access memory (RAM) 1102 and a read-only memory (ROM) 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The server 1100 also includes a basic input / output system (I / O system) 1106 for transferring information between various components within the computer, and a mass storage device 1107 for storing an operating system 1113, application programs 1114, and other program modules 1115.
[0249] The basic input / output system 1106 includes a display 1108 for displaying information and input devices 1109 such as a mouse and a keyboard for user input of information. The display 1108 and the input devices 1109 are both connected to the central processing unit 1101 through an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may also include an input / output controller 1110 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1110 also provides output to a display screen, a printer, or other types of output devices.
[0250] The mass storage device 1107 is connected to the central processing unit 1101 through a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1107 and its associated computer-readable medium provide non-volatile storage for the server 1100. That is, the mass storage device 1107 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM drive.
[0251] Without loss of generality, computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM, EEPROM, flash memory or other solid state storage technologies, CD-ROM, DVD or other optical storage, magnetic tape cartridges, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that computer storage media is not limited to the above several types. The above-mentioned system memory 1104 and mass storage device 1107 may be collectively referred to as memory.
[0252] According to various embodiments of the present application, the server 1100 may also be run by a remote computer on the network connected through a network such as the Internet. That is, the server 1100 may be connected to the network 1112 through the network interface unit 1111 connected to the system bus 1105, or in other words, the network interface unit 1111 may also be used to connect to other types of networks or remote computer systems (not shown).
[0253] The memory also includes one or more programs. One or more programs are stored in the memory, and one or more programs include instructions for performing content recommendation provided in the embodiments of the present application.
[0254] Those of ordinary skill in the art can understand that all or part of the steps in the content recommendation method of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0255] Those of ordinary skill in the art can understand that all or part of the steps in the content recommendation method of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0256] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0257] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A content recommendation method, characterized in that, it includes: Determine the user group to which the target user belongs; Determine the content evaluation index information corresponding to the user group, where the content evaluation index information is calculated based on the historical behavior data of the users in the user group for the recommended content; the content evaluation index information includes at least two content evaluation indexes and the weight of each content evaluation index, and the weight corresponding to each content evaluation index is calculated based on the sample content obtained from the historical behavior data of the users in the user group for the recommended content; Evaluate each content to be recommended for the target user according to the content evaluation index information; Determine the recommended content to be recommended to the target user from all the content to be recommended according to the evaluation result; wherein, the weight corresponding to the content evaluation index is determined according to the following method: At least obtain the estimated value of the content evaluation index of the recommended content and the final objective function corresponding to the user group, where the final objective function includes the weight parameter of the content evaluation index of the recommended content; Determine the gradient of the weight parameter of the content evaluation index for the final objective function; Use the estimated value of the content evaluation index of the recommended content, and perform iterative calculation for the gradient according to the gradient descent method. When the difference between two adjacent iterations is less than the preset threshold or the number of iterations is reached, determine the weight of the corresponding content evaluation index.
2. The method according to claim 1, characterized in that, The evaluating each content to be recommended for the target user according to the content evaluation index information includes: For each content to be recommended, calculate the recommendation score of the content to be recommended by using the estimated value of the content evaluation index of the content to be recommended and the weight corresponding to each content evaluation index; The estimated value of the content evaluation index of the content to be recommended is estimated based on the content feature value of the content to be recommended and the user feature value of the target user.
3. The method according to claim 2, characterized in that, The estimated value of the content evaluation index of the content to be recommended is determined according to the following method: Input the content feature value of the content to be recommended and the user feature value of the target user into the trained deep neural network model to obtain the estimated value of the content evaluation index of the content to be recommended; The deep neural network model is trained according to the content feature value, user feature value of the sample content and the interaction behavior data of the user for the sample content to obtain the corresponding model parameters.
4. The method according to claim 1, characterized in that, The estimated value of the content evaluation index of the recommended content is determined according to the following method: Input the content feature value of the recommended content and the user feature value of the recommended content into the trained deep neural network model, and calculate to obtain the estimated value of the content evaluation index of the recommended content; The deep neural network model is trained according to the content feature value, user feature value of the sample content and the interaction behavior data of the user for the sample content to obtain the corresponding model parameters.
5. The method according to claim 1, characterized in that, The final objective function corresponding to the user grouping is determined as follows: Determine the constraint conditions and the initial objective function corresponding to the user grouping; Determine the final objective function according to the constraint conditions and the initial objective function.
6. The method according to claim 5, wherein, the determining the final objective function according to the constraint conditions and the initial objective function includes: Combining the constraint conditions and the initial objective function into a transitional objective function; Converting the non-differentiable terms in the transitional objective function into differentiable terms to obtain the final objective function.
7. The method according to claim 1, wherein, before determining the recommended content to be recommended to the target user from all the content to be recommended according to the evaluation result, further includes: Determining all relevant content corresponding to the user grouping according to the content evaluation index information corresponding to the user grouping; Determining the content to be recommended from all the relevant content according to the filtering rules.
8. A content recommendation device, wherein, includes: A grouping unit for determining the user grouping to which the target user belongs; An index unit for determining the content evaluation index information corresponding to the user grouping, the content evaluation index information being calculated based on the historical behavior data of the users in the user grouping for the recommended content; the content evaluation index information includes at least two content evaluation indexes and the weight of each content evaluation index, and the weight corresponding to each content evaluation index is calculated based on the sample content obtained from the historical behavior data of the users in the user grouping for the recommended content; An evaluation unit for evaluating each piece of content to be recommended for the target user according to the content evaluation index information; A determination unit for determining the recommended content to be recommended to the target user from all the content to be recommended according to the evaluation result; The index unit is further configured to: at least obtain the estimated value of the content evaluation index of the recommended content and the final objective function corresponding to the user grouping, where the final objective function includes the weight parameter of the content evaluation index of the recommended content; determine the gradient of the weight parameter of the final objective function with respect to the content evaluation index; use the estimated value of the content evaluation index of the recommended content, and perform iterative calculation on the gradient according to the gradient descent method, and when the difference between two adjacent iterations is less than a preset threshold or the number of iterations is reached, determine the weight of the corresponding content evaluation index.
9. The device according to claim 8, wherein, the evaluation unit is specifically configured to: For each piece of content to be recommended, calculate the recommendation score of the content to be recommended by using the estimated value of the content evaluation index of the content to be recommended and the weight corresponding to each content evaluation index; The estimated value of the content evaluation index of the content to be recommended is estimated based on the content feature value of the content to be recommended and the user feature value of the target user.
10. A computer device, wherein, includes: At least one processor, and A memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the at least one processor implements the method according to any one of claims 1 to 7 by executing the instructions stored in the memory.
11. A computer-readable storage medium, characterized in that the computer-readable storage medium stores computer instructions, and when the computer instructions run on a computer, the computer is caused to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Information recommending method and device
CN104504098A
Shared account detection method and related device
CN110175438A
Information flow recommendation method and device, based on deep network, equipment and storage medium
CN110266745A