Content recommendation method and device, model training method and device, electronic equipment and medium

By using a multi-task learning model to predict click-through rates and browsing duration, and combining this with operational information, the problem of low exposure for high-quality content in existing technologies has been solved, enabling widespread recommendation of high-quality content and ensuring user preference.

CN120950768APending Publication Date: 2025-11-14BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511068920.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing content recommendation systems rely too heavily on users' historical interests, resulting in low exposure of high-quality content and an inability to recommend it widely.

Method used

A multi-task learning model is used to combine content information, operational information, and user information of candidate recommended content to predict click-through rate, browsing time, and exposure probability. The recommendation score is determined through the multi-task learning model, and high-quality content is selected and recommended.

Benefits of technology

It increases the exposure of high-quality content, ensuring that recommended content is more likely to be liked by users, and making the recommended content richer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950768A_ABST
    Figure CN120950768A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a content recommendation method and device, a model training method and device, electronic equipment and a medium, and relates to the technical field of data analysis, and the technical scheme comprises the following steps: obtaining content information and operation information of a plurality of candidate recommendation contents and user information of a user to be recommended; and based on the content information of the candidate recommendation content, the operation information and the user information, utilizing a multi-task learning model to determine a predicted click rate, a predicted browsing duration and a predicted warranty probability. And determining recommendation scores of the candidate recommendation contents according to the predicted click rate, the predicted browsing duration and the predicted warranty probability of the candidate recommendation contents, selecting a plurality of candidate recommendation contents as to-be-recommended contents according to a sequence of the recommendation scores of the candidate recommendation contents from high to low, and recommending the to-be-recommended contents to a to-be-recommended user. When the content is recommended to the user, excessive dependence on the historical interest of the user is avoided, and the exposure of the high-quality content is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis technology, and in particular to a content recommendation method, a model training method, an apparatus, an electronic device, and a medium. Background Technology

[0002] In content delivery platforms, to ensure a good user experience, content that users are more likely to be interested in is typically recommended. For example, when a user enters a video website, videos that the user is more likely to be interested in are usually retrieved and recommended based on the user's information, including the user's interest in previously recommended content. The recommended videos are then shown to the user to increase click-through rates.

[0003] However, this method relies too heavily on users' historical interests when determining recommended content, which can easily lead to a "Matthew effect" where the "stronger get stronger." This results in increasingly limited content types being displayed to users, and some high-quality content, such as newly launched content, exclusive copyrighted content, and self-produced content, receives low exposure and cannot be widely recommended to users. Summary of the Invention

[0004] The purpose of this application is to provide a content recommendation method, model training method, apparatus, electronic device, and medium to avoid over-reliance on users' historical interests when recommending content to users, thereby ensuring the exposure of high-quality content. The specific technical solution is as follows:

[0005] A first aspect of this application provides a content recommendation method, the method comprising:

[0006] The system acquires content and operational information of multiple candidate recommended content, as well as user information of users to be recommended. The operational information is used to indicate the required exposure of the candidate recommended content, and the user information is used to indicate the degree of interest of the users to be recommended in the recommended content.

[0007] For each candidate recommendation content, based on the content information, operational information, and user information of the candidate recommendation content, a multi-task learning model is used to determine the predicted click-through rate, predicted browsing duration, and predicted probability of maintaining exposure. The predicted click-through rate represents the probability that the user to be recommended will click on the candidate recommendation content if it is recommended to the user to be recommended. The predicted browsing duration represents the duration that the user to be recommended will browse the candidate recommendation content if it is recommended to the user to be recommended. The predicted probability of maintaining exposure represents the probability of recommending the candidate recommendation content to the user to be recommended in order to achieve the required exposure of the candidate recommendation content.

[0008] The recommendation score for the candidate content is determined based on its predicted click-through rate, predicted browsing duration, and predicted retention rate.

[0009] Based on the recommendation scores of each candidate content in descending order, select multiple candidate content as the content to be recommended;

[0010] The content to be recommended is recommended to the user to be recommended.

[0011] In some embodiments of this application, the multi-task learning model includes multiple expert networks, a click-through rate (CTR) gating network, a browsing duration gating network, a user retention gating network, a CTR task tower, a browsing duration task tower, and a user retention task tower; the step of determining the predicted CTR, predicted browsing duration, and predicted user retention probability using the multi-task learning model based on the content information, operational information, and user information of the candidate recommendation content includes:

[0012] The content information, operational information, and user information of the candidate recommendation content are concatenated to obtain the concatenated information.

[0013] The spliced ​​information is input into each expert network and each gate network of the multi-task learning model to obtain the spliced ​​features output by each expert network after feature extraction of the spliced ​​information, and the weights of each expert network output by each gate network based on the spliced ​​information.

[0014] The multi-task learning model multiplies the weights of each expert network output by the click-through rate gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the click-through rate task tower to obtain the predicted click-through rate output by the click-through rate task tower.

[0015] The multi-task learning model multiplies the weights of each expert network output by the browsing duration gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the browsing duration task tower to obtain the predicted browsing duration output by the browsing duration task tower.

[0016] The multi-task learning model multiplies the weights of each expert network output by the quantity-preserving gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the quantity-preserving task tower to obtain the predicted quantity-preserving probability output by the quantity-preserving task tower.

[0017] In some embodiments of this application, the multi-task learning model includes a first expert network, a second expert network, a click-through rate (CTR) gating network, a browsing duration gating network, a user retention gating network, a CTR task tower, a browsing duration task tower, and a user retention task tower; the step of determining the predicted CTR, predicted browsing duration, and predicted user retention probability using the multi-task learning model based on the content information, operational information, and user information of the candidate recommendation content includes:

[0018] The content information of the candidate recommendation and the user information are concatenated to obtain the first concatenation information;

[0019] The first spliced ​​information is input into the first expert network, the click-through rate gating network, and the browsing duration gating network respectively to obtain the features output by the first expert network after feature extraction of the first spliced ​​information. The click-through rate gating network outputs the weights of each expert network based on the first spliced ​​information, and the browsing duration gating network outputs the weights of each expert network based on the first spliced ​​information.

[0020] The content information, operational information, and user information of the candidate recommendation content are concatenated to obtain the second concatenated information.

[0021] The second splicing information is input into the second expert network and the quantity-preserving gating network respectively to obtain the features output by the second expert network after feature extraction of the second splicing information, and the weights of each expert network output by the quantity-preserving gating network based on the second splicing information.

[0022] The multi-task learning model multiplies the weights of each expert network output by the click-through rate gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the click-through rate task tower to obtain the predicted click-through rate output by the click-through rate task tower.

[0023] The multi-task learning model multiplies the weights of each expert network output by the browsing duration gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is then input into the browsing duration task tower to obtain the predicted browsing duration output by the browsing duration task tower.

[0024] The multi-task learning model multiplies the weights of each expert network output by the quantity-preserving gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the quantity-preserving task tower to obtain the predicted quantity-preserving probability output by the quantity-preserving task tower.

[0025] In some embodiments of this application, obtaining the operational information of the candidate recommended content includes:

[0026] For each candidate recommended content, a preset content priority coefficient is obtained, whereby the content priority coefficient represents the priority at which the content is recommended;

[0027] Determine the target exposure for this candidate recommendation content;

[0028] Based on the content priority coefficient and the target exposure, the operational information of the candidate recommended content is constructed.

[0029] In some embodiments of this application, determining the recommendation score of the candidate recommended content based on its predicted click-through rate, predicted browsing duration, and predicted retention probability includes:

[0030] Determine the exposure requirements for the candidate recommended content;

[0031] Based on the preset mapping relationship between exposure demand and weight, the weight mapped to the exposure demand of the candidate recommended content is used as the weight of the predicted volume guarantee probability, wherein the exposure demand represented by the preset mapping relationship is positively correlated with the weight.

[0032] The weighted sum of the predicted click-through rate, predicted browsing duration, and predicted retention probability of the candidate content is determined as the recommendation score for that candidate content.

[0033] In some embodiments of this application, determining the exposure requirement of the candidate recommended content includes:

[0034] The difference between the preset target exposure for the candidate recommended content and the actual exposure of the candidate recommended content is determined, and the ratio of the difference to the target exposure is determined to obtain the exposure gap intensity, which is used as the exposure requirement;

[0035] Alternatively, the predicted quantity retention probability can be used as the exposure requirement.

[0036] A second aspect of this application provides a model training method, the method comprising:

[0037] Obtain content and operational information of multiple sample contents recommended to sample users in history, as well as user information of the sample users, wherein the operational information is used to indicate the required exposure of the sample content;

[0038] For each sample content, based on the content information, operational information, and user information of the sample content, a multi-task learning model is used to determine the predicted click-through rate, predicted browsing duration, and predicted retention probability. The predicted click-through rate represents the probability that the sample user will click on the sample content if it is recommended to the sample user; the predicted browsing duration represents the duration that the sample user will browse the sample content if it is recommended to the sample user; and the predicted retention probability represents the probability that the sample content will be recommended to the sample user to achieve the required exposure.

[0039] The joint loss value is calculated based on the predicted click-through rate, predicted browsing duration, and predicted retention probability of each sample content, as well as the actual click-through rate, actual browsing duration, and standard retention probability of each sample content.

[0040] The network parameters of the multi-task learning model are adjusted using the joint loss value. The steps of determining the predicted click-through rate, predicted browsing duration, and predicted volume retention probability using the multi-task learning model for each sample content based on the content information, operational information, and user information of the sample content are returned until the multi-task learning model converges, at which point the training is considered complete.

[0041] In some embodiments of this application, the standard exposure retention probability is the ratio of the exposure gap to the target exposure volume preset for the sample content, where the exposure gap is the difference between the target exposure volume and the actual exposure volume of the sample content; the calculation of the joint loss value based on the predicted click-through rate, predicted browsing duration, and predicted exposure retention probability of each sample content, as well as the actual click-through rate, actual browsing duration, and standard exposure retention probability of each sample content, includes:

[0042] Calculate the click-through rate loss value based on the predicted click-through rate and the actual click-through rate of each sample content;

[0043] Based on the predicted browsing time and actual browsing time of each sample content, calculate the browsing time loss value;

[0044] Based on the predicted probability of maintaining quantity and the standard probability of maintaining quantity for each sample, calculate the loss value for maintaining quantity.

[0045] Based on the preset mapping relationship between the quantity preservation probability and the weight, the weight of the average standard quantity preservation probability mapping of each sample content is determined, and used as the weight corresponding to the quantity preservation loss value. The quantity preservation probability represented by the preset mapping relationship is positively correlated with the weight.

[0046] The weighted sum of the click-through rate loss value, the browsing time loss value, and the retention loss value is taken as the joint loss value.

[0047] In some embodiments of this application, before determining the predicted click-through rate, predicted browsing duration, and predicted retention probability using a multi-task learning model for each sample content based on the content information, operational information, and user information of that sample content, the method further includes:

[0048] For each sample content, based on the content information of the sample content and the user information, a multi-task learning model is used to determine the predicted click-through rate and the predicted browsing duration.

[0049] The pre-training loss value is calculated based on the predicted click-through rate and predicted browsing duration of each sample content, as well as the actual click-through rate and actual browsing duration of each sample content.

[0050] The network parameters of the multi-task learning model are adjusted using the pre-training loss value. The steps of determining the predicted click-through rate and predicted browsing duration based on the content information of the sample content and the user information using the multi-task learning model are returned for each sample content. This process continues until the multi-task learning model converges. Then, the steps of determining the predicted click-through rate, predicted browsing duration, and predicted volume retention probability based on the content information, operational information, and user information of the sample content using the multi-task learning model are executed.

[0051] A third aspect of the embodiments of this application provides a content recommendation device, the device comprising:

[0052] The acquisition module is used to acquire content information and operational information of multiple candidate recommended content, as well as user information of the user to be recommended. The operational information is used to indicate the exposure required for the candidate recommended content, and the user information is used to indicate the degree of interest of the user to be recommended in the recommended content.

[0053] The prediction module is used to determine the predicted click-through rate, predicted browsing duration, and predicted volume retention probability for each candidate recommendation content based on the content information, operational information, and user information obtained by the acquisition module. The predicted click-through rate represents the probability that the user to be recommended will click on the candidate recommendation content if it is recommended to them; the predicted browsing duration represents the duration that the user to be recommended will browse the candidate recommendation content if it is recommended to them; and the predicted volume retention probability represents the probability that the candidate recommendation content will be recommended to the user to achieve the required exposure level.

[0054] The determination module is used to determine the recommendation score of the candidate recommended content based on the predicted click-through rate, predicted browsing time, and predicted retention probability of the candidate recommended content predicted by the prediction module.

[0055] The selection module is used to select multiple candidate recommended contents as recommended contents in descending order of their recommendation scores as determined by the determination module.

[0056] The recommendation module is used to recommend the content selected by the selection module to the user to be recommended.

[0057] In some embodiments of this application, the multi-task learning model includes multiple expert networks, a click-through rate (CTR) gating network, a browsing duration gating network, a page view retention (PV) gating network, a CTR task tower, a browsing duration task tower, and a PV retention task tower; the prediction module is specifically used for:

[0058] The content information, operational information, and user information of the candidate recommendation content are concatenated to obtain the concatenated information.

[0059] The spliced ​​information is input into each expert network and each gate network of the multi-task learning model to obtain the spliced ​​features output by each expert network after feature extraction of the spliced ​​information, and the weights of each expert network output by each gate network based on the spliced ​​information.

[0060] The multi-task learning model multiplies the weights of each expert network output by the click-through rate gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the click-through rate task tower to obtain the predicted click-through rate output by the click-through rate task tower.

[0061] The multi-task learning model multiplies the weights of each expert network output by the browsing duration gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the browsing duration task tower to obtain the predicted browsing duration output by the browsing duration task tower.

[0062] The multi-task learning model multiplies the weights of each expert network output by the quantity-preserving gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the quantity-preserving task tower to obtain the predicted quantity-preserving probability output by the quantity-preserving task tower.

[0063] In some embodiments of this application, the multi-task learning model includes a first expert network, a second expert network, a click-through rate (CTR) gating network, a browsing duration gating network, a page view retention (PV) gating network, a CTR task tower, a browsing duration task tower, and a PV retention task tower; the prediction module is specifically used for:

[0064] The content information of the candidate recommendation and the user information are concatenated to obtain the first concatenation information;

[0065] The first spliced ​​information is input into the first expert network, the click-through rate gating network, and the browsing duration gating network respectively to obtain the features output by the first expert network after feature extraction of the first spliced ​​information. The click-through rate gating network outputs the weights of each expert network based on the first spliced ​​information, and the browsing duration gating network outputs the weights of each expert network based on the first spliced ​​information.

[0066] The content information, operational information, and user information of the candidate recommendation content are concatenated to obtain the second concatenated information.

[0067] The second splicing information is input into the second expert network and the quantity-preserving gating network respectively to obtain the features output by the second expert network after feature extraction of the second splicing information, and the weights of each expert network output by the quantity-preserving gating network based on the second splicing information.

[0068] The multi-task learning model multiplies the weights of each expert network output by the click-through rate gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the click-through rate task tower to obtain the predicted click-through rate output by the click-through rate task tower.

[0069] The multi-task learning model multiplies the weights of each expert network output by the browsing duration gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is then input into the browsing duration task tower to obtain the predicted browsing duration output by the browsing duration task tower.

[0070] The multi-task learning model multiplies the weights of each expert network output by the quantity-preserving gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the quantity-preserving task tower to obtain the predicted quantity-preserving probability output by the quantity-preserving task tower.

[0071] In some embodiments of this application, the acquisition module is specifically used for:

[0072] For each candidate recommended content, a preset content priority coefficient is obtained, whereby the content priority coefficient represents the priority at which the content is recommended;

[0073] Determine the target exposure for this candidate recommendation content;

[0074] Based on the content priority coefficient and the target exposure, the operational information of the candidate recommended content is constructed.

[0075] In some embodiments of this application, the determining module is specifically used for:

[0076] Determine the exposure requirements for the candidate recommended content;

[0077] Based on the preset mapping relationship between exposure demand and weight, the weight mapped to the exposure demand of the candidate recommended content is used as the weight of the predicted volume guarantee probability, wherein the exposure demand represented by the preset mapping relationship is positively correlated with the weight.

[0078] The weighted sum of the predicted click-through rate, predicted browsing duration, and predicted retention probability of the candidate content is determined as the recommendation score for that candidate content.

[0079] In some embodiments of this application, the determining module is specifically used for:

[0080] The difference between the preset target exposure for the candidate recommended content and the actual exposure of the candidate recommended content is determined, and the ratio of the difference to the target exposure is determined to obtain the exposure gap intensity, which is used as the exposure requirement;

[0081] Alternatively, the predicted quantity retention probability can be used as the exposure requirement.

[0082] A fourth aspect of this application provides a model training apparatus, the apparatus comprising:

[0083] The acquisition module is used to acquire content information and operational information of multiple sample contents recommended to sample users in history, as well as user information of the sample users. The operational information is used to indicate the exposure required for the sample content.

[0084] The prediction module is used to determine the predicted click-through rate, predicted browsing duration, and predicted retention probability for each sample content based on the content information, operational information, and user information obtained by the acquisition module. The predicted click-through rate represents the probability that a sample user will click on the sample content if it is recommended to them; the predicted browsing duration represents the duration a sample user will browse the sample content if it is recommended to them; and the predicted retention probability represents the probability of recommending the sample content to the sample user to achieve the required exposure level.

[0085] The calculation module is used to calculate the joint loss value based on the predicted click-through rate, predicted browsing time and predicted volume retention probability of each sample content predicted by the prediction module, as well as the actual click-through rate, actual browsing time and standard volume retention probability of each sample content.

[0086] The adjustment module is used to adjust the network parameters of the multi-task learning model using the joint loss value calculated by the calculation module, and to call the prediction module to perform the steps of determining the predicted click-through rate, predicted browsing duration, and predicted volume retention probability for each sample content based on the content information, operational information, and user information of the sample content using the multi-task learning model, until the multi-task learning model converges, at which point training is determined to be complete.

[0087] In some embodiments of this application, the standard exposure probability is the ratio of the exposure gap to the target exposure for the sample content, where the exposure gap is the difference between the target exposure and the actual exposure of the sample content; the calculation module is specifically used for:

[0088] Calculate the click-through rate loss value based on the predicted click-through rate and the actual click-through rate of each sample content;

[0089] Based on the predicted browsing time and actual browsing time of each sample content, calculate the browsing time loss value;

[0090] Based on the predicted probability of maintaining quantity and the standard probability of maintaining quantity for each sample, calculate the loss value for maintaining quantity.

[0091] Based on the preset mapping relationship between the quantity preservation probability and the weight, the weight of the average standard quantity preservation probability mapping of each sample content is determined, and used as the weight corresponding to the quantity preservation loss value. The quantity preservation probability represented by the preset mapping relationship is positively correlated with the weight.

[0092] The weighted sum of the click-through rate loss value, the browsing time loss value, and the retention loss value is taken as the joint loss value.

[0093] In some embodiments of this application, the prediction module is further configured to determine the predicted click-through rate, predicted browsing duration, and predicted volume retention probability for each sample content using a multi-task learning model, before determining the predicted click-through rate and predicted browsing duration for each sample content based on the content information, operational information, and user information of the sample content.

[0094] The calculation module is also used to calculate the pre-training loss value based on the predicted click-through rate and predicted browsing time of each sample content, as well as the actual click-through rate and actual browsing time of each sample content.

[0095] The adjustment module is further configured to adjust the network parameters of the multi-task learning model using the pre-trained loss value, and call the prediction module to execute the steps of determining the predicted click-through rate and predicted browsing duration for each sample content based on the content information of the sample content and the user information using the multi-task learning model, until the multi-task learning model converges, and then call the prediction module to execute the steps of determining the predicted click-through rate, predicted browsing duration, and predicted volume retention probability for each sample content based on the content information, operational information, and user information using the multi-task learning model.

[0096] A fifth aspect of the embodiments of this application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0097] Memory, used to store computer programs;

[0098] When a processor executes a program stored in memory, it implements any of the content recommendation methods or model training methods described above.

[0099] A sixth aspect of the embodiments of this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the content recommendation method and model training method described in any of the preceding claims.

[0100] A seventh aspect of the embodiments of this application also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the content recommendation method or model training method described in any of the above claims.

[0101] The content recommendation method, model training method, apparatus, electronic device, and medium provided in this application can, for each candidate recommended content, determine the predicted click-through rate, predicted browsing duration, and predicted exposure probability using a multi-task learning model based on the content information, operational information, and user information of the user to be recommended. This achieves the prediction of the user's click-through rate and browsing duration for the candidate recommended content, as well as the probability of recommending the candidate recommended content to the user to achieve the required exposure. Therefore, candidate recommended content with higher predicted click-through rate, predicted browsing duration, and predicted exposure probability can be selected as the recommended content and recommended to the user. In other words, by considering predicted click-through rate and predicted browsing duration when determining the recommended content, this application increases the likelihood that the recommended content will be liked by the user, and by considering the predicted exposure probability, it makes the exposure of the candidate recommended content closer to the required exposure. Thus, this application achieves a richer selection of recommended content while ensuring user preference, and guarantees the exposure of each piece of content. Attached Figure Description

[0102] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0103] Figure 1 A flowchart illustrating the first content recommendation method provided in this application embodiment;

[0104] Figure 2 This is a schematic diagram of the structure of a multi-task learning model provided in an embodiment of this application;

[0105] Figure 3 A flowchart illustrating the second content recommendation method provided in this application embodiment;

[0106] Figure 4 This is a schematic diagram of the structure of another multi-task learning model provided in an embodiment of this application;

[0107] Figure 5 A flowchart illustrating the third content recommendation method provided in this application embodiment;

[0108] Figure 6 A flowchart illustrating a model training method provided in this application embodiment;

[0109] Figure 7 A flowchart illustrating another model training method provided in this application embodiment;

[0110] Figure 8 An exemplary schematic diagram of a content recommendation process provided in an embodiment of this application;

[0111] Figure 9 This is a schematic diagram of the structure of a content recommendation device provided in an embodiment of this application;

[0112] Figure 10 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application;

[0113] Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0114] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0115] To avoid over-reliance on users' historical interests when recommending content and to ensure the exposure of high-quality content, this application provides a content recommendation method. This method is applied to electronic devices, such as servers, desktop computers, or laptops—devices with data processing capabilities. See also... Figure 1 The content recommendation method provided in this application includes the following steps:

[0116] S101. Obtain content and operational information of multiple candidate recommended content, as well as user information of users to be recommended.

[0117] The candidate recommended content can be: long videos, short videos, advertisements, product links, audio, images, news, or documents, etc., and this application embodiment does not specifically limit it. For example, long videos include: movies, TV series, and variety shows, etc.

[0118] Content information includes: content title, description, and historical statistical parameters. Historical statistical parameters include: click-through rate, cumulative viewing time, number of user interactions, and actual exposure of candidate recommended content in the previous statistical period. Number of user interactions includes: number of likes, shares, and comments; actual exposure includes the actual number of times the content was recommended, i.e., the actual number of times it was displayed on the user's device.

[0119] Operational information is used to indicate the amount of exposure required for candidate recommended content.

[0120] User information is used to indicate the degree of interest of the user to be recommended in the recommended content. For example, user information includes: click-through rate, cumulative browsing time, and number of interactions for each recommended piece of content. User information may also include other information, such as user profile information, etc., which are not specifically limited in this embodiment of the application.

[0121] S102. For each candidate recommended content, based on the content information, operational information and user information of the candidate recommended content, use a multi-task learning model to determine the predicted click-through rate, predicted browsing duration and predicted volume retention probability.

[0122] The predicted click-through rate (CTR) represents the probability that a user will click on the recommended content if it is presented to them. CTR is further denoted as the click-through rate.

[0123] Predicted browsing time indicates the estimated time a user will spend browsing the recommended content if it is suggested to them.

[0124] The predicted probability of ensuring exposure represents the probability that a candidate content will be recommended to users to achieve the required exposure volume. In other words, the predicted probability of ensuring exposure reflects the intensity of the exposure demand for the candidate content in this instance, i.e., the urgency of exposing the candidate content to users in this instance.

[0125] Among them, the multi-task learning (MTL) model can be a multi-gate mixture-of-experts (MMoE) model or a progressive layered extraction (PLE) model, etc.

[0126] S103. Based on the predicted click-through rate, predicted browsing duration, and predicted retention probability of the candidate recommended content, determine the recommendation score of the candidate recommended content.

[0127] The predicted click-through rate, predicted browsing duration, and predicted probability of maintaining views for a candidate recommendation can be summed to obtain the recommendation score for that candidate recommendation.

[0128] Alternatively, the recommendation score for a candidate content can be the weighted sum of its predicted click-through rate (CTR), predicted viewing time, and predicted retention rate. The weights for each of these factors can be pre-set based on the specific business needs of the scenario. For example, the weight for predicted CTR could be 0.4, predicted viewing time 0.3, and predicted retention rate 0.3.

[0129] Alternatively, the method for determining the recommendation score of candidate content can be found in the description below.

[0130] S104. Select multiple candidate recommended contents as recommended contents in descending order of their recommendation scores.

[0131] You can select a preset number of candidate recommended content items as the content to be recommended, based on their recommendation scores from highest to lowest. The preset number can be pre-set according to the actual business needs of the scenario, for example, a preset number of 20.

[0132] Alternatively, candidate recommended content can be selected from those with recommendation scores exceeding a preset score, in descending order of their recommendation scores. The preset score can be pre-set based on the specific business needs of the actual scenario.

[0133] Alternatively, other methods can be used to select content to be recommended, and this application embodiment does not specifically limit this method.

[0134] S105. Recommend content to be recommended to users.

[0135] Electronic devices can send content to be recommended to the user's device, so that the user's device can display the content to be recommended.

[0136] In addition, when sending content to be recommended, electronic devices can also send sorting information for the content to be recommended. The sorting information indicates the ranking of each content to be recommended in descending order of recommendation score. This allows user devices to display the content to be recommended in the order it is presented, thereby increasing the click-through rate and browsing time of content that users are more likely to be interested in, as well as giving priority exposure to high-quality content, thus increasing the click-through rate and browsing time of high-quality content.

[0137] The content recommendation method provided in this application can, for each candidate recommended content, determine the predicted click-through rate, predicted browsing duration, and predicted exposure probability using a multi-task learning model based on the content information, operational information, and user information of the user to be recommended. This achieves the prediction of the user's click-through rate and browsing duration for the candidate recommended content, as well as the probability of recommending the candidate recommended content to the user to achieve the required exposure. Therefore, candidate recommended content with higher predicted click-through rate, predicted browsing duration, and predicted exposure probability can be selected as the recommended content and recommended to the user. In other words, by considering predicted click-through rate and predicted browsing duration when determining the recommended content, this application increases the likelihood that the recommended content will be liked by the user, and by considering the predicted exposure probability, it makes the exposure of the candidate recommended content closer to the required exposure. Thus, this application achieves a richer selection of recommended content while ensuring user preference, and guarantees the exposure of each piece of content.

[0138] The following provides a detailed description of the content recommendation method provided in the embodiments of this application.

[0139] In some embodiments of this application, the electronic device may execute the above-described S101 when it detects that the recommended conditions are met.

[0140] The recommended conditions may include: the electronic device receiving a page retrieval request from the user terminal for a specified page. The specified page can be the homepage of the content provider platform or a specified channel page, such as a movie channel page or a TV series channel page. The user terminal can access the content provider platform's page through applications (APPs), mini-programs, or browsers.

[0141] The recommended conditions may also include other conditions, which are not specifically limited in this embodiment. For example, the recommended conditions may also include being in a preset test period at the current time.

[0142] When recommendation criteria are met, the electronic device can identify the user belonging to the user terminal to which the recommendation criteria apply as the user to be recommended. For example, the electronic device can obtain user identification information from the page retrieval request and use the user identified by the user identification information as the user to be recommended.

[0143] In this embodiment of the application, the multiple candidate recommended contents in S101 include: multiple interest contents that are highly likely to be of interest to the user to be recommended.

[0144] When filtering candidate recommended content, electronic devices can select multiple pieces of content from a content library based on user information, choosing those with a probability higher than a threshold of interest for the user. For example, user information can be converted into user vectors, and the vector similarity between the user vector and the content vectors of each piece of content in the content library can be determined. Multiple pieces of content can then be selected as interest content in descending order of vector similarity. Here, the content vector can be a vector obtained by vector transformation of the content information.

[0145] The multiple candidate recommendations in S101 can also include: multiple high-quality content items that require a high level of exposure.

[0146] Electronic devices can filter multiple high-quality content items based on operational information. For example, they can first filter content in the content library where the difference between the target exposure and the actual exposure is greater than a preset difference. Here, the target exposure represents the required exposure for the content, which can be preset according to the needs of the actual business scenario. Then, from the filtered content, multiple items can be selected as high-quality content in descending order of the difference, or vice versa, from the filtered content in descending order of the target exposure.

[0147] This application also supports other methods for filtering candidate recommended content, but this application does not specifically limit these methods.

[0148] In this embodiment of the application, the operational information of the candidate recommended content in S101 may include: the intensity of the exposure gap and the content priority coefficient. Accordingly, the method of obtaining the operational information of the candidate recommended content in S101 includes the following steps:

[0149] Step 1: For each candidate recommended content, obtain the preset content priority coefficient for that candidate recommended content.

[0150] The content priority coefficient represents the priority at which content is recommended. That is, the higher the content priority coefficient, the higher the priority of the content being recommended to the user, and vice versa.

[0151] The content priority coefficient can be set according to the exposure required by the content in the actual application scenario. The content priority coefficient is positively correlated with the exposure required by the content. For example, the value range of the content priority coefficient is [0.2, 1.0]. The content priority coefficient of each piece of content is a value within this range. The higher the content priority coefficient, the higher the exposure required by the content. Conversely, the lower the content priority coefficient, the lower the exposure required by the content.

[0152] Step 2: Determine the target exposure for the candidate recommended content.

[0153] The target exposure for this candidate recommendation content represents the amount of exposure required for the candidate recommendation content based on the needs of the actual application scenario. The target exposure can be pre-configured by staff.

[0154] Step 3: Based on the content priority coefficient and target exposure, construct the operational information for the candidate recommended content.

[0155] Electronic devices can combine content priority coefficients and target exposure levels to create operational information for the candidate recommended content.

[0156] Electronic devices can also construct operational information based on other information, such as the actual exposure of candidate recommended content, but this application does not specifically limit this.

[0157] This application's embodiments can determine the target exposure volume of candidate recommended content, thereby reflecting the required exposure volume of the candidate recommended content. Furthermore, this application's embodiments can also obtain the content priority coefficient of the candidate recommended content, thereby reflecting the priority differences in how different content is recommended to users, allowing higher-priority content to be recommended to users first. Based on the content priority coefficient and the target exposure volume, this application's embodiments construct operational information for candidate recommended content, enabling the operational information to reflect the content's exposure needs and its recommendation priority, thus allowing for more accurate prediction of the content's retention probability in the future.

[0158] In this application embodiment, the content priority coefficient and target exposure of each content support minute-level updates. When constructing operational information, electronic devices can obtain the most recently updated information to ensure the timeliness of the constructed operational information and adapt to the real-time requirements of trending content.

[0159] After obtaining various information in S101, the above-mentioned S102 makes predictions based on this information in the following two ways:

[0160] Before introducing the first implementation method of S102, let's first explain the structure of the multi-task learning model. See [link / reference] Figure 2 The multi-task learning model includes: multiple expert networks, a click-through rate (CTR) gate network, a browsing duration gate network, a page view maintenance (PV) gate network, a CTR task tower, a browsing duration task tower, and a PV maintenance task tower. Input data can be fed separately into each expert network and each gate network. Figure 2 The circle in the diagram represents: a weighted sum is calculated based on the outputs of each expert network and the weights of the outputs of each expert network from a gating network, and this weighted sum is used as input data for a task tower.

[0161] Figure 2 The example shows three expert networks. In real-world applications, the number of expert networks included in the multi-task learning model used in this application is not limited to these.

[0162] In the first implementation of S102, the multi-task learning model can be an MMoE model, see [link / reference]. Figure 3 S102 includes the following steps:

[0163] S1021. The content information, operational information and user information of the candidate recommended content are spliced ​​together to obtain spliced ​​information.

[0164] In this process, electronic devices can first preprocess content information, operational information, and user information. Preprocessing includes, for example, removing duplicate data, filling in missing values, handling outlier data, and feature encoding, thereby transforming the information into continuous features that can be processed by a multi-task learning model. Then, the preprocessed information is concatenated to obtain the concatenated information.

[0165] S1022. Input the splicing information into each expert network and each gated network of the multi-task learning model to obtain the splicing features output by each expert network after extracting features from the splicing information, and the weights of each expert network output by each gated network based on the splicing information.

[0166] Although the input information is the same for each expert network, each expert network can learn to pay different attention to each piece of information in the spliced ​​information during the pre-training process. Therefore, different expert networks can output different splicing features.

[0167] For example, a multi-task learning model includes three expert networks. The first expert network focuses more on the content information within the concatenated information during feature extraction; the second expert network focuses more on both the content information and user information within the concatenated information; and the third expert network focuses more on operational information. This results in different outputs from the three expert networks.

[0168] Similarly, although the input information is the same for each gating network, during pre-training, each gating network can learn to pay different amounts of attention to various information items in the spliced ​​information. Therefore, different gating networks can output different weight combinations. Specifically, the gating networks can process the spliced ​​information through an attention mechanism, thereby outputting the corresponding weights for each expert network.

[0169] S1023. The weights of each expert network output by the click-through rate gating network are multiplied by the concatenated features output by the expert network through the multi-task learning model to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature. This weighted sum feature is then input into the click-through rate task tower to obtain the predicted click-through rate output by the click-through rate task tower.

[0170] The predicted click-through rate (CTR) ranges from [0,1]. A higher CTR indicates a higher likelihood that the user will click on the recommended content if it is recommended to them. Conversely, a lower CTR indicates a lower likelihood that the user will click on the recommended content if it is recommended to them.

[0171] S1024. Using a multi-task learning model, the weights corresponding to each expert network output by the browsing duration gating network are multiplied by the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature. This weighted sum feature is then input into the browsing duration task tower to obtain the predicted browsing duration output by the browsing duration task tower.

[0172] The predicted browsing time ranges from [0,1], representing the normalized value of the browsing time of the candidate recommended content if it is recommended to the user to be recommended.

[0173] S1025. Through the multi-task learning model, the weights corresponding to each expert network output by the quantity-preserving gating network are multiplied by the concatenated features output by the expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature. This weighted sum feature is input into the quantity-preserving task tower to obtain the predicted quantity-preserving probability output by the quantity-preserving task tower.

[0174] The predicted probability of maintaining the quantity ranges from [0,1], representing the probability of recommending the candidate content to the user to be recommended in order to achieve the required exposure quantity for the candidate content. In other words, it represents the probability of recommending the candidate content to the user to complete the exposure quantity task for the candidate content.

[0175] S1023 to S1025 can be executed in parallel.

[0176] In the multi-task learning model of this application embodiment, not only are click-through rate (CTR) prediction and browsing duration prediction branches set up, but a volume preservation task branch is also innovatively extended, constructing a three-task model that includes CTR prediction, browsing duration prediction, and volume preservation prediction. In the multi-task model, each task tower can share the output results of each expert network. By adjusting the weights of the output results of each expert network through different gating networks, the features extracted by different expert networks are dynamically selected to adapt to different prediction tasks. This approach simplifies the network structure of the multi-task model while achieving multi-task prediction, reduces the model's computational load, and improves the efficiency of multi-task prediction.

[0177] Before introducing the second implementation method of S102, let's first explain the structure of the multi-task learning model. See [link / reference] Figure 4 The multi-task learning model includes: a first expert network, a second expert network, a click-through rate (CTR) gating network, a browsing duration gating network, a user retention gating network, a CTR task tower, a browsing duration task tower, and a user retention task tower. Content information and user information can be concatenated and then input into the respective first expert network, browsing duration gating network, and CTR gating network. Similarly, content information, operational information, and user information can be concatenated and then input into the second expert network and the user retention gating network, respectively. Figure 4The circle in the diagram represents a weighted sum based on the outputs of multiple expert networks and the weights of the outputs of each expert network from a gating network. This weighted sum is then used as input data for a task tower.

[0178] Figure 4 The example shows two first expert networks and one second expert network. In practical application scenarios, the number of expert networks included in the multi-task learning model used in this application embodiment is not limited to this.

[0179] In the second implementation of S102, the multi-task learning model can be a PLE model, see [link / reference]. Figure 5 S102 includes the following steps:

[0180] S1026. The content information and user information of the candidate recommendation content are concatenated to obtain the first concatenated information. The first concatenated information is then input into the first expert network, the click-through rate gating network, and the browsing time gating network to obtain the features output by the first expert network after feature extraction of the first concatenated information, the weights of each expert network output by the click-through rate gating network based on the first concatenated information, and the weights of each expert network output by the browsing time gating network based on the first concatenated information.

[0181] The specific implementation of S1026 can be found in S1021 and S1022 above, and will not be repeated here.

[0182] Since both click-through rate (CTR) prediction and browsing duration prediction tasks have low correlation with operational information but high correlation with content and user information, the weights of the second expert network output by the CTR gating network and browsing duration gating network can be smaller than the weights of the first expert network. For example, the weight of the second expert network can be 0. This reduces the impact of operational information on the prediction results of the CTR and browsing duration prediction tasks, thereby improving the accuracy of CTR and browsing duration prediction.

[0183] S1027. The content information, operation information and user information of the candidate recommended content are concatenated to obtain the second concatenated information. The second concatenated information is then input into the second expert network and the quantity-preserving gating network to obtain the concatenated features output by the second expert network after feature extraction of the second concatenated information, and the weights of each expert network output by the quantity-preserving gating network based on the second concatenated information.

[0184] The specific implementation of S1027 can be referred to S1021 and S1022 above, and will not be repeated here. S1026 and S1027 can be executed sequentially or in parallel, and this embodiment does not specifically limit this.

[0185] Since the volume retention prediction task is more closely related to the operational information of the content, the weight of the second expert network output by the volume retention gating network can be higher than that of the first expert network. This allows the volume retention task tower to focus more on operational information and less on content and user information when predicting the probability of volume retention, thus improving the accuracy of the prediction of the probability of volume retention.

[0186] S1028. Through the multi-task learning model, the weights corresponding to each expert network output by the click-through rate gating network are multiplied by the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature. This weighted sum feature is then input into the click-through rate task tower to obtain the predicted click-through rate output by the click-through rate task tower.

[0187] The specific implementation of S1028 can be referred to S1023 above, and will not be repeated here.

[0188] S1029. Using a multi-task learning model, the weights corresponding to each expert network output by the browsing duration gating network are multiplied by the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature. This weighted sum feature is then input into the browsing duration task tower to obtain the predicted browsing duration output by the browsing duration task tower.

[0189] The specific implementation of S1029 can be referred to S1024 above, and will not be repeated here.

[0190] S10210. Through the multi-task learning model, the weights corresponding to each expert network output by the quantity-preserving gating network are multiplied by the concatenated features output by the expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature. This weighted sum feature is input into the quantity-preserving task tower to obtain the predicted quantity-preserving probability output by the quantity-preserving task tower.

[0191] The specific implementation of S10210 can be referred to S1025 above, and will not be repeated here.

[0192] S1028 to S10210 can be executed in parallel.

[0193] In this embodiment, the multi-task learning model not only includes click-through rate (CTR) prediction and browsing duration prediction branches, but also innovatively extends this to include a volume preservation task branch, constructing a three-task model encompassing CTR prediction, browsing duration prediction, and volume preservation prediction. In this multi-task model, each task leader can share the output results of its respective expert network. By adjusting the weights of each expert network's output results through different gating networks, the features extracted by different expert networks are dynamically selected to adapt to different prediction tasks. This approach simplifies the network structure of the multi-task model while reducing computational load and improving multi-task prediction efficiency.

[0194] Furthermore, the embodiments of this application can input different information into different expert networks, making the information input to the downstream task tower more flexible. This can reduce the impact of operational information on click-through rate prediction tasks and browsing time prediction tasks, and can also balance the attention of quantity prediction tasks to content information, user information, and operational information, thereby making the prediction accuracy of each prediction task higher.

[0195] After obtaining the predicted click-through rate, predicted browsing duration, and predicted retention probability of the candidate recommended content, the method for determining the recommendation score of the candidate recommended content in S103 above includes the following steps:

[0196] Step 1: Determine the exposure requirements for the candidate recommended content.

[0197] In one implementation, the difference between the preset target exposure for the candidate recommended content and the actual exposure of the candidate recommended content can be determined, and the ratio of the difference to the target exposure can be determined to obtain the exposure gap intensity, which is used as the exposure requirement. That is, exposure gap intensity = (target exposure - actual exposure) / target exposure.

[0198] The actual exposure of the candidate recommended content refers to the actual cumulative exposure after the candidate recommended content goes live. In this embodiment of the application, the actual exposure of each content supports minute-level updates. When determining the intensity of the exposure gap, the electronic device can obtain the actual exposure after the most recent update to ensure the timeliness of the exposure gap intensity.

[0199] In another implementation, the predicted quantity probability output by the multi-task learning model can be used as the exposure requirement.

[0200] Using the above method, the embodiments of this application can determine the difference between the target exposure and the actual exposure, thereby reflecting the current exposure gap of the candidate recommended content, and determine the ratio of the difference to the target exposure to obtain the exposure gap intensity, that is, the proportion of the current unrealized exposure to the total exposure required. This can reflect the urgency of the candidate recommended content for exposure. Therefore, the exposure gap intensity can be used as the exposure demand to ensure that the exposure demand is more in line with the actual exposure demand of the candidate recommended content.

[0201] Alternatively, since the predicted probability of maintaining exposure determined by the multi-task learning model can reflect the intensity of the exposure demand for the candidate recommended content in this instance, that is, the urgency of exposing the candidate recommended content to the users to be recommended, the predicted probability of maintaining exposure output by the multi-task learning model can be directly used as the exposure demand, which improves the efficiency of determining the exposure demand and thus improves the efficiency of content recommendation.

[0202] Step II: Based on the preset mapping relationship between exposure demand and weight, the weight mapped to the exposure demand of the candidate recommended content is used as the weight for predicting the probability of maintaining the volume.

[0203] Among them, the exposure demand represented by the preset mapping relationship is positively correlated with the weight.

[0204] For example, assuming the exposure requirement ranges from [0,1], the default value of the weight corresponding to the exposure requirement is 0.1. When the exposure requirement is greater than 0.3 and less than or equal to 0.5, the corresponding weight is 0.2, and when the exposure requirement is greater than 0.5, the corresponding weight is 0.3.

[0205] Step III: Determine the weighted sum of the predicted click-through rate, predicted browsing duration, and predicted retention probability of the candidate recommended content, and use this sum as the recommendation score for the candidate recommended content.

[0206] Using the above method, this application embodiment can determine the weight of the predicted volume retention probability based on exposure demand, and the exposure demand is positively correlated with the weight. That is, the higher the exposure demand, the higher the weight of the predicted volume retention probability. This achieves dynamic adjustment of the importance of the predicted volume retention probability in calculating the recommendation score based on the magnitude of exposure demand, reducing the cost of manually adjusting the weight. Moreover, this application embodiment can use a default value for the weight of the predicted volume retention probability when the exposure demand is low, and only use a larger value for the weight of the predicted volume retention probability when the exposure demand is high. This ensures that when recommending content to users, attention is increased only when the exposure demand is high, reducing the situation where the predicted volume retention probability is overemphasized when determining recommended content, while ignoring user interests.

[0207] Based on the same inventive concept, embodiments of this application also provide a model training method, which is applied to an electronic device, such as a server, desktop computer, or laptop computer, or other device with data processing capabilities. The electronic device used for the content recommendation method and the electronic device used for the model training method may be the same electronic device or different electronic devices; this application does not specifically limit this.

[0208] See Figure 6 The model training method also provided in this application includes the following steps:

[0209] S601. Obtain content and operational information of multiple sample contents recommended to sample users in the past, as well as user information of the sample users.

[0210] The operational information is used to indicate the amount of exposure required for the sample content.

[0211] The method for obtaining the content information and operational information of the sample content in S601 can refer to the method for obtaining the content information and operational information of the candidate recommended content in S101 above. The method for obtaining the user information of the sample users in S601 can refer to the method for obtaining the user information of the users to be recommended in S101 above. It will not be repeated here.

[0212] S602. For each sample content, based on the content information, operational information and user information of the sample content, use a multi-task learning model to determine the predicted click-through rate, predicted browsing duration and predicted volume retention probability.

[0213] Here, predicted click-through rate (CTR) represents the probability that a sample user will click on the sample content if it is recommended to them. Predicted viewing time represents the duration a sample user will view the sample content if it is recommended to them. Predicted reach probability represents the probability of recommending the sample content to the sample user to achieve the required exposure.

[0214] For details on the specific implementation of S602, please refer to the relevant description of S102 above, which will not be repeated here.

[0215] S603. Calculate the joint loss value based on the predicted click-through rate, predicted browsing duration, and predicted retention probability of each sample content, as well as the actual click-through rate, actual browsing duration, and standard retention probability of each sample content.

[0216] The standard quantity retention probability can be the quantity retention probability set by the user for the sample content, or it can be the exposure gap intensity of the sample content. The specific implementation of S603 is described below.

[0217] S604. Adjust the network parameters of the multi-task learning model using the joint loss value, and return to the steps in S602 for each sample content, based on the content information, operation information and user information of the sample content, to determine the predicted click-through rate, predicted browsing duration and predicted volume retention probability using the multi-task learning model, until the multi-task learning model converges, and the training is considered complete.

[0218] Electronic devices can employ gradient descent to adjust the network parameters of a multi-task learning model using the joint loss value and determine whether the multi-task learning model has converged. For example, it can determine whether the number of iterations of the multi-task learning model has reached a preset number; if so, the multi-task learning model is considered converged; otherwise, it is considered non-converged. Another example is whether the currently calculated joint loss value is less than a preset loss value; if so, the multi-task learning model is considered converged; otherwise, it is considered non-converged. Yet another example is whether the difference between the joint loss values ​​calculated in the previous preset number of iterations is less than a preset difference; if so, the multi-task learning model is considered converged; otherwise, it is considered non-converged. Alternatively, other methods can be used to determine whether the multi-task learning model has converged; this embodiment does not specifically limit this method. When the multi-task learning model is determined to have converged, the model training is considered complete; when the multi-task learning model is determined to have not converged, the process returns to S602, thus entering the next iteration.

[0219] In this process, each iteration can use a different batch of sample content, thereby improving the generalization of the multi-task learning model.

[0220] The model training method provided in this application can, for each sample content, determine the predicted click-through rate (CTR), predicted browsing duration, and predicted retention probability using a multi-task learning model based on the content information, operational information, and user information of the sample content. This achieves the prediction of the user's CTR, browsing duration, and the probability of recommending the sample content to the user to achieve the required exposure. Therefore, a joint loss value can be calculated based on the differences between the predicted CTR and actual CTR, the predicted browsing duration and actual browsing duration, and the predicted retention probability and standard retention probability for each sample content. This loss value is then used to train the multi-task learning model, making the predicted CTR, browsing duration, and retention probability more accurate. This allows the trained multi-task learning model to more accurately predict the CTR, browsing duration, and retention probability of content, thereby recommending content that users are more likely to be interested in and ensuring that the content's exposure is closer to the required exposure. Therefore, this application embodiment achieves a balance between ensuring that recommended content is liked by users, enriching the recommended content, and guaranteeing the exposure of each piece of content.

[0221] The model training method provided in the embodiments of this application will be described in detail below:

[0222] The above-mentioned S603 calculates the joint loss value based on the predicted click-through rate, predicted browsing duration, and predicted retention probability of each sample content, as well as the actual click-through rate, actual browsing duration, and standard retention probability of each sample content. The method includes the following steps:

[0223] Step 1: Calculate the click-through rate loss value based on the predicted click-through rate and the actual click-through rate of each sample content.

[0224] The actual click-through rate (CTR) represents the actual click-through rate of sample users on sample content. The CTR loss can be calculated using methods such as the binary cross-entropy loss function, the cross-entropy loss function, or the mean squared error loss function.

[0225] Step 2: Calculate the browsing time loss value based on the predicted browsing time and actual browsing time of each sample content.

[0226] The actual browsing time refers to the cumulative browsing time of sample users on the sample content. The browsing time loss can be calculated using the mean squared error loss function, the binary cross-loss function, or the cross-entropy loss function, among others.

[0227] Step 3: Calculate the loss value for maintaining volume based on the predicted probability and standard probability of maintaining volume for each sample.

[0228] The standard retention probability is defined as the ratio of the exposure gap to the target exposure for the sample content, where the exposure gap is the difference between the target exposure and the actual exposure of the sample content. In other words, the standard retention probability is the intensity of the aforementioned exposure gap.

[0229] The loss value for maintaining quantity can be calculated using the mean squared error loss function, the binary cross loss function, or the cross-entropy loss function, etc.

[0230] Step 4: Based on the preset mapping relationship between the quantity preservation probability and the weight, determine the weight of the average standard quantity preservation probability mapping for each sample content, and use it as the weight corresponding to the quantity preservation loss value.

[0231] The average standard probability of preserving quantity can be calculated for a batch of sample content used in this iteration, and then the weights mapped by the average standard probability of preserving quantity can be used as the weights corresponding to the preserving quantity loss values.

[0232] The pre-defined mapping relationship represents a positive correlation between the probability of maintaining volume and the weight. For example, assuming the average probability of maintaining volume ranges from [0,1], the default value of the weight corresponding to the average probability of maintaining volume is 0.1. When the average standard probability of maintaining volume is greater than 0.3 and less than or equal to 0.5, the corresponding weight is 0.2, and when the average standard probability of maintaining volume is greater than 0.5, the corresponding weight is 0.3.

[0233] Step 5: Take the weighted sum of the click-through rate loss, browsing time loss, and retention loss as the joint loss value.

[0234] The weights corresponding to the click-through rate loss value and the browsing time loss value are preset fixed values.

[0235] For example, if the weight of the click-through rate loss value is 0.6, the weight of the browsing time loss value is 0.3, and the weight of the retention loss value is 0.1, then the combined loss value = click-through rate loss value × 0.6 + browsing time loss value × 0.3 + retention loss value × 0.1.

[0236] This application embodiment utilizes click-through rate (CTR) loss, browsing duration loss, and retention loss to jointly train a multi-task learning model. During model training, the weights of the retention loss values ​​are flexibly adjusted based on the standard retention probability. This dynamically adjusts the importance of the retention task in the joint training task according to the size of the exposure gap, reducing the cost of manually adjusting weights. This application embodiment only increases the multi-task learning model's focus on the retention task when the exposure gap is large, reducing the model's long-term overemphasis on the retention task and minimizing the neglect of the primary CTR and browsing duration tasks.

[0237] Understandably, in another implementation of S603 above, the click-through rate (CTR) loss value, browsing time loss value, and retention loss value each have their own preset weights. Based on these preset weights, a weighted sum of the CTR loss value, browsing time loss value, and retention loss value can be calculated as the joint loss value. This method can improve the efficiency of determining the joint loss value.

[0238] In some embodiments of this application, before determining the predicted click-through rate, predicted browsing duration, and predicted retention probability using the multi-task learning model in S602 above, the electronic device may also pre-train the multi-task learning model, see [link to relevant documentation]. Figure 7 The pre-training phase includes the following steps:

[0239] S605. For each sample content, based on the content information and user information of the sample content, use a multi-task learning model to determine the predicted click-through rate and predicted browsing duration.

[0240] The methods used by multi-task learning models to determine predicted click-through rate and predicted browsing duration can be found in the description above, and will not be repeated here.

[0241] S606. Calculate the pre-training loss value based on the predicted click-through rate and predicted browsing duration of each sample content, as well as the actual click-through rate and actual browsing duration of each sample content.

[0242] Electronic devices can calculate the click-through rate loss value based on the predicted click-through rate and the actual click-through rate of each sample content, and calculate the browsing time loss value based on the predicted browsing time and the actual browsing time of each sample content. Then, the weighted sum of the click-through rate loss value and the browsing time loss value is calculated as the pre-training loss value.

[0243] The calculation methods for click-through rate loss and browsing time loss can be found in the above description and will not be repeated here.

[0244] S607. Adjust the network parameters of the multi-task learning model using the pre-trained loss value, and return to S605. For each sample content, based on the content information and user information of the sample content, use the multi-task learning model to determine the predicted click-through rate and the predicted browsing time, until the multi-task learning model converges, and then execute S602.

[0245] The specific implementation of S607 is the same as that of S604 above, and can be referred to S604 above. It will not be repeated here.

[0246] This application's embodiments can first pre-train the multi-task learning model based on click-through rate (CTR) prediction and browsing duration prediction tasks, and then jointly train it based on CTR prediction, browsing duration prediction, and volume preservation prediction tasks, forming a progressive training process. The pre-training process involves fewer tasks, thus achieving high training efficiency. Furthermore, during pre-training, the multi-task learning model can learn to predict CTR and browsing duration, thereby shortening the time required for subsequent joint training and improving the overall efficiency of model training.

[0247] Furthermore, in the training process of the multi-task learning model, this application embodiment can also introduce adversarial training. Adversarial training is a training method that improves the robustness of the model by introducing adversarial examples. Adversarial examples are samples that cause the model to make incorrect predictions by adding small perturbations to the training samples. By introducing adversarial training, this application embodiment enables the multi-task learning model to learn more robust feature representations, thereby improving the predictive performance of the multi-task learning model in the face of noise or malicious attacks.

[0248] See Figure 8 The following example illustrates the overall process of content recommendation involved in the embodiments of this application, based on practical application scenarios:

[0249] S801. Obtain content and operational information of multiple sample contents recommended to sample users in the past, as well as user information of the sample users.

[0250] S802. Use the content information and user information of each sample to pre-train the multi-task learning model.

[0251] S803. Utilize the content and operational information of each sample, as well as the user information of the sample users, to jointly train the multi-task learning model.

[0252] S804. Upon receiving a page retrieval request from a user terminal, obtain content information and operational information of multiple candidate recommended content, as well as user information of the user to be recommended to the user terminal.

[0253] S805. For each candidate recommended content, based on the content information, operational information and user information of the candidate recommended content, a multi-task learning model is used to determine the predicted click-through rate, predicted browsing duration and predicted volume retention probability, and the weighted sum of the predicted click-through rate, predicted browsing duration and predicted volume retention probability of the candidate recommended content is used as the recommendation score of the candidate recommended content.

[0254] S806. Select multiple candidate recommended contents as recommended contents in descending order of their recommendation scores.

[0255] S807: Send the recommended content to the user terminal.

[0256] Figure 8 The specific implementation methods for each step can be found in the relevant descriptions above, and will not be repeated here.

[0257] As can be seen, this application provides an end-to-end optimization scheme that combines operational goals with user interests. It avoids the traditional recommendation algorithm's need to first recall content based on user interests, and then manually intervene in the recalled content. This involves forcibly inserting content requiring increased exposure into the recalled content according to manually set rules, disrupting the overall ranking of recommended content and reducing user experience. Furthermore, this method relies excessively on manual intervention, and manually set rules depend on A / B testing and experience summarization, taking hours or even days, making it unsuitable for real-time operational needs. In contrast, this application's embodiment can predict the probability of content retention through a multi-task learning model and use this probability for content recommendation. This avoids using manually set rules, does not disrupt the ranking of recommended content, and the operational information used can be updated in a timely manner. Combined with a lightweight multi-task learning network, it enables rapid content recommendation, improving user experience, enhancing the intelligent operational capabilities of the recommendation system, and adapting to real-time operational needs, making it suitable for recommending trending content.

[0258] Furthermore, due to the insufficient initial exposure of newly launched content, such as new TV series and movies, and the sparse user behavior data on this content, traditional recommendation algorithms tend to recommend content with higher popularity—that is, content with greater user interaction. This results in newly launched content not being widely recommended to users, easily creating a "virtual exposure loop" where content with higher exposure is recommended more widely, and the more widely recommended content receives, the higher its exposure becomes. However, the embodiments of this application, when recommending content to users, can consider the required exposure level of the content, thereby avoiding content recommendations solely based on users' historical interests and avoiding excessive focus on content popularity. This solves the cold start problem for newly launched content and ensures sufficient exposure.

[0259] Furthermore, the embodiments of this application can dynamically adjust the weights of the volume-preserving loss value when training the multi-task learning model, dynamically balancing the importance of different prediction tasks, and achieving a balance between content exposure and user experience through Pareto optimization. This allows for a balance between user interests and commercial value when subsequently using the multi-task learning model for content delivery. For example, in advertising scenarios, it can balance the traffic allocation between performance ads with higher user preference and brand ads with higher commercial value; or in content recommendation scenarios, it can coordinate the exposure ratio between quality content with higher user preference and commercial content with higher commercial value, maximizing the overall revenue of the content recommendation system and achieving a win-win situation for both commercial goals and user experience.

[0260] Based on the same inventive concept, and corresponding to the above method embodiments, this application provides a content recommendation device applied to electronic devices. See [link to relevant documentation]. Figure 9 The device includes: an acquisition module 901, a prediction module 902, a determination module 903, a selection module 904, and a recommendation module 905;

[0261] The acquisition module 901 is used to acquire content information and operational information of multiple candidate recommended content, as well as user information of users to be recommended. The operational information is used to indicate the exposure required for the candidate recommended content, and the user information is used to indicate the degree of interest of users to be recommended in the recommended content.

[0262] The prediction module 902 is used to determine the predicted click-through rate, predicted browsing duration, and predicted volume retention probability for each candidate recommendation content based on the content information, operational information, and user information of the candidate recommendation content obtained by the acquisition module 901, using a multi-task learning model. The predicted click-through rate represents the probability that the user to be recommended will click on the candidate recommendation content if it is recommended to the user to be recommended. The predicted browsing duration represents the duration that the user to be recommended will browse the candidate recommendation content if it is recommended to the user to be recommended. The predicted volume retention probability represents the probability that the candidate recommendation content will be recommended to the user to be recommended in order to achieve the required exposure of the candidate recommendation content.

[0263] The determination module 903 is used to determine the recommendation score of the candidate recommended content based on the predicted click-through rate, predicted browsing time and predicted retention probability of the candidate recommended content predicted by the prediction module 902.

[0264] The selection module 904 is used to select multiple candidate recommended contents as recommended contents in descending order of their recommendation scores as determined by the determination module 903.

[0265] The recommendation module 905 is used to recommend the content selected by the selection module 904 to the user to be recommended.

[0266] In some embodiments of this application, the multi-task learning model includes multiple expert networks, a click-through rate (CTR) gating network, a browsing duration gating network, a page view retention (PV) gating network, a CTR task tower, a browsing duration task tower, and a PV retention task tower; the prediction module 902 is specifically used for:

[0267] The content information, operational information, and user information of the candidate recommendation are combined to obtain the combined information.

[0268] The spliced ​​information is input into each expert network and each gate network of the multi-task learning model to obtain the spliced ​​features output by each expert network after feature extraction of the spliced ​​information, and the weights of each expert network output by each gate network based on the spliced ​​information.

[0269] The multi-task learning model multiplies the weights of each expert network output by the click rate gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is then input into the click rate task tower to obtain the predicted click rate output by the click rate task tower.

[0270] The multi-task learning model multiplies the weights of each expert network output by the browsing duration gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is then input into the browsing duration task tower to obtain the predicted browsing duration output by the browsing duration task tower.

[0271] The multi-task learning model multiplies the weights of each expert network output by the quantity-preserving gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the quantity-preserving task tower to obtain the predicted quantity-preserving probability output by the quantity-preserving task tower.

[0272] In some embodiments of this application, the multi-task learning model includes a first expert network, a second expert network, a click-through rate gating network, a browsing duration gating network, a page view maintenance gating network, a click-through rate task tower, a browsing duration task tower, and a page view maintenance task tower; the prediction module 902 is specifically used for:

[0273] The content information and user information of the candidate recommendation are concatenated to obtain the first concatenated information;

[0274] The first spliced ​​information is input into the first expert network, the click-through rate gating network, and the browsing time gating network respectively to obtain the features output by the first expert network after feature extraction of the first spliced ​​information, the weights of each expert network output by the click-through rate gating network based on the first spliced ​​information, and the weights of each expert network output by the browsing time gating network based on the first spliced ​​information.

[0275] The content information, operational information, and user information of the candidate recommendation are combined to obtain the second combined information.

[0276] The second splicing information is input into the second expert network and the quantity-preserving gating network respectively to obtain the features output by the second expert network after feature extraction of the second splicing information, and the weights of each expert network output by the quantity-preserving gating network based on the second splicing information.

[0277] The multi-task learning model multiplies the weights of each expert network output by the click rate gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature. This weighted sum feature is then input into the click rate task tower to obtain the predicted click rate output by the click rate task tower.

[0278] The multi-task learning model multiplies the weights of each expert network output by the browsing duration gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature. This weighted sum feature is then input into the browsing duration task tower to obtain the predicted browsing duration output by the browsing duration task tower.

[0279] The multi-task learning model multiplies the weights of each expert network output by the quantity-preserving gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature. This weighted sum feature is then input into the quantity-preserving task tower to obtain the predicted quantity-preserving probability output by the quantity-preserving task tower.

[0280] In some embodiments of this application, the acquisition module 901 is specifically used for:

[0281] For each candidate recommended content, obtain the content priority coefficient preset for that candidate recommended content. The content priority coefficient indicates the priority at which the content is recommended.

[0282] Determine the target exposure for this candidate recommendation content;

[0283] Based on the content priority coefficient and target exposure, construct the operational information for the candidate recommended content.

[0284] In some embodiments of this application, the determining module 903 is specifically used for:

[0285] Determine the exposure requirements for the candidate recommended content;

[0286] Based on the preset mapping relationship between exposure demand and weight, the weight mapped to the exposure demand of the candidate recommended content is used as the weight for predicting the probability of maintaining the volume. The preset mapping relationship represents a positive correlation between exposure demand and weight.

[0287] The weighted sum of the predicted click-through rate, predicted browsing duration, and predicted retention probability of the candidate content is determined as the recommendation score for that candidate content.

[0288] In some embodiments of this application, the determining module 903 is specifically used for:

[0289] The difference between the preset target exposure for the candidate recommended content and the actual exposure of the candidate recommended content is determined, and the ratio of the difference to the target exposure is determined to obtain the exposure gap intensity, which is used as the exposure requirement;

[0290] Alternatively, the probability of maintaining a certain volume can be predicted as the exposure requirement.

[0291] Based on the same inventive concept, and corresponding to the above method embodiments, this application provides a model training device applied to electronic devices. See [link to relevant documentation]. Figure 10 The device includes: an acquisition module 1001, a prediction module 1002, a calculation module 1003, and an adjustment module 1004;

[0292] The acquisition module 1001 is used to acquire content information and operational information of multiple sample contents recommended to sample users in history, as well as user information of sample users. The operational information is used to indicate the exposure required for the sample content.

[0293] The prediction module 1002 is used to determine the predicted click-through rate, predicted browsing duration, and predicted retention probability for each sample content based on the content information, operational information, and user information of the sample content obtained by the acquisition module 1001, using a multi-task learning model. The predicted click-through rate represents the probability that a sample user will click on the sample content if it is recommended to the sample user; the predicted browsing duration represents the duration that a sample user will browse the sample content if it is recommended to the sample user; and the predicted retention probability represents the probability that the sample content will be recommended to the sample user to achieve the required exposure for the sample content.

[0294] The calculation module 1003 is used to calculate the joint loss value based on the predicted click-through rate, predicted browsing time and predicted volume retention probability of each sample content predicted by the prediction module 1002, as well as the actual click-through rate, actual browsing time and standard volume retention probability of each sample content.

[0295] The adjustment module 1004 is used to adjust the network parameters of the multi-task learning model using the joint loss value calculated by the calculation module 1003. It calls the prediction module 1002 to perform the following steps for each sample content: based on the content information, operation information and user information of the sample content, the multi-task learning model is used to determine the predicted click-through rate, predicted browsing duration and predicted volume retention probability, until the multi-task learning model converges, at which point the training is considered complete.

[0296] In some embodiments of this application, the standard exposure probability is the ratio of the exposure gap to the target exposure for the sample content, where the exposure gap is the difference between the target exposure and the actual exposure of the sample content; the calculation module 1003 is specifically used for:

[0297] Calculate the click-through rate loss value based on the predicted click-through rate and the actual click-through rate of each sample content;

[0298] Based on the predicted browsing time and actual browsing time of each sample content, calculate the browsing time loss value;

[0299] Based on the predicted and standard probability of maintaining volume for each sample, calculate the loss value for maintaining volume.

[0300] Based on the preset mapping relationship between the probability of maintaining quantity and the weight, the weight of the average standard probability of maintaining quantity for each sample content is determined and used as the weight corresponding to the loss value of maintaining quantity. The probability of maintaining quantity represented by the preset mapping relationship is positively correlated with the weight.

[0301] The weighted sum of the click-through rate loss, browsing time loss, and retention loss is used as the joint loss value.

[0302] In some embodiments of this application, the prediction module 1002 is further configured to determine the predicted click-through rate, predicted browsing duration, and predicted volume retention probability for each sample content using a multi-task learning model before determining the predicted click-through rate and predicted browsing duration based on the content information, operational information, and user information of the sample content.

[0303] The calculation module 1003 is also used to calculate the pre-training loss value based on the predicted click-through rate and predicted browsing time of each sample content, as well as the actual click-through rate and actual browsing time of each sample content.

[0304] The adjustment module 1004 is also used to adjust the network parameters of the multi-task learning model using the pre-trained loss value, and to call the prediction module 1002 to perform the steps of determining the predicted click-through rate and predicted browsing duration for each sample content based on the content information and user information of the sample content using the multi-task learning model, until the multi-task learning model converges. Then, the prediction module 1002 is called to perform the steps of determining the predicted click-through rate, predicted browsing duration, and predicted volume retention probability for each sample content based on the content information, operational information, and user information of the sample content using the multi-task learning model.

[0305] This application also provides an electronic device, such as... Figure 11 As shown, it includes a processor 1101, a communication interface 1102, a memory 1103, and a communication bus 1104. The processor 1101, communication interface 1102, and memory 1103 communicate with each other via the communication bus 1104.

[0306] Memory 1103 is used to store computer programs;

[0307] The processor 1101 is used to execute the program stored in the memory 1103 to implement the method steps in the above method embodiments.

[0308] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0309] The communication interface is used for communication between the aforementioned terminal and other devices.

[0310] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0311] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0312] In another embodiment provided in this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, it implements any of the content recommendation methods and model training methods described in the above embodiments.

[0313] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the content recommendation methods and model training methods described in the above embodiments.

[0314] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0315] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0316] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0317] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A content recommendation method, characterized in that, The method includes: The system acquires content and operational information of multiple candidate recommended content, as well as user information of users to be recommended. The operational information is used to indicate the required exposure of the candidate recommended content, and the user information is used to indicate the degree of interest of the users to be recommended in the recommended content. For each candidate recommendation content, based on the content information, operational information, and user information of the candidate recommendation content, a multi-task learning model is used to determine the predicted click-through rate, predicted browsing duration, and predicted probability of maintaining exposure. The predicted click-through rate represents the probability that the user to be recommended will click on the candidate recommendation content if it is recommended to the user to be recommended. The predicted browsing duration represents the duration that the user to be recommended will browse the candidate recommendation content if it is recommended to the user to be recommended. The predicted probability of maintaining exposure represents the probability of recommending the candidate recommendation content to the user to be recommended in order to achieve the required exposure of the candidate recommendation content. The recommendation score for the candidate content is determined based on its predicted click-through rate, predicted browsing duration, and predicted retention rate. Based on the recommendation scores of each candidate content in descending order, select multiple candidate content as the content to be recommended; The content to be recommended is recommended to the user to be recommended.

2. The method according to claim 1, characterized in that, The multi-task learning model includes multiple expert networks, click-through rate gating networks, browsing time gating networks, volume retention gating networks, click-through rate task towers, browsing time task towers, and volume retention task towers. The process of determining the predicted click-through rate, predicted browsing duration, and predicted volume retention probability using a multi-task learning model based on the content information, operational information, and user information of the candidate recommended content includes: The content information, operational information, and user information of the candidate recommendation content are concatenated to obtain the concatenated information. The spliced ​​information is input into each expert network and each gate network of the multi-task learning model to obtain the spliced ​​features output by each expert network after feature extraction of the spliced ​​information, and the weights of each expert network output by each gate network based on the spliced ​​information. The multi-task learning model multiplies the weights of each expert network output by the click-through rate gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the click-through rate task tower to obtain the predicted click-through rate output by the click-through rate task tower. The multi-task learning model multiplies the weights of each expert network output by the browsing duration gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the browsing duration task tower to obtain the predicted browsing duration output by the browsing duration task tower. The multi-task learning model multiplies the weights of each expert network output by the quantity-preserving gating network with the concatenated features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the quantity-preserving task tower to obtain the predicted quantity-preserving probability output by the quantity-preserving task tower.

3. The method according to claim 1, characterized in that, The multi-task learning model includes a first expert network, a second expert network, a click-rate gating network, a browsing duration gating network, a volume retention gating network, a click-rate task tower, a browsing duration task tower, and a volume retention task tower. The process of determining the predicted click-through rate, predicted browsing duration, and predicted volume retention probability using a multi-task learning model based on the content information, operational information, and user information of the candidate recommended content includes: The content information of the candidate recommendation and the user information are concatenated to obtain the first concatenation information; The first spliced ​​information is input into the first expert network, the click-through rate gating network, and the browsing duration gating network respectively to obtain the features output by the first expert network after feature extraction of the first spliced ​​information. The click-through rate gating network outputs the weights of each expert network based on the first spliced ​​information, and the browsing duration gating network outputs the weights of each expert network based on the first spliced ​​information. The content information, operational information, and user information of the candidate recommendation content are concatenated to obtain the second concatenated information. The second splicing information is input into the second expert network and the quantity-preserving gating network respectively to obtain the features output by the second expert network after feature extraction of the second splicing information, and the weights of each expert network output by the quantity-preserving gating network based on the second splicing information. The multi-task learning model multiplies the weights of each expert network output by the click-through rate gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the click-through rate task tower to obtain the predicted click-through rate output by the click-through rate task tower. The multi-task learning model multiplies the weights of each expert network output by the browsing duration gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is then input into the browsing duration task tower to obtain the predicted browsing duration output by the browsing duration task tower. The multi-task learning model multiplies the weights of each expert network output by the quantity-preserving gating network with the features output by that expert network to obtain weighted features. The weighted features are then summed to obtain a weighted sum feature, which is input into the quantity-preserving task tower to obtain the predicted quantity-preserving probability output by the quantity-preserving task tower.

4. The method according to any one of claims 1-3, characterized in that, Obtain the operational information of the candidate recommended content, including: For each candidate recommended content, a preset content priority coefficient is obtained, whereby the content priority coefficient represents the priority at which the content is recommended; Determine the target exposure for this candidate recommendation content; Based on the content priority coefficient and the target exposure, the operational information of the candidate recommended content is constructed.

5. The method according to any one of claims 1-3, characterized in that, The process of determining the recommendation score of the candidate content based on its predicted click-through rate, predicted browsing duration, and predicted retention probability includes: Determine the exposure requirements for the candidate recommended content; Based on the preset mapping relationship between exposure demand and weight, the weight mapped to the exposure demand of the candidate recommended content is used as the weight of the predicted volume guarantee probability, wherein the exposure demand represented by the preset mapping relationship is positively correlated with the weight. The weighted sum of the predicted click-through rate, predicted browsing duration, and predicted retention probability of the candidate content is determined as the recommendation score for that candidate content.

6. The method according to claim 5, characterized in that, Determining the exposure requirements for the candidate recommended content includes: The difference between the preset target exposure for the candidate recommended content and the actual exposure of the candidate recommended content is determined, and the ratio of the difference to the target exposure is determined to obtain the exposure gap intensity, which is used as the exposure requirement; Alternatively, the predicted quantity retention probability can be used as the exposure requirement.

7. A model training method, characterized in that, The method includes: Obtain content and operational information of multiple sample contents recommended to sample users in history, as well as user information of the sample users, wherein the operational information is used to indicate the required exposure of the sample content; For each sample content, based on the content information, operational information, and user information of the sample content, a multi-task learning model is used to determine the predicted click-through rate, predicted browsing duration, and predicted retention probability. The predicted click-through rate represents the probability that the sample user will click on the sample content if it is recommended to the sample user; the predicted browsing duration represents the duration that the sample user will browse the sample content if it is recommended to the sample user; and the predicted retention probability represents the probability that the sample content will be recommended to the sample user to achieve the required exposure. The joint loss value is calculated based on the predicted click-through rate, predicted browsing duration, and predicted retention probability of each sample content, as well as the actual click-through rate, actual browsing duration, and standard retention probability of each sample content. The network parameters of the multi-task learning model are adjusted using the joint loss value. The steps of determining the predicted click-through rate, predicted browsing duration, and predicted volume retention probability using the multi-task learning model for each sample content based on the content information, operational information, and user information of the sample content are returned until the multi-task learning model converges, at which point the training is considered complete.

8. The method according to claim 7, characterized in that, The standard exposure retention probability is the ratio of the exposure gap to the target exposure volume preset for the sample content, where the exposure gap is the difference between the target exposure volume and the actual exposure volume of the sample content; the calculation of the joint loss value based on the predicted click-through rate, predicted browsing duration, and predicted exposure retention probability of each sample content, as well as the actual click-through rate, actual browsing duration, and standard exposure retention probability of each sample content, includes: Calculate the click-through rate loss value based on the predicted click-through rate and the actual click-through rate of each sample content; Based on the predicted browsing time and actual browsing time of each sample content, calculate the browsing time loss value; Based on the predicted probability of maintaining quantity and the standard probability of maintaining quantity for each sample, calculate the loss value for maintaining quantity. Based on the preset mapping relationship between the quantity preservation probability and the weight, the weight of the average standard quantity preservation probability mapping of each sample content is determined, and used as the weight corresponding to the quantity preservation loss value. The quantity preservation probability represented by the preset mapping relationship is positively correlated with the weight. The weighted sum of the click-through rate loss value, the browsing time loss value, and the retention loss value is taken as the joint loss value.

9. The method according to claim 7 or 8, characterized in that, Before determining the predicted click-through rate, predicted browsing duration, and predicted retention probability using a multi-task learning model for each sample content based on the content information, operational information, and user information of that sample content, the method further includes: For each sample content, based on the content information of the sample content and the user information, a multi-task learning model is used to determine the predicted click-through rate and the predicted browsing duration. The pre-training loss value is calculated based on the predicted click-through rate and predicted browsing duration of each sample content, as well as the actual click-through rate and actual browsing duration of each sample content. The network parameters of the multi-task learning model are adjusted using the pre-training loss value. The steps of determining the predicted click-through rate and predicted browsing duration based on the content information of the sample content and the user information using the multi-task learning model are returned for each sample content. This process continues until the multi-task learning model converges. Then, the steps of determining the predicted click-through rate, predicted browsing duration, and predicted volume retention probability based on the content information, operational information, and user information of the sample content using the multi-task learning model are executed.

10. A content recommendation device, characterized in that, The device includes: The acquisition module is used to acquire content information and operational information of multiple candidate recommended content, as well as user information of the user to be recommended. The operational information is used to indicate the exposure required for the candidate recommended content, and the user information is used to indicate the degree of interest of the user to be recommended in the recommended content. The prediction module is used to determine the predicted click-through rate, predicted browsing duration, and predicted volume retention probability for each candidate recommendation content based on the content information, operational information, and user information obtained by the acquisition module. The predicted click-through rate represents the probability that the user to be recommended will click on the candidate recommendation content if it is recommended to them; the predicted browsing duration represents the duration that the user to be recommended will browse the candidate recommendation content if it is recommended to them; and the predicted volume retention probability represents the probability that the candidate recommendation content will be recommended to the user to achieve the required exposure level. The determination module is used to determine the recommendation score of the candidate recommended content based on the predicted click-through rate, predicted browsing time, and predicted retention probability of the candidate recommended content predicted by the prediction module. The selection module is used to select multiple candidate recommended contents as recommended contents in descending order of their recommendation scores as determined by the determination module. The recommendation module is used to recommend the content selected by the selection module to the user to be recommended.

11. A model training device, characterized in that, The device includes: The acquisition module is used to acquire content information and operational information of multiple sample contents recommended to sample users in history, as well as user information of the sample users. The operational information is used to indicate the exposure required for the sample content. The prediction module is used to determine the predicted click-through rate, predicted browsing duration, and predicted retention probability for each sample content based on the content information, operational information, and user information obtained by the acquisition module. The predicted click-through rate represents the probability that a sample user will click on the sample content if it is recommended to them; the predicted browsing duration represents the duration a sample user will browse the sample content if it is recommended to them; and the predicted retention probability represents the probability of recommending the sample content to the sample user to achieve the required exposure level. The calculation module is used to calculate the joint loss value based on the predicted click-through rate, predicted browsing time and predicted volume retention probability of each sample content predicted by the prediction module, as well as the actual click-through rate, actual browsing time and standard volume retention probability of each sample content. The adjustment module is used to adjust the network parameters of the multi-task learning model using the joint loss value calculated by the calculation module, and to call the prediction module to perform the steps of determining the predicted click-through rate, predicted browsing duration, and predicted volume retention probability for each sample content based on the content information, operational information, and user information of the sample content using the multi-task learning model, until the multi-task learning model converges, at which point training is determined to be complete.

12. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-9.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-9.