Content recommendation method and apparatus, electronic device, and computer-readable storage medium

CN119025743BActive Publication Date: 2026-08-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310595601.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-24
Publication Date
2026-08-18
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

[0003]在对现有技术的研究和实践过程中,本申请的发明人发现将交互持续时间作为特征引入时,由于模型的训练往往包括很多的特征,特征经过多层神经网络作用到最后的预估值,作用不够直接,使得信息增益小,而采用交互持续时间进行加权修正就需要增加部署一个交互持续时间的预测模型,增加模型部署的资源,另外,采用交互持续时间对推荐概率进行加权修正往往会依赖人工经验,修正的效果会有折损,因此,导致内容推荐的推荐效率不足

Benefits of technology

[0033]本发明实施例在获取内容推荐模型的当前训练集后,该当前训练集包括对象集合中每一对象的交互内容样本和交互内容样本对应的交互持续时间,采用内容推荐模型,预测交互内容样本针对对象的推荐概率,得到预测推荐结果,然后,根据交互持续时间,对交互内容样本进行分类,并将不同类别的交互内容样本对应的预测推荐结果进行对比,以得到交互置信度损失,然后,基于交互置信度损失和预测推荐结果,确定内容推荐模型的目标推荐损失,然后,根据目标推荐损失对内容推荐模型进行更新,并采用更新后的内容推荐模型,将待推荐内容推荐至对象集合中的目标对象;由于该方案可以通过交互持续时间来衡量交互行为的置信度,并通过交互持续时间对交互内容样本进行分类,然后,将不同类别的交互内容样本对应的预测推荐结果进行对比,从而得到交互置信度损失,通过交互置信度损失使得内容推荐模型可以显式的学习基于交互持续时间的偏序关系,可以精准的刻画对象与内容之间的交互行为的真实意图,提升内容推荐模型的精度,进而提升内容推荐的推荐效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119025743B_ABST
    Figure CN119025743B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a content recommendation method and device, electronic equipment and a computer readable storage medium; an embodiment of the present application obtains a current training set of a content recommendation model, the current training set including an interaction content sample of each object in an object set and an interaction duration corresponding to the interaction content sample, then, a recommendation probability of the interaction content sample for the object is predicted by using the content recommendation model to obtain a predicted recommendation result, the interaction content sample is classified according to the interaction duration, and the predicted recommendation results corresponding to the interaction content samples of different categories are compared to obtain an interaction confidence loss, a target recommendation loss of the content recommendation model is determined based on the interaction confidence loss and the predicted recommendation result, the content recommendation model is updated according to the target recommendation loss, and the updated content recommendation model is used to recommend to-be-recommended content to a target object in the object set; the scheme can improve the recommendation efficiency of content recommendation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of content recommendation, and more specifically to a content recommendation method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] In recent years, with the rapid development of internet technology, an increasing number of diverse types of content have emerged online. To increase the effective exposure or conversion of this content, it can be recommended to different audiences. The duration of interaction varies depending on the content being discussed. For example, taking a click as an interaction, the dwell time after clicking on recommended content can differ depending on the specific content. To improve the efficiency of content recommendation, the interaction duration factor needs to be considered. Existing content recommendation methods often use interaction duration as a training feature for the content recommendation model, or they use the predicted interaction duration to correct the predicted recommendation probability and then recommend content based on the corrected probability.

[0003] In the process of researching and practicing existing technologies, the inventors of this application discovered that when interaction duration is introduced as a feature, the model training often includes many features, and the features are processed through multiple layers of neural networks to reach the final prediction value. This results in insufficient directness and low information gain. Furthermore, using interaction duration for weighted correction requires the deployment of an additional interaction duration prediction model, increasing the resources available for model deployment. In addition, using interaction duration to weighted correct the recommendation probability often relies on human experience, which can diminish the effectiveness of the correction. Therefore, the recommendation efficiency of content recommendation is insufficient. Summary of the Invention

[0004] This invention provides a content recommendation method, apparatus, electronic device, and computer-readable storage medium, which can improve the efficiency of content recommendation.

[0005] A content recommendation method includes:

[0006] Obtain the current training set of the content recommendation model, wherein the current training set includes the interaction content sample of each object in the object set and the interaction duration corresponding to the interaction content sample;

[0007] Using the content recommendation model, the recommendation probability of the interactive content sample for the object is predicted, and the prediction recommendation result is obtained;

[0008] Based on the duration of the interaction, the interaction content samples are classified, and the prediction and recommendation results corresponding to the interaction content samples of different categories are compared to obtain the interaction confidence loss.

[0009] Based on the interaction confidence loss and the predicted recommendation results, the target recommendation loss of the content recommendation model is determined;

[0010] The content recommendation model is updated based on the target recommendation loss, and the updated content recommendation model is used to recommend the content to be recommended to the target object in the object set.

[0011] Accordingly, embodiments of the present invention provide a content recommendation device, including:

[0012] The acquisition unit is used to acquire the current training set of the content recommendation model, wherein the current training set includes the interaction content sample of each object in the object set and the interaction duration corresponding to the interaction content sample;

[0013] The prediction unit is used to use the content recommendation model to predict the recommendation probability of the interactive content sample for the object, and obtain the prediction recommendation result.

[0014] The comparison unit is used to classify the interaction content samples according to the interaction duration, and compare the prediction recommendation results corresponding to the interaction content samples of different categories to obtain the interaction confidence loss.

[0015] The determining unit is used to determine the target recommendation loss of the content recommendation model based on the interaction confidence loss and the prediction recommendation result;

[0016] The recommendation unit is used to update the content recommendation model based on the target recommendation loss, and use the updated content recommendation model to recommend the content to be recommended to the target objects in the object set.

[0017] In some embodiments, the comparison unit may be specifically used to obtain the interaction duration threshold corresponding to the current training set; and to classify the interaction content samples based on the interaction duration threshold and the interaction duration.

[0018] In some embodiments, the comparison unit may be specifically used to obtain the historical interaction duration of each object interacting with at least one historical interaction content; determine the time mean of the historical interaction duration, and use the time mean as the initial interaction duration threshold corresponding to each object; and use the initial interaction duration threshold as the interaction duration threshold corresponding to the current training set.

[0019] In some embodiments, the comparison unit may be specifically used to compare the interaction duration with the interaction duration threshold; to take the interaction content samples whose interaction duration is greater than or equal to the interaction duration threshold as first category samples to obtain a first category sample set; and to take the interaction content samples whose interaction duration is less than the interaction duration threshold as second category samples to obtain a second category sample set.

[0020] In some embodiments, the comparison unit may be specifically used to construct a recommendation result pair based on the first category sample set and the second category sample set; calculate the difference between the recommendation results in the recommendation result pair to obtain the recommendation difference; and correct the recommendation difference to obtain the interaction confidence loss.

[0021] In some embodiments, the comparison unit may be specifically used to select one interactive content sample from the first category sample set and one from the second category sample set to obtain a target sample pair; to filter out the predicted recommendation results corresponding to the interactive content samples in the target sample pair from the predicted recommendation results to obtain the current recommendation results; and to construct the recommendation result pair based on the current recommendation results and the target sample pair.

[0022] In some embodiments, the comparison unit can be specifically used to obtain historical recommendation results corresponding to candidate objects, wherein the candidate objects include objects corresponding to interactive content samples in the target sample pair; when the target sample pair contains interactive content samples of the same object, the current recommendation result and the historical recommendation result are combined to obtain the recommendation result pair; when the target sample pair contains interactive content samples of different objects, the current recommendation result is adjusted based on the historical recommendation result to obtain the recommendation result pair.

[0023] In some embodiments, the comparison unit may be specifically used to filter out candidate objects corresponding to interactive content samples in the target sample pair from the object set; obtain at least one historical recommendation probability of the candidate object, wherein the historical recommendation probability includes the recommendation probability of historical interactive content for the candidate object predicted by the content recommendation model; determine the probability mean of the historical recommendation probability, and use the probability mean as the historical recommendation result corresponding to the candidate object.

[0024] In some embodiments, the comparison unit may be specifically used to calculate the difference between the historical recommendation result and the current recommendation result to obtain the adjusted current recommendation result; and to combine the adjusted current recommendation results to obtain the recommendation result pair.

[0025] In some embodiments, the comparison unit may be specifically used to combine the current recommendation results to obtain the recommendation result pair when the target sample pair contains interactive content samples of the same object.

[0026] In some embodiments, the comparison unit may be specifically used to sort the recommendation results in the recommendation result pair based on the interaction duration and the interaction duration threshold; and calculate the difference between the recommendation results in the recommendation result pair according to the sorting result to obtain the recommendation difference.

[0027] In some embodiments, the comparison unit may be specifically used to obtain the training penalty value corresponding to the current training set; fuse the training penalty value with the recommendation difference to obtain the target recommendation difference; compare the target recommendation difference with the preset recommendation difference, and determine the interaction confidence loss based on the comparison result.

[0028] In some embodiments, the determining unit may be specifically used to obtain the labeled recommendation result corresponding to the interactive content sample; compare the labeled recommendation result with the corresponding predicted recommendation result to obtain an initial recommendation loss; and fuse the initial recommendation loss with the interaction confidence loss to obtain the target recommendation loss.

[0029] In some embodiments, the recommendation unit may be used to obtain content to be recommended, and use the updated content recommendation model to predict the current recommendation probability of the content to be recommended for each object in the object set; based on the current recommendation probability, filter out the target object corresponding to the content to be recommended in the object set, and recommend the content to be recommended to the target object.

[0030] Furthermore, embodiments of the present invention also provide an electronic device, including a processor and a memory, wherein the memory stores an application program, and the processor is used to run the application program in the memory to implement the content recommendation method provided in embodiments of the present invention.

[0031] Furthermore, embodiments of the present invention also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the content recommendation methods provided in embodiments of the present invention.

[0032] Furthermore, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the content recommendation method provided in embodiments of this application.

[0033] In this embodiment of the invention, after obtaining the current training set of the content recommendation model, the current training set includes interaction content samples of each object in the object set and the interaction duration corresponding to the interaction content samples. The content recommendation model is used to predict the recommendation probability of the interaction content samples for the objects, obtaining a predicted recommendation result. Then, based on the interaction duration, the interaction content samples are classified, and the predicted recommendation results corresponding to different categories of interaction content samples are compared to obtain an interaction confidence loss. Then, based on the interaction confidence loss and the predicted recommendation result, the target recommendation loss of the content recommendation model is determined. Then, the content recommendation model is updated according to the target recommendation loss, and the updated content recommendation model is used to recommend the content to be recommended to the target objects in the object set. Since this scheme can measure the confidence of interaction behavior through interaction duration and classify interaction content samples through interaction duration, and then compare the predicted recommendation results corresponding to different categories of interaction content samples to obtain the interaction confidence loss, the interaction confidence loss allows the content recommendation model to explicitly learn the partial order relationship based on interaction duration, accurately characterizing the true intent of the interaction behavior between objects and content, improving the accuracy of the content recommendation model, and thus improving the recommendation efficiency of content recommendation. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a schematic diagram of a scenario for the content recommendation method provided in an embodiment of the present invention;

[0036] Figure 2 This is a flowchart illustrating the content recommendation method provided in an embodiment of the present invention;

[0037] Figure 3 This is a schematic diagram comparing the training methods provided in the embodiments of the present invention;

[0038] Figure 4 This is a schematic diagram illustrating the display of content to be recommended, provided in an embodiment of the present invention;

[0039] Figure 5 This is another flowchart illustrating the content recommendation method provided in this embodiment of the invention;

[0040] Figure 6 This is a schematic diagram of the structure of the content recommendation device provided in an embodiment of the present invention;

[0041] Figure 7This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0043] This invention provides a content recommendation method, apparatus, and computer-readable storage medium. The content recommendation apparatus can be integrated into an electronic device, which may be a server or a terminal, etc.

[0044] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0045] For example, see Figure 1 Taking the integration of a content recommendation device into an electronic device as an example, the electronic device obtains the current training set of the content recommendation model. This current training set may include the interaction content samples of each object in the object set and the interaction duration corresponding to the interaction content samples. Then, the content recommendation model is used to predict the recommendation probability of the interaction content samples for the objects, and the predicted recommendation result is obtained. According to the interaction duration, the interaction content samples are classified, and the predicted recommendation results corresponding to the interaction content samples of different categories are compared to obtain the interaction confidence loss. Then, based on the interaction confidence loss and the predicted recommendation result, the target recommendation loss of the content recommendation model is determined. The content recommendation model is updated according to the target recommendation loss, and the updated content recommendation model is used to recommend the content to be recommended to the target objects in the object set, thereby improving the recommendation efficiency of the content recommendation.

[0046] The content recommendation method provided in this application relates to machine learning within artificial intelligence. This application can adjust the partial order relationship of the predicted recommendation results of interactive content samples by adjusting the interaction duration, thereby enabling the model to explicitly learn this partial order relationship and improve the training accuracy of the content recommendation model.

[0047] Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0048] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0049] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0050] It is understood that, in the specific embodiments of this application, data related to the interaction content samples, interaction duration, historical interaction content or historical interaction duration of the object are involved. When the following embodiments of this application are applied to specific products or technologies, permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0051] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0052] This embodiment will be described from the perspective of a content recommendation device, which can be integrated into an electronic device, such as a server or a terminal. The terminal can include tablet computers, laptops, personal computers (PCs), wearable devices, virtual reality devices, or other smart devices capable of content recommendation.

[0053] A content recommendation method includes:

[0054] Obtain the current training set of the content recommendation model. The current training set includes interaction content samples and corresponding interaction durations for each object in the object set. Using the content recommendation model, predict the recommendation probability of the interaction content samples for the object and obtain the prediction recommendation result. Classify the interaction content samples according to the interaction duration and compare the prediction recommendation results corresponding to the interaction content samples of different categories to obtain the interaction confidence loss. Based on the interaction confidence loss and the prediction recommendation result, determine the target recommendation loss of the content recommendation model. Update the content recommendation model according to the target recommendation loss and use the updated content recommendation model to recommend the content to be recommended to the target object in the object set.

[0055] like Figure 2 As shown, the specific process of this content recommendation method is as follows:

[0056] 101. Obtain the current training set of the content recommendation model.

[0057] The content recommendation model can be a model that recommends content to objects in a set of objects. This model can include an initial, untrained content recommendation model or a model trained using a historical training set corresponding to the set of objects.

[0058] The current training set can include interaction content samples for each object in the object set and the corresponding interaction duration. The object set can include at least one user using a content recommendation model for content recommendation. The corresponding interaction content samples are the content samples that interact with the user. There are various types of interactions, such as clicks and conversions. Clicks can include opening, forwarding, liking, commenting, saving, screenshotting, copying, or content selection, etc. Conversions can include registering an account, downloading an app, following a public account, or submitting various materials, etc. It should be noted that interaction content samples can include positive and negative samples. Positive samples are content samples that have been interacted with by the object, and negative samples are content samples that have not been interacted with by the object. The content types of interaction content samples can be various, such as text, audio / video, or other types of content.

[0059] The interaction duration corresponding to the interactive content sample can be the duration of the interaction between the object and the positive sample. Taking the interaction as a click as an example, the interaction duration can include the time the object stays on the content sample after clicking, which can also be called the landing page dwell time. For example, if object A clicks on ad A and exits after 0.1 seconds, the dwell time for ad A can be 0.1 seconds.

[0060] There are several ways to obtain the current training set of the content recommendation model, as follows:

[0061] For example, it can directly receive the current training set of the content recommendation model uploaded by the terminal or client. Alternatively, it can obtain the interaction records of each object in the object set, filter the interaction content samples of each object in the network or content database based on the interaction records, and count the interaction duration corresponding to the interaction content samples in the interaction records to obtain the current training set of the content recommendation model. Or, it can filter the interaction content corresponding to at least one object in the network or database, obtain the interaction content samples, and determine the interaction duration corresponding to the interaction content samples to obtain the current training set of the content recommendation model. Or, when the memory of the current training set is large or the number of samples is large, it can also receive update requests for the content recommendation model. The update request may include the storage address of the current training set, and obtain the current training set based on the storage address, and so on.

[0062] 102. Using a content recommendation model, predict the recommendation probability of interactive content samples for specific objects, and obtain the predicted recommendation results.

[0063] Among them, the prediction and recommendation results can characterize the probability or degree of recommending interactive content samples to the corresponding objects.

[0064] There are several ways to use a content recommendation model to predict the probability of recommending interactive content samples to an object, as follows:

[0065] For example, a content recommendation model can be used to extract features from interactive content samples to obtain multi-dimensional content features. These features can then be fused to obtain recommendation features for the interactive content sample relative to the corresponding object in the object set. Based on these recommendation features, the recommendation probability of the interactive content sample relative to that object can be determined, thereby obtaining the predicted recommendation result for the interactive content sample.

[0066] The recommendation probability can include the probability that the interactive content sample will be recommended to the corresponding object, and that the object will perform various interactive operations such as clicking or converting the interactive content. For example, when the interaction is click, the recommendation probability can be the probability that the object clicks the interactive content, which can also be called the click rate for the interactive content sample.

[0067] 103. Based on the duration of the interaction, classify the interaction content samples and compare the prediction and recommendation results corresponding to the interaction content samples of different categories to obtain the interaction confidence loss.

[0068] Interaction confidence loss characterizes the loss of confidence an object has in interacting with a sample of interactive content. Interaction confidence is used to evaluate the authenticity of an interaction. A good metric for measuring interaction confidence can be the interaction duration. Taking a click as an example, the interaction duration can be the dwell time after the click or the dwell time on the landing page. The dwell time after the click can measure the interaction confidence. For instance, if an object *u* clicks on ad *a* and exits after 0.1 seconds, and clicks on ad *b* and exits after 5 seconds, then ad *a* might have been a false click, meaning the confidence of this click is low, or the click behavior does not convey the credibility of the interaction intent.

[0069] The process of classifying interaction content samples based on interaction duration and comparing the predicted recommendation results corresponding to different categories of interaction content samples to obtain the interaction confidence loss can be described as follows:

[0070] S1. Classify the interaction content samples according to the duration of the interaction.

[0071] For example, the interaction duration threshold corresponding to the current training set can be obtained, and the interaction content samples can be classified based on the interaction duration threshold and the interaction duration.

[0072] The interaction duration threshold serves as a criterion for classifying the length of an interaction. For example, a value greater than or equal to the interaction duration threshold indicates a long interaction duration for the given content sample, while a value less than the threshold indicates a short interaction duration. There are several ways to obtain the interaction duration threshold for the current training set. For instance, one can obtain the historical interaction durations of each object interacting with at least one historical content sample, determine the average historical interaction duration, and use this average as the initial interaction duration threshold for each object. This initial threshold is then used as the interaction duration threshold for the current training set. Alternatively, a preset interaction duration threshold can be obtained and used as the interaction duration threshold for the current training set. Another approach is to obtain the content types of the interaction content samples in the current training set, filter out the corresponding time thresholds from a preset time threshold set, and thus obtain the interaction duration threshold for the current training set. Yet another approach is to filter out the corresponding time thresholds from a preset time threshold set based on the object identifier of each object in the object set, and so on.

[0073] The historical interaction content may include at least one piece of content that each object interacted with within a historical time period.

[0074] After obtaining the interaction duration threshold corresponding to the current training set, the interaction content samples can be classified based on the interaction duration threshold and the interaction duration. There are several ways to classify them. For example, the interaction duration can be compared with the interaction duration threshold, and the interaction content samples with an interaction duration greater than or equal to the interaction duration threshold can be regarded as the first category samples to obtain the first category sample set, and the interaction content samples with an interaction duration less than the interaction duration threshold can be regarded as the second category samples to obtain the second category sample set.

[0075] S2. Compare the prediction and recommendation results corresponding to different categories of interactive content samples to obtain the interaction confidence loss.

[0076] For example, recommendation result pairs can be constructed based on the first category sample set and the second category sample set. The difference between the recommendation results in the recommendation result pair can be calculated to obtain the recommendation difference. The recommendation difference can then be corrected to obtain the interaction confidence loss, as follows:

[0077] C1. Construct recommendation result pairs based on the first category sample set and the second category sample set.

[0078] The recommendation result pair can include recommendation results from different categories of interactive content samples. These recommendation results can include predicted recommendations, historical predictions, revised predictions, and so on. Recommendation result pairs, also known as "pair pairs," are used to represent the partial order relationship of interaction duration.

[0079] There are several ways to construct recommendation result pairs based on the first category sample set and the second category sample set, as follows:

[0080] For example, one interactive content sample can be selected from each of the first and second category sample sets to obtain a target sample pair. The predicted recommendation results corresponding to the interactive content samples in the target sample pair can be selected from the predicted recommendation results to obtain the current recommendation results. Based on the current recommendation results and the target sample pair, a recommendation result pair can be constructed.

[0081] The target sample pairs can include interactive content samples of different categories.

[0082] The current recommendation result can be the predicted recommendation result for each interaction content sample in the target sample pair. Based on the current recommendation result and the target sample pair, there are several ways to construct recommendation result pairs. For example, historical recommendation results for candidate objects can be obtained. When the target sample pair contains interaction content samples of the same object, the current recommendation result and historical recommendation results are combined to obtain a recommendation result pair. When the target sample pair contains interaction content samples of different objects, the current recommendation result is adjusted based on historical recommendation results to obtain a recommendation result pair. Specifically, it can be as follows:

[0083] (1) Obtain the historical recommendation results corresponding to the candidate objects.

[0084] Candidate objects include objects corresponding to the interactive content samples in the target sample pair.

[0085] There are several ways to obtain the historical recommendation results corresponding to the candidate objects, as follows:

[0086] For example, candidate objects corresponding to interactive content samples in the target sample pair can be selected from the object set, at least one historical recommendation probability of the candidate object can be obtained, the probability mean of the historical recommendation probability can be determined, and the probability mean can be used as the historical recommendation result corresponding to the candidate object.

[0087] The historical recommendation probability can include the recommendation probability of candidate objects predicted by using a content recommendation model based on historical interaction content. There are several ways to obtain at least one historical recommendation probability of a candidate object. For example, one can obtain at least one historical interaction content corresponding to the candidate object, and use a content recommendation model to predict the recommendation probability corresponding to the historical interaction content to obtain the historical recommendation probability.

[0088] It's important to note that the historical interaction content here can be the same as or different from the historical interaction content used to determine the interaction duration threshold. However, the interaction time for these two historical interactions can be the same interaction time period. This ensures that the interaction duration threshold determined based on the historical interaction content and the historical recommendation probability are on the same dimension, thereby guaranteeing the update accuracy of the content recommendation model. It should also be noted that each object in the object set corresponds to one historical recommendation result, and the recommendation time period for historical recommendation results of different objects can be the same.

[0089] (2) When the target sample pair contains interactive content samples of the same object, the current recommendation result and the historical recommendation result are combined to obtain the recommendation result pair.

[0090] For example, when a target sample pair contains interactive content samples of the same object, the current recommendation result of each interactive content sample in the target sample pair is combined with the historical recommendation result to obtain a recommendation result pair.

[0091] Since the target sample pair contains two interactive content samples, the recommendation result pair can include each interactive content sample's own recommendation result pair with its historical recommendation result. For example, if the target sample pair includes interactive content sample A and interactive content sample B corresponding to object u, with the current recommendation result for interactive content sample A being recommendation result A, the current recommendation result for interactive content sample B being recommendation result B, and the historical recommendation result for object u being recommendation result C, then the recommendation result pair can include either the recommendation result pair consisting of recommendation result A and recommendation result C, or it can include the recommendation result pair consisting of recommendation result B and recommendation result C.

[0092] (3) When the target sample pair contains interactive content samples of different objects, the current recommendation result is adjusted based on the historical recommendation result to obtain the recommendation result pair.

[0093] For example, the difference between historical recommendation results and current recommendation results can be calculated to obtain the adjusted current recommendation results. The adjusted current recommendation results can then be combined to obtain recommendation result pairs.

[0094] There are several ways to calculate the difference between the historical recommendation result and the current recommendation result. For example, the current recommendation result can be subtracted from the historical recommendation result to obtain the difference between the historical recommendation result and the current recommendation result, and this difference can be used as the current recommendation result. Specifically, it can be shown in formula (1):

[0095]

[0096] in, The adjusted current recommendation result is represented by p(u,a), which is the current recommendation result. For historical recommendation results, u is the object identifier, and a is the interaction content. The interaction content of the current recommendation result can be the same as or different from the interaction content in the historical recommendation results.

[0097] It's important to note that since the target sample pair contains interaction content samples from different objects, and the differences in behavioral habits among these objects can lead to significant discrepancies in the corresponding predicted recommendation results, the subsequent interaction confidence loss needs to compare the predicted recommendation results corresponding to the interaction content samples of different objects. If the predicted recommendation results for different objects are already significantly different, comparison will be impossible, or the differences will be too large. Therefore, when constructing the recommendation result pairs for different objects, normalization can be used to smooth out the behavioral differences between different objects, ensuring that the predicted recommendation results for different objects are within the same computational dimension. This normalization can be understood as normalizing the predicted recommendation results of the interaction content samples according to the average estimated value (historical recommendation results), making the current recommendation results for different objects in the target sample pair comparable, thereby obtaining a more accurate interaction confidence loss.

[0098] Optionally, in some embodiments, the method of constructing a recommendation result pair based on the current recommendation result and the target sample pair may also include: when the target sample pair contains interactive content samples of the same object, the current recommendation results are combined to obtain a recommendation result pair. For example, taking two different interactive content samples A and B that contain the same object u in the target sample pair, with the current recommendation result corresponding to interactive content sample A being recommendation result A and the current recommendation result corresponding to interactive content sample B being recommendation result B, recommendation result A and recommendation result B can be combined to obtain a recommendation result pair {recommendation result A, recommendation result B}, and so on.

[0099] C2. Calculate the difference between the recommendations in the recommendation result pair to obtain the recommendation difference.

[0100] The recommendation difference can be the difference between recommendation results. For example, taking an interaction as a click, the recommendation result can be the click-through rate estimate, and the corresponding recommendation difference can be the difference between the click-through rate estimates.

[0101] There are several ways to calculate the difference between recommendation results in a pair, as follows:

[0102] For example, the recommended results in a pair can be sorted based on the interaction duration and the interaction duration threshold. Based on the sorting results, the difference between the recommended results in the pair can be calculated to obtain the recommendation difference.

[0103] There are several ways to sort the recommendations in a pair of recommendations based on interaction duration and an interaction duration threshold. For example, the recommendations can be sorted according to the interaction duration corresponding to each recommendation in the pair. For instance, when a pair contains the current recommendation, it can be sorted according to the interaction duration corresponding to the current recommendation. When a pair contains an adjusted current recommendation, it can be sorted according to the interaction duration corresponding to the adjusted current recommendation before the adjustment. When a pair contains both the current and historical recommendations, the interaction duration threshold can be used as the interaction duration corresponding to the historical recommendation, and then the current and historical recommendations can be sorted according to their respective interaction durations.

[0104] After sorting the recommendations in a pair, the difference between the recommendations in the pair can be calculated based on the sorting results. There are several ways to calculate the difference between recommendations in a pair. For example, based on the sorting results, a target recommendation with a short interaction duration can be selected from the pair. This target recommendation can be used as the minuend, and the other recommendation in the pair can be used as the subtrahend. The recommendation difference can then be calculated based on the minuend and the subtrahend.

[0105] For example, if a pair of recommendation results contains recommendation results A and B corresponding to the same object, and the interaction duration of recommendation result A is less than the interaction duration of recommendation result B, then the recommendation difference can be the difference between recommendation result A and recommendation result B. Similarly, if a pair of recommendation results contains recommendation result A and historical recommendation result B corresponding to the same object, and the interaction duration of recommendation result A is less than the threshold of the interaction duration of historical recommendation result B, then the recommendation difference can be the difference between recommendation result A and historical recommendation result B. Or, if a pair of recommendation results contains adjusted recommendation result A and adjusted recommendation result B corresponding to different objects, and the interaction duration of adjusted recommendation result A before adjustment is less than the interaction duration of adjusted recommendation result B before adjustment, then the recommendation difference can be the difference between adjusted recommendation result A and adjusted recommendation result B, and so on.

[0106] C3. Correct the recommendation difference to obtain the interaction confidence loss.

[0107] For example, the training penalty value corresponding to the current training set can be obtained, the training penalty value can be fused with the recommendation difference to obtain the target recommendation difference, the target recommendation difference can be compared with the preset recommendation difference, and the interaction confidence loss can be determined based on the comparison result.

[0108] The training penalty value can be understood as a penalty term during model training, or as a supplementary adjustment to the loss function, primarily used to alleviate overfitting. The training penalty value can be a pre-defined penalty term or a randomly generated penalty term. The training penalty value can be 0 or any other arbitrary value. The training penalty values ​​for different training sets can be the same or different. There are several ways to fuse the training penalty value with the recommendation difference; for example, the training penalty value can be directly added to the recommendation difference to obtain the target recommendation difference.

[0109] The preset recommendation difference can be a pre-defined value used to compare with the calculated target recommendation difference, thereby selecting the recommendation difference used to determine the interaction confidence loss. After comparing the target recommendation difference with the preset recommendation difference, the interaction confidence loss can be determined based on the comparison result in several ways. For example, when the target recommendation difference exceeds the preset recommendation difference, the target recommendation difference is fused with a pre-set confidence parameter to obtain the interaction confidence loss; when the target recommendation difference does not exceed the preset recommendation difference, the preset recommendation difference is fused with a pre-set confidence parameter to obtain the interaction confidence loss.

[0110] In this case, taking the recommendation results A and B corresponding to two different interactive content samples of the same object u as an example, and the interaction duration corresponding to recommendation result A is less than the interaction duration corresponding to recommendation result B, with the preset recommendation difference value of 0, the interaction confidence loss of the recommendation result pair can be shown in formula (2), and can be as follows:

[0111] L1=λ|max(f clicj (v - ,w c )-f click (v + ,w c )+τ),0| (2)

[0112] Where L1 is the interaction confidence loss, f click (v _ ,w c ) represents the recommended result A, f click (v + ,w c ) represents recommendation result B, v - v represents the sample of interactive content whose interaction duration is less than the interaction duration threshold. + w represents the interactive content samples whose interaction duration is greater than or equal to an interaction duration threshold. c τ represents the training weights, and τ is the training penalty value.

[0113] Taking the current recommendation result A and the historical recommendation result B of the same object u as an example, with a preset recommendation difference of 0, when the interaction duration corresponding to the current recommendation result A is less than the interaction duration threshold, the interaction confidence loss corresponding to the recommendation result pair can be as shown in formula (3), specifically as follows:

[0114]

[0115] Where L1 is the interaction confidence loss, f click (u,w c ( ) represents the current recommendation result A. Let B be the historical recommendation result, u be the corresponding object, τ be the training penalty value, and w be the value of the object. c These are the training weights.

[0116] Optionally, when the interaction duration corresponding to the current recommendation result A is greater than the interaction duration threshold, the confidence loss of the corresponding interaction for this recommendation result can be as shown in formula (4), specifically as follows:

[0117]

[0118] Where L1 is the interaction confidence loss, f click(u,w c ( ) represents the current recommendation result A. Let B be the historical recommendation result, u be the corresponding object, τ be the training penalty value, and w be the value of the object. c These are the training weights.

[0119] Among them, taking the adjusted recommendation results A and B corresponding to the interaction content samples of different objects, the interaction duration corresponding to the adjusted recommendation result A before adjustment is less than the interaction duration corresponding to the adjusted recommendation result B before adjustment, with the preset recommendation difference value of 0 as an example, the interaction confidence loss of the corresponding recommendation result can be shown in formula (5), and can be specifically as follows:

[0120]

[0121] Where L1 is the interaction confidence loss, The adjusted recommendation result is A. For the adjusted recommendation result B, u - For the object corresponding to recommendation result A, u + Let a1 be the object corresponding to recommendation result B, a2 be the interaction content sample corresponding to recommendation result A, and a3 be the interaction content sample corresponding to recommendation result B.

[0122] 104. Based on the interaction confidence loss and the predicted recommendation results, determine the target recommendation loss of the content recommendation model.

[0123] The target recommendation loss can include the loss generated by the content recommendation model when making content recommendations on interactive content samples in the current training set.

[0124] There are several ways to determine the target recommendation loss of a content recommendation model based on interaction confidence loss and predicted recommendation results, as follows:

[0125] For example, the labeled recommendation results corresponding to the interactive content samples can be obtained, and the labeled recommendation results can be compared with the corresponding predicted recommendation results to obtain the initial recommendation loss. The initial recommendation loss and the interaction confidence loss can be fused to obtain the target recommendation loss.

[0126] The labeled recommendation result can be the recommendation result labeled on the interactive content sample. The labeled recommendation result indicates whether the interactive content sample was actually recommended to the corresponding object, or it can indicate whether the object actually clicked on the interactive content sample or converted based on it after the interactive content sample was pushed to the corresponding object. The labeled recommendation result can be a binary label, for example, 1 if the object clicked the interactive content sample, and 0 if it did not click. The labeled recommendation result can also be a multivariate label, using multivariate recommendation probability values ​​to characterize the degree to which the object actually clicked or converted the interactive content sample. For example, it can be any value between [0,1], where a larger value indicates a higher degree of actual click or conversion by the user, and so on.

[0127] The initial recommendation loss can be obtained by using clicks or conversions as supervisory signals; it is a supervised learning loss. There are several ways to compare the labeled recommendation results with the corresponding predicted recommendation results to obtain the initial recommendation loss. For example, the cross-entropy loss function can be used, or other types of loss functions can be employed.

[0128] After comparing the labeled recommendation results with the corresponding predicted recommendation results to obtain the initial recommendation loss, the initial recommendation loss can be fused with the interaction confidence loss to obtain the target recommendation loss. There are several ways to fuse the initial recommendation loss and the interaction confidence loss. For example, one can select the initial recommendation loss for each interactive content sample in the target sample pair corresponding to the interaction confidence loss from the initial recommendation loss, obtain the candidate recommendation loss corresponding to the interaction confidence loss, add the candidate recommendation loss to the interaction confidence loss to obtain the current recommendation loss corresponding to the interactive content sample in the target sample pair or the intersection of the target sample pair, and accumulate the current recommendation loss to obtain the target recommendation loss corresponding to the current training set. Alternatively, one can select the initial recommendation loss for each interactive content sample in the target sample pair corresponding to the interaction confidence loss from the initial recommendation loss, obtain the candidate recommendation loss corresponding to the interaction confidence loss, obtain the training weights, and weight the candidate recommendation loss and the interaction confidence loss based on the training weights, and add the weighted candidate recommendation loss and the weighted interaction confidence loss to obtain the current recommendation loss corresponding to the target sample pair, and accumulate the current recommendation loss to obtain the target recommendation loss corresponding to the current training set, and so on.

[0129] The current recommendation loss can include the initial recommendation loss corresponding to the interactive content samples in the target sample pair and the interaction confidence loss of the corresponding recommendation result pair. Taking the cross-entropy loss as the loss function corresponding to the initial recommendation loss as an example, it can be specifically shown in formula (6), and can be further described as follows:

[0130] L = L CE (y click ,f click (v,w c ))+L1(6)

[0131] Where L is the current recommended loss, L CE (y click ,f click (v,w c )) represents the initial recommendation loss, y click To label the recommendation results, f click (v,w c ) represents the predicted recommendation result, v represents the content features (embedding) of the interactive content sample, and w represents the prediction result. c L1 is the training weight, and L1 is the interaction confidence loss for the target sample pair.

[0132] It should be noted that, for the initial recommendation loss, taking the target sample pair corresponding to the interaction confidence loss as including interaction content sample A and interaction content sample B as an example, the initial recommendation loss corresponding to the interaction confidence loss can include the initial recommendation losses corresponding to these two interaction content samples. Furthermore, it should be noted that a target sample pair can correspond to one recommendation result pair or two recommendation result pairs; therefore, there can be one or multiple interaction confidence losses corresponding to a target sample pair. For example, consider a target sample pair including interactive content sample A and interactive content sample B, with recommendation result A corresponding to interactive content sample A and recommendation result B corresponding to interactive content sample B. The corresponding recommendation result pair can include {recommendation result A, recommendation result B}, or it can include {recommendation result A, historical recommendation result} and {recommendation result B, historical recommendation result}. When the recommendation result pair is {recommendation result A, recommendation result B}, the corresponding initial recommendation loss can include the initial recommendation loss corresponding to interactive content sample A and interactive content sample B respectively. When the recommendation result pair is {recommendation result A, historical recommendation result} and {recommendation result B, historical recommendation result}, then the initial recommendation loss corresponding to {recommendation result A, historical recommendation result} can be the initial recommendation loss corresponding to interactive content sample A, and the initial recommendation loss corresponding to {recommendation result B, historical recommendation result} can be the initial recommendation loss corresponding to interactive content sample B. When determining the current recommendation loss, the initial recommendation loss of the same interactive content sample can be added to the corresponding interaction confidence loss to obtain the current recommendation loss for that interactive content sample. Alternatively, the initial recommendation loss of the interactive content sample in the same target sample pair can be added to the corresponding interaction confidence loss to obtain the current recommendation loss for the target sample pair.

[0133] After determining the current recommendation loss, the current recommendation loss can be accumulated to obtain the target recommendation loss of the current training set, and this target recommendation loss can be used as the target recommendation loss of the content recommendation model.

[0134] In training content recommendation models, supervised learning loss functions (i.e., the initial recommendation loss corresponding to the cross-entropy loss function) are typically used. This approach, however, adds a new contrastive learning / self-supervised learning (SL) loss function (i.e., the corresponding interaction confidence loss). Taking a DNN-based content recommendation model as an example, a comparison between existing training methods and this approach can be seen as follows: Figure 3As shown, by explicitly learning the partial order relationship based on the interaction duration, the prediction and recommendation results of interaction content samples corresponding to different interaction durations can be adjusted. This makes the prediction and recommendation results of interaction content samples with shorter interaction durations smaller than those of interaction content samples with longer interaction durations. This allows for a more accurate portrayal of the true intent of the object's interaction behavior, improves the training accuracy of the content recommendation model, and ultimately enhances the recommendation efficiency of the content recommendation.

[0135] 105. Update the content recommendation model based on the target recommendation loss, and use the updated content recommendation model to recommend the content to be recommended to the target object in the object set.

[0136] The content to be recommended can be any content that needs to be recommended. The types of content to be recommended can be various, such as text, audio and video, hyperlinks, or other content that can be pushed or exposed, etc.

[0137] There are several ways to update the content recommendation model based on the target recommendation loss, as follows:

[0138] For example, a gradient descent algorithm can be used to update the network parameters of the content recommendation model based on the target recommendation loss, resulting in an updated content recommendation model. This updated model is then used as the final content recommendation model, and the process of retrieving the current training set for the content recommendation model is repeated until the model converges, yielding the updated target recommendation model. Alternatively, other network parameter update algorithms can be used to update the network parameters of the content recommendation model based on the target recommendation loss, resulting in an updated content recommendation model. This updated model is then used as the final content recommendation model, and the process of retrieving the current training set for the content recommendation model is repeated until the model converges, yielding the updated target recommendation model.

[0139] After updating the content recommendation model, the updated model can be used to recommend the content to be recommended to the target objects in the object set. There are several ways to recommend content to the target objects in the object set. For example, the content to be recommended can be obtained, and the updated content recommendation model can be used to predict the current recommendation probability of the content to be recommended for each object in the object set. Based on the current recommendation probability, the target objects corresponding to the content to be recommended can be selected from the object set, and the content to be recommended can be recommended to the target objects.

[0140] There are several ways to obtain the content to be recommended. For example, you can directly obtain at least one piece of content uploaded by the terminal or client to obtain the content to be recommended. Alternatively, you can filter at least one piece of content from the network or content library to obtain the content to be recommended. Or, you can obtain at least one piece of published content from the content interaction platform to obtain the content to be recommended. Or, when there is a large number of content to be recommended or the memory is large, you can also receive a content recommendation request, which carries the storage address of the content to be recommended. Based on the storage address, you can obtain the content to be recommended, and so on.

[0141] After acquiring the content to be recommended, and using the updated content recommendation model, the current recommendation probability of the content to be recommended for each object in the object set is predicted. The method for predicting the current recommendation probability is similar to the method for predicting the recommendation probability of interactive content samples for objects, as described above, and will not be repeated here.

[0142] After predicting the recommendation probability, the target object corresponding to the content to be recommended can be selected from the object set based on the current recommendation probability. There are several ways to select the target object. For example, at least one object with a current recommendation probability exceeding a preset probability threshold can be selected from the object set to obtain the target object. Alternatively, the objects in the object set can be classified based on the current recommendation probability, and at least one target object corresponding to a preset ranking range (e.g., Top N, where N is a positive integer) can be selected from the object set based on the classification results to obtain the target object, and so on.

[0143] After identifying the target audience, the content to be recommended can be presented to them. There are several ways to present the content to the target audience; for example, the content can be directly pushed to the target audience's corresponding device, or the content can be displayed on the target audience's corresponding device. Specifically... Figure 4 As shown, etc.

[0144] There are several ways to display the recommended content on the target's terminal or client. For example, it can be displayed as a patch on the page currently being viewed by the target, or it can be displayed directly in a preset area on the target's terminal or client. Alternatively, the content can be displayed as a hyperlink on the target's terminal or client, and the terminal or client can respond to the trigger of the hyperlink to display the recommended content. Or, the recommended content can be displayed on the target's terminal or client. If a preset time threshold is exceeded and the terminal or client does not receive a trigger operation from the target, the recommended content can be withdrawn, set to transparent, or the display of the recommended content can be stopped.

[0145] After receiving the content to be recommended, the target audience can interact with it through a terminal or client. There are various ways to interact, such as clicking or converting.

[0146] As can be seen from the above, in this embodiment, after obtaining the current training set of the content recommendation model, the current training set includes the interaction content samples of each object in the object set and the interaction duration corresponding to the interaction content samples. The content recommendation model is used to predict the recommendation probability of the interaction content samples for the objects, obtaining the predicted recommendation result. Then, based on the interaction duration, the interaction content samples are classified, and the predicted recommendation results corresponding to different categories of interaction content samples are compared to obtain the interaction confidence loss. Then, based on the interaction confidence loss and the predicted recommendation result, the target recommendation loss of the content recommendation model is determined. Finally, the content recommendation is performed according to the target recommendation loss. The recommendation model is updated, and the updated content recommendation model is used to recommend the content to be recommended to the target objects in the object set. Since this scheme can measure the confidence of the interaction behavior by the interaction duration and classify the interaction content samples by the interaction duration, the predicted recommendation results corresponding to the interaction content samples of different categories are compared to obtain the interaction confidence loss. Through the interaction confidence loss, the content recommendation model can explicitly learn the partial order relationship based on the interaction duration, which can accurately characterize the true intent of the interaction behavior between objects and content, improve the accuracy of the content recommendation model, and thus improve the recommendation efficiency.

[0147] Based on the method described in the above embodiments, the following examples will provide further detailed explanations.

[0148] In this embodiment, the content recommendation device is specifically integrated into an electronic device, the electronic device is a server, the object set is a user set, the object is a user, the interactive content sample is an advertisement sample, the interaction is a click, the interaction duration is the click dwell time, the content to be recommended is an advertisement, the content recommendation model is an advertisement recommendation model, the recommendation probability is the click-through rate, and the predicted recommendation result is the click-through rate estimate.

[0149] like Figure 5 As shown, a content recommendation method has the following specific process:

[0150] 201. The server obtains the current training set of the advertising recommendation model.

[0151] For example, the server can directly receive the current training set of the advertising recommendation model uploaded by the terminal or client. Alternatively, it can obtain the click records of each user in the user set, filter the corresponding advertising samples for each user based on the click records, and calculate the click dwell time corresponding to the advertising samples in the click records to obtain the current training set of the advertising recommendation model. Or, it can filter the advertisements corresponding to at least one user in the network or database, obtain advertising samples, and determine the click dwell time corresponding to the advertising samples to obtain the current training set of the advertising recommendation model. Or, when the memory of the current training set is large or the number of samples is large, it can also receive update requests for the advertising recommendation model. The update request may include the storage address of the current training set, and obtain the current training set based on the storage address, and so on.

[0152] 202. Using an advertising recommendation model, predict the click-through rate of advertising samples for users and obtain the estimated click-through rate.

[0153] For example, the server can use an advertising recommendation model to extract features from advertising samples, obtain content features in multiple dimensions, fuse the content features to obtain the click features of the advertising sample for the corresponding user in the user set, and determine the click-through rate of the advertising sample for that user based on the click features, thereby obtaining the estimated click-through rate of the advertising sample.

[0154] 203. The server categorizes the ad samples based on the click-through time.

[0155] For example, the server can obtain the historical click-and-dwell time of each user's clicks with at least one historical ad sample, determine the average historical click-and-dwell time, and use the average time as the initial click-and-dwell time threshold for each user. This initial click-and-dwell time threshold is then used as the click-and-dwell time threshold for the current training set. Alternatively, it can obtain a preset click-and-dwell time threshold and use it as the click-and-dwell time threshold for the current training set. Or, it can obtain the ad types of the ad samples in the current training set, filter out the time thresholds corresponding to the content types from the preset time threshold set, and thus obtain the click-and-dwell time threshold for the current training set. Alternatively, it can also filter out the time thresholds corresponding to the user identifiers from the preset time threshold set based on the user identifiers of each user in the user set, and so on.

[0156] The server compares the click-and-dwell time with the click-and-dwell time threshold, and selects the ad samples with click-and-dwell time greater than or equal to the threshold as the first category sample, thus obtaining the first category sample set. The ad samples with click-and-dwell time less than the threshold are selected as the second category sample, thus obtaining the second category sample set.

[0157] 204. The server constructs the estimated value pairs based on the first category sample set and the second category sample set.

[0158] For example, the server can select one ad sample from the first category sample set and one from the second category sample set to obtain a target sample pair. Then, it can filter out the estimated click-through rate (CTR) of the ad sample in the target sample pair from the estimated CTR, and obtain the current estimated CTR.

[0159] The server can filter candidate users corresponding to ad samples in the target sample pair from the user set, obtain at least one historical ad sample corresponding to the candidate user, and use an ad recommendation model to predict the click-through rate (CTR) corresponding to the historical ad sample to obtain the historical CTR. The probability mean of the historical CTR is determined and used as the historical estimated value for the candidate user.

[0160] When the target sample pair contains ad samples from the same user, the server combines the current estimated value of each ad sample in the target sample pair with the historical estimated value to obtain an estimated value pair. Alternatively, when the target sample pair contains ad samples from different users, the server calculates the difference between the historical estimated value and the current estimated value to obtain an adjusted current estimated value, and combines the adjusted current estimated values ​​to obtain an estimated value pair.

[0161] 205. The server calculates the difference between the estimated values ​​in the estimated value pair to obtain the estimated difference.

[0162] For example, when a prediction pair contains the current prediction, the server can sort the current predictions according to the click-and-dwell time corresponding to the current prediction. When a prediction pair contains the adjusted current prediction, the server can sort the adjusted current prediction according to the click-and-dwell time corresponding to the adjusted current prediction before the adjustment. When a prediction pair contains both the current prediction and historical predictions, the click-and-dwell time threshold can be used as the click-and-dwell time corresponding to the historical prediction. Then, the current prediction and the historical prediction can be sorted according to the click-and-dwell time corresponding to the current prediction and the historical prediction, respectively.

[0163] Based on the sorting results, the server can filter out the target estimated value with the shorter click dwell time from the estimated value pair, use the target estimated value as the minuend, use the other estimated value in the estimated value pair as the subtrahend, and calculate the estimated difference between the estimated values ​​based on the minuend and the subtrahend.

[0164] 206. The server corrects the estimated difference to obtain the interaction confidence loss.

[0165] For example, the server can obtain the training penalty value corresponding to the current training set, add the training penalty value to the estimated difference, and thus obtain the target estimated difference. The target estimated difference is compared with the preset estimated difference. When the target estimated difference exceeds the preset estimated difference, the target estimated difference is fused with the preset confidence parameter to obtain the interactive confidence loss. When the target estimated difference does not exceed the preset estimated difference, the preset estimated difference is fused with the preset confidence parameter to obtain the interactive confidence loss. Specifically, it can be shown in formulas (2), (3), (4) and (5), as described above, and will not be repeated here.

[0166] 207. The server determines the target click loss of the advertising recommendation model based on the interaction confidence loss and the click-through rate prediction.

[0167] For example, the server can obtain the labeled estimated values ​​corresponding to the ad samples. Using the cross-entropy loss function, the labeled results are compared with the corresponding estimated click-through rate (CTR) values ​​to obtain the initial recommendation loss. Alternatively, other types of loss functions can be used to compare the labeled results with the corresponding estimated CTR values ​​to obtain the initial recommendation loss.

[0168] The server selects the initial recommendation loss for each ad sample in the target sample pair corresponding to the interaction confidence loss from the initial recommendation loss, obtains the candidate recommendation loss corresponding to the interaction confidence, adds the candidate recommendation loss to the interaction confidence loss, and obtains the current recommendation loss corresponding to the ad sample in the target sample pair or the target sample pair. The current recommendation loss is accumulated to obtain the target recommendation loss corresponding to the current training set. Alternatively, the server can select the initial recommendation loss for each ad sample in the target sample pair corresponding to the interaction confidence loss from the initial recommendation loss, obtain the candidate recommendation loss corresponding to the interaction confidence, obtain the training weights, and weight the candidate recommendation loss and the interaction confidence loss respectively based on the training weights. The weighted candidate recommendation loss and the weighted interaction confidence loss are added to obtain the current recommendation loss corresponding to the target sample pair. The current recommendation loss is accumulated to obtain the target recommendation loss corresponding to the current training set. The specific details are shown in formula (6), as described above, and will not be repeated here.

[0169] 208. The server updates the ad recommendation model based on the target recommendation loss.

[0170] For example, the server can use the gradient descent algorithm to update the network parameters of the advertising recommendation model based on the target recommendation loss, obtaining an updated advertising recommendation model. This updated model is then used as the final advertising recommendation model, and the server returns to the step of obtaining the current training set for the advertising recommendation model. This process continues until the advertising recommendation model converges, resulting in the updated target recommendation model. Alternatively, other network parameter update algorithms can be used to update the network parameters of the advertising recommendation model based on the target recommendation loss, obtaining an updated advertising recommendation model. This updated model is then used as the final advertising recommendation model, and the server returns to the step of obtaining the current training set for the advertising recommendation model. This process continues until the advertising recommendation model converges, resulting in the updated target recommendation model.

[0171] 209. The server uses the updated advertising recommendation model to recommend advertisements to target users in the user set.

[0172] For example, the server can directly obtain at least one piece of content uploaded by the terminal or client to obtain an advertisement; or, it can filter at least one piece of content from the network or content library to obtain an advertisement; or, it can obtain at least one piece of published content from the content interaction platform to obtain an advertisement; or, when there are many advertisements or a large amount of memory, it can also receive content recommendation requests that carry the storage address of the advertisement, and obtain the advertisement based on the storage address, and so on.

[0173] The server uses an updated ad recommendation model to predict the current click-through rate (CTR) of ads for each user in the user set. It then selects at least one user whose CTR exceeds a preset probability threshold to obtain the target user. Alternatively, it can classify the users in the user set based on the current CTR, and based on the classification results, select at least one target user corresponding to a preset ranking range (e.g., Top N, where N is a positive integer) to obtain the target user, and so on.

[0174] The server can directly push advertisements to the target user's corresponding terminal, or it can display advertisements on the target user's corresponding terminal, and so on.

[0175] After receiving the advertisement, the target user can click on or convert the advertisement through a terminal or client.

[0176] As described above, in this embodiment, after the server obtains the current training set of the advertising recommendation model, this current training set includes the advertising samples of each user in the user set and the corresponding click-through time of the advertising samples. The advertising recommendation model is used to predict the click-through rate (CTR) of the advertising samples for their target audience, obtaining a CTR prediction value. Then, based on the click-through time, the advertising samples are classified, and the CTR prediction values ​​corresponding to different categories of advertising samples are compared to obtain the interaction confidence loss. Then, based on the interaction confidence loss and the predicted recommendation results, the target recommendation loss of the advertising recommendation model is determined. Finally, the advertising recommendation model is adjusted according to the target recommendation loss. The system is updated, and the updated advertising recommendation model is used to recommend content to target users in the object set. Since this approach measures the confidence of click behavior based on click-dwell time and classifies advertising samples by click-dwell time, the predicted click values ​​for different categories of advertising samples are compared to obtain the interaction confidence loss. This interaction confidence loss allows the advertising recommendation model to explicitly learn the partial order relationship based on click-dwell time, accurately characterizing the true intent of user click behavior with ads, improving the accuracy of the advertising recommendation model, and thus improving the recommendation efficiency.

[0177] To better implement the above methods, embodiments of the present invention also provide a content recommendation device, which can be integrated into an electronic device, such as a server or terminal, and the terminal may include a tablet computer, a laptop computer, and / or a personal computer.

[0178] For example, such as Figure 6 As shown, the content recommendation device may include an acquisition unit 301, a prediction unit 302, a comparison unit 303, a determination unit 304, and a recommendation unit 305, as follows:

[0179] (1) Obtain unit 301;

[0180] The acquisition unit 301 is used to acquire the current training set of the content recommendation model. The current training set includes the interaction content samples of each object in the object set and the interaction duration corresponding to the interaction content samples.

[0181] For example, the acquisition unit 301 can be used to receive the current training set of the content recommendation model uploaded by the terminal or client, or it can acquire the interaction record of each object in the object set, and based on the interaction record, filter the interaction content samples of each object in the network or content database, and count the interaction duration corresponding to the interaction content samples in the interaction record to obtain the current training set of the content recommendation model. Alternatively, it can filter the interaction content corresponding to at least one object in the network or database, obtain the interaction content samples, and determine the interaction duration corresponding to the interaction content samples to obtain the current training set of the content recommendation model. Or, when the memory of the current training set is large or the number of samples is large, it can also receive an update request for the content recommendation model. The update request may include the storage address of the current training set, and based on the storage address, acquire the current training set, and so on.

[0182] (2) Prediction unit 302;

[0183] The prediction unit 302 is used to use a content recommendation model to predict the recommendation probability of interactive content samples for an object, and obtain the prediction recommendation result.

[0184] For example, prediction unit 302 can be used to extract features from interactive content samples using a content recommendation model, obtain content features in multiple dimensions, fuse the content features to obtain recommendation features for the interactive content sample for the corresponding object in the object set, determine the recommendation probability of the interactive content sample for the object based on the recommendation features, and thus obtain the prediction recommendation result corresponding to the interactive content sample.

[0185] (3) Comparison unit 303;

[0186] The comparison unit 303 is used to classify the interaction content samples according to the interaction duration and compare the prediction recommendation results corresponding to the interaction content samples of different categories to obtain the interaction confidence loss.

[0187] For example, the comparison unit 303 can be used to obtain the interaction duration threshold corresponding to the current training set, classify the interaction content samples based on the interaction duration threshold and the interaction duration, construct recommendation result pairs based on the classified first category sample set and second category sample set, calculate the difference between the recommendation results in the recommendation result pair, obtain the recommendation difference, and correct the recommendation difference to obtain the interaction confidence loss.

[0188] (4) Determine unit 304;

[0189] Unit 304 is used to determine the target recommendation loss of the content recommendation model based on the interaction confidence loss and the predicted recommendation results.

[0190] For example, the determination unit 304 can be used to obtain the labeled recommendation results corresponding to the interactive content samples, compare the labeled recommendation results with the corresponding predicted recommendation results to obtain the initial recommendation loss, and fuse the initial recommendation loss with the interaction confidence loss to obtain the target recommendation loss.

[0191] (5) Recommended Unit 305;

[0192] Recommendation unit 305 is used to update the content recommendation model based on the target recommendation loss, and then use the updated content recommendation model to recommend the content to be recommended to the target object in the object set.

[0193] For example, recommendation unit 305 can be used to update the content recommendation model based on the target recommendation loss, obtain the content to be recommended, and use the updated content recommendation model to predict the current recommendation probability of the content to be recommended for each object in the object set. Based on the current recommendation probability, the target object corresponding to the content to be recommended is selected from the object set, and the content to be recommended is recommended to the target object.

[0194] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.

[0195] As can be seen from the above, in this embodiment, after the acquisition unit 301 acquires the current training set of the content recommendation model, the current training set includes the interaction content samples of each object in the object set and the interaction duration corresponding to the interaction content samples. The prediction unit 302 uses the content recommendation model to predict the recommendation probability of the interaction content samples for the objects, and obtains the prediction recommendation result. Then, the comparison unit 303 classifies the interaction content samples according to the interaction duration and compares the prediction recommendation results corresponding to the interaction content samples of different categories to obtain the interaction confidence loss. Then, the determination unit 304 determines the target recommendation loss of the content recommendation model based on the interaction confidence loss and the prediction recommendation result. Finally, the recommendation unit 304 makes the recommendation of the target recommendation of the content recommendation model. Yuan305 updates the content recommendation model based on the target recommendation loss and uses the updated model to recommend content to target objects in the object set. Since this approach measures the confidence of interaction behavior by interaction duration and classifies interaction content samples by duration, it compares the predicted recommendation results for different categories of interaction content samples to obtain the interaction confidence loss. This interaction confidence loss allows the content recommendation model to explicitly learn the partial order relationship based on interaction duration, accurately characterizing the true intent of the interaction between objects and content, improving the accuracy of the content recommendation model, and ultimately enhancing the recommendation efficiency.

[0196] This invention also provides an electronic device, such as... Figure 7 As shown, it illustrates a structural schematic diagram of the electronic device involved in an embodiment of the present invention, specifically:

[0197] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0198] The processor 401 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes software programs and / or modules stored in the memory 402, and calls data stored in the memory 402, to perform various functions and process data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0199] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0200] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0201] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0202] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:

[0203] Obtain the current training set of the content recommendation model. The current training set includes interaction content samples and corresponding interaction durations for each object in the object set. Using the content recommendation model, predict the recommendation probability of the interaction content samples for the object and obtain the prediction recommendation result. Classify the interaction content samples according to the interaction duration and compare the prediction recommendation results corresponding to the interaction content samples of different categories to obtain the interaction confidence loss. Based on the interaction confidence loss and the prediction recommendation result, determine the target recommendation loss of the content recommendation model. Update the content recommendation model according to the target recommendation loss and use the updated content recommendation model to recommend the content to be recommended to the target object in the object set.

[0204] For example, an electronic device can acquire the current training set of a content recommendation model. This current training set includes interaction content samples for each object in the object set and the corresponding interaction duration. The content recommendation model extracts features from the interaction content samples, obtaining multi-dimensional content features. These features are then fused to obtain recommendation features for the interaction content sample relative to the corresponding object in the object set. Based on these recommendation features, the recommendation probability for the interaction content sample relative to that object is determined, thus obtaining the predicted recommendation result for the interaction content sample. The interaction duration threshold corresponding to the current training set is acquired. Based on the interaction duration threshold and the interaction duration, the interaction content samples are classified. Based on the first and second category sample sets, recommendation result pairs are constructed. The difference between the recommendation results in the recommendation result pair is calculated to obtain the recommendation difference. The recommendation difference is then corrected to obtain the interaction confidence loss. The labeled recommendation results corresponding to the interaction content samples are acquired and compared with the corresponding predicted recommendation results to obtain the initial recommendation loss. The initial recommendation loss and the interaction confidence loss are then fused to obtain the target recommendation loss. The content recommendation model is updated based on the target recommendation loss. The content to be recommended is obtained, and the updated content recommendation model is used to predict the current recommendation probability of the content to be recommended for each object in the object set. Based on the current recommendation probability, the target object corresponding to the content to be recommended is selected from the object set, and the content to be recommended is recommended to the target object.

[0205] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0206] As can be seen from the above, in this embodiment, after obtaining the current training set of the content recommendation model, the current training set includes the interaction content samples of each object in the object set and the interaction duration corresponding to the interaction content samples. The content recommendation model is used to predict the recommendation probability of the interaction content samples for the objects, obtaining the predicted recommendation result. Then, based on the interaction duration, the interaction content samples are classified, and the predicted recommendation results corresponding to different categories of interaction content samples are compared to obtain the interaction confidence loss. Then, based on the interaction confidence loss and the predicted recommendation result, the target recommendation loss of the content recommendation model is determined. Finally, the content recommendation is performed according to the target recommendation loss. The recommendation model is updated, and the updated content recommendation model is used to recommend the content to be recommended to the target objects in the object set. Since this scheme can measure the confidence of the interaction behavior by the interaction duration and classify the interaction content samples by the interaction duration, the predicted recommendation results corresponding to the interaction content samples of different categories are compared to obtain the interaction confidence loss. Through the interaction confidence loss, the content recommendation model can explicitly learn the partial order relationship based on the interaction duration, which can accurately characterize the true intent of the interaction behavior between objects and content, improve the accuracy of the content recommendation model, and thus improve the recommendation efficiency.

[0207] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0208] To this end, embodiments of the present invention provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the content recommendation methods provided in the embodiments of the present invention. For example, the instructions can execute the following steps:

[0209] Obtain the current training set of the content recommendation model. The current training set includes interaction content samples and corresponding interaction durations for each object in the object set. Using the content recommendation model, predict the recommendation probability of the interaction content samples for the object and obtain the prediction recommendation result. Classify the interaction content samples according to the interaction duration and compare the prediction recommendation results corresponding to the interaction content samples of different categories to obtain the interaction confidence loss. Based on the interaction confidence loss and the prediction recommendation result, determine the target recommendation loss of the content recommendation model. Update the content recommendation model according to the target recommendation loss and use the updated content recommendation model to recommend the content to be recommended to the target object in the object set.

[0210] For example, an electronic device can acquire the current training set of a content recommendation model. This current training set includes interaction content samples for each object in the object set and the corresponding interaction duration. The content recommendation model extracts features from the interaction content samples, obtaining multi-dimensional content features. These features are then fused to obtain recommendation features for the interaction content sample relative to the corresponding object in the object set. Based on these recommendation features, the recommendation probability for the interaction content sample relative to that object is determined, thus obtaining the predicted recommendation result for the interaction content sample. The interaction duration threshold corresponding to the current training set is acquired. Based on the interaction duration threshold and the interaction duration, the interaction content samples are classified. Based on the first and second category sample sets, recommendation result pairs are constructed. The difference between the recommendation results in the recommendation result pair is calculated to obtain the recommendation difference. The recommendation difference is then corrected to obtain the interaction confidence loss. The labeled recommendation results corresponding to the interaction content samples are acquired and compared with the corresponding predicted recommendation results to obtain the initial recommendation loss. The initial recommendation loss and the interaction confidence loss are then fused to obtain the target recommendation loss. The content recommendation model is updated based on the target recommendation loss. The content to be recommended is obtained, and the updated content recommendation model is used to predict the current recommendation probability of the content to be recommended for each object in the object set. Based on the current recommendation probability, the target object corresponding to the content to be recommended is selected from the object set, and the content to be recommended is recommended to the target object.

[0211] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0212] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0213] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the content recommendation methods provided in the embodiments of the present invention, the beneficial effects that any of the content recommendation methods provided in the embodiments of the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0214] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the content recommendation or content push aspects described above.

[0215] The foregoing has provided a detailed description of a content recommendation method, apparatus, electronic device, and computer-readable storage medium provided by embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A content recommendation method, characterized in that, include: Obtain the current training set of the content recommendation model, wherein the current training set includes the interaction content sample of each object in the object set and the interaction duration corresponding to the interaction content sample; Using the content recommendation model, the recommendation probability of the interactive content sample for the object is predicted, and the prediction recommendation result is obtained; Obtain the interaction duration threshold corresponding to the current training set; compare the interaction duration with the interaction duration threshold; Interaction content samples with an interaction duration greater than or equal to the interaction duration threshold are designated as first category samples to obtain a first category sample set; interaction content samples with an interaction duration less than the interaction duration threshold are designated as second category samples to obtain a second category sample set. Based on the first category sample set and the second category sample set, construct recommendation result pairs; Calculate the difference between the recommendation results in the pair to obtain the recommendation difference; The recommendation difference is corrected to obtain the interaction confidence loss; Based on the interaction confidence loss and the predicted recommendation results, the target recommendation loss of the content recommendation model is determined; The content recommendation model is updated based on the target recommendation loss, and the updated content recommendation model is used to recommend the content to be recommended to the target object in the object set.

2. The content recommendation method according to claim 1, characterized in that, The step of obtaining the interaction duration threshold corresponding to the current training set includes: Get the duration of historical interactions between each object and at least one historical interaction content; Determine the average duration of the historical interactions, and use the average duration as the initial interaction duration threshold for each object; The initial interaction duration threshold is used as the interaction duration threshold corresponding to the current training set.

3. The content recommendation method according to claim 1, characterized in that, The step of constructing recommendation result pairs based on the first category sample set and the second category sample set includes: One interactive content sample is selected from each of the first category sample set and the second category sample set to obtain the target sample pair; The predicted recommendation results are then filtered out from the predicted recommendation results to obtain the current recommendation results; Based on the current recommendation results and the target sample pair, construct the recommendation result pair.

4. The content recommendation method according to claim 3, characterized in that, The step of constructing the recommendation result pair based on the current recommendation result and the target sample pair includes: Obtain the historical recommendation results corresponding to the candidate objects, wherein the candidate objects include the objects corresponding to the interactive content samples in the target sample pair; When the target sample pair contains interactive content samples of the same object, the current recommendation result and the historical recommendation result are combined to obtain the recommendation result pair; When the target sample pair contains interactive content samples of different objects, the current recommendation result is adjusted based on the historical recommendation results to obtain the recommendation result pair.

5. The content recommendation method according to claim 4, characterized in that, The step of obtaining the historical recommendation results corresponding to the candidate object includes: From the object set, candidate objects corresponding to the interactive content samples in the target sample pair are selected; Obtain at least one historical recommendation probability of the candidate object, wherein the historical recommendation probability includes the recommendation probability of the candidate object based on the content recommendation model predicted by the content recommendation model for historical interaction content; The mean probability of the historical recommendation probabilities is determined, and the mean probability is used as the historical recommendation result corresponding to the candidate object.

6. The content recommendation method according to claim 4, characterized in that, The step of adjusting the current recommendation result based on the historical recommendation results to obtain the recommendation result pair includes: Calculate the difference between the historical recommendation results and the current recommendation results to obtain the adjusted current recommendation results; The adjusted current recommendation results are combined to obtain the recommendation result pair.

7. The content recommendation method according to claim 3, characterized in that, The step of constructing the recommendation result pair based on the current recommendation result and the target sample pair includes: When the target sample pair contains interactive content samples of the same object, the current recommendation results are combined to obtain the recommendation result pair.

8. The content recommendation method according to claim 1, characterized in that, The step of calculating the difference between the recommendation results in the recommendation result pair to obtain the recommendation difference includes: Based on the interaction duration and the interaction duration threshold, the recommendation results in the recommendation result pair are sorted. Based on the sorting results, the difference between the recommended results in the recommended result pair is calculated to obtain the recommendation difference.

9. The content recommendation method according to claim 1, characterized in that, The step of correcting the recommendation difference to obtain the interaction confidence loss includes: Obtain the training penalty value corresponding to the current training set; The training penalty value and the recommendation difference are fused together to obtain the target recommendation difference; The target recommendation difference is compared with the preset recommendation difference, and the interaction confidence loss is determined based on the comparison result.

10. The content recommendation method according to any one of claims 1 to 9, characterized in that, The step of determining the target recommendation loss of the content recommendation model based on the interaction confidence loss and the predicted recommendation result includes: Obtain the labeled recommendation results corresponding to the interactive content samples; The labeled recommendation results are compared with the corresponding predicted recommendation results to obtain the initial recommendation loss; The initial recommendation loss is fused with the interaction confidence loss to obtain the target recommendation loss.

11. The content recommendation method according to claim 10, characterized in that, The step of using the updated content recommendation model to recommend content to target objects in the object set includes: Obtain the content to be recommended, and use the updated content recommendation model to predict the current recommendation probability of the content to be recommended for each object in the object set; Based on the current recommendation probability, the target object corresponding to the content to be recommended is selected from the object set, and the content to be recommended is recommended to the target object.

12. A content recommendation device, characterized in that, include: The acquisition unit is used to acquire the current training set of the content recommendation model, wherein the current training set includes the interaction content sample of each object in the object set and the interaction duration corresponding to the interaction content sample; The prediction unit is used to use the content recommendation model to predict the recommendation probability of the interactive content sample for the object, and obtain the prediction recommendation result. The comparison unit is used to obtain the interaction duration threshold corresponding to the current training set, compare the interaction duration with the interaction duration threshold, take the interaction content samples whose interaction duration is greater than or equal to the interaction duration threshold as the first category samples to obtain the first category sample set, take the interaction content samples whose interaction duration is less than the interaction duration threshold as the second category samples to obtain the second category sample set, construct recommendation result pairs based on the first category sample set and the second category sample set, calculate the difference between the recommendation results in the recommendation result pair to obtain the recommendation difference, and correct the recommendation difference to obtain the interaction confidence loss; The determining unit is used to determine the target recommendation loss of the content recommendation model based on the interaction confidence loss and the prediction recommendation result; The recommendation unit is used to update the content recommendation model based on the target recommendation loss, and use the updated content recommendation model to recommend the content to be recommended to the target objects in the object set.

13. An electronic device, characterized in that, It includes a processor and a memory, the memory storing an application program, and the processor running the application program within the memory to perform the steps of the content recommendation method according to any one of claims 1 to 11.

14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the content recommendation method according to any one of claims 1 to 11.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the content recommendation method according to any one of claims 1 to 11.