Video recommendation method and corresponding model training method, device, equipment and medium
By using a debiasing revenue prediction model, which predicts the debiasing revenue after video recommendation based on the relevant features of users and videos, the problem of over-recommendation of popular resources in existing technologies is solved, thereby improving the accuracy of video recommendation.
Patent Information
- Application Number
- CN202311083557.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-25
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-08-25
AI Technical Summary
In existing video recommendation technologies, strategies based on duration targets are prone to over-recommending popular resources, resulting in poor recommendation accuracy. Furthermore, external signal correction methods may inadvertently harm high-quality videos.
A pre-trained debiasing revenue prediction model is used to predict the debiasing revenue after video recommendation based on the relevant features of users and videos, and then video recommendation is performed based on the debiasing revenue.
It improves the accuracy of video recommendations, reduces false recommendations of high-quality videos, and meets users' video viewing needs.
Smart Images

Figure CN117251594B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, specifically to the fields of video recommendation and artificial intelligence, and particularly to a video recommendation method and corresponding model training method, apparatus, device and medium. Background Technology
[0002] To provide a better immersive short video viewing experience and better meet user needs, short video applications (Apps) often consider multiple objectives when recommending videos, such as swiping speed, duration, following, and commenting.
[0003] The time-based target not only considers the revenue generated by a single video's duration but also the total revenue generated by that video over time. This long-term revenue-driven recommendation strategy means that if a user watches many videos, the short video app's recommendation system will suggest related and interesting videos, thereby increasing user viewing time and engagement. This effectively enhances user dwell time and better meets users' video viewing needs. Summary of the Invention
[0004] This disclosure provides a video recommendation method and corresponding model training methods, devices, equipment, and media.
[0005] According to one aspect of this disclosure, a video recommendation method is provided, comprising:
[0006] Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended in the multiple videos to be recommended, obtain the relevant features of the user to be recommended and each video to be recommended;
[0007] For each of the videos to be recommended, based on the corresponding relevant features, a pre-trained debiasing revenue prediction model is used to predict the debiasing revenue of the video to be recommended after recommending it to the user to be recommended; the debiasing revenue of the video to be recommended is used to identify the total consumption time of the video to be recommended to the user to be recommended within a preset time period after the debias is eliminated.
[0008] Based on the anti-biasing benefits corresponding to each of the videos to be recommended, video recommendations are made to the users to be recommended.
[0009] According to another aspect of this disclosure, a training method for a debiasing profit prediction model is provided, comprising:
[0010] Based on the attribute information of the first training user and the attribute information of the first training video, obtain the first correlation feature between the first training user and the first training video;
[0011] Obtain the actual debiasing benefit of the first training video after recommending the first training video to the first training user;
[0012] Based on the first relevant feature and the actual debiasing gain of the first training video, the debiasing gain prediction model is trained.
[0013] According to another aspect of this disclosure, a video recommendation device is provided, comprising:
[0014] The feature acquisition module is used to acquire the relevant features of the user to be recommended and each of the videos to be recommended based on the attribute information of the user to be recommended and the attribute information of each video to be recommended in a plurality of videos to be recommended;
[0015] The prediction module is used to predict the bias reduction benefit of each video to be recommended after recommending it to the user, based on the corresponding relevant features and using a pre-trained bias reduction benefit prediction model. The bias reduction benefit of the video to be recommended is used to identify the total consumption time of the video to be recommended after bias reduction for the user within a preset time period.
[0016] The recommendation module is used to recommend videos to the users to be recommended based on the anti-bias benefits corresponding to each of the videos to be recommended.
[0017] According to another aspect of this disclosure, a training apparatus for a debiased return prediction model is provided, comprising:
[0018] The first feature acquisition module is used to acquire the first correlation feature between the first training user and the first training video based on the attribute information of the first training user and the attribute information of the first training video.
[0019] The debiasing benefit acquisition module is used to acquire the actual debiasing benefit of the first training video after recommending the first training video to the first training user.
[0020] The first training module is used to train the debiasing gain prediction model based on the first relevant features and the real debiasing gains of the first training video.
[0021] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0022] At least one processor; and
[0023] A memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described above and any possible implementations.
[0025] According to yet another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described above and any possible implementation thereof.
[0026] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the aspects and any possible implementations described above.
[0027] The technology disclosed herein can effectively improve the accuracy of video recommendations.
[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0029] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0030] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;
[0031] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;
[0032] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure;
[0033] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure;
[0034] Figure 5 This is a schematic diagram according to the fifth embodiment of the present disclosure;
[0035] Figure 6 This is a schematic diagram according to the sixth embodiment of the present disclosure;
[0036] Figure 7 This is a schematic diagram according to the seventh embodiment of the present disclosure;
[0037] Figure 8 This is a schematic diagram according to the eighth embodiment of the present disclosure;
[0038] Figure 9 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation
[0039] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0040] Obviously, the described embodiments are only some, not all, of the embodiments disclosed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0041] It should be noted that the terminal devices involved in the embodiments of this disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers; the display devices may include, but are not limited to, personal computers, televisions, and other devices with display functions.
[0042] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0043] Existing video recommendation strategies based on duration targets can easily lead to more exposure for popular content. In some cases, videos are recommended to users not because they are of interest, but because of their high popularity. To overcome this problem and improve the accuracy of video recommendations, existing technologies can also rely on external signals for correction. For example, by statistically analyzing video impressions, resources with excessively high impressions can be suppressed, or signals such as click-through rate and completion rate can be used to suppress videos with high impressions but low user engagement.
[0044] However, the above-mentioned method of correction using external signals is prone to misrepresenting high-quality, trending videos due to the lack of precision of the external signals, resulting in poor accuracy in video recommendations.
[0045] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure; as shown Figure 1This embodiment provides a video recommendation method that can be applied to the video recommendation system of a short video app, and specifically includes the following steps:
[0046] S101. Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended in multiple videos to be recommended, obtain the relevant features of the user to be recommended and each video to be recommended.
[0047] In this embodiment, the attribute information of the user to be recommended may include the user's gender, age range, occupation, interests, spending level, and spending time characteristics, etc. For example, the spending level can identify whether the user to be recommended is a new user, an active user, or an inactive user. The spending time characteristics can identify the time distribution of the user's historical spending.
[0048] Alternatively, the attribute information of the user to be recommended may also include the attribute information of the terminal device used by the user. For example, it may include the type of mobile operating system used by the user, the network type of the mobile phone used by the user, etc.
[0049] The attribute information for videos to be recommended can include the video's type, tags, and language. The video's type identifies its genre, such as action, comedy, horror, romance, etc. Tags are keywords that identify key information about the video's content. For example, they could include humor, family, or gaming. The language indicates whether the video is in Chinese, English, or another language.
[0050] S102. For each video to be recommended, based on the corresponding relevant features, a pre-trained debiasing revenue prediction model is used to predict the debiasing revenue of the video to be recommended after recommending it to the user.
[0051] In this embodiment, a pre-trained debiasing revenue prediction model is used. This model can predict the debiasing revenue of the recommended video after it is recommended to the user. The debiasing revenue of the recommended video is used to identify the total consumption time of the recommended video after eliminating bias for the user within a preset time period.
[0052] The preset time period in this embodiment can be set to one day, two days, one week or other time lengths according to actual needs, and is not limited here.
[0053] For example, in a real-world application scenario, after recommending a video to a user, at least one additional trigger video will be recommended to that user within a preset time period, triggered by the video to be recommended. In this case, the de-biasing benefit of the video to be recommended includes the user's viewing time of the video to be recommended, and theoretically, the sum of the viewing time of each trigger video to be recommended, influenced by the video to be recommended.
[0054] For example, for each triggering video, the viewing time of the user to be recommended can be divided into two parts: one part is the viewing time of the user to be recommended who is not affected by the triggering video; this part corresponds to the viewing time of the user voluntarily watching the triggering video. The other part is the viewing time of the user to be recommended who is affected by the triggering video. For example, the total viewing time of the user to be recommended is t; if it is determined that the viewing time of the user to be recommended who voluntarily watches the triggering video is t1, then theoretically, the viewing time of the user to be recommended who is affected by the triggering video is t-t1.
[0055] In this embodiment, the debiasing benefit of the recommended video includes the consumption time of the recommended video, as well as the sum of the consumption time of the recommended user watching each triggering video, which is theoretically affected by the recommended video. This can eliminate the consumption time of the recommended user spontaneously watching the triggering video and improve the accuracy of the debiasing benefit of the recommended video.
[0056] S103. Based on the anti-biasing benefits corresponding to each video to be recommended, recommend videos to the users to be recommended.
[0057] By following the above method, the anti-bias benefit of each video to be recommended can be accurately obtained, and then more accurate video recommendations can be made to users based on the anti-bias benefit of each video to be recommended.
[0058] The video recommendation method in this embodiment employs a pre-trained debiasing revenue prediction model. Based on the correlation features between the user to be recommended and each video to be recommended, it can accurately predict the debiasing revenue of the recommended videos after recommending them to the user. Therefore, based on the debiasing revenue corresponding to each video to be recommended, more accurate video recommendations can be made to the user, which can effectively improve the accuracy of video recommendations.
[0059] Furthermore, the anti-bias revenue prediction model used in this embodiment can significantly improve the accuracy of the anti-bias revenue prediction for the recommended video, thus better meeting the user's video recommendation needs.
[0060] Further, optionally, in one embodiment of this disclosure, step S101 may include the following two implementation methods:
[0061] The first implementation method involves obtaining the characteristics of the user to be recommended and the characteristics of each video to be recommended based on the attribute information of the user to be recommended and the attribute information of each video to be recommended among multiple videos to be recommended.
[0062] The second implementation method involves obtaining the characteristics of the user to be recommended, the characteristics of each video to be recommended, and the interaction combination characteristics between the user to be recommended and each video to be recommended, based on the attribute information of the user to be recommended and the attribute information of each video to be recommended.
[0063] In the second implementation, compared to the first implementation, an interaction combination feature between the user to be recommended and each video to be recommended is added. For example, the interaction combination feature between the same user to be recommended and the same video to be recommended can include one, two, or more features. This interaction combination feature can include a combination of a feature of the user to be recommended and a feature of the video to be recommended. For example, if the user to be recommended is male and the video to be recommended carries the "funny" tag, this constitutes an interaction combination feature. If both the user is male and the video to be recommended carries the "funny" tag, this interaction combination feature is 1; otherwise, it is 0. Similarly, if the user to be recommended is female and the video to be recommended is of the "romance" type, this can also constitute an interaction combination feature. If both the user is female and the video to be recommended is of the "romance" type, this interaction combination feature is 1; otherwise, it is 0. Multiple interaction combination features can be set in a similar manner; further examples are not provided here.
[0064] Both of the above implementation methods can accurately and efficiently obtain the relevant features of the users to be recommended and each video to be recommended. In particular, the second implementation method, by adding interactive combination features, can further highlight the correlation features between the features of the users to be recommended and the features of the videos to be recommended, thereby enabling a more accurate prediction of the debiasing benefits of the videos to be recommended.
[0065] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure; this embodiment provides a training method for a debiased return prediction model, applied in a training device for a debiased return prediction model, and specifically may include the following steps:
[0066] S201. Based on the attribute information of the first training user and the attribute information of the first training video, obtain the first correlation feature between the first training user and the first training video;
[0067] S202. Obtain the actual debiasing benefit of the first training video after recommending it to the first training user.
[0068] S203. Based on the first relevant feature and the real debiasing gain of the first training video, train the debiasing gain prediction model.
[0069] In this embodiment, the training data is based on data about the first training user and the first training video. In practical applications, similar methods can be used to obtain relevant data about multiple first training users and first training videos, and the bias reduction profit prediction model can be trained in the manner described in this embodiment.
[0070] During specific training, we can first obtain the first relevant features that can identify the first training user and the first training video based on the attribute information of the first training user and the attribute information of the first training video.
[0071] In practical applications, after recommending the first training video to the first training user, at least one trigger video will be recommended to the first training user within a preset time period, triggered by the first training video. At this time, the debiasing benefit of the first training video includes the viewing time of the first training user and, theoretically, the sum of the viewing time of each trigger video influenced by the first training video. In this embodiment, the true debiasing benefit of the first training video refers to the actual debiasing benefit of the first training video obtained through statistics or calculation. This true debiasing benefit of the first training video is used as label data during training to supervise the training of the debiasing benefit prediction model.
[0072] Referring to the relevant descriptions in the above embodiments, the debiasing gain of the first training video in this embodiment can eliminate the consumption time of the first training user spontaneously watching the triggered video, thereby improving the accuracy of the debiasing gain of the first training video. Based on this concept, the accuracy of the true debiasing gain of the first training video obtained in this embodiment can be ensured.
[0073] In this embodiment, when training the debiasing gain prediction model based on the first relevant feature and the actual debiasing gain of the first training video, the first relevant feature can be used as input data, and the actual debiasing gain of the first training video can be used as supervision data to train the debiasing gain prediction model. Specifically, the first relevant feature is input into the debiasing gain prediction model, which can predict and output the predicted debiasing gain of the first training video based on the input information. Then, a loss function is constructed based on the actual debiasing gain and the predicted debiasing gain of the first training video; and the parameters of the debiasing gain model are adjusted with the direction of convergence of the loss function as the adjustment target. Following the above method, the debiasing gain prediction model is trained using data from multiple groups of first training users and corresponding first training videos until the model converges, the parameters of the debiasing gain prediction model are determined, and thus the debiasing gain prediction model is finalized.
[0074] The training method for the debiased revenue prediction model in this embodiment, by adopting the above method, can accurately and efficiently train the debiased revenue prediction model, thereby ensuring the accuracy of the trained debiased revenue prediction model; furthermore, when using the debiased revenue prediction model for video recommendation, it can effectively improve the accuracy of video recommendation.
[0075] Further, optionally, in this embodiment, step S201 can be implemented in the following two ways:
[0076] First implementation method: Based on the attribute information of the first training user and the attribute information of the first training video, obtain the features of the first training user and the features of the first training video.
[0077] The second implementation method is to obtain the features of the first training user, the features of the first training video, and the interaction combination features of the first training user and the first training video based on the attribute information of the first training user and the attribute information of the first training video.
[0078] Compared to the first implementation, the second implementation adds the interaction combination features of the first training user and the first training video. Specifically, refer to the above-mentioned related embodiments for two implementation methods of obtaining the relevant features of the user to be recommended and each video to be recommended, which will not be elaborated further here.
[0079] Similarly, both of the above implementation methods can efficiently and accurately obtain the primary correlation features between the first training user and the first training video. In particular, the second implementation method, by adding interactive combination features, can further highlight the correlation features between the features of the first training user and the first training video, thereby increasing the accuracy and rationality of the primary correlation features between the first training user and the first training video, and thus improving the training effect of the debiasing profit prediction model.
[0080] Figure 3 This is a schematic diagram based on the third embodiment of the present disclosure; in the above... Figure 2 Based on the technical solutions of the illustrated embodiments, this embodiment further describes the technical solutions in more detail. Figure 2 The specific implementation of step S202 in the illustrated embodiment: The technical solution of this disclosure may specifically include the following steps:
[0081] S301. Based on the first relevant feature and a pre-trained long-term revenue prediction model, predict the long-term revenue of the first training video after recommending it to the first training user.
[0082] The long-term benefits of the first training video include the consumption time of the first training video within a preset time period for the first training user, and the sum of the consumption time of each of the first triggered videos in at least one first triggered video recommended in response to the first training video.
[0083] The long-term benefit of the first training video in this embodiment differs from the debiasing benefit of the first training video described above in that the consumption duration of each first trigger video in the latter half is different. The long-term benefit of the first training video is the accumulated consumption duration of each first trigger video. However, the debiasing benefit of the first training video is the accumulated consumption duration of each first trigger video watched by the first training user, theoretically influenced by the first training video. In other words, compared to the debiasing benefit of the first training video, the long-term benefit of the first training video also includes the consumption duration of each first trigger video watched independently by the first training user, unaffected by the first training video.
[0084] Specifically, in this embodiment, a pre-trained long-term revenue prediction model can be used to accurately predict the long-term revenue of the first training video. During prediction, first relevant features are input into the long-term revenue prediction model. Based on these first relevant features, such as the features of the first training user, the features of the first training video, and the interaction combination features between the first training user and the first training video, the long-term revenue prediction model can accurately and efficiently predict the long-term revenue of the first training video.
[0085] S302. Based on the first video browsing behavior information of the first training user within a preset time period collected in advance, obtain the information of each first triggered video;
[0086] For example, the first video browsing behavior information of the first training user within a preset time period can include the identifiers of all videos viewed or watched by the first training user within that preset time period, the viewing time, the attribute information of each video, and the identifier of the source triggering video corresponding to each video, etc. That is, when recommending videos, if the video recommendation system recommends video 2 to the user based on video 1 watched by the user, the source triggering video identifier for video 2 will be recorded as video 1 in the first video browsing behavior information. If the recommendation system further recommends video 3 to the user based on video 2, the source triggering video identifier for video 3 will be recorded as video 2 in the first video browsing behavior information. The first video recommended to the first training user within the preset time period may not have a source triggering video identifier; this first video may be recommended by the recommendation system based on other recommendation rules such as user attribute information or video popularity. For non-first videos recommended to the first training user within a preset time period, such as one day, a source triggering video identifier is usually present.
[0087] For example, in this embodiment, all information about each triggered video other than the first video can be obtained from the first video browsing behavior information of the first training user, such as identifiers, attribute information, etc. Each first triggered video may include a video directly or indirectly triggered by the first video.
[0088] S303. Based on the information of each first trigger video, obtain the single benefit of each first trigger video when recommending each first trigger video to the first training user;
[0089] The single revenue from the first triggered video includes the consumption time of the first triggered video for the first training user when it is not triggered by the first training video; that is, the consumption time of the first training user when watching or browsing the first triggered video on their own without being triggered by other videos.
[0090] For example, the specific implementation of step S303 may include the following steps:
[0091] (1) Based on the attribute information of the first training user and the information of each first trigger video, obtain the second correlation feature between the first training user and each first trigger video;
[0092] The acquisition of the second related features of the first training user and each first trigger video is the same as the acquisition of the first related features of the first training user and the first training video in the above embodiments, and the acquisition of the related features of the user to be recommended and the video to be recommended in the above embodiments. The content is also the same, so it will not be repeated here.
[0093] (2) For each first trigger video, based on the corresponding second related features, a pre-trained single-benefit prediction model is used to predict the single benefit of the first trigger video after it is recommended to the first training user.
[0094] In this embodiment, a pre-trained single-revenue prediction model can be used to predict the single revenue of the first triggering video. In specific use, the second relevant features are input into the single-revenue prediction model. The single-revenue prediction model can accurately and efficiently predict the single revenue of the first triggering video based on the features of the first training user and the features of the first triggering video in the input second relevant features, or it can also include the interaction combination features of the first training user and the first triggering video.
[0095] S304. Based on the long-term benefit of the first training video and the single benefit of each first trigger video, obtain the true debiasing benefit of the first training video.
[0096] Specifically, combining the concepts of long-term gain of the first training video, single gain of each first triggering video, and debiasing gain of the first training video introduced above, the long-term gain of the first training video minus the single gain of each first triggering video can be used as the true debiasing gain of the first training video.
[0097] In practical applications, other methods can also be used to obtain the true debiasing gain of the first training video. For example, the true debiasing gain of the first training video can be predicted through statistical analysis, which will not be elaborated here.
[0098] In this embodiment, by adopting the above method, the true debiasing gains of the first training video can be obtained accurately and efficiently, providing effective data support for the training of the debiasing gains prediction model.
[0099] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure; this embodiment is used to provide Figure 3 The training method for the single-revenue prediction model in the illustrated embodiment may specifically include the following steps:
[0100] S401. Collect the second video browsing behavior information of the second training user within a preset time period;
[0101] In this embodiment, the second video browsing behavior information includes information about the second training video. The second training video is the first video that the second training user browses within a preset time period without being triggered by other videos. The second video browsing behavior information also includes the actual consumption time of the second training user when browsing the second training video.
[0102] The second video browsing behavior information is used to provide training data for the single-revenue prediction model. Therefore, unlike the first video browsing behavior information, the second video browsing behavior information may only include the first video viewed by the second training user within a preset time period, i.e., the second training video. This first video is not triggered by other videos. Therefore, the actual consumption time of the second training user browsing the second training video recorded in the second video browsing behavior information is the consumption time of the user's independent browsing, not triggered by any other videos, and belongs to the actual single revenue of the second training user browsing the second training video.
[0103] S402. Based on the second video browsing behavior information of the second training user, obtain the real single revenue of the second training video;
[0104] That is, the actual consumption time of the second training user in browsing the second training video is obtained from the second training user's second video browsing behavior information, and is used as the actual single revenue of the second training video.
[0105] S403. Based on the attribute information of the second training user and the attribute information of the second training video, obtain the third related feature between the second training user and the second training video.
[0106] Specifically, the acquisition of the second related features of the second related features of the first training user and the first trigger video mentioned above, as well as the acquisition of the first related features of the first training user and the first training video in the above embodiments, are the same and include the same content, so they will not be repeated here.
[0107] S404. Based on the third relevant feature and the real single benefit of the second training video, the single benefit prediction model is trained.
[0108] The training method in this embodiment is supervised training. Specifically, the third relevant feature is first input into the single-revenue prediction model. This model can predict and output the predicted single revenue of the second training video based on the input third relevant feature. Then, based on the predicted single revenue and the actual single revenue of the second training video, a loss function is constructed, and the parameters of the single-revenue prediction model are adjusted with the direction of loss function convergence as the adjustment target. Following the above method, the single-revenue prediction model is trained using data from multiple groups of second training users and corresponding second training videos until the model converges, the parameters of the single-revenue prediction model are determined, and thus the single-revenue prediction model is finalized.
[0109] The training method for the single-income prediction model in this embodiment can accurately and efficiently train the single-income prediction model, thereby ensuring the accuracy of the trained single-income prediction model; furthermore, it can provide reliable support for the training of the debiased income prediction model and improve the training effect of the debiased income prediction model.
[0110] Figure 5 This is a schematic diagram according to the fifth embodiment of the present disclosure; this embodiment is used to provide Figure 3 The training method for the long-term return prediction model in the illustrated embodiment may specifically include the following steps:
[0111] S501. Collect the third video browsing behavior information of the third training user within a preset time period;
[0112] The third video browsing behavior information includes information about the third training video, which is the first video that the third training user browses within a preset time period without being triggered by other videos. The third video browsing behavior information also includes information about at least one second trigger video recommended to the third training user after being triggered by the third training video, as well as the consumption time of the third training user in browsing the third training video and each second trigger video.
[0113] The third video browsing behavior information of the third training user in this embodiment is the same as the first video browsing behavior information of the first training user in the above embodiment. It includes not only the identifier, viewing time, consumption duration, and attribute information of the first video viewed by the third training user within a preset time period, i.e., the third training video, but also the identifier, viewing time, consumption duration, attribute information, and the identifier of the source triggering video corresponding to each second triggering video recommended to the third training user within a preset time period, which is directly or indirectly triggered by the third training video.
[0114] The consumption time of the third training user browsing each of the second-triggered videos recorded in the third video browsing behavior information is the actual consumption time of the third training user browsing the corresponding second-triggered videos in practical applications. However, theoretically, the consumption time of the third training user browsing the corresponding second-triggered videos is a comprehensive consumption time, including the time when the third training user browses the corresponding second-triggered videos independently without being triggered by other videos, i.e., the single benefit of the third training user browsing the second-triggered videos; and also including the time when the third training user browses the corresponding second-triggered videos triggered by other videos.
[0115] S502. Based on the third video browsing behavior information of the third training user, obtain the real long-term benefits of the third training video.
[0116] Specifically, in this embodiment, the consumption time of the third training user browsing the third training video can be obtained from the third training user's third video browsing behavior information, as well as the consumption time of other second-triggered videos browsed by the third training user within a preset time period, which are directly or indirectly triggered by the third training video. Then, the consumption time of the third training video and the consumption time of each second-triggered video are added together as the true long-term benefit of the third training video.
[0117] S503. Based on the attribute information of the third training user and the attribute information of the third training video, obtain the fourth related feature between the third user and the third training video.
[0118] Specifically, the acquisition methods for the second related features are the same as those for the third related features of the second training user and the second related features of the first training user and the first triggering video, and the content is also the same, so they will not be repeated here.
[0119] S504. Based on the fourth relevant feature and the real long-term returns of the third training video, train the long-term return prediction model.
[0120] The training method in this embodiment is supervised training. Specifically, the fourth relevant feature is first input into the long-term return prediction model. This model can predict and output the predicted long-term return of the third training video based on the input fourth relevant feature. Then, based on the predicted long-term return and the actual long-term return of the third training video, a loss function is constructed, and the parameters of the long-term return prediction model are adjusted with the direction of loss function convergence as the adjustment target. Following the above method, the long-term return prediction model is trained using data from multiple groups of third training users and corresponding third training videos until the model converges, the parameters of the long-term return prediction model are determined, and thus the long-term return prediction model is finalized.
[0121] The training method for the long-term return prediction model in this embodiment can accurately and efficiently train the long-term return prediction model, thereby ensuring the accuracy of the trained long-term return prediction model; furthermore, it can provide reliable support for the training of the debiased return prediction model and improve the training effect of the debiased return prediction model.
[0122] Figure 6 This is a schematic diagram according to the sixth embodiment of this disclosure; as shown Figure 6 As shown, this embodiment provides a video recommendation device 600, including:
[0123] The feature acquisition module 601 is used to acquire the relevant features of the user to be recommended and each of the videos to be recommended based on the attribute information of the user to be recommended and the attribute information of each video to be recommended in a plurality of videos to be recommended;
[0124] The prediction module 602 is used to predict the bias reduction benefit of each video to be recommended after recommending it to the user based on the corresponding relevant features and a pre-trained bias reduction benefit prediction model. The bias reduction benefit of the video to be recommended is used to identify the total consumption time after bias reduction brought by the video to be recommended to the user within a preset time period.
[0125] The recommendation module 603 is used to recommend videos to the users to be recommended based on the anti-bias benefits corresponding to each of the videos to be recommended.
[0126] The video recommendation device 600 in this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0127] Further optionally, in one embodiment of this disclosure, the feature acquisition module 601 is configured to:
[0128] Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended, the features of the user to be recommended and the features of each video to be recommended are obtained.
[0129] Further optionally, in one embodiment of this disclosure, the feature acquisition module 601 is configured to:
[0130] Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended, the features of the user to be recommended, the features of each video to be recommended, and the interaction combination features of the user to be recommended and each video to be recommended are obtained.
[0131] Figure 7 This is a schematic diagram according to the seventh embodiment of the present disclosure; as shown Figure 7 As shown, this embodiment provides a training device 700 for a debiased profit prediction model, including:
[0132] The first feature acquisition module 701 is used to acquire the first correlation feature between the first training user and the first training video based on the attribute information of the first training user and the attribute information of the first training video.
[0133] The debiasing benefit acquisition module 702 is used to acquire the actual debiasing benefit of the first training video after recommending the first training video to the first training user.
[0134] The first training module 703 is used to train the debiasing revenue prediction model based on the first relevant features and the real debiasing revenue of the first training video.
[0135] The training device 700 for the debiased return prediction model in this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0136] Figure 8 This is a schematic diagram based on the eighth embodiment of the present disclosure; as shown Figure 9 As shown, this embodiment provides a training device 800 for a debiased profit prediction model, including the above-mentioned... Figure 7 The modules with the same name and function shown are: First Feature Acquisition Module 801, Bias Reduction Gain Acquisition Module 802, and First Training Module 803.
[0137] In one embodiment of this disclosure, the first feature acquisition module 801 is configured to:
[0138] Based on the attribute information of the first training user and the attribute information of the first training video, the features of the first training user and the features of the first training video are obtained.
[0139] Optionally, in one embodiment of this disclosure, the first feature acquisition module 801 is configured to:
[0140] Based on the attribute information of the first training user and the attribute information of the first training video, the features of the first training user, the features of the first training video, and the interaction combination features of the first training user and the first training video are obtained.
[0141] Further optional, such as Figure 8 As shown, in one embodiment of this disclosure, the debiasing benefit acquisition module 802 includes:
[0142] The prediction unit 8021 is used to predict the long-term benefit of the first training video after recommending the first training video to the first training user based on the first relevant features and a pre-trained long-term benefit prediction model. The long-term benefit of the first training video includes the consumption time of the first training video for the first training user within a preset time period, and the sum of the consumption time of each of the first trigger videos in at least one first trigger video recommended due to the trigger of the first training video.
[0143] The video acquisition unit 8022 is used to acquire information of each of the first triggered videos based on the first video browsing behavior information of the first training user within a preset time period that has been collected in advance.
[0144] The revenue acquisition unit 8023 is used to acquire, based on the information of each of the first trigger videos, the single revenue of each of the first trigger videos when recommending each of the first trigger videos to the first training user; the single revenue of the first trigger video includes the consumption time of the first trigger video for the first training user when it is not triggered by the first training video.
[0145] The debiasing revenue acquisition unit 8024 is used to acquire the true debiasing revenue of the first training video based on the long-term revenue of the first training video and the single revenue of each of the first triggering videos.
[0146] Further optionally, in one embodiment of this disclosure, the revenue-generating unit 8023 is used for:
[0147] Based on the attribute information of the first training user and the information of each of the first triggering videos, a second correlation feature between the first training user and each of the first triggering videos is obtained.
[0148] For each of the first triggered videos, based on the corresponding second related features, a pre-trained single-revenue prediction model is used to predict the single revenue of the first triggered video after it is recommended to the first training user.
[0149] Further optional, such as Figure 8 As shown, in one embodiment of this disclosure, the training device 800 for the debiasing profit prediction model further includes:
[0150] The first acquisition module 804 is used to acquire the second video browsing behavior information of the second training user within a preset time period. The second video browsing behavior information includes information about the second training video. The second training video is the first video that the second training user browses within the preset time period without being triggered by other videos. The second video browsing behavior information also includes the actual consumption time of the second training user when browsing the second training video.
[0151] The single revenue acquisition module 805 is used to acquire the real single revenue of the second training video based on the second video browsing behavior information of the second training user.
[0152] The second feature acquisition module 806 is used to acquire a third related feature between the second training user and the second training video based on the attribute information of the second training user and the attribute information of the second training video.
[0153] The second training module 807 is used to train the single-income prediction model based on the third relevant feature and the real single income of the second training video.
[0154] Further optional, such as Figure 8 As shown, in one embodiment of this disclosure, the training device 800 for the debiasing profit prediction model further includes:
[0155] The second acquisition module 808 is used to acquire third video browsing behavior information of a third training user within a preset time period. The third video browsing behavior information includes information about the third training video, which is the first video viewed by the third training user within the preset time period without being triggered by other videos. The third video browsing behavior information also includes information about at least one second trigger video recommended to the third training user based on the trigger of the third training video, as well as the consumption time of the third training user in viewing the third training video and each of the second trigger videos.
[0156] The long-term revenue acquisition module 809 is used to acquire the real long-term revenue of the third training video based on the third video browsing behavior information of the third training user.
[0157] The third feature acquisition module 810 is used to acquire a fourth related feature between the third user and the third training video based on the attribute information of the third training user and the attribute information of the third training video.
[0158] The third training module 811 is used to train the long-term return prediction model based on the fourth relevant feature and the real long-term return of the third training video.
[0159] The training device 800 for the debiased return prediction model in this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0160] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0161] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0162] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0163] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0164] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0165] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the methods of this disclosure. For example, in some embodiments, the methods of this disclosure may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the methods of this disclosure described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the methods of this disclosure by any other suitable means (e.g., by means of firmware).
[0166] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0167] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0168] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0169] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0170] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0171] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0172] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0173] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A training method for a debiased return prediction model, comprising: Based on the attribute information of the first training user and the attribute information of the first training video, obtain the first correlation feature between the first training user and the first training video; Based on the first relevant features and the pre-trained long-term revenue prediction model, the long-term revenue of the first training video is predicted after it is recommended to the first training user. The long-term revenue of the first training video includes the consumption time of the first training video within a preset time period for the first training user, and the sum of the consumption time of each of the first trigger videos in at least one first trigger video recommended due to the trigger of the first training video. Based on the first video browsing behavior information of the first training user within a preset time period, information of each of the first triggered videos is obtained; Based on the information of each of the first triggered videos, obtain the single benefit of each of the first triggered videos when recommending each of the first triggered videos to the first training user; the single benefit of the first triggered video includes the consumption time of the first triggered video for the first training user when it is not triggered by the first training video; Based on the long-term benefit of the first training video and the single benefit of each of the first triggering videos, the true debiasing benefit of the first training video is obtained. Based on the first relevant feature and the actual debiasing gains of the first training video, a debiasing gains prediction model is trained, and the debiasing gains prediction model is used for video recommendation.
2. The method according to claim 1, wherein, Based on the attribute information of the first training user and the attribute information of the first training video, the first correlation feature between the first training user and the first training video is obtained, including: Based on the attribute information of the first training user and the attribute information of the first training video, the features of the first training user and the features of the first training video are obtained.
3. The method according to claim 1, wherein, Based on the attribute information of the first training user and the attribute information of the first training video, the first correlation feature between the first training user and the first training video is obtained, including: Based on the attribute information of the first training user and the attribute information of the first training video, the features of the first training user, the features of the first training video, and the interaction combination features of the first training user and the first training video are obtained.
4. The method according to claim 1, wherein, Based on the information of each of the first triggered videos, obtain the single benefit of each of the first triggered videos when recommending each of the first triggered videos to the first training user, including: Based on the attribute information of the first training user and the information of each of the first triggering videos, a second correlation feature between the first training user and each of the first triggering videos is obtained. For each of the first triggered videos, based on the corresponding second related features, a pre-trained single-revenue prediction model is used to predict the single revenue of the first triggered video after it is recommended to the first training user.
5. The method according to claim 4, wherein, For each of the first triggered videos, based on the corresponding second relevant features, a pre-trained single-revenue prediction model is used to predict the single-revenue of the first triggered video after recommending it to the first training user. Before this prediction, the method further includes: Collect second video browsing behavior information of the second training user within a preset time period. The second video browsing behavior information includes information about the second training video. The second training video is the first video that the second training user browses within the preset time period without being triggered by other videos. The second video browsing behavior information also includes the actual consumption time of the second training user when browsing the second training video. Based on the second video browsing behavior information of the second training user, obtain the real single revenue of the second training video; Based on the attribute information of the second training user and the attribute information of the second training video, a third correlation feature between the second training user and the second training video is obtained. The single-reward prediction model is trained based on the third relevant feature and the real single reward of the second training video.
6. The method according to claim 1, wherein, Based on the first relevant feature and a pre-trained long-term return prediction model, before predicting the long-term return of the first training video after recommending it to the first training user, the method further includes: Collect third video browsing behavior information of a third training user within a preset time period. The third video browsing behavior information includes information about the third training video, which is the first video that the third training user browses within the preset time period without being triggered by other videos. The third video browsing behavior information also includes information about at least one second trigger video recommended to the third training user based on the third training video, as well as the consumption time of the third training user in browsing the third training video and each of the second trigger videos. Based on the third video browsing behavior information of the third training user, the real long-term benefits of the third training video are obtained. Based on the attribute information of the third training user and the attribute information of the third training video, a fourth related feature of the third training user and the third training video is obtained. The long-term return prediction model is trained based on the fourth relevant feature and the real long-term returns of the third training video.
7. A video recommendation method, comprising: Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended in the multiple videos to be recommended, obtain the relevant features of the user to be recommended and each video to be recommended; For each of the videos to be recommended, based on the corresponding relevant features, a pre-trained debiasing revenue prediction model is used to predict the debiasing revenue of the video to be recommended after recommending it to the user to be recommended; the debiasing revenue of the video to be recommended is used to identify the total consumption time of the video to be recommended to the user to be recommended within a preset time period after eliminating the bias; the debiasing revenue prediction model is the debiasing revenue prediction model trained by the method described in any one of claims 1-6 above. Based on the anti-biasing benefits corresponding to each of the videos to be recommended, video recommendations are made to the users to be recommended.
8. The method according to claim 7, wherein, Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended among multiple videos to be recommended, the relevant features between the user to be recommended and each of the videos to be recommended are obtained, including: Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended, the features of the user to be recommended and the features of each video to be recommended are obtained.
9. The method according to claim 7, wherein, Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended among multiple videos to be recommended, the relevant features between the user to be recommended and each of the videos to be recommended are obtained, including: Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended, the features of the user to be recommended, the features of each video to be recommended, and the interaction combination features of the user to be recommended and each video to be recommended are obtained.
10. A training device for a debiased return prediction model, comprising: The first feature acquisition module is used to acquire the first correlation feature between the first training user and the first training video based on the attribute information of the first training user and the attribute information of the first training video. The debiasing revenue acquisition module includes: The prediction unit is used to predict the long-term benefit of the first training video after recommending the first training video to the first training user, based on the first relevant features and a pre-trained long-term benefit prediction model. The long-term benefit of the first training video includes the consumption time of the first training video for the first training user within a preset time period, and the sum of the consumption time of each of the first trigger videos in at least one first trigger video recommended due to the triggering of the first training video. The video acquisition unit is used to acquire information about each of the first triggered videos based on the first video browsing behavior information of the first training user within a preset time period that has been collected in advance. The revenue acquisition unit is used to acquire, based on the information of each of the first trigger videos, the single revenue of each of the first trigger videos when recommending each of the first trigger videos to the first training user; the single revenue of the first trigger video includes the consumption time of the first trigger video for the first training user when it is not triggered by the first training video. The debiasing benefit acquisition unit is used to acquire the true debiasing benefit of the first training video based on the long-term benefit of the first training video and the single benefit of each of the first triggering videos. The first training module is used to train a debiasing revenue prediction model based on the first relevant features and the real debiasing revenue of the first training video. The debiasing revenue prediction model is used for video recommendation.
11. The apparatus according to claim 10, wherein, The first feature acquisition module is used for: Based on the attribute information of the first training user and the attribute information of the first training video, the features of the first training user and the features of the first training video are obtained.
12. The apparatus according to claim 10, wherein, The first feature acquisition module is used for: Based on the attribute information of the first training user and the attribute information of the first training video, the features of the first training user, the features of the first training video, and the interaction combination features of the first training user and the first training video are obtained.
13. The apparatus according to claim 10, wherein, The revenue-generating unit is used for: Based on the attribute information of the first training user and the information of each of the first triggering videos, a second correlation feature between the first training user and each of the first triggering videos is obtained. For each of the first triggered videos, based on the corresponding second related features, a pre-trained single-revenue prediction model is used to predict the single revenue of the first triggered video after it is recommended to the first training user.
14. The apparatus according to claim 13, wherein, The device further includes: The first acquisition module is used to acquire the second video browsing behavior information of the second training user within a preset time period. The second video browsing behavior information includes information about the second training video. The second training video is the first video that the second training user browses within the preset time period without being triggered by other videos. The second video browsing behavior information also includes the actual consumption time of the second training user when browsing the second training video. The single revenue acquisition module is used to acquire the real single revenue of the second training video based on the second video browsing behavior information of the second training user. The second feature acquisition module is used to acquire a third related feature between the second training user and the second training video based on the attribute information of the second training user and the attribute information of the second training video. The second training module is used to train the single-revenue prediction model based on the third relevant feature and the real single revenue of the second training video.
15. The apparatus according to claim 10, wherein, The device further includes: The second acquisition module is used to acquire the third video browsing behavior information of the third training user within a preset time period. The third video browsing behavior information includes information about the third training video, which is the first video that the third training user browses within the preset time period without being triggered by other videos. The third video browsing behavior information also includes information about at least one second trigger video recommended to the third training user based on the trigger of the third training video, as well as the consumption time of the third training user in browsing the third training video and each of the second trigger videos. The long-term revenue acquisition module is used to acquire the real long-term revenue of the third training video based on the third training user's third video browsing behavior information. The third feature acquisition module is used to acquire the fourth related feature between the third training user and the third training video based on the attribute information of the third training user and the attribute information of the third training video. The third training module is used to train the long-term return prediction model based on the fourth relevant feature and the real long-term return of the third training video.
16. A video recommendation device, comprising: The feature acquisition module is used to acquire the relevant features of the user to be recommended and each of the videos to be recommended based on the attribute information of the user to be recommended and the attribute information of each video to be recommended in a plurality of videos to be recommended; The prediction module is used to predict, for each of the videos to be recommended, the anti-bias revenue of the video to be recommended after recommending it to the user, based on the corresponding relevant features and using a pre-trained anti-bias revenue prediction model; the anti-bias revenue of the video to be recommended is used to identify the total consumption time of the video to be recommended to the user after eliminating bias within a preset time period; the anti-bias revenue prediction model is an anti-bias revenue prediction model trained by the device described in any one of claims 10-15 above. The recommendation module is used to recommend videos to the users to be recommended based on the anti-bias benefits corresponding to each of the videos to be recommended.
17. The apparatus according to claim 16, wherein, The feature acquisition module is used for: Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended, the features of the user to be recommended and the features of each video to be recommended are obtained.
18. The apparatus according to claim 16, wherein, The feature acquisition module is used for: Based on the attribute information of the user to be recommended and the attribute information of each video to be recommended, the features of the user to be recommended, the features of each video to be recommended, and the interaction combination features of the user to be recommended and each video to be recommended are obtained.
19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-6 or 7-9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6 or 7-9.
21. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-6 or 7-9.
Citation Information
Patent Citations
Recommendation method, training method of recommendation model and related device
CN114661999A