Training method of video recommendation model and video recommendation method

By constructing a training sample set and deep learning model of multi-dimensional recommended values, the video recommendation model is optimized, and the low accuracy problem caused by relying on the current recommended values in the existing technology is solved, and more accurate video recommendation is achieved.

CN120296203APending Publication Date: 2025-07-11阿里巴巴(上海)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510321588.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing video recommendation model mainly relies on the current recommendation value, resulting in low accuracy of video recommendations and failure to effectively utilize the long-term potential of video to affect user behavior.

Method used

By constructing a training sample set, including the feature set of sample videos and multi-dimensional recommendation values, the influence between videos is calculated using the guidance duration, content association and author association, and iterative training is used for deep learning models to optimize the video recommendation model.

Benefits of technology

It improves the accuracy of video recommendations, can better understand the connection between videos and the complexity of user behavior, recommends matching users' current interests and guides viewing more related videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296203A_ABST
    Figure CN120296203A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of a video recommendation model and a video recommendation method. Relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a training sample set which at least comprises a plurality of sample videos, a feature set corresponding to a sample video in the plurality of sample videos, and a first recommendation value set corresponding to the sample video; processing the training sample set through the initial video recommendation model to obtain a prediction recommendation value set; and training the initial video recommendation model according to the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model, and determining a to-be-recommended video based on a video recommendation value output by the target video recommendation model. According to the method and the device, the technical problem of relatively low accuracy of video recommendation caused by performing video recommendation according to the current recommendation value corresponding to the video by a video recommendation model in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method for training a video recommendation model and a video recommendation method. Background Art

[0002] In the field of video recommendation, a key technology is to effectively model the video recommendation value. Usually, by modeling different tasks and finally combining them to measure the overall recommendation value, such as the watch time and completion rate tasks. However, in the prior art, only the current recommendation value of the video is often concerned. In fact, the video has the potential to have a long-term impact on users. Therefore, in the related technology, the video recommendation model recommends videos according to the current recommendation value corresponding to the video, and there is a problem that the accuracy of video recommendation is relatively low.

[0003] For the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] The embodiments of the present application provide a method for training a video recommendation model and a video recommendation method, so as to at least solve the technical problem that the video recommendation model in the related technology recommends videos according to the current recommendation value corresponding to the video, resulting in relatively low accuracy of video recommendation.

[0005] According to one aspect of the embodiments of the present application, a method for training a video recommendation model is provided, including: obtaining a training sample set, where the training sample set at least includes: a plurality of sample videos, a feature set corresponding to the sample videos in the plurality of sample videos, and a first recommendation value set corresponding to the sample videos, and the first recommendation value set at least includes: a first recommendation value, a second recommendation value, and a third recommendation value, the first recommendation value is obtained based on the guiding duration of the sample video, the second recommendation value is obtained based on the watch time of a first video that satisfies a first association relationship with the sample video, and the third recommendation value is obtained based on the watch time of a second video that satisfies a second association relationship with the sample video; processing the training sample set through an initial video recommendation model to obtain a predicted recommendation value set; training the initial video recommendation model according to the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model, where the video to be recommended is determined based on the video recommendation value output by the target video recommendation model.

[0006] Further, the first recommended value set is obtained through the following method: obtaining the guiding duration corresponding to the sample video in the multiple sample videos, and obtaining the first recommended value based on the guiding duration; determining the first video that satisfies the first association relationship with the sample video in the multiple sample videos, and calculating based on the viewing duration of the first video to obtain the second recommended value; determining the second video that satisfies the second association relationship with the sample video in the multiple sample videos, and calculating based on the viewing duration of the second video to obtain the third recommended value.

[0007] Further, obtaining the guiding duration corresponding to the sample video in the multiple sample videos includes: obtaining the third video that the target object watches within the first time period after watching the target sample video, where the target sample video is any one of the multiple sample videos; obtaining the viewing duration corresponding to the third video; calculating based on the viewing duration to obtain the guiding duration.

[0008] Further, obtaining the first recommended value based on the guiding duration includes: obtaining the position information corresponding to the sample video in the multiple sample videos, and performing grouping processing on the multiple sample videos based on the position information to obtain multiple groups of sample videos; for any group of sample videos, calculating based on the guiding duration corresponding to the sample video in the group of sample videos to obtain the equal-frequency quantile corresponding to the sample video in the group of sample videos; obtaining the first recommended value based on the equal-frequency quantile and the guiding duration.

[0009] Further, obtaining the first recommended value based on the equal-frequency quantile and the guiding duration includes: determining the target guiding duration based on the guiding duration corresponding to the sample video in the group of sample videos; performing bucketing processing on the sample videos in the group of sample videos according to the equal-frequency quantile to obtain the bucket value corresponding to the equal-frequency quantile; calculating based on the bucket value and the target guiding duration to obtain the first recommended value.

[0010] Further, determining the first video that satisfies the first association relationship with the sample video in the multiple sample videos includes: obtaining multiple first candidate videos that the target object watches within the second time period after watching the target sample video, where the target sample video is any one of the multiple sample videos; determining whether the first candidate video in the multiple first candidate videos satisfies the first association relationship to obtain the first target judgment result; determining the first video from the multiple first candidate videos based on the first target judgment result.

[0011] Further, determining whether the first candidate video among the multiple first candidate videos meets the first association relationship to obtain a first target determination result includes: determining whether the content of the first candidate video among the multiple first candidate videos is similar to that of the target sample video to obtain a first determination result; determining whether the access behavior of the first candidate video among the multiple first candidate videos is similar to that of the target sample video to obtain a second determination result; determining whether the rank of the first candidate video among the multiple first candidate videos is the target rank to obtain a third determination result; and determining whether the first candidate video among the multiple first candidate videos meets the first association relationship according to the first determination result, the second determination result, and the third determination result to obtain the first target determination result.

[0012] Further, determining the second video that meets the second association relationship with the sample video among the multiple sample videos includes: obtaining multiple second candidate videos watched by the target object within a second time period after watching the target sample video and multiple third candidate videos watched by the target object within a third time period after watching the target sample video; determining whether the second candidate video among the multiple second candidate videos meets the second association relationship to obtain a second target determination result, and determining whether the third candidate video among the multiple third candidate videos meets the second association relationship to obtain a third target determination result; and determining the second video from the multiple second candidate videos and the multiple third candidate videos according to the second target determination result and the third target determination result.

[0013] Further, determining whether the second candidate video among the multiple second candidate videos meets the second association relationship to obtain a second target determination result includes: determining whether the publishing object corresponding to the second candidate video among the multiple second candidate videos is the same as that of the target sample video; in the case where the publishing object corresponding to the second candidate video among the multiple second candidate videos is the same as that of the target sample video, determining that the second target determination result is that the second candidate video meets the second association relationship; and in the case where the publishing object corresponding to the second candidate video among the multiple second candidate videos is different from that of the target sample video, determining that the second target determination result is that the second candidate video does not meet the second association relationship.

[0014] Further, training the initial video recommendation model based on the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model includes: calculating a loss based on the first predicted recommendation value and the first recommendation value in the predicted recommendation value set to obtain a first loss function; calculating a loss based on the second predicted recommendation value and the second recommendation value in the predicted recommendation value set to obtain a second loss function; calculating a loss based on the third predicted recommendation value and the third recommendation value in the predicted recommendation value set to obtain a third loss function; training the initial video recommendation model based on the first loss function, the second loss function, and the third loss function to obtain the target video recommendation model.

[0015] Further, calculating a loss based on the second predicted recommendation value and the second recommendation value in the predicted recommendation value set to obtain a second loss function includes: calculating a mean square error between the second predicted recommendation value and the second recommendation value to obtain a first loss sub-function; calculating a loss based on the exponential distribution family for the second predicted recommendation value and the second recommendation value to obtain a second loss sub-function; obtaining the second loss function based on the first loss sub-function and the second loss sub-function.

[0016] According to another aspect of the embodiments of the present application, a video recommendation method is further provided, including: obtaining target feature information corresponding to multiple initial videos; processing the target feature information through a target video recommendation model to obtain a second recommendation value set corresponding to an initial video among the multiple initial videos, where the target video recommendation model is trained by using the training method of the video recommendation model described in any one of the above; determining a target video to be recommended from the multiple initial videos according to the second recommendation value set.

[0017] According to another aspect of the embodiments of the present application, a video recommendation method is further provided, including: obtaining target feature information corresponding to multiple initial videos uploaded by a client; processing the target feature information through a target video recommendation model in a cloud server to obtain a second recommendation value set corresponding to an initial video among the multiple initial videos, where the target video recommendation model is trained by using the training method of the video recommendation model described in any one of the above; determining a target video to be recommended from the multiple initial videos according to the second recommendation value set; returning the target video to be recommended to the client.

[0018] According to another aspect of the embodiments of the present application, there is also provided a video recommendation method, including: a first acquisition unit, configured to acquire a training sample set, where the training sample set at least includes: a plurality of sample videos, a feature set corresponding to the sample videos in the plurality of sample videos, and a first recommendation value set corresponding to the sample videos, and the first recommendation value set at least includes: a first recommendation value, a second recommendation value, and a third recommendation value, the first recommendation value is obtained based on the guiding duration of the sample video, the second recommendation value is obtained based on the viewing duration of a first video that satisfies a first association relationship with the sample video, and the third recommendation value is obtained based on the viewing duration of a second video that satisfies a second association relationship with the sample video; a first processing unit, configured to process the training sample set through an initial video recommendation model to obtain a predicted recommendation value set; a training unit, configured to train the initial video recommendation model according to the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model, where the video to be recommended is determined based on the video recommendation value output by the target video recommendation model.

[0019] Further, the recommendation value set is obtained through the following device: a second acquisition unit, configured to acquire the guiding duration corresponding to the sample videos in the plurality of sample videos, and obtain the first recommendation value according to the guiding duration; a first determination unit, configured to determine a first video that satisfies a first association relationship with the sample videos in the plurality of sample videos, and calculate according to the viewing duration of the first video to obtain the second recommendation value; a second determination unit, configured to determine a second video that satisfies a second association relationship with the sample videos in the plurality of sample videos, and calculate according to the viewing duration of the second video to obtain the third recommendation value.

[0020] Further, the second acquisition unit includes: a first acquisition module, configured to acquire a third video viewed by a target object within a first time period after viewing a target sample video, where the target sample video is any one of the plurality of sample videos; a second acquisition module, configured to acquire the viewing duration corresponding to the third video; a first calculation module, configured to calculate according to the viewing duration to obtain the guiding duration.

[0021] Further, the second acquisition unit includes: a third acquisition module, configured to acquire the position information corresponding to the sample videos in the plurality of sample videos, and perform grouping processing on the plurality of sample videos based on the position information to obtain multiple groups of sample videos; a second calculation module, configured to calculate, for any group of sample videos, according to the guiding duration corresponding to the sample videos in the group to obtain the equal-frequency quantile corresponding to the sample videos in the group; a first determination module, configured to obtain the first recommendation value according to the equal-frequency quantile and the guiding duration.

[0022] Further, the determination module includes: a first determination sub-module, configured to determine a target guidance duration according to the guidance duration corresponding to the sample video in the group of sample videos; a processing sub-module, configured to perform bucketing processing on the sample videos in the group of sample videos according to the equal-frequency quantiles to obtain the bucket values corresponding to the equal-frequency quantiles; and a first calculation sub-module, configured to perform a calculation according to the bucket values and the target guidance duration to obtain the first recommended value.

[0023] Further, the first determination unit includes: a fourth acquisition module, configured to acquire a plurality of first candidate videos that the target object watches after watching the target sample video and within a second time period, where the target sample video is any one of the plurality of sample videos; a first judgment module, configured to judge whether the first candidate videos in the plurality of first candidate videos satisfy the first association relationship to obtain a first target judgment result; and a second determination module, configured to determine the first video from the plurality of first candidate videos according to the first target judgment result.

[0024] Further, the judgment module includes: a first judgment sub-module, configured to judge whether the content of the first candidate videos in the plurality of first candidate videos is similar to the target sample video to obtain a first judgment result; a second judgment sub-module, configured to judge whether the access behavior of the first candidate videos in the plurality of first candidate videos is similar to the target sample video to obtain a second judgment result; a third judgment sub-module, configured to judge whether the sequence position of the first candidate videos in the plurality of first candidate videos is the target sequence position to obtain a third judgment result; and a fourth judgment sub-module, configured to judge whether the first candidate videos in the plurality of first candidate videos satisfy the first association relationship according to the first judgment result, the second judgment result, and the third judgment result to obtain the first target judgment result.

[0025] Further, the second determination unit includes: a fifth acquisition module, configured to acquire a plurality of second candidate videos that the target object watches after watching the target sample video and within a second time period and a plurality of third candidate videos that the target object watches after watching the target sample video and within a third time period; a second judgment module, configured to judge whether the second candidate videos in the plurality of second candidate videos satisfy the second association relationship to obtain a second target judgment result, and judge whether the third candidate videos in the plurality of third candidate videos satisfy the second association relationship to obtain a third target judgment result; and a second determination module, configured to determine the second video from the plurality of second candidate videos and the plurality of third candidate videos according to the second target judgment result and the third target judgment result.

[0026] Further, the second judgment module includes: a fifth judgment sub-module, configured to judge whether the release object corresponding to the second candidate video in the multiple second candidate videos is the same as the target sample video; a second determination sub-module, configured to, when the release object corresponding to the second candidate video in the multiple second candidate videos is the same as the target sample video, determine that the second target judgment result is that the second candidate video meets the second association relationship; a third determination sub-module, configured to, when the release object corresponding to the second candidate video in the multiple second candidate videos is different from the target sample video, determine that the second target judgment result is that the second candidate video does not meet the second association relationship.

[0027] Further, the training unit includes: a second calculation module, configured to perform loss calculation based on the first predicted recommendation value and the first recommendation value in the predicted recommendation value set to obtain a first loss function; a third calculation module, configured to perform loss calculation based on the second predicted recommendation value and the second recommendation value in the predicted recommendation value set to obtain a second loss function; a fourth calculation module, configured to perform loss calculation based on the third predicted recommendation value and the third recommendation value in the predicted recommendation value set to obtain a third loss function; a training module, configured to train the initial video recommendation model based on the first loss function, the second loss function, and the third loss function to obtain the target video recommendation model.

[0028] Further, the third calculation module includes: a second calculation sub-module, configured to calculate the mean square error between the second predicted recommendation value and the second recommendation value to obtain a first loss sub-function; a third calculation sub-module, configured to perform loss calculation on the second predicted recommendation value and the second recommendation value according to the exponential distribution family to obtain a second loss sub-function; a fourth determination sub-module, configured to obtain the second loss function based on the first loss sub-function and the second loss sub-function.

[0029] According to another aspect of the embodiments of the present application, there is also provided a video recommendation device, including: a third acquisition unit, configured to acquire target feature information corresponding to multiple initial videos; a first processing unit, configured to process the target feature information through a target video recommendation model to obtain a second recommendation value set corresponding to the initial video in the multiple initial videos, where the target video recommendation model is trained by using the training method of the video recommendation model described in any one of the above; a third determination unit, configured to determine a target video to be recommended from the multiple initial videos according to the second recommendation value set.

[0030] According to another aspect of the embodiments of the present invention, an electronic device is further provided, including: a memory storing an executable program; a processor for running the program, wherein when the program runs, it executes the training method of the video recommendation model in any one of the above, or the video recommendation method.

[0031] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is further provided, and the storage medium stores a program, wherein when the program runs, it controls the device where the storage medium is located to execute the training method of the video recommendation model in any one of the above, or the video recommendation method.

[0032] According to another aspect of the embodiments of the present invention, a computer program product is further provided, including a computer program or instruction, and when the computer program or instruction is executed by a processor, it implements the training method of the video recommendation model in any one of the above, or the video recommendation method.

[0033] In the embodiments of the present application, the following steps are adopted: obtaining a training sample set, wherein the training sample set at least includes: a plurality of sample videos, a feature set corresponding to the sample videos in the plurality of sample videos, and a first recommended value set corresponding to the sample videos, and the first recommended value set at least includes: a first recommended value, a second recommended value, and a third recommended value, the first recommended value is obtained based on the guiding duration of the sample video, the second recommended value is obtained based on the viewing duration of a first video that satisfies a first association relationship with the sample video, and the third recommended value is obtained based on the viewing duration of a second video that satisfies a second association relationship with the sample video; processing the training sample set through an initial video recommendation model to obtain a predicted recommended value set; training the initial video recommendation model according to the predicted recommended value set and the first recommended value set to obtain a target video recommendation model, wherein the video to be recommended is determined based on the video recommendation value output by the target video recommendation model, which solves the technical problem in the related art that the video recommendation model recommends videos according to the current recommended value corresponding to the video, resulting in relatively low accuracy of video recommendation.

[0034] In this solution, a training sample set is constructed through multiple sample videos, the feature sets corresponding to the sample videos in the multiple sample videos, and the first recommendation value set corresponding to the sample videos. Then, the initial video recommendation model is iteratively trained through the training sample set to obtain the target video recommendation model. The first recommendation value set includes at least a first recommendation value, a second recommendation value, and a third recommendation value. The first recommendation value is calculated through the guiding duration of the sample video and can accurately evaluate the influence degree of the sample video on subsequent videos. The second recommendation value and the third recommendation value are respectively calculated based on the association relationship between videos and the viewing duration of subsequent videos. Such multi-dimensional recommendation values enable the video recommendation model to more accurately understand the connections between videos and the complexity of user behavior, so as to recommend videos that not only match the user's current interests but also can guide the user to watch more relevant videos, thereby achieving the technical effect of improving the accuracy of video recommendations. Description of the Drawings

[0035] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0036] Figure 1 is a block diagram of the hardware structure of a computer terminal provided in Embodiment 1 of the present application;

[0037] Figure 2 is a flowchart of a method for training a video recommendation model provided in Embodiment 1 of the present application;

[0038] Figure 3 is a schematic diagram of the guiding duration distribution in the prior art Figure 1 ;

[0039] Figure 4 is a schematic diagram of the guiding duration distribution in the prior art Figure 2 ;

[0040] Figure 5 is a schematic diagram of the recommendation value of a video provided in Embodiment 1 of the present application;

[0041] Figure 6 is a schematic diagram of a sample video provided in Embodiment 1 of the present application;

[0042] Figure 7 is a schematic diagram of a method for training a video recommendation model provided in Embodiment 1 of the present application;

[0043] Figure 8 is a schematic diagram of the effect of a video recommendation model provided in Embodiment 1 of the present application;

[0044] Figure 9It is a flowchart of the video recommendation method provided in the second embodiment of the present application;

[0045] Figure 10 It is a flowchart of the video recommendation method provided in the third embodiment of the present application;

[0046] Figure 11 It is a schematic diagram of the training device of the video recommendation model provided in the fourth embodiment of the present application;

[0047] Figure 12 It is a schematic diagram of the video recommendation device provided in the fifth embodiment of the present application;

[0048] Figure 13 It is a block diagram of the structure of the electronic device provided in the sixth embodiment of the present application. Detailed implementation manners

[0049] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0050] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0051] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:

[0052] Guiding duration: It refers to the total time spent by a video to guide a user to watch a subsequent video sequence.

[0053] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in the relevant region, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0054] Embodiment 1

[0055] According to an embodiment of the present application, there is also provided a method for training a video recommendation model. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0056] The method embodiment provided in the first embodiment of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing the method for training a video recommendation model is shown. As Figure 1 shown, the computer terminal (or mobile device) 10 may include a set of processors 102 (the set of processors 102 may include but is not limited to processing devices such as a microprocessor MCU or a programmable logic device FPGA, and the set of processors 102 may include a set of processors, Figure 1 which are shown as 102a, 102b,..., 102n in the figure), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.

[0057] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).

[0058] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the training method of the video recommendation model in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned training method of the video recommendation model. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0059] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0060] The display can be a touch-screen liquid crystal display, which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0061] Under the above operating environment, the present application provides a training method of a video recommendation model as Figure 2 shown. Figure 2 It is a flowchart of the training method of the video recommendation model according to Embodiment 1 of the present application. The training method includes:

[0062] Step S201: Obtain a training sample set, where the training sample set includes at least: multiple sample videos, a feature set corresponding to the sample videos in the multiple sample videos, and a first recommendation value set corresponding to the sample videos. The first recommendation value set includes at least: a first recommendation value, a second recommendation value, and a third recommendation value. The first recommendation value is obtained based on the guiding duration of the sample video. The second recommendation value is obtained based on the viewing duration of a first video that has a first association relationship with the sample video. The third recommendation value is obtained based on the viewing duration of a second video that has a second association relationship with the sample video.

[0063] Optionally, extract a series of videos from the user's historical browsing records as samples. These videos can be the videos that the user has watched in the past period of time, or the videos recommended by the system to the user. For the sample videos, obtain the feature set corresponding to the sample videos. The feature set may include user features (such as age, gender, viewing history, click preferences, etc.), video features (such as video type, duration, upload time, number of views, etc.), and context features (such as the time of recommendation, location, the user's current activity, etc.).

[0064] Then, obtain the first recommendation value according to the quantile corresponding to the guiding duration of the sample video. The guiding duration refers to the total duration that the user is guided to watch other videos after watching a video. Using the quantile as the recommendation value can eliminate the influence of position bias. For example, if the quantile corresponding to the guiding duration of a video is the 75th percentile, then the first recommendation value can be the average guiding duration corresponding to the 75th percentile.

[0065] Obtain the second recommendation value according to the viewing duration of the first video that has a first association relationship with the sample video. It should be noted that the first video that has a first association relationship with the sample video can be a video that has a similar theme or style to the sample video. By considering the viewing durations of these associated videos, the model can evaluate the direct influence of the sample video and its influence on the user's subsequent behavior, and help identify the internal connections between videos.

[0066] Obtain the above-mentioned third recommendation value according to the viewing duration of the second video that has a second association relationship with the sample video. It should be noted that the second video that has a second association relationship with the sample video can be a video with the same author as the sample video.

[0067] It should be noted that the above-mentioned first recommendation value, second recommendation value, and third recommendation value are all real recommendation values, which are the real labels for the subsequent iterative training of the initial video recommendation model.

[0068] By constructing the first recommendation value set, the video recommendation system can more accurately understand the value of the videos and the user's behavior patterns, so as to provide higher-quality, more relevant, and more personalized video recommendations.

[0069] Step S202: Process the training sample set through the initial video recommendation model to obtain a set of predicted recommendation values.

[0070] Optionally, input the feature set in the above training sample set into the initial video recommendation model, and process the feature set through the initial recommendation model to obtain the above set of predicted recommendation values. It should be noted that the initial video recommendation model can be a deep learning model, such as a neural network model. For example, the initial video recommendation model consists of an Embedding layer and multiple feature extraction modules. The Embedding module performs vector transformation on the feature set in the training sample set to obtain corresponding feature vectors, and the feature extraction module corresponding to the first recommendation value processes the feature vectors to obtain the first predicted recommendation value, as well as the second and third predicted recommendation values.

[0071] Step S203: Train the initial video recommendation model based on the set of predicted recommendation values and the first set of recommendation values to obtain a target video recommendation model, where the video to be recommended is determined based on the video recommendation value output by the target video recommendation model.

[0072] Optionally, adjust the parameters and weights in the initial video recommendation model by comparing the differences between the set of predicted recommendation values and the set of recommendation values to obtain the final target video recommendation model. The target video recommendation model trained through the set of predicted recommendation values and the set of recommendation values can significantly improve the accuracy and efficiency of video recommendation, reduce position deviation and target underestimation, and at the same time enhance the prediction ability of users' long-term behavior.

[0073] In summary, a training sample set is constructed through multiple sample videos, the feature sets corresponding to the sample videos in the multiple sample videos, and the first set of recommendation values corresponding to the sample videos. Then, the initial video recommendation model is iteratively trained through the training sample set to obtain a target video recommendation model. The first set of recommendation values includes at least the first recommendation value, the second recommendation value, and the third recommendation value. The first recommendation value is calculated through the guiding duration of the sample video and can accurately evaluate the influence degree of the sample video on subsequent videos. The second and third recommendation values are calculated based on the correlation relationship between videos and the viewing duration of subsequent videos respectively. Such multi-dimensional recommendation values enable the video recommendation model to more accurately understand the connections between videos and the complexity of user behavior, so as to recommend videos that not only match the user's current interests but also can guide the user to watch more relevant videos, thereby achieving the technical effect of improving the accuracy of video recommendation.

[0074] To improve the accuracy of the recommended value set, in the training method of the video recommendation model provided in the first embodiment of this application, the first recommended value set is obtained through the following method: Obtain the guiding duration corresponding to the sample video in multiple sample videos, and obtain the first recommended value based on the guiding duration; Determine the first video that satisfies the first association relationship with the sample video in multiple sample videos, and calculate based on the viewing duration of the first video to obtain the second recommended value; Determine the second video that satisfies the second association relationship with the sample video in multiple sample videos, and calculate based on the viewing duration of the second video to obtain the third recommended value.

[0075] Optionally, the guiding duration refers to the total viewing duration of the subsequent videos that the user is guided to watch after watching a video. Therefore, the guiding duration can be calculated based on the time of the videos that the user watches after the sample video and before ending the video viewing.

[0076] Then, to determine the first video that satisfies the first association relationship with the sample video in multiple sample videos, it should be noted that the first association relationship usually refers to a certain direct association between videos, such as content similarity or the same category, etc. After determining the first video that satisfies the first association relationship with the sample video, calculate the second recommended value based on the viewing duration of these videos. Through the second recommended value, the influence of the sample video on the associated first video can be effectively evaluated. By considering the viewing duration of these associated videos, the video recommendation model can learn the internal connection between videos and predict the attractiveness of the associated videos that a video may lead to viewing.

[0077] Furthermore, to determine the second video that satisfies the second association relationship with the sample video in multiple sample videos, it should be noted that the second video that satisfies the second association relationship can be a video with the same release object as the sample video within a preset time. After determining the second video that satisfies the second association relationship with the sample video, calculate the third recommended value based on the viewing duration of these videos. For example, accumulate and sum the viewing duration of the second videos to obtain the third recommended value.

[0078] Training the model with the first recommended value, the second recommended value, and the third recommended value can enhance the video recommendation model's understanding of the user's personalized preferences, thereby providing recommended content that better meets the user's interests and expectations.

[0079] To improve the accuracy of calculating the guiding duration, in the training method of the video recommendation model provided in the first embodiment of this application, obtaining the guiding duration corresponding to the sample video in multiple sample videos includes: Obtain the third video that the target object watches after watching the target sample video and within the first time period, where the target sample video is any one of the multiple sample videos; Obtain the viewing duration corresponding to the third video; Calculate based on the viewing duration to obtain the guiding duration.

[0080] Optionally, for any one of the multiple sample videos (i.e., the above-mentioned target sample video), obtain the third video that the user (i.e., the above-mentioned target object) watches within the first time period after watching the target sample video. It should be noted that the first time period can be the period when the user finishes watching the video after watching the target sample video, or it can be a time period set according to the user's behavior pattern and task requirements.

[0081] Obtain the viewing duration corresponding to the third video, and sum these viewing times to obtain the guidance duration. For example, the guidance duration of the sample video at a certain position n is defined as slide time = t1 + t2 + … + t k , where t k represents the viewing time of the video at position n + k. To consider the upper limit Q of the influence of the current card on subsequent cards, the final calculation result of the guidance duration can be used as y = min(t1 + t2 + … + t k , Q). It should be noted that the upper limit Q can be set according to requirements. After calculating the guidance duration, the guidance duration can be directly used as the above-mentioned first recommended value.

[0082] By calculating the guidance duration, the coherence and depth of the user's viewing behavior can be understood more deeply, and the accuracy of the subsequent recommended video model can be improved.

[0083] To improve the accuracy of calculating the first recommended value, in the training method of the video recommendation model provided in Embodiment 1 of the present application, obtaining the first recommended value based on the guidance duration includes: obtaining the position information corresponding to the sample videos in the multiple sample videos, and performing grouping processing on the multiple sample videos based on the position information to obtain multiple groups of sample videos; for any group of sample videos, calculating based on the guidance duration corresponding to the sample videos in the group of sample videos to obtain the equal-frequency quantiles corresponding to the sample videos in the group of sample videos; obtaining the first recommended value based on the equal-frequency quantiles and the guidance duration.

[0084] Optionally, there is a position bias in the initial guidance time method. The recommendation system often shows unfair preference for videos in relatively later display positions. As Figure 3 shown, the guidance duration of the video is significantly affected by its position. When the average position of the video reaches a certain threshold, its guidance duration begins to increase. This is because the videos at the end of the sequence are usually watched by active users, and these users are more likely to watch more videos. As Figure 4As shown, it presents the three-dimensional distribution corresponding to different quantiles of the guiding duration on different pages (i.e., the above-mentioned location information). It can be observed that as the page increases, the quantile distribution changes significantly. At the same quantile, the larger the page, the longer the corresponding guiding duration, and the lighter the color indicates a longer guiding duration. To overcome this phenomenon, with the page as the main grouping unit, each page request will include four videos.

[0085] The sample videos are grouped according to the page to obtain multiple groups of sample videos. For example, they are grouped into M equal parts according to the page. Then, for any group of sample videos, the equal-frequency quantiles corresponding to the sample videos in this group of sample videos are calculated based on the guiding durations corresponding to the sample videos in this group of sample videos. For example, the equal-frequency quantiles are calculated based on the guiding time distribution within the group. Any group of sample videos is discretized into T quantiles, and k represents the kth group. The equal-frequency quantile means that after the data is sorted by value, it is divided into equal numbers of parts, and each part contains the same number of data points. For example, if there are 100 videos in a group, the guiding duration values at the 10th percentile, 20th percentile, etc. can be calculated until the 100th percentile.

[0086] Finally, according to the equal-frequency quantiles and the guiding duration, the first recommendation value of each video is calculated. For example, the guiding duration of the video is converted into a relative attraction index relative to other videos within the location group, that is, the equal-frequency quantile ranking, and then the ranking is used as the first recommendation value.

[0087] By grouping based on location information and calculating equal-frequency quantiles, the influence of video location on its guiding duration can be effectively eliminated. Eliminating location bias and more accurately evaluating video value helps to provide more personalized and high-quality recommended content for users, improve users' viewing satisfaction and engagement, and ultimately increase user stickiness and platform activity.

[0088] To further improve the accuracy of calculating the first recommendation value, in the training method of the video recommendation model provided in the first embodiment of this application, obtaining the first recommendation value based on the equal-frequency quantiles and the guiding duration includes: determining the target guiding duration based on the guiding duration corresponding to the sample videos in this group of sample videos; performing bucketing processing on the sample videos in this group of sample videos according to the equal-frequency quantiles to obtain the bucket values corresponding to the equal-frequency quantiles; calculating based on the bucket values and the target guiding duration to obtain the first recommendation value.

[0089] Optionally, since when the page type is 0, the actual lead time values below the 32nd percentile will be set to zero. To solve this problem, a dynamic starting point is designed, that is, the target lead time is determined according to the lead time corresponding to the sample videos in this group of sample videos. For example, the lead time that is less than all other lead times in this group of sample videos is used as the target lead time.

[0090] Then, the sample videos in this group of sample videos are bucketed according to equal-frequency percentiles. Finally, calculations are performed based on the bucket values and the target lead time to obtain the first recommended value. By calculating the equal-frequency percentiles, the boundaries of the buckets can be determined to ensure that the distribution of the video lead times within the buckets is roughly uniform. For example, assuming there are lead time data of 100 videos, they can be divided into 4 buckets, and each bucket contains lead time data of 25 videos. Through this bucketing, even if there are significant differences in the lead times of the videos, they can be classified according to their relative positions in the overall distribution, reducing the bias of the position on the video value evaluation. Estimating percentiles is simpler than estimating specific lead times because their value range is from 0 to 1 and they change from continuous values to discrete values.

[0091] For example, the first recommended value is B(s i ,D i* ) is the above-mentioned bucket value, s i* is the target lead time, i represents the i-th sample video, and s i is the lead time corresponding to the i-th sample video.

[0092] By combining the lead time and the equal-frequency percentiles to obtain the first recommended value, the attractiveness and value of the video can be more comprehensively reflected, thereby improving the prediction accuracy of the recommendation model.

[0093] To more accurately determine the first video, in the training method of the video recommendation model provided in the first embodiment of this application, determining the first video that satisfies the first association relationship with the sample videos among multiple sample videos includes: obtaining multiple first candidate videos that the target object watches after watching the target sample video and within the second time period, where the target sample video is any one of the multiple sample videos; determining whether the first candidate videos among the multiple first candidate videos satisfy the first association relationship to obtain the first target judgment result; and determining the first video from the multiple first candidate videos based on the first target judgment result.

[0094] Optionally, after the user finishes watching a video (the target sample video), the viewing history of the user in the subsequent second time period will be tracked to identify all subsequent videos watched by the user. These subsequent videos are the first candidate videos, and the first candidate videos represent all the content that the user may be interested in after watching the target sample video. It should be noted that the second time period can be the period when the user finishes watching the video after watching the target sample video, or a time period set according to the user's behavior pattern and task requirements.

[0095] Then, it is determined whether the first candidate videos among the multiple first candidate videos satisfy the first association relationship. It should be noted that the first association relationship refers to the degree of association between the first candidate video and the target sample video, and it can be evaluated from multiple dimensions, such as the author dimension, the category dimension, the multi-modal similarity, the co-occurrence frequency, etc. For example, it is determined whether the first candidate videos among the multiple first candidate videos are released by the same author, whether they belong to the same category, and whether their contents are highly similar.

[0096] Finally, according to the first target judgment result, the multiple first candidate videos are screened. For example, the first candidate videos that are indeed associated with the target sample video are determined as the first videos.

[0097] By screening the first videos associated with the target sample video, the video recommendation model can learn which video types, themes, or authors are more likely to attract the user's continuous attention, thereby improving the prediction ability of the user's future viewing behavior.

[0098] To improve the accuracy of determining whether the first association relationship is satisfied, in the training method of the video recommendation model provided in the first embodiment of the present application, determining whether the first candidate videos among the multiple first candidate videos satisfy the first association relationship and obtaining the first target judgment result includes: determining whether the content of the first candidate videos among the multiple first candidate videos is similar to the target sample video to obtain the first judgment result; determining whether the access behavior of the first candidate videos among the multiple first candidate videos is similar to the target sample video to obtain the second judgment result; determining whether the rank of the first candidate videos among the multiple first candidate videos is the target rank to obtain the third judgment result; and determining whether the first candidate videos among the multiple first candidate videos satisfy the first association relationship based on the first judgment result, the second judgment result, and the third judgment result to obtain the first target judgment result.

[0099] Optionally, in video recommendation, evaluating the relevance between the subsequent videos (the first candidate videos) watched by a user after watching a video (the target sample video) and this video is an important part of ensuring the quality of recommended content, improving user satisfaction and engagement. Therefore, determining whether the first candidate video among multiple first candidate videos meets the first relevance relationship and obtaining the first target judgment result includes: determining whether the content of the first candidate video is similar to the target sample video. In this part, the similarity between the first candidate video and the target sample video in terms of content is mainly determined, and multimodal similarity calculation methods such as image similarity, text similarity, and audio similarity can be used to quantify the content relevance between videos. For example, by comparing the titles, descriptions, image frames, and audio features of videos, the similarity score at the content level can be obtained. If the score is high, it indicates that the two videos have a high similarity in content and may thus meet the first relevance relationship.

[0100] For example, judge whether the content of the first candidate video is similar to the target sample video from three dimensions: multimodal similarity, same author, and same first-level category. For multimodal similarity: calculate the multimodal similarity pairwise for the videos, and determine the videos with multimodal similarity greater than 0.9 as similar to the target sample video; for the same author, if the publishing author of the first candidate video is the same as that of the target sample video, it is determined to be similar to the target sample video; for the same first-level category: if the category of the first candidate video is the same as that of the target sample video, it is determined to be similar to the target sample video.

[0101] Then, determine whether the access behavior of the first candidate video is similar to the target sample video. For example, whether the access behavior is similar to the target sample video can be the same recall or v2v co-occurrence. Same recall means whether the first candidate video and the target sample video are the same video. V2v co-occurrence means pairwise calculation of the videos under the same session (the behavior of currently watching a video), and determining the videos with v2v left table similarity greater than 0.5 as the videos similar to the target sample video.

[0102] V2V co-occurrence (Video-to-Video Co-occurrence) is a method for analyzing and measuring the relevance between videos, mainly applied in recommendation systems, especially in the video recommendation scenario. It infers whether there is a certain form of connection or similarity between videos by statistically analyzing the co-occurrence patterns of videos in the user's viewing history. The V2V left table refers to a specific data structure or query method used in V2V similarity or relevance analysis, that is, the left join table. Left join is a database operation used to combine two tables based on a common key, ensuring that all records in the left table will appear in the result set, even if there are no matching records in the right table.

[0103] For example, the V2V left table may include metadata information of videos, such as video ID, title, upload time, category label, author ID, etc. It serves as the main collection of video data and is the basis for performing V2V correlation or similarity calculations. When conducting V2V similarity or correlation analysis, the left table is used as the initiator of the query. Based on each video in the left table, other videos with co-occurrence relationships or similar features are searched for. For example, other videos that frequently co-occur with video A in the user's viewing sequence can be queried to identify potential associated videos of video A. By using the V2V left table and the left join operation, the correlation and similarity between videos can be analyzed more comprehensively.

[0104] Secondly, it is determined whether the ordinal position of the first candidate video is the target ordinal position. For example, if 4 videos are returned due to 1 request, the 6 videos (i.e., the first 6 sorted first candidate videos) subsequently led out by the current video (i.e., the above-mentioned target sample video) are considered as the first videos that satisfy the first association relationship with the target sample video.

[0105] Finally, according to the first judgment result, the second judgment result, and the third judgment result, it is determined whether the first candidate video among the multiple first candidate videos satisfies the first association relationship, and the first target judgment result is obtained.

[0106] By comprehensively considering content similarity, access behavior similarity, and ordinal position continuity to determine whether the first candidate video is truly associated with the target sample video, it is possible to more accurately identify which subsequent videos are truly led to be watched by the target sample video, which helps to reduce the noise of irrelevant videos and improve the accuracy of subsequent video recommendations.

[0107] To improve the accuracy of determining the second video, in the training method of the video recommendation model provided in the first embodiment of the present application, determining the second video that satisfies the second association relationship with the sample video among multiple sample videos includes: obtaining multiple second candidate videos that the target object watches within the second time period after watching the target sample video and multiple third candidate videos that the target object watches within the third time period after watching the target sample video; determining whether the second candidate video among the multiple second candidate videos satisfies the second association relationship to obtain the second target judgment result, and determining whether the third candidate video among the multiple third candidate videos satisfies the second association relationship to obtain the third target judgment result; and determining the second video from the multiple second candidate videos and the multiple third candidate videos based on the second target judgment result and the third target judgment result.

[0108] Optionally, after a user finishes watching a certain video (i.e., the target sample video), information about other videos watched by the user within a set time period (the second time period or the third time period) is collected. Among them, the second candidate video and the third candidate video are both videos that may be associated with the target sample video. For example, the second candidate video may refer to those videos watched within the same video viewing behavior, while the third candidate video refers to videos watched across sessions and across days.

[0109] In an optional embodiment, the second candidate video is the video watched on the same day after watching the target sample video. The third candidate video is the video watched within 7 days after watching the target sample video.

[0110] Then, it is determined whether the second candidate video among the multiple second candidate videos satisfies the second association relationship and whether the third candidate video among the multiple third candidate videos satisfies the second association relationship.

[0111] In the short video field, high-quality creators have a huge influence on users. Users often visit their favorite authors frequently and even watch repeatedly for several consecutive days. In the re-ranking stage, the diversity of the entire sequence usually needs to be considered comprehensively. The diversity of authors in a single session limits the cumulative exposure time of users to a specific author, thus restricting the in-depth understanding of user preferences. Therefore, in order to more accurately capture user preferences and interests, a broader time perspective should be adopted to deeply model and analyze the authors favored by users. Specifically, by extending the observation period to seven days, the long-term attractiveness and potential value of each author to readers can be evaluated more comprehensively. In this way, the possible impact of the current video on users in the next few days can be simulated.

[0112] That is, it can be determined whether the second candidate video satisfies the second association relationship and whether the third candidate video satisfies the second association relationship by determining whether the publishing object corresponding to the second candidate video among the multiple second candidate videos is the same as the target sample video and whether the publishing object corresponding to the third candidate video among the multiple third candidate videos is the same as the target sample video. For example, when the publishing object corresponding to the second candidate video among the multiple second candidate videos is the same as the target sample video, it is determined that the second target judgment result is that the second candidate video satisfies the second association relationship; when the publishing object corresponding to the second candidate video among the multiple second candidate videos is different from the target sample video, it is determined that the second target judgment result is that the second candidate video does not satisfy the second association relationship.

[0113] By identifying the second video and the third video that have the second association relationship with the target sample video, the continuity and evolution of user viewing behavior can be understood more comprehensively, and the accuracy and personalization level of recommendations can be improved.

[0114] How to train a video recommendation model is crucial. Therefore, in the training method of the video recommendation model provided in Embodiment 1 of this application, training the initial video recommendation model based on the predicted recommendation value set and the first recommendation value set to obtain the target video recommendation model includes: calculating the loss based on the first predicted recommendation value and the first recommendation value in the predicted recommendation value set to obtain the first loss function; calculating the loss based on the second predicted recommendation value and the second recommendation value in the predicted recommendation value set to obtain the second loss function; calculating the loss based on the third predicted recommendation value and the third recommendation value in the predicted recommendation value set to obtain the third loss function; training the initial video recommendation model based on the first loss function, the second loss function, and the third loss function to obtain the target video recommendation model.

[0115] Optionally, calculating the loss based on the first predicted recommendation value and the first recommendation value in the predicted recommendation value set to obtain the first loss function. For example, where θ(u i ) is the first predicted recommendation value, and u i is the feature set of the sample video.

[0116] Then, calculating the loss based on the second predicted recommendation value and the second recommendation value in the predicted recommendation value set to obtain the second loss function, and calculating the loss based on the third predicted recommendation value and the third recommendation value in the predicted recommendation value set to obtain the third loss function. It should be noted that the second loss function and the third loss function can also use MSE (mean squared error) as the loss function. Finally, training the initial video recommendation model based on the first loss function, the second loss function, and the third loss function to obtain the target video recommendation model. In the training stage, considering the first, second, and third loss functions simultaneously, adjusting the model parameters through the backpropagation algorithm, so that the video recommendation model can consider the long-term value and deep user interests, thereby improving the overall recommendation accuracy.

[0117] Since the distribution of the second recommendation value is not a normal distribution, but more conforms to the tweed ie (exponential distribution family) distribution. MSE Loss is proven to be more suitable for learning samples with a normal distribution. In addition, through the analysis of the estimated scores, it is found that there is an underestimation phenomenon for this target. Therefore, in the training method of the video recommendation model provided in Embodiment 1 of this application, calculating the loss based on the second predicted recommendation value and the second recommendation value in the predicted recommendation value set to obtain the second loss function includes: calculating the mean squared error between the second predicted recommendation value and the second recommendation value to obtain the first loss sub-function; calculating the loss based on the exponential distribution family for the second predicted recommendation value and the second recommendation value to obtain the second loss sub-function; obtaining the second loss function based on the first loss sub-function and the second loss sub-function.

[0118] Optionally, calculate the mean squared error between the second predicted recommendation value and the second recommendation value to obtain the first loss sub-function. Mean Squared Error (MSE) is a common loss function used to measure the difference between the predicted value and the true value. It is suitable for samples with a normal distribution but may not perform well on non-normal distribution data, especially when dealing with the second recommendation value with a large number of zeros. Therefore, calculate the loss between the second predicted recommendation value and the second recommendation value based on Tweedie to obtain the second loss sub-function. The Tweedie loss function calculates the loss by directly optimizing the statistic related to the Tweedie distribution, which can more accurately reflect the difference between the predicted value and the true value, especially on non-normal distribution data. Finally, obtain the second loss function according to the first loss sub-function and the second loss sub-function.

[0119] For example, Loss = Loss MSE + rLoss Tweedie , where Loss is the second loss function, Loss MSE is the first loss sub-function, Loss Tweedie is the second loss sub-function, and r is a preset weight. y i is the second recommendation value, μ i is the second predicted recommendation value, and ρ is a preset hyperparameter.

[0120] By using the Tweedie loss function to process the non-normal distribution of the second recommendation value, the model can more accurately fit the true distribution of the data and improve the accuracy of prediction.

[0121] In an optional embodiment, as shown in the schematic diagram Figure 5 , the recommendation values of the video are divided into the current recommendation value, the session-level recommendation value (i.e., the first recommendation value and the second recommendation value above), and the daily-level recommendation value (the third recommendation value). The user's behavior towards the current video (such as following, commenting, and favoriting) indicates the current recommendation value; the leading duration of the current video indicates the session-level recommendation value; the user's repeated viewing of a certain author's content over multiple days indicates the daily-level value.

[0122] Regarding the calculation of session-level recommendation values through the guiding duration of the current video, the existing calculation methods have the following problems: Modeling only using the cumulative play time at the current position of the current video will lead to significant biases. Simply summing up the play times of subsequent videos is too simplistic and may introduce noise. Therefore, in response to the problems of the existing technology, the session-level recommendation value is split into the above-mentioned first recommendation value and second recommendation value. The first recommendation value: It is determined based on the quantiles of the guiding duration across each video page; the second recommendation value: Based on three dimensions of context, behavioral similarity, and content similarity, excluding irrelevant video content, the second recommendation value is calculated based on the filtered videos.

[0123] For the daily recommendation value, the observation period is extended to seven days to more comprehensively evaluate the long-term attractiveness of each author to readers and their potential value. The definition of its daily-level value is as follows:

[0124]

[0125] Among them, V mj represents the daily-level value of the j-th video, represents the sum of the cumulative viewing durations of α videos on the m-th day. C mα takes a value of 0 or 1. When video j and the α-th video are created by the same author, it is 1; otherwise, it is 0. represents the sum of the cumulative viewing durations of α videos on the m-th day. C dα takes a value of 0 or 1. When video j and the α-th video are created by the same author, it is 1; otherwise, it is 0.

[0126] It should be noted that since the aggregation process spans several days in the future, this process cannot be tagged on the same day. Therefore, the samples will be assembled with a delay of A days, as Figure 6 shown, while obtaining the normal samples for the A-th day and the daily-level value samples from A days.

[0127] Since the samples cover up to A days, initially separate and model this task separately to avoid affecting the update of the online model. However, this method leads to an increase in the online response time and requires additional resources. Therefore, it is necessary to integrate the model so that it can be trained on different samples from t - A days and t - 1 days simultaneously. Adopt the co-train strategy to jointly model the long-term value model of the author and the daily-level update model. Multiple tasks share the underlying embedding. When training the daily-level value task, apply stop-gradient to prevent it from affecting the update of the shared embedding. For other tasks, the shared embedding is updated normally. This method enables the model to be effectively updated and trained in multiple loops and tasks. As Figure 7As shown in the figure, the video recommendation model includes a shared embedding module and multiple feature extraction modules (the first feature extraction module, the second feature extraction module, and the third feature extraction model). Among them, the embedding module consists of a SetNet (deep convolutional network), an embedding layer, and a target attention layer (Target Attention). The feature extraction modules in the multiple feature extraction modules are composed of a feature processing sub-model and a parameter personalization network. The feature processing sub-model is composed of multiple feature processing layers (such as Figure 7 Layer 1, Layer 2, and Layer n in Figure 7 ), and the parameter personalization network is composed of multiple gated neural networks (i.e.,

[0128] GateNN1, GateNN2, and GateNNn in

[0129]

[0130]

[0131] Method XAUC MSE MAE PCOC XAUC-2 Prior Art 0.6252 4.9847 1.7434 0.8485 - This Application 0.6378(+0.0126) 0.0946 0.2402 0.8481 0.6894

[0132]

[0133] The first feature extraction module is used to obtain the daily recommendation value, the second feature extraction module is used to obtain the first recommendation value in the session-level recommendation value, and the third feature extraction module is used to obtain the second recommendation value in the session-level recommendation value.

[0133] In an optional embodiment, a third feature extraction module may also be set in the video recommendation model, and the current value of the video is obtained through the third feature extraction module.

[0129] In an optional implementation, a series of ablation studies were conducted to evaluate the effectiveness of the method in this application on the first recommendation value in the original and session-level recommendation values. Experiments show that the method of calculating the first recommendation value in this application improves the prediction accuracy of the sliding time sorting. In addition, this application significantly reduces the MSE error. XAUC-2 is calculated based on the grouped quantile label of the bootstrap duration and exceeds its own XAUC score. Overall, the method in this application performs better than the baseline in most metrics. As shown in Table 1.

[0130] Table 1

[0131] Method XAUC MSE MAE PCOC XAUC-2 Prior Art 0.6252 4.9847 1.7434 0.8485 - This Application 0.6378(+0.0126) 0.0946 0.2402 0.8481 0.6894

[0132] Among them, MSE is the mean square error, MAE is the mean absolute error, XAUC is the inverse order pair, and the consistency between the prediction order and the true order; PCOC (Predict Click Over Click) is a metric used to evaluate the effect of prediction calibration, and the closer it is to 1, the closer the predicted value is to the true value.

[0133] The effectiveness of this application on the original and attributed sliding time (i.e., the second recommendation value) is shown in Table 2. Among them, it is observed that the XAUC increases by 0.0118, while the MSE decreases by 0.8755. Through the integration of Tweedie loss optimization, the performance is further improved. This adjustment effectively alleviates the underestimation problem and makes the model prediction more accurate.

[0134] Table 2

[0135]

[0136]

[0137] Two methods are adopted to train the author duration task (i.e., the third recommended value). One method is to start a new model (such as the single model in Figure 8 ) specifically for this task and perform warm-up (learning rate adjustment strategy) based on the existing model. This method ensures that the latency of this task will not affect other tasks. Alternatively, the co-train framework can be used to train the author duration task together with other tasks. The offline evaluation shows the comparison of the effects of the single model and co-train on MSE, MAE, XAUC, and PCOC as shown in Figure 8 . Compared with the single model, the co-training method provides an additional improvement in the PCOC metric.

[0138] In the training method of the video recommendation model provided in the first embodiment of this application, by obtaining a training sample set, where the training sample set at least includes: a plurality of sample videos, a feature set corresponding to the sample videos in the plurality of sample videos, and a first recommended value set corresponding to the sample videos, the first recommended value set at least includes: a first recommended value, a second recommended value, and a third recommended value, the first recommended value is obtained based on the guiding duration of the sample video, the second recommended value is obtained based on the viewing duration of a first video that satisfies a first association relationship with the sample video, and the third recommended value is obtained based on the viewing duration of a second video that satisfies a second association relationship with the sample video; by processing the training sample set through an initial video recommendation model, a predicted recommended value set is obtained; and by training the initial video recommendation model according to the predicted recommended value set and the first recommended value set, a target video recommendation model is obtained, where the video to be recommended is determined based on the video recommendation value output by the target video recommendation model, which solves the technical problem in the related art that the video recommendation model performs video recommendation according to the current recommended value corresponding to the video, resulting in relatively low accuracy of video recommendation.

[0139] In this solution, a training sample set is constructed through multiple sample videos, the feature sets corresponding to the sample videos in the multiple sample videos, and the first recommendation value sets corresponding to the sample videos. Then, the initial video recommendation model is iteratively trained through the training sample set to obtain the target video recommendation model. The first recommendation value set includes at least a first recommendation value, a second recommendation value, and a third recommendation value. The first recommendation value is calculated through the guiding duration of the sample video and can accurately evaluate the influence degree of the sample video on subsequent videos. The second recommendation value and the third recommendation value are respectively calculated based on the association relationship between videos and the viewing duration of subsequent videos. Such multi-dimensional recommendation values enable the video recommendation model to more accurately understand the connections between videos and the complexity of user behavior, so as to recommend videos that not only match the user's current interests but also can guide the user to watch more relevant videos, thereby achieving the technical effect of improving the accuracy of video recommendations.

[0140] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of this application.

[0142] Embodiment 2

[0143] According to an embodiment of the present application, a video recommendation method is further provided, as Figure 9 shown. The video recommendation method includes:

[0144] Step S901, obtaining target feature information corresponding to multiple initial videos;

[0145] Step S902: Process the target feature information through the target video recommendation model to obtain a set of second recommendation values corresponding to the initial videos in multiple initial videos, where the target video recommendation model is trained by using the training method of the video recommendation model in any one of Embodiment 1.

[0146] Step S903: Determine the target video to be recommended from multiple initial videos according to the set of second recommendation values.

[0147] Optionally, collect the target feature information corresponding to multiple initial videos. The target feature information at least includes video features: including the title, description, duration, upload time, category label, author information, user interaction data (such as likes, comments, shares), etc. User features: interest preferences, activity patterns; Context features: location information of the video, device information, etc.; Sequence features: the user's viewing history videos, etc. Input the target feature information collected in Step S901 into the target video recommendation model to obtain a set of second recommendation values corresponding to the initial videos in multiple initial videos. It should be noted that the set of second recommendation values may include the current value and long-term value of the initial video. Based on the obtained set of second recommendation values, select the target video to be recommended therefrom. For example, sort the set of second recommendation values from high to low according to the recommendation value, and select the target video to be recommended according to the sorting result.

[0148] Through these steps, the target video recommendation model can provide personalized and high-quality video recommendations according to the user's interest preferences and video content. This not only increases the possibility of users watching videos, but also promotes the content distribution efficiency and user satisfaction of the platform.

[0149] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0150] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0151] Embodiment 3

[0152] According to an embodiment of the present application, a video recommendation method is further provided. As Figure 10 shown, the video recommendation method includes:

[0153] Step S1001: Obtain target feature information corresponding to multiple initial videos uploaded by a client;

[0154] Step S1002: Process the target feature information through a target video recommendation model in a cloud server to obtain a second recommendation value set corresponding to the initial video in the multiple initial videos, where the target video recommendation model is trained by using the training method of the video recommendation model in any one of Embodiment 1; determine a target video to be recommended from the multiple initial videos according to the second recommendation value set;

[0155] Step S1003: Return the target video to be recommended to the client.

[0156] It should be noted that the specific method of the video recommendation method in the cloud server is the same as that in Embodiment 2 and will not be described in detail here.

[0157] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0158] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present application.

[0159] Embodiment 4

[0160] According to an embodiment of the present application, there is also provided a training device for a video recommendation model for implementing the training method of the above video recommendation model, as Figure 11 shown. The device includes: a first acquisition unit 1101, a first processing unit 1102, and a training unit 1103.

[0161] The first acquisition unit 1101 is configured to acquire a training sample set, where the training sample set at least includes: a plurality of sample videos, a feature set corresponding to the sample videos in the plurality of sample videos, and a first recommendation value set corresponding to the sample videos. The first recommendation value set at least includes: a first recommendation value, a second recommendation value, and a third recommendation value. The first recommendation value is obtained based on the guiding duration of the sample video. The second recommendation value is obtained based on the viewing duration of a first video that satisfies a first association relationship with the sample video. The third recommendation value is obtained based on the viewing duration of a second video that satisfies a second association relationship with the sample video.

[0162] The first processing unit 1102 is configured to process the training sample set through an initial video recommendation model to obtain a predicted recommendation value set.

[0163] The training unit 1103 is configured to train the initial video recommendation model based on the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model, where the video to be recommended is determined based on the video recommendation value output by the target video recommendation model.

[0164] In the training device of the video recommendation model provided in the fourth embodiment of the present application, a training sample set is obtained through the first acquisition unit 1101. The training sample set at least includes: a plurality of sample videos, a feature set corresponding to the sample videos in the plurality of sample videos, and a first recommendation value set corresponding to the sample videos. The first recommendation value set at least includes: a first recommendation value, a second recommendation value, and a third recommendation value. The first recommendation value is obtained based on the guiding duration of the sample video. The second recommendation value is obtained based on the viewing duration of the first video that satisfies the first association relationship with the sample video. The third recommendation value is obtained based on the viewing duration of the second video that satisfies the second association relationship with the sample video. The first processing unit 1102 processes the training sample set through the initial video recommendation model to obtain a predicted recommendation value set. The training unit 1103 trains the initial video recommendation model based on the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model. Among them, the video to be recommended is determined based on the video recommendation value output by the target video recommendation model, which solves the technical problem in the related art that the video recommendation model recommends videos according to the current recommendation value corresponding to the video, resulting in relatively low accuracy of video recommendation.

[0165] In this solution, a training sample set is constructed through a plurality of sample videos, a feature set corresponding to the sample videos in the plurality of sample videos, and a first recommendation value set corresponding to the sample videos. Then, the initial video recommendation model is iteratively trained through the training sample set to obtain a target video recommendation model. The first recommendation value set at least includes a first recommendation value, a second recommendation value, and a third recommendation value. The first recommendation value is calculated through the guiding duration of the sample video and can accurately evaluate the influence degree of the sample video on subsequent videos. The second recommendation value and the third recommendation value are respectively calculated based on the association relationship between videos and the viewing duration of subsequent videos. Such multi-dimensional recommendation values enable the video recommendation model to more accurately understand the connections between videos and the complexity of user behavior, so as to recommend videos that not only match the user's current interests but also can guide the user to watch more relevant videos, thereby achieving the technical effect of improving the accuracy of video recommendation.

[0166] Optionally, in the training device of the video recommendation model provided in the fourth embodiment of the present application, the recommendation value set is obtained through the following device: a second acquisition unit, configured to obtain the guiding duration corresponding to the sample videos in the plurality of sample videos, and obtain the first recommendation value based on the guiding duration; a first determination unit, configured to determine the first video that satisfies the first association relationship with the sample videos in the plurality of sample videos, and calculate the second recommendation value based on the viewing duration of the first video; a second determination unit, configured to determine the second video that satisfies the second association relationship with the sample videos in the plurality of sample videos, and calculate the third recommendation value based on the viewing duration of the second video.

[0167] Optionally, in the training device of the video recommendation model provided in the fourth embodiment of this application, the second acquisition unit includes: a first acquisition module, configured to acquire a third video that the target object watches within a first time period after watching the target sample video, where the target sample video is any one of a plurality of sample videos; a second acquisition module, configured to acquire the viewing duration corresponding to the third video; a first calculation module, configured to perform a calculation based on the viewing duration to obtain a guiding duration.

[0168] Optionally, in the training device of the video recommendation model provided in the fourth embodiment of this application, the second acquisition unit includes: a third acquisition module, configured to acquire the position information corresponding to the sample videos in a plurality of sample videos, and perform grouping processing on the plurality of sample videos based on the position information to obtain multiple groups of sample videos; a second calculation module, configured to, for any group of sample videos, perform a calculation based on the guiding duration corresponding to the sample videos in this group of sample videos to obtain the equal-frequency quantile corresponding to the sample videos in this group of sample videos; a first determination module, configured to obtain a first recommendation value based on the equal-frequency quantile and the guiding duration.

[0169] Optionally, in the training device of the video recommendation model provided in the fourth embodiment of this application, the determination module includes: a first determination sub-module, configured to determine a target guiding duration based on the guiding duration corresponding to the sample videos in this group of sample videos; a processing sub-module, configured to perform bucketing processing on the sample videos in this group of sample videos according to the equal-frequency quantile to obtain a bucketing value corresponding to the equal-frequency quantile; a first calculation sub-module, configured to perform a calculation based on the bucketing value and the target guiding duration to obtain a first recommendation value.

[0170] Optionally, in the training device of the video recommendation model provided in the fourth embodiment of this application, the first determination unit includes: a fourth acquisition module, configured to acquire a plurality of first candidate videos that the target object watches within a second time period after watching the target sample video, where the target sample video is any one of a plurality of sample videos; a first judgment module, configured to judge whether the first candidate videos in the plurality of first candidate videos satisfy a first association relationship to obtain a first target judgment result; a second determination module, configured to determine a first video from the plurality of first candidate videos according to the first target judgment result.

[0171] Optionally, in the training device of the video recommendation model provided in the fourth embodiment of the present application, the judgment module includes: a first judgment sub-module, configured to judge whether the content of the first candidate video among the multiple first candidate videos is similar to the target sample video, and obtain a first judgment result; a second judgment sub-module, configured to judge whether the access behavior of the first candidate video among the multiple first candidate videos is similar to the target sample video, and obtain a second judgment result; a third judgment sub-module, configured to judge whether the sequence number of the first candidate video among the multiple first candidate videos is the target sequence number, and obtain a third judgment result; a fourth judgment sub-module, configured to judge whether the first candidate video among the multiple first candidate videos meets the first association relationship according to the first judgment result, the second judgment result, and the third judgment result, and obtain a first target judgment result.

[0172] Optionally, in the training device of the video recommendation model provided in the fourth embodiment of the present application, the second determination unit includes: a fifth acquisition module, configured to acquire multiple second candidate videos watched by the target object within a second time period after watching the target sample video and multiple third candidate videos watched by the target object within a third time period after watching the target sample video; a second judgment module, configured to judge whether the second candidate video among the multiple second candidate videos meets the second association relationship, and obtain a second target judgment result, and judge whether the third candidate video among the multiple third candidate videos meets the second association relationship, and obtain a third target judgment result; a second determination module, configured to determine a second video from the multiple second candidate videos and the multiple third candidate videos according to the second target judgment result and the third target judgment result.

[0173] Optionally, in the training device of the video recommendation model provided in the fourth embodiment of the present application, the second judgment module includes: a fifth judgment sub-module, configured to judge whether the publishing object corresponding to the second candidate video among the multiple second candidate videos is the same as the target sample video; a second determination sub-module, configured to determine that the second target judgment result is that the second candidate video meets the second association relationship when the publishing object corresponding to the second candidate video among the multiple second candidate videos is the same as the target sample video; a third determination sub-module, configured to determine that the second target judgment result is that the second candidate video does not meet the second association relationship when the publishing object corresponding to the second candidate video among the multiple second candidate videos is not the same as the target sample video.

[0174] Optionally, in the training device of the video recommendation model provided in the fourth embodiment of the present application, the training unit includes: a second calculation module, configured to calculate a loss based on a first predicted recommendation value and a first recommendation value in a set of predicted recommendation values to obtain a first loss function; a third calculation module, configured to calculate a loss based on a second predicted recommendation value and a second recommendation value in the set of predicted recommendation values to obtain a second loss function; a fourth calculation module, configured to calculate a loss based on a third predicted recommendation value and a third recommendation value in the set of predicted recommendation values to obtain a third loss function; and a training module, configured to train an initial video recommendation model based on the first loss function, the second loss function, and the third loss function to obtain a target video recommendation model.

[0175] Optionally, in the training device of the video recommendation model provided in the fourth embodiment of the present application, the third calculation module includes: a second calculation sub-module, configured to calculate a mean square error between the second predicted recommendation value and the second recommendation value to obtain a first loss sub-function; a third calculation sub-module, configured to calculate a loss based on the exponential distribution family for the second predicted recommendation value and the second recommendation value to obtain a second loss sub-function; and a fourth determination sub-module, configured to obtain a second loss function based on the first loss sub-function and the second loss sub-function.

[0176] It should be noted here that the above-mentioned first acquisition unit 1101, first processing unit 1102, and training unit 1103 correspond to steps S201 to S203 in Embodiment 1. The functions of the three units are the same as those of the corresponding steps in terms of implementation examples and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.

[0177] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.

[0178] Embodiment 5

[0179] According to an embodiment of the present application, there is also provided a video recommendation device for implementing the above video recommendation method, as Figure 12 shown. The device includes: a third acquisition unit 1201, a first processing unit 1202, and a third determination unit 1203.

[0180] The third acquisition unit 1201 is configured to acquire target feature information corresponding to multiple initial videos;

[0181] The first processing unit 1202 is configured to process the target feature information through a target video recommendation model to obtain a second recommendation value set corresponding to an initial video in multiple initial videos, where the target video recommendation model is trained by using the video recommendation model training device in any one of the first embodiments;

[0182] The third determination unit 1203 is configured to determine a target video to be recommended from multiple initial videos according to the second recommendation value set. It should be noted here that the above-mentioned third acquisition unit 1201, the first processing unit 1202, and the third determination unit 1203 correspond to steps S901 to S903 in the second embodiment. The functions of the three units are the same as those of the corresponding steps in terms of implementation examples and application scenarios, but are not limited to the content disclosed in the above-mentioned second embodiment.

[0183] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as those provided in the second embodiment in terms of application scenarios and implementation processes, but are not limited to the solutions provided in the second embodiment.

[0184] Embodiment 6

[0185] An embodiment of the present application can provide an electronic device, and the electronic device can be any one of the electronic device terminals in the electronic device terminal group. Optionally, in this embodiment, the above-mentioned electronic device can also be replaced with a terminal device such as a mobile terminal.

[0186] Optionally, in this embodiment, the above-mentioned electronic device can be at least one of multiple network devices in a computer network.

[0187] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the video recommendation model training method or the video recommendation method: obtaining a training sample set, where the training sample set at least includes: multiple sample videos, a feature set corresponding to the sample video in the multiple sample videos, and a first recommendation value set corresponding to the sample video, and the first recommendation value set at least includes: a first recommendation value, a second recommendation value, and a third recommendation value, the first recommendation value is obtained based on the guiding duration of the sample video, the second recommendation value is obtained based on the viewing duration of a first video that satisfies a first association relationship with the sample video, and the third recommendation value is obtained based on the viewing duration of a second video that satisfies a second association relationship with the sample video; processing the training sample set through an initial video recommendation model to obtain a predicted recommendation value set; training the initial video recommendation model according to the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model, where the video to be recommended is determined based on the video recommendation value output by the target video recommendation model.

[0188] The above computer terminal can execute the program code for the following steps in the video recommendation model training method or the video recommendation method: The first recommendation value set is obtained through the following method: Obtain the guiding duration corresponding to the sample video in multiple sample videos, and based on the guiding duration, obtain the first recommendation value; Determine the first video that satisfies the first association relationship with the sample video in multiple sample videos, and calculate based on the viewing duration of the first video to obtain the second recommendation value; Determine the second video that satisfies the second association relationship with the sample video in multiple sample videos, and calculate based on the viewing duration of the second video to obtain the third recommendation value.

[0189] The above computer terminal can execute the program code for the following steps in the video recommendation model training method or the video recommendation method: Obtaining the guiding duration corresponding to the sample video in multiple sample videos includes: Obtain the third video that the target object watches after watching the target sample video and within the first time period, where the target sample video is any one of the multiple sample videos; Obtain the viewing duration corresponding to the third video; Calculate based on the viewing duration to obtain the guiding duration.

[0190] The above computer terminal can execute the program code for the following steps in the video recommendation model training method or the video recommendation method: Obtaining the first recommendation value based on the guiding duration includes: Obtain the position information corresponding to the sample video in multiple sample videos, and perform grouping processing on the multiple sample videos based on the position information to obtain multiple groups of sample videos; For any group of sample videos, calculate based on the guiding duration corresponding to the sample video in the group to obtain the equal-frequency quantile corresponding to the sample video in the group; Based on the equal-frequency quantile and the guiding duration, obtain the first recommendation value.

[0191] The above computer terminal can execute the program code for the following steps in the video recommendation model training method or the video recommendation method: Obtaining the first recommendation value based on the equal-frequency quantile and the guiding duration includes: Based on the guiding duration corresponding to the sample video in the group, determine the target guiding duration; Perform bucketing processing on the sample video in the group based on the equal-frequency quantile to obtain the bucket value corresponding to the equal-frequency quantile; Calculate based on the bucket value and the target guiding duration to obtain the first recommendation value.

[0192] The above computer terminal can execute the program code for the following steps in the video recommendation model training method or the video recommendation method: Determining the first video that satisfies the first association relationship with the sample video among multiple sample videos includes: obtaining multiple first candidate videos that the target object watches after watching the target sample video and within the second time period, where the target sample video is any one of the multiple sample videos; determining whether the first candidate video among the multiple first candidate videos satisfies the first association relationship to obtain a first target judgment result; and determining the first video from the multiple first candidate videos based on the first target judgment result.

[0193] The above computer terminal can execute the program code for the following steps in the video recommendation model training method or the video recommendation method: Determining whether the first candidate video among the multiple first candidate videos satisfies the first association relationship to obtain a first target judgment result includes: determining whether the content of the first candidate video among the multiple first candidate videos is similar to the target sample video to obtain a first judgment result; determining whether the access behavior of the first candidate video among the multiple first candidate videos is similar to the target sample video to obtain a second judgment result; determining whether the sequence position of the first candidate video among the multiple first candidate videos is the target sequence position to obtain a third judgment result; and determining whether the first candidate video among the multiple first candidate videos satisfies the first association relationship based on the first judgment result, the second judgment result, and the third judgment result to obtain a first target judgment result.

[0194] The above computer terminal can execute the program code for the following steps in the video recommendation model training method or the video recommendation method: Determining the second video that satisfies the second association relationship with the sample video among multiple sample videos includes: obtaining multiple second candidate videos that the target object watches after watching the target sample video and within the second time period and multiple third candidate videos that the target object watches after watching the target sample video and within the third time period; determining whether the second candidate video among the multiple second candidate videos satisfies the second association relationship to obtain a second target judgment result, and determining whether the third candidate video among the multiple third candidate videos satisfies the second association relationship to obtain a third target judgment result; and determining the second video from the multiple second candidate videos and the multiple third candidate videos based on the second target judgment result and the third target judgment result.

[0195] The above computer terminal can execute the program code of the following steps in the video recommendation model training method or the video recommendation method: determining whether a second candidate video among a plurality of second candidate videos satisfies a second association relationship, and obtaining a second target determination result, including: determining whether the publishing object corresponding to the second candidate video among the plurality of second candidate videos is the same as the target sample video; in the case where the publishing object corresponding to the second candidate video among the plurality of second candidate videos is the same as the target sample video, determining that the second target determination result is that the second candidate video satisfies the second association relationship; in the case where the publishing object corresponding to the second candidate video among the plurality of second candidate videos is different from the target sample video, determining that the second target determination result is that the second candidate video does not satisfy the second association relationship.

[0196] The above computer terminal can execute the program code of the following steps in the video recommendation model training method or the video recommendation method: training an initial video recommendation model based on a predicted recommendation value set and a first recommendation value set to obtain a target video recommendation model, including: calculating a loss based on a first predicted recommendation value and a first recommendation value in the predicted recommendation value set to obtain a first loss function; calculating a loss based on a second predicted recommendation value and a second recommendation value in the predicted recommendation value set to obtain a second loss function; calculating a loss based on a third predicted recommendation value and a third recommendation value in the predicted recommendation value set to obtain a third loss function; training the initial video recommendation model based on the first loss function, the second loss function, and the third loss function to obtain a target video recommendation model.

[0197] The above computer terminal can execute the program code of the following steps in the video recommendation model training method or the video recommendation method: calculating a loss based on a second predicted recommendation value and a second recommendation value in the predicted recommendation value set to obtain a second loss function, including: calculating a mean square error between the second predicted recommendation value and the second recommendation value to obtain a first loss sub-function; calculating a loss based on the exponential distribution family for the second predicted recommendation value and the second recommendation value to obtain a second loss sub-function; obtaining the second loss function based on the first loss sub-function and the second loss sub-function.

[0198] According to another aspect of the embodiments of the present application, there is also provided a video recommendation method, including: obtaining target feature information corresponding to a plurality of initial videos; processing the target feature information through a target video recommendation model to obtain a second recommendation value set corresponding to an initial video among the plurality of initial videos, where the target video recommendation model is trained by using the video recommendation model training method of any one of the above; determining a target video to be recommended from the plurality of initial videos based on the second recommendation value set.

[0199] Optionally, Figure 13 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 13As shown, the electronic device 130 may include: one or more ( Figure 13 only one is shown in the figure) processors 1302 and a memory 1304. The electronic device 130 may further include a storage controller for controlling and managing the memory 1304; the electronic device 130 may further include a peripheral interface for connecting a radio frequency module, an audio module, a display screen, etc.

[0200] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the training method of the video recommendation model or the video recommendation method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned training method of the video recommendation model or the video recommendation method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely set relative to the processor, and these remote memories can be connected to the electronic device 130 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0201] The processor can call the information and application programs stored in the memory through a transmission device to execute the following steps: obtaining a training sample set, where the training sample set at least includes: a plurality of sample videos, a feature set corresponding to the sample videos in the plurality of sample videos, and a first recommendation value set corresponding to the sample videos, and the first recommendation value set at least includes: a first recommendation value, a second recommendation value, and a third recommendation value. The first recommendation value is obtained based on the leading duration of the sample video, the second recommendation value is obtained based on the viewing duration of a first video that satisfies a first association relationship with the sample video, and the third recommendation value is obtained based on the viewing duration of a second video that satisfies a second association relationship with the sample video; processing the training sample set through an initial video recommendation model to obtain a predicted recommendation value set; training the initial video recommendation model according to the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model, where the video to be recommended is determined based on the video recommendation value output by the target video recommendation model.

[0202] Optionally, the above processor may further execute the program code of the following steps: the first recommendation value set is obtained by the following method: obtaining the leading duration corresponding to the sample video in the plurality of sample videos, and obtaining the first recommendation value according to the leading duration; determining a first video that satisfies a first association relationship with the sample video in the plurality of sample videos, and calculating according to the viewing duration of the first video to obtain the second recommendation value; determining a second video that satisfies a second association relationship with the sample video in the plurality of sample videos, and calculating according to the viewing duration of the second video to obtain the third recommendation value.

[0203] Optionally, the above-mentioned processor may also execute the program code of the following steps: Obtaining the guidance duration corresponding to the sample video in multiple sample videos includes: obtaining a third video that the target object watches within the first time period after watching the target sample video, where the target sample video is any one of the multiple sample videos; obtaining the viewing duration corresponding to the third video; and calculating based on the viewing duration to obtain the guidance duration.

[0204] Optionally, the above-mentioned processor may also execute the program code of the following steps: Obtaining the first recommended value based on the guidance duration includes: obtaining the position information corresponding to the sample video in multiple sample videos, and performing grouping processing on the multiple sample videos based on the position information to obtain multiple groups of sample videos; for any group of sample videos, calculating based on the guidance duration corresponding to the sample video in the group of sample videos to obtain the equal-frequency quantile corresponding to the sample video in the group of sample videos; and obtaining the first recommended value based on the equal-frequency quantile and the guidance duration.

[0205] Optionally, the above-mentioned processor may also execute the program code of the following steps: Obtaining the first recommended value based on the equal-frequency quantile and the guidance duration includes: determining the target guidance duration based on the guidance duration corresponding to the sample video in the group of sample videos; performing bucketing processing on the sample videos in the group of sample videos according to the equal-frequency quantile to obtain the bucket value corresponding to the equal-frequency quantile; and calculating based on the bucket value and the target guidance duration to obtain the first recommended value.

[0206] Optionally, the above-mentioned processor may also execute the program code of the following steps: Determining the first video that satisfies the first association relationship with the sample video in multiple sample videos includes: obtaining multiple first candidate videos that the target object watches within the second time period after watching the target sample video, where the target sample video is any one of the multiple sample videos; determining whether the first candidate video in the multiple first candidate videos satisfies the first association relationship to obtain the first target judgment result; and determining the first video from the multiple first candidate videos based on the first target judgment result.

[0207] Optionally, the above-mentioned processor may also execute the program code of the following steps: determining whether a first candidate video among multiple first candidate videos satisfies a first association relationship, and obtaining a first target judgment result, including: determining whether the content of the first candidate video among multiple first candidate videos is similar to that of a target sample video, and obtaining a first judgment result; determining whether the access behavior of the first candidate video among multiple first candidate videos is similar to that of the target sample video, and obtaining a second judgment result; determining whether the order of the first candidate video among multiple first candidate videos is the target order, and obtaining a third judgment result; and determining whether the first candidate video among multiple first candidate videos satisfies the first association relationship according to the first judgment result, the second judgment result, and the third judgment result, and obtaining a first target judgment result.

[0208] Optionally, the above-mentioned processor may also execute the program code of the following steps: determining a second video that satisfies a second association relationship with a sample video among multiple sample videos, including: obtaining multiple second candidate videos watched by a target object within a second time period after watching the target sample video and multiple third candidate videos watched by the target object within a third time period after watching the target sample video; determining whether a second candidate video among multiple second candidate videos satisfies the second association relationship, and obtaining a second target judgment result, and determining whether a third candidate video among multiple third candidate videos satisfies the second association relationship, and obtaining a third target judgment result; and determining the second video from multiple second candidate videos and multiple third candidate videos according to the second target judgment result and the third target judgment result.

[0209] Optionally, the above-mentioned processor may also execute the program code of the following steps: determining whether a second candidate video among multiple second candidate videos satisfies the second association relationship, and obtaining a second target judgment result, including: determining whether the publishing object corresponding to the second candidate video among multiple second candidate videos is the same as that of the target sample video; in the case where the publishing object corresponding to the second candidate video among multiple second candidate videos is the same as that of the target sample video, determining that the second target judgment result is that the second candidate video satisfies the second association relationship; and in the case where the publishing object corresponding to the second candidate video among multiple second candidate videos is not the same as that of the target sample video, determining that the second target judgment result is that the second candidate video does not satisfy the second association relationship.

[0210] Optionally, the above-mentioned processor may also execute the program code of the following steps: training the initial video recommendation model based on the predicted recommendation value set and the first recommendation value set to obtain the target video recommendation model, including: calculating the loss based on the first predicted recommendation value and the first recommendation value in the predicted recommendation value set to obtain the first loss function; calculating the loss based on the second predicted recommendation value and the second recommendation value in the predicted recommendation value set to obtain the second loss function; calculating the loss based on the third predicted recommendation value and the third recommendation value in the predicted recommendation value set to obtain the third loss function; training the initial video recommendation model based on the first loss function, the second loss function, and the third loss function to obtain the target video recommendation model.

[0211] Optionally, the above-mentioned processor may also execute the program code of the following steps: calculating the loss based on the second predicted recommendation value and the second recommendation value in the predicted recommendation value set to obtain the second loss function, including: calculating the mean square error between the second predicted recommendation value and the second recommendation value to obtain the first loss sub-function; calculating the loss based on the exponential distribution family for the second predicted recommendation value and the second recommendation value to obtain the second loss sub-function; obtaining the second loss function based on the first loss sub-function and the second loss sub-function.

[0212] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtaining the target feature information corresponding to multiple initial videos; processing the target feature information through the target video recommendation model to obtain the second recommendation value set corresponding to the initial video among the multiple initial videos, where the target video recommendation model is trained by using the training method of the video recommendation model in any of the above items; determining the target video to be recommended from the multiple initial videos based on the second recommendation value set.

[0213] Those of ordinary skill in the art can understand that Figure 13 the structure shown is only schematic, and the electronic device 130 may also be a smart phone (such as an Android phone, an iOS phone), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 13 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device 130 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown Figure 13 in the figure, or have a different configuration from that shown Figure 13 in the figure.

[0214] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and this program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.

[0215] Embodiment 7

[0216] The embodiment of the present application also provides a computer program product. Optionally, in this embodiment, the above computer program product can be used to store the program code executed by the training method or the video recommendation method of the video recommendation model provided in the first embodiment above.

[0217] Optionally, in this embodiment, the above computer program product can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0218] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0219] In the above embodiments of the present application, the descriptions of the various embodiments each have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0220] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0221] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0222] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0223] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0224] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A training method for a video recommendation model, characterized in that, Including: Obtain a training sample set, where the training sample set at least includes: a plurality of sample videos, a feature set corresponding to the sample videos in the plurality of sample videos, and a first recommendation value set corresponding to the sample videos. The first recommendation value set at least includes: a first recommendation value, a second recommendation value, and a third recommendation value. The first recommendation value is obtained based on the guiding duration of the sample video, the second recommendation value is obtained based on the viewing duration of a first video that satisfies a first association relationship with the sample video, and the third recommendation value is obtained based on the viewing duration of a second video that satisfies a second association relationship with the sample video; Process the training sample set through an initial video recommendation model to obtain a predicted recommendation value set; Train the initial video recommendation model according to the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model, where the video to be recommended is determined based on the video recommendation value output by the target video recommendation model.

2. The method according to claim 1, characterized in that, The first recommendation value set is obtained through the following method: Obtain the guiding duration corresponding to the sample videos in the plurality of sample videos, and obtain the first recommendation value based on the guiding duration; Determine a first video that satisfies a first association relationship with the sample videos in the plurality of sample videos, and calculate based on the viewing duration of the first video to obtain the second recommendation value; Determine a second video that satisfies a second association relationship with the sample videos in the plurality of sample videos, and calculate based on the viewing duration of the second video to obtain the third recommendation value.

3. The method according to claim 2, wherein Obtaining the guiding duration corresponding to the sample videos in the plurality of sample videos includes: Obtain a third video that the target object watches after watching the target sample video and within a first time period, where the target sample video is any one of the plurality of sample videos; Obtain the viewing duration corresponding to the third video; Calculate based on the viewing duration to obtain the guiding duration.

4. The method according to claim 2, wherein Obtaining the first recommendation value based on the guiding duration includes: Obtain the position information corresponding to the sample videos in the plurality of sample videos, and perform grouping processing on the plurality of sample videos based on the position information to obtain multiple groups of sample videos; For any group of sample videos, calculate based on the guiding duration corresponding to the sample videos in the group to obtain the equal-frequency quantile corresponding to the sample videos in the group; Obtain the first recommendation value according to the equal-frequency quantile and the guiding duration.

5. The method according to claim 4, wherein Obtaining the first recommendation value according to the equal-frequency quantile and the guiding duration includes: Determine the target guiding duration according to the guiding duration corresponding to the sample videos in the group; Perform bucketing processing on the sample videos in the group according to the equal-frequency quantile to obtain the bucket value corresponding to the equal-frequency quantile; Calculate based on the bucket value and the target guiding duration to obtain the first recommendation value.

6. The method according to claim 2, wherein Determining a first video that satisfies a first association relationship with the sample videos in the plurality of sample videos includes: Obtain multiple first candidate videos that the target object watches after watching the target sample video and within the second time period, where the target sample video is any one of the multiple sample videos; Determine whether a first candidate video among the multiple first candidate videos satisfies the first association relationship to obtain a first target judgment result; Determine the first video from the multiple first candidate videos according to the first target judgment result.

7. The method according to claim 6, characterized in that, Determining whether a first candidate video among the multiple first candidate videos satisfies the first association relationship to obtain a first target judgment result includes: Determine whether the content of a first candidate video among the multiple first candidate videos is similar to the target sample video to obtain a first judgment result; Determine whether the access behavior of a first candidate video among the multiple first candidate videos is similar to the target sample video to obtain a second judgment result; Determine whether the sequence position of a first candidate video among the multiple first candidate videos is the target sequence position to obtain a third judgment result; According to the first judgment result, the second judgment result, and the third judgment result, determine whether a first candidate video among the multiple first candidate videos satisfies the first association relationship to obtain the first target judgment result.

8. The method according to claim 2, characterized in that Determining the second video that satisfies the second association relationship with the sample video among the multiple sample videos includes: Obtain multiple second candidate videos that the target object watches after watching the target sample video and within the second time period, and multiple third candidate videos that the target object watches after watching the target sample video and within the third time period; Determine whether a second candidate video among the multiple second candidate videos satisfies the second association relationship to obtain a second target judgment result, and determine whether a third candidate video among the multiple third candidate videos satisfies the second association relationship to obtain a third target judgment result; Determine the second video from the multiple second candidate videos and the multiple third candidate videos according to the second target judgment result and the third target judgment result.

9. The method according to claim 8, wherein Determining whether a second candidate video among the multiple second candidate videos satisfies the second association relationship to obtain a second target judgment result includes: Determine whether the publishing object corresponding to the second candidate video among the multiple second candidate videos is the same as the target sample video; When the publishing object corresponding to the second candidate video among the multiple second candidate videos is the same as the target sample video, determine that the second target judgment result is that the second candidate video satisfies the second association relationship; When the publishing object corresponding to the second candidate video among the multiple second candidate videos is different from the target sample video, determine that the second target judgment result is that the second candidate video does not satisfy the second association relationship.

10. The method according to claim 1, wherein Training the initial video recommendation model according to the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model includes: Calculate a loss according to the first predicted recommendation value in the predicted recommendation value set and the first recommendation value to obtain a first loss function; Calculate a second loss function based on the second predicted recommendation value and the second recommendation value in the predicted recommendation value set; Calculate a third loss function based on the third predicted recommendation value and the third recommendation value in the predicted recommendation value set; Train the initial video recommendation model based on the first loss function, the second loss function, and the third loss function to obtain the target video recommendation model.

11. The method according to claim 10, wherein Calculating a second loss function based on the second predicted recommendation value and the second recommendation value in the predicted recommendation value set includes: Calculate a first loss sub-function by calculating the mean square error between the second predicted recommendation value and the second recommendation value; Calculate a second loss sub-function based on the exponential distribution family for the second predicted recommendation value and the second recommendation value; Obtain the second loss function based on the first loss sub-function and the second loss sub-function.

12. A video recommendation method, characterized in that, including: Obtain the target feature information corresponding to multiple initial videos; Process the target feature information through a target video recommendation model to obtain a second recommendation value set corresponding to an initial video in the multiple initial videos, where the target video recommendation model is trained by using the training method of the video recommendation model according to any one of claims 1 to 11; Determine a target video to be recommended from the multiple initial videos according to the second recommendation value set.

13. A video recommendation method, characterized in that, including: Obtain the target feature information corresponding to multiple initial videos uploaded by the client; Process the target feature information through a target video recommendation model in a cloud server to obtain a second recommendation value set corresponding to an initial video in the multiple initial videos, where the target video recommendation model is trained by using the training method of the video recommendation model according to any one of claims 1 to 11; determine a target video to be recommended from the multiple initial videos according to the second recommendation value set; Return the target video to be recommended to the client.

14. A training device for a video recommendation model, characterized in that, including: A first acquisition unit for acquiring a training sample set, where the training sample set at least includes: multiple sample videos, a feature set corresponding to a sample video in the multiple sample videos, and a first recommendation value set corresponding to the sample video, and the first recommendation value set at least includes: a first recommendation value, a second recommendation value, and a third recommendation value, the first recommendation value is obtained based on the guiding duration of the sample video, the second recommendation value is obtained based on the viewing duration of a first video that satisfies a first association relationship with the sample video, and the third recommendation value is obtained based on the viewing duration of a second video that satisfies a second association relationship with the sample video; A first processing unit for processing the training sample set through an initial video recommendation model to obtain a predicted recommendation value set; A training unit for training the initial video recommendation model based on the predicted recommendation value set and the first recommendation value set to obtain a target video recommendation model, where a video to be recommended is determined based on a video recommendation value output by the target video recommendation model.

15. A video recommendation device, characterized in that, including: A third acquisition unit, configured to acquire target feature information corresponding to a plurality of initial videos; A first processing unit, configured to process the target feature information through a target video recommendation model to obtain a second recommendation value set corresponding to an initial video in the plurality of initial videos, where the target video recommendation model is trained by using the training method of the video recommendation model according to any one of claims 1 to 11; A third determination unit, configured to determine a target video to be recommended from the plurality of initial videos according to the second recommendation value set.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where when the program runs, it controls the device where the storage medium is located to execute the training method of the video recommendation model according to any one of claims 1 to 11, or the video recommendation method according to any one of claims 12 to 13.

17. An electronic device, characterized in that, Comprising: A memory, storing an executable program; A processor, configured to run the program, where when the program runs, it executes the training method of the video recommendation model according to any one of claims 1 to 11, or the video recommendation method according to any one of claims 12 to 13.

18. A computer program product, characterized in that, Comprising a computer program or instruction, where when the computer program or instruction is executed by a processor, it implements the training method of the video recommendation model according to any one of claims 1 to 11, or the video recommendation method according to any one of claims 12 to 13.