Model training method and device, video recommendation method and device, equipment and storage medium
By using a tree structure to model the playback time prediction model, combining the extraction and learning of user features and video features, the problem of inaccurate existing models is solved, and more accurate prediction of video playback time is achieved.
Patent Information
- Application Number
- CN202510351640.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-13
AI Technical Summary
The existing video playback time prediction model is not accurate enough to effectively meet users' fragmented entertainment needs.
The video playback time is modeled using a tree structure. By obtaining user features and video features in the training sample, historical sequence feature extraction and feature importance learning, input tree playback time full-connection network for progressive playback time learning, and adjust the parameters of the playback time prediction model.
Through the tree structure, the order relationship and size dependence of playback time are fully considered, and the prediction ability of the playback time prediction model is improved, making the video playback time prediction more accurate.
Smart Images

Figure CN120151601A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to technical fields such as deep learning, data processing, and recommendation algorithms in the field of artificial intelligence, and particularly relates to a model training method, a video recommendation method, an apparatus, a device, and a storage medium. Background Art
[0002] With the development of the mobile Internet and the popularization of smart phones, videos, especially short videos, have gradually become one of the important ways to obtain entertainment and relaxation, and can better meet the fragmented entertainment needs.
[0003] Currently, when making video (such as short video) recommendations, the video playback duration of candidate videos is usually predicted by a pre-trained playback duration prediction model, and then, based on the video playback duration, target candidate videos to be recommended to users are determined from the candidate videos. Among them, the playback duration prediction model is obtained by modeling the playback duration percentile in a manner of multi-scale playback duration and resource duration bucketing.
[0004] However, the video playback duration predicted by the above playback duration prediction model is not accurate enough. Summary of the Invention
[0005] The present disclosure provides a model training method, a video recommendation method, an apparatus, a device, and a storage medium for more accurately predicting the video playback duration.
[0006] According to a first aspect of the present disclosure, there is provided a model training method for training a playback duration prediction model, where the playback duration prediction model includes a tree-shaped playback duration fully connected network, and the model training method includes:
[0007] Obtain training samples, where the training samples include user features of a sample user, video features of the historical videos watched by the sample user, and the reference playback duration corresponding to the historical videos watched;
[0008] Extract historical sequence features based on the user features and the video features to obtain a first feature vector;
[0009] Perform feature importance learning based on the first feature vector to obtain a second feature vector;
[0010] Input the second feature vector into the tree-shaped playback duration fully connected network for progressive playback duration learning to obtain the predicted playback duration corresponding to the historical videos watched, where the tree-shaped playback duration fully connected network is obtained by modeling the playback duration using a tree structure;
[0011] Adjust the parameters of the playback duration prediction model based on the predicted playback duration and the reference playback duration.
[0012] According to a second aspect of the present disclosure, there is provided a video recommendation method, including:
[0013] Obtaining user characteristics of a target user, video characteristics of historical videos viewed by the target user, and candidate video characteristics respectively corresponding to a plurality of candidate videos;
[0014] For each candidate video among the plurality of candidate videos, according to the user characteristics, video characteristics, and candidate video characteristics, predicting the playing duration of the candidate video through a playing duration prediction model to obtain a target playing duration corresponding to the candidate video, where the playing duration prediction model is obtained by using the model training method described in the first aspect of the present disclosure;
[0015] Based on the target playing duration, determining a target candidate video recommended for the target user.
[0016] According to a third aspect of the present disclosure, there is provided a model training apparatus for training a playing duration prediction model, where the playing duration prediction model includes a tree-shaped playing duration fully-connected network, and the model training apparatus includes:
[0017] An obtaining unit for obtaining training samples, where the training samples include user characteristics of a sample user, video characteristics of historical videos viewed by the sample user, and a reference playing duration corresponding to the historical video;
[0018] A first feature extraction unit for performing historical sequence feature extraction according to the user characteristics and video characteristics to obtain a first feature vector;
[0019] A second feature extraction unit for performing feature importance learning based on the first feature vector to obtain a second feature vector;
[0020] A prediction unit for inputting the second feature vector into the tree-shaped playing duration fully-connected network for progressive playing duration learning to obtain a predicted playing duration corresponding to the historical video, where the tree-shaped playing duration fully-connected network is obtained by modeling the playing duration using a tree structure;
[0021] An adjustment unit for adjusting parameters of the playing duration prediction model based on the predicted playing duration and the reference playing duration.
[0022] According to a fourth aspect of the present disclosure, there is provided a video recommendation apparatus, including:
[0023] An obtaining unit for obtaining user characteristics of a target user, video characteristics of historical videos viewed by the target user, and candidate video characteristics respectively corresponding to a plurality of candidate videos;
[0024] A prediction unit, configured to, for each of a plurality of candidate videos, predict the playback duration of the candidate video through a playback duration prediction model according to user characteristics, video characteristics, and candidate video characteristics, so as to obtain a target playback duration corresponding to the candidate video, where the playback duration prediction model is obtained by using the model training method described in the first aspect of the present disclosure;
[0025] A determination unit, configured to determine a target candidate video recommended for a target user based on the target playback duration.
[0026] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the model training method described in the first aspect or execute the video recommendation method described in the second aspect.
[0027] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the model training method described in the first aspect or execute the video recommendation method described in the second aspect.
[0028] According to a seventh aspect of the present disclosure, there is provided a computer program product, including: a computer program, the computer program is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and when the at least one processor executes the computer program, the electronic device is caused to execute the model training method described in the first aspect or execute the video recommendation method described in the second aspect.
[0029] The technology according to the present disclosure solves the problem that the video playback duration predicted by the current playback duration prediction model is not accurate enough. The playback duration prediction model of the present disclosure includes a tree-shaped playback duration fully connected network; by obtaining training samples, the training samples include user characteristics of sample users, video characteristics of historical videos viewed by sample users, and reference playback durations corresponding to the historical videos viewed; extracting historical sequence features according to the user characteristics and video characteristics to obtain a first feature vector; performing feature importance learning based on the first feature vector to obtain a second feature vector; inputting the second feature vector into the tree-shaped playback duration fully connected network for progressive playback duration learning to obtain a predicted playback duration corresponding to the historical video viewed. The tree-shaped playback duration fully connected network is obtained by modeling the playback duration using a tree structure. Through the tree structure, the order relationship of the predicted values of the playback duration and the dependence of the playback duration magnitude can be fully considered, adding prior knowledge of the order relationship of the predicted playback duration values to the playback duration prediction model, which can effectively improve the prediction ability of the playback duration prediction model; furthermore, based on the predicted playback duration and the reference playback duration, the parameters of the playback duration prediction model can be adjusted to obtain a trained playback duration prediction model. The obtained playback duration prediction model can more accurately predict the video playback duration and has good versatility.
[0030] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0032] Figure 1 is a schematic structural diagram of a multi-objective model for video playback duration prediction provided by the related art;
[0033] Figure 2 is a schematic diagram of an application scenario applicable to the model training method of the present disclosure;
[0034] Figure 3 is a schematic diagram according to the first embodiment of the present disclosure;
[0035] Figure 4 is a schematic diagram according to the second embodiment of the present disclosure;
[0036] Figure 5 is a schematic diagram according to the third embodiment of the present disclosure;
[0037] Figure 6 is a schematic diagram according to the fourth embodiment of the present disclosure;
[0038] Figure 7is a schematic diagram according to the fifth embodiment of the present disclosure;
[0039] Figure 8 is a schematic diagram according to the sixth embodiment of the present disclosure;
[0040] Figure 9 is a schematic diagram according to the seventh embodiment of the present disclosure;
[0041] Figure 10 is a schematic diagram according to the eighth embodiment of the present disclosure;
[0042] Figure 11 is a schematic block diagram of an exemplary electronic device 1100 that can be used to implement the embodiments of the present disclosure. Detailed implementation manners
[0043] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0044] The present disclosure provides a model training method, a video recommendation method, an apparatus, a device, and a storage medium, which are applied to technical fields such as deep learning, data processing, and recommendation algorithms in the field of artificial intelligence, so as to achieve the purpose of more accurately predicting the video playback duration.
[0045] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0046] Currently, when recommending videos (such as short videos), the video playback duration of candidate videos is usually predicted by a pre-trained playback duration prediction model, and then, based on the video playback duration, the target video to be recommended to the user is determined from the candidate videos. Among them, the playback duration prediction model is obtained by modeling the playback duration percentile in a manner of multi-scale playback duration and resource duration bucketing. However, the video playback duration predicted by the above playback duration prediction model is not accurate enough.
[0047] Exemplarily, Figure 1 is a schematic structural diagram of a multi-objective model provided by the related art for predicting the short video playback duration, as Figure 1 shown. For the playback duration-related objectives, the model models the playback duration in a manner of multi-scale playback duration and resource duration bucketing. The specific playback duration-related objectives include:
[0048] (1) 10 - second play duration: Samples with a play duration greater than 10 s are positive samples, and samples with a play duration less than or equal to 10 s are negative samples;
[0049] (2) 20 - second play duration: Samples with a play duration greater than 20 s are positive samples, and samples with a play duration less than or equal to 20 s are negative samples;
[0050] (3) 30 - second play duration: Samples with a play duration greater than 30 s are positive samples, and samples with a play duration less than or equal to 30 s are negative samples;
[0051] (4) 50 - second play duration: Samples with a play duration greater than 50 s are positive samples, and samples with a play duration less than or equal to 50 s are negative samples;
[0052] (5) Resource - duration - binned play duration: First, group the distributed videos according to the resource duration, then perform equal - frequency aggregation on adjacent resource durations according to the distribution ratio. After that, use the unbiased within - bucket thousand - percentile as the label for predicting the play duration.
[0053] The model can also include targets such as completion rate, satisfaction, full - play, and fast - skip. Among them, for the completion - rate target, it can be obtained by dividing the user's viewing duration of the video by the resource duration of the video; for the satisfaction target, after binning according to the resource duration, sort the completion rates in each bin from smallest to largest. Samples with a completion rate ranked, for example, in the top 20% are negative samples, and samples with a completion rate ranked, for example, in the bottom 20% are positive samples; for the full - play target, samples with a completion rate of about 95% can be used as positive samples; for the fast - skip target, samples with a play duration less than the play - duration threshold (such as 3 s) can be used as positive samples, and samples with a play duration greater than or equal to the play - duration threshold are used as negative samples. Each target of the model corresponds to an independent sub - network, sharing the expert network to utilize information or experts at different scales to process data and sharing the underlying network architecture.
[0054] Reference Figure 1The model shown can finely characterize users' consumption habits by modeling multi-scale play durations, and fit the play durations through the resource duration bucketing method to alleviate the problem of the model being biased by long videos. However, both the above methods of modeling play durations by enumeration or debiasing do not consider the sequential relationship between the predicted values of play durations. For example, watching a video to completion can only occur on the premise of watching half of the video. In the video recommendation scenario, the accuracy of play duration prediction is crucial. The modeling of multi-scale play durations models the play duration-related objectives by enumeration, but this method is coarse-grained and exhaustive; although the play duration modeling of resource duration bucketing can eliminate the influence of video duration on video exposure, it does not consider the sequential relationship between the predicted values of play durations.
[0055] To solve the above problems, the present disclosure uses a tree structure to model the video play duration to obtain a play duration prediction model. The play duration prediction model aims at the progressive play duration, and considers the sequential relationship of the target predicted values of the video play duration and the dependence of the video play duration size through the tree structure to perform progressive play duration learning, which can more accurately predict the video play duration and improve the prediction ability of the play duration prediction model.
[0056] Figure 2 FIG. 7 is a schematic diagram of an application scenario applicable to the model training method of the present disclosure. The application scenario may include: a server cluster 21 and a terminal 22. Among them, the server cluster 21 includes multiple servers 211 and a memory 212. The terminal 22 may be a tablet computer, a laptop computer, a desktop computer, a smart home appliance, etc. The server 211 is used to obtain training samples based on the interaction data (such as user logs) obtained by the user watching videos through the terminal 22, and train the play duration prediction model with the training samples. During the training process, data is obtained from the memory 212 and the generated data is stored in the memory 212. In addition, during the training process, communication is performed with the terminal 22 through a wireless network or a wired network.
[0057] In addition, the embodiments of the present disclosure can be applied in the video recommendation scenario. For example, for multiple candidate videos, the play duration prediction model predicts the play duration of each candidate video to obtain the target play duration corresponding to the candidate video, and then based on the target play duration, determines the target candidate video recommended for the target user from the multiple candidate videos.
[0058] Next, specific embodiments will be used to describe in detail the technical solutions of the present disclosure and how the technical solutions of the present disclosure solve the above technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0059] Figure 3 is a schematic diagram according to the first embodiment of the present disclosure. As Figure 3 shown, the model training method provided by the first embodiment of the present disclosure is used to train a playback duration prediction model, and the playback duration prediction model includes a tree-shaped fully connected network for playback duration. The model training method provided by the first embodiment of the present disclosure includes:
[0060] S301. Obtain training samples, where the training samples include user characteristics of the sample user, video characteristics of the historical videos watched by the sample user, and the reference playback duration corresponding to the historical videos watched.
[0061] In the embodiments of the present disclosure, exemplarily, relevant feedback on the videos watched by the user can be collected through the behavior collector of the server, and combined with the historical consumption situation of the user for splicing to obtain the consumption log of the user, that is, the user log; furthermore, feature extraction required for the playback duration prediction model can be performed based on the user log to obtain training samples. The training samples include user characteristics of the sample user, video characteristics of the historical videos watched by the sample user, and the reference playback duration corresponding to the historical videos watched. The reference playback duration corresponding to the historical videos watched can be understood as a label, which is the actual playback duration of the videos watched by the sample user. Among them, the user characteristics of the sample user can include, for example, viewing behavior characteristics such as the time, progress, viewing frequency, and viewing duration (i.e., the playback duration of the video) of the sample user watching each video, as well as interaction behavior characteristics such as the like, attention, sharing, and comment of the sample user on each video. The video characteristics of the historical videos watched by the sample user can include, for example, video metadata characteristics such as the title, description, tags, classification, release time, and creator of the historical videos watched, as well as video quality characteristics such as resolution, bit rate, and frame rate.
[0062] S302. Extract historical sequence features according to the user characteristics and video characteristics to obtain a first feature vector.
[0063] In this step, for each training sample, historical sequence features can be extracted according to the user characteristics and video characteristics to obtain a first feature vector in the time dimension. For how to specifically extract historical sequence features according to the user characteristics and video characteristics to obtain a first feature vector, reference can be made to the subsequent embodiments.
[0064] S303. Perform feature importance learning based on the first feature vector to obtain a second feature vector.
[0065] In this step, after obtaining the first feature vector, feature importance learning can be performed based on the first feature vector to determine the weight or importance of each feature in the first feature vector, so as to perform weighted fusion processing on the first feature vector to obtain the second feature vector. For the specific method of performing feature importance learning based on the first feature vector to obtain the second feature vector, reference can be made to the subsequent embodiments.
[0066] S304. Input the second feature vector into a tree-shaped play duration fully connected network for progressive play duration learning to obtain the predicted play duration corresponding to the historical watched video. The tree-shaped play duration fully connected network is obtained by modeling the play duration using a tree structure.
[0067] It can be understood that the tree-shaped play duration fully connected network is obtained by modeling the play duration using a tree structure. With the progressive play duration as the goal, the order relationship of the predicted values of the video play duration and the dependence of the video play duration size are considered through the tree structure. For example, for a video with a sample user viewing duration of 100s, it is necessary to first view from 0s to 50s. Between 50s and 100s, it is necessary to first view 75s. Between 75s and 100s, it is necessary to first view 95s, and finally view 100s. Among them, different time ranges can be understood as different levels of the tree structure. The tree-shaped play duration fully connected network learns the above progressive relationship of the play duration through the tree structure, so as to more accurately predict the video play duration. For the specific tree structure of the tree-shaped play duration fully connected network, reference can be made to the subsequent embodiments.
[0068] In this step, after obtaining the second feature vector, the second feature vector can be input into a tree-shaped play duration fully connected network for progressive play duration learning to obtain the predicted play duration corresponding to the historical watched video. For the specific method of obtaining the predicted play duration corresponding to the historical watched video through the tree-shaped play duration fully connected network, reference can be made to the subsequent embodiments.
[0069] S305. Adjust the parameters of the play duration prediction model based on the predicted play duration and the reference play duration.
[0070] Exemplarily, after obtaining the predicted play duration corresponding to the historical watched video, the predicted play duration and the reference play duration can be compared to determine the loss value. Among them, the loss value is, for example, the logloss value. Adjust the parameters of the play duration prediction model according to the loss value until the loss value meets the convergence condition, such as the loss value converges to a ten-thousandth difference, so as to obtain the trained play duration prediction model.
[0071] In the embodiments of the present disclosure, the playback duration prediction model includes a tree-shaped full connection network for playback duration; by obtaining training samples, the training samples include user characteristics of sample users, video characteristics of the historical videos watched by the sample users, and the reference playback duration corresponding to the historical videos watched; extracting historical sequence features based on the user characteristics and video characteristics to obtain a first feature vector; performing feature importance learning based on the first feature vector to obtain a second feature vector; inputting the second feature vector into the tree-shaped full connection network for playback duration for progressive playback duration learning to obtain the predicted playback duration corresponding to the historical videos watched. The tree-shaped full connection network for playback duration is obtained by modeling the playback duration using a tree structure. Through the tree structure, the order relationship of the predicted values of the playback duration and the dependence of the playback duration magnitude can be fully considered, adding prior knowledge of the order relationship of the predicted playback duration values to the playback duration prediction model, which can effectively improve the prediction ability of the playback duration prediction model; furthermore, based on the predicted playback duration and the reference playback duration, the parameters of the playback duration prediction model can be adjusted to obtain a trained playback duration prediction model. The obtained playback duration prediction model can more accurately predict the video playback duration and has good versatility.
[0072] Figure 4 It is a schematic diagram according to the second embodiment of the present disclosure. On the basis of the above embodiments, the embodiments of the present disclosure further illustrate the model training method. As Figure 4 shown, the model training method provided by the second embodiment of the present disclosure includes:
[0073] S401. Obtain training samples, where the training samples include user characteristics of sample users, video characteristics of the historical videos watched by the sample users, and the reference playback duration corresponding to the historical videos watched.
[0074] Among them, the implementation principle and technical effects of S401 can be referred to the foregoing embodiments and will not be elaborated here.
[0075] Considering that the playback duration prediction model further includes a playback duration expert network, therefore, in the embodiments of the present disclosure, Figure 3 step S302 in can further include the following S402 step:
[0076] S402. Input the user characteristics and video characteristics into the playback duration expert network for historical sequence feature extraction to obtain the first feature vector output by the playback duration expert network.
[0077] It can be understood that the expert network (Experts) is usually a feedforward neural network (multi-layer perceptron). The expert network can independently process the input data set, extract features in the time dimension, and generate an output for the data set. In this step, the user features and video features are input into the play duration expert network for historical sequence feature extraction, and the first feature vector output by the play duration expert network can be obtained.
[0078] Considering that the play duration prediction model also includes a play duration gating network, therefore, in the embodiments of the present disclosure, Figure 3 step S303 in can further include the following step S403:
[0079] S403. Input the first feature vector into the play duration gating network for feature importance learning, and obtain the second feature vector output by the play duration gating network.
[0080] It can be understood that the gating network (Gating Network) can dynamically determine the weight or importance of each input feature according to the features of the input data. The gating network is usually composed of one or two layers of feedforward neural networks (multi-layer perceptrons). In this step, the second feature vector is input into the play duration gating network for feature importance learning, and then feature selection is performed, so that the second feature vector output by the play duration gating network can be obtained.
[0081] In the embodiments of the present disclosure, Figure 3 step S304 in can further include the following three steps S404 to S406:
[0082] S404. Input the second feature vector into the tree-shaped play duration fully connected network. The tree-shaped play duration fully connected network determines the target node path and the target leaf node from the root node to the leaf node in the tree structure corresponding to the historical watched video according to the second feature vector. The tree-shaped play duration fully connected network is obtained by modeling the play duration using a tree structure.
[0083] In this step, after obtaining the second feature vector, the second feature vector can be input into the tree-shaped play duration fully connected network for further feature extraction to determine the target node path and the target leaf node from the root node to the leaf node in the tree structure corresponding to the historical watched video.
[0084] Optionally, the tree structure of the tree-shaped play duration fully connected network includes multiple nodes at a set level, and each node in the multiple nodes includes the play duration range of the historical watched video passing through the node; the non-leaf nodes in the multiple nodes are binary classifiers, and the non-leaf nodes perform binary classification processing on the play duration of the historical watched video according to the play duration range of the non-leaf nodes to determine the child nodes of the non-leaf nodes corresponding to the historical watched video; the leaf nodes in the multiple nodes correspond to a set play duration.
[0085] Exemplarily, Figure 5 is a schematic diagram according to the third embodiment of the present disclosure. As Figure 5 shown, in the third embodiment of the present disclosure, the output dimension of the tree-shaped play duration fully connected network is a set dimension of 15 dimensions. By expanding the 15 dimensions, a tree structure including 5 levels (i.e., the set level) can be obtained. The tree structure includes non-leaf nodes (including the root node) and leaf nodes. Among them, each node includes the play duration range (i.e., the upper and lower bounds of the play duration) of the historical watched video (i.e., the sample) passing through the node. For example, taking the leaf node LN8 as an example, the play duration range of the leaf node LN8 is (90, 100), and the unit is seconds (s). Each leaf node corresponds to a set play duration (such as represented by pl i ), indicating the play duration of the leaf node during prediction. The set play duration can be, for example, the average value of the upper and lower bounds of the play duration of the leaf node. For example, Figure 5 the set play duration pl 8 of the leaf node LN8 in i is 95s. Each leaf node corresponds to a probability value (such as represented by prob
[0086] ), indicating the probability that the historical watched video (i.e., the sample) falls into the node. The non-leaf nodes are binary classifiers. For example, node 11 is a non-leaf node, and the play duration range is (80, 100), and the unit is seconds (s). The classification task is to determine whether the play duration of the historical watched video (i.e., the sample) is greater than 90s. If the play duration is greater than 90s, it is determined that the next node (i.e., the child node of node 11) corresponding to the historical watched video is the right child node of node 11; if the play duration is less than or equal to 90s, it is determined that the next node (i.e., the child node of node 11) corresponding to the historical watched video is the left child node of node 11. Figure 5 It can be understood that the play duration of the historical watched video (i.e., the sample) is judged layer by layer starting from the root node until it falls into a certain leaf node. For example,
[0087] S405. Obtain the target probability value corresponding to the target leaf node according to the probability value corresponding to the target node path.
[0088] Exemplarily, referring to Figure 5 , taking the target node path where the historical viewing video (sample) falls from the root node (node 1) to the leaf node LN8 as an example, assuming that the classification logit (used to predict the probability of going right in the next step) of each non-leaf node is represented as p i , then the target probability value corresponding to the target leaf node LN8 can be determined according to this target node path as: (1 - p 1 ) * p 2 * p 5 * p 11 , where p 1 represents the probability value from node 1 to node 2, p 2 represents the probability value from node 2 to node 5, p 5 represents the probability value from node 5 to node 11, p 11 represents the probability value from node 11 to the target leaf node LN8.
[0089] S406. Obtain the predicted playback duration corresponding to the historical viewing video based on the set playback duration and the target probability value corresponding to the target leaf node.
[0090] Exemplarily, referring to Figure 5 , each of the 16 leaf nodes corresponds to a set playback duration pl i and a probability value prob i , so the predicted playback duration (such as represented by q) corresponding to the historical viewing video can be obtained based on the set playback duration and the target probability value corresponding to the target leaf node as the weighted sum of the playback durations of the leaf nodes, that is:
[0091] In the embodiments of the present disclosure, Figure 3 step S305 can further include the following three steps of S407 to S409:
[0092] S407. Obtain the first loss value corresponding to each node in the target node path passed by the historical viewing video; obtain the second loss value according to the predicted playback duration and the reference playback duration.
[0093] Exemplarily, referring to Figure 5, the first loss value corresponding to each node in the target node path passed by the historical watched video can be obtained. The first loss value can also be understood as the binary classification loss of non-leaf nodes. The second loss value can be obtained according to the predicted playback duration and the reference playback duration corresponding to the historical watched video. The second loss value is, for example, the mean squared error (MSE) between the predicted playback duration and the reference playback duration, or the log loss (logloss).
[0094] S408. Perform weighted fusion processing on the first loss value and the second loss value to obtain the target loss value.
[0095] In this step, the target loss value can be obtained by performing weighted fusion processing according to the first loss value, the first weight corresponding to the first loss value, the second loss value, and the second weight corresponding to the second loss value.
[0096] S409. Adjust the parameters of the playback duration prediction model according to the target loss value.
[0097] In this step, after obtaining the target loss value, the parameters of the playback duration prediction model can be adjusted according to the target loss value until the target loss value meets the convergence condition. For example, the target loss value converges to a ten-thousandth difference, so as to obtain the trained playback duration prediction model.
[0098] In the embodiments of the present disclosure, the playing duration prediction model includes a playing duration expert network, a playing duration gating network, and a tree-shaped playing duration fully connected network. By obtaining training samples, where the training samples include the user characteristics of the sample user, the video characteristics of the historical videos watched by the sample user, and the reference playing duration corresponding to the historical videos watched; inputting the user characteristics and the video characteristics into the playing duration expert network for historical sequence feature extraction to obtain a first feature vector output by the playing duration expert network; inputting the first feature vector into the playing duration gating network for feature importance learning to obtain a second feature vector output by the playing duration gating network; inputting the second feature vector into the tree-shaped playing duration fully connected network, and the tree-shaped playing duration fully connected network determines the target node path and the target leaf node from the root node to the leaf node in the tree structure corresponding to the historical video watched according to the second feature vector, and the tree-shaped playing duration fully connected network is obtained by modeling the playing duration using a tree structure; obtaining the target probability value corresponding to the target leaf node according to the probability value corresponding to the target node path, and obtaining the predicted playing duration corresponding to the historical video watched based on the set playing duration and the target probability value corresponding to the target leaf node; the order relationship of the predicted values of the playing duration and the dependence of the size of the playing duration can be fully considered through the tree structure, adding prior knowledge of the order relationship of the predicted values of the playing duration to the playing duration prediction model, and effectively improving the prediction ability of the playing duration prediction model; obtaining the first loss value corresponding to each node in the target node path when the historical video watched passes through; obtaining the second loss value according to the predicted playing duration and the reference playing duration; performing weighted fusion processing on the first loss value and the second loss value to obtain the target loss value, and adjusting the parameters of the playing duration prediction model according to the target loss value, so as to more efficiently obtain the trained playing duration prediction model. The obtained playing duration prediction model can more accurately predict the video playing duration, has good versatility, and can reduce the calculation amount.
[0099] Based on the above embodiments, Figure 6 is a schematic diagram according to the fourth embodiment of the present disclosure. As Figure 6 shown, the playing duration prediction model in the fourth embodiment of the present disclosure includes a playing duration expert network, a playing duration gating network, and a tree-shaped playing duration fully connected network. Refer to Figure 4Example, where the playing duration expert network is used to extract historical sequence features based on the user features of the sample users included in the training samples and the video features of the historical watched videos of the sample users, to obtain a first feature vector; the playing duration gating network is used to perform feature importance learning based on the first feature vector to obtain a second feature vector; the tree-shaped playing duration fully connected network is used to perform progressive playing duration learning based on the second feature vector to obtain the predicted playing duration corresponding to the historical watched video. Furthermore, based on the predicted playing duration and the reference playing duration corresponding to the historical watched video, the parameters of the playing duration prediction model can be adjusted to obtain a trained playing duration prediction model.
[0100] Based on the above embodiments, Figure 7 is a schematic diagram according to the fifth embodiment of the present disclosure. As Figure 7 shown, the video recommendation method provided by the fifth embodiment of the present disclosure includes:
[0101] S701. Obtain the user features of the target user, the video features of the historical watched videos of the target user, and the candidate video features corresponding to multiple candidate videos respectively.
[0102] Exemplarily, the user features of the target user may include viewing behavior features such as the time, progress, viewing frequency, and viewing duration (i.e., the playing duration of the video) of the target user for each video, as well as interaction behavior features such as the like, follow, share, and comment of the target user on each video. The video features of the historical watched videos of the target user may include video metadata features such as the title, description, tags, classification, release time, and creator of the historical watched videos, as well as video quality features such as resolution, bit rate, and frame rate. The candidate video features corresponding to multiple candidate videos may include video metadata features such as the identifier, title, description, tags, classification, release time, and creator of the candidate videos, as well as video quality features such as resolution, bit rate, and frame rate.
[0103] S702. For each candidate video among the multiple candidate videos, predict the playing duration of the candidate video through the playing duration prediction model according to the user features, video features, and candidate video features, to obtain the target playing duration corresponding to the candidate video, where the playing duration prediction model is obtained by using the model training method in any of the above method embodiments.
[0104] In this step, for each candidate video among the multiple candidate videos, the user features, video features, and candidate video features may be input into the playing duration prediction model, and the playing duration prediction model predicts the playing duration of the candidate video to obtain the target playing duration corresponding to the candidate video.
[0105] S703. Based on the target playing duration, determine the target candidate video recommended for the target user.
[0106] Exemplarily, after obtaining the target play duration corresponding to each candidate video among multiple candidate videos, the target play durations corresponding to the multiple candidate videos can be sorted according to the magnitudes of the target play durations, so as to determine that the target candidate videos recommended for the target user are a preset number of candidate videos with larger target play durations. The target candidate videos can be output to the user to facilitate the user to watch the target candidate videos in a timely manner, and the target candidate videos are videos that the target user may be satisfied with.
[0107] Optionally, the consumption behavior data of the target user for the recommended target candidate videos can also be obtained, and then the play duration prediction model can be continuously trained and optimized based on the consumption behavior data.
[0108] In the embodiments of the present disclosure, by obtaining the user characteristics of the target user, the video characteristics of the videos historically watched by the target user, and the candidate video characteristics corresponding to the multiple candidate videos respectively; for each candidate video among the multiple candidate videos, according to the user characteristics, video characteristics, and candidate video characteristics, the play duration of the candidate video is predicted through the play duration prediction model to obtain the target play duration corresponding to the candidate video; since the play duration prediction model in the embodiments of the present disclosure is trained by using the model training method in any of the above method embodiments, therefore, through the play duration prediction model, the target play duration corresponding to the candidate video can be obtained more accurately, and then based on the target play duration, the target candidate videos recommended for the target user can be determined, realizing more accurate and efficient video recommendation for the target user.
[0109] Based on the above embodiments, Figure 8 is a schematic diagram according to the sixth embodiment of the present disclosure. As Figure 8 shown, in the sixth embodiment of the present disclosure, when performing offline training on the play duration prediction model, the features required for the play duration prediction model can be extracted based on the user logs to obtain training samples, and the play duration prediction model is constructed offline, and the play duration prediction model is trained with the progressive play duration as the target. For example, taking logloss as the target loss, when the loss converges to a ten-thousandth difference, the play duration prediction model is basically convergent, thereby obtaining the trained play duration prediction model. Among them, the user logs can be collected by the behavior collector of the server for the relevant feedback of the videos watched by the user, and spliced in combination with the historical consumption situation of the user to obtain.
[0110] After obtaining the trained playback duration prediction model, when performing online prediction through the playback duration prediction model, in response to a video viewing request of a target user, based on the user information of the target user, the video information of the historical videos viewed by the target user, and the video information of multiple candidate videos, feature extraction is performed to obtain the user features of the target user, the video features of the historical videos viewed by the target user, and the candidate video features corresponding to the multiple candidate videos respectively. For each candidate video among the multiple candidate videos, the playback duration of the candidate video can be predicted through the playback duration prediction model according to the user features, video features, and candidate video features to obtain the target playback duration (i.e., the progressive playback duration) corresponding to the candidate video; thus, the target playback durations can be sorted, and the target candidate videos recommended for the target user are determined to be a preset number of candidate videos with larger target playback durations, and the target candidate videos are sent to the target user. After the target user views the recommended target candidate videos, feedback data of the target user on the recommended target candidate videos can also be obtained. The feedback data may include, for example, the implicit feedback data and explicit feedback data of the target user. The implicit feedback data may include, for example, data such as the playback duration and playback completion rate of the target user, and the explicit feedback data may include data such as the likes and follows of the target user for the video. The playback duration prediction model can be continuously trained and optimized based on the feedback data.
[0111] Based on the above embodiments, recommending videos to users according to the video playback durations predicted by the playback duration prediction model provided in the embodiments of the present disclosure can significantly improve the new user retention rate of video products, thereby gradually expanding the scale of video products. The solution of the embodiments of the present disclosure has good generality and can be directly migrated to other user products for application.
[0112] The following is an embodiment of the apparatus of the present disclosure, which can be used to execute the method embodiment of the present disclosure. For details not disclosed in the embodiment of the apparatus of the present disclosure, please refer to the method embodiment of the present disclosure.
[0113] Figure 9 is a schematic diagram according to the seventh embodiment of the present disclosure. As Figure 9As shown in the figure, the model training device 900 provided in the seventh embodiment of the present disclosure is used to train a playback duration prediction model, and the playback duration prediction model includes a tree-shaped full connection network for playback duration. The model training device 900 includes: an acquisition unit 901, configured to acquire training samples, where the training samples include user features of a sample user, video features of historical videos watched by the sample user, and a reference playback duration corresponding to the historical videos watched; a first feature extraction unit 902, configured to perform historical sequence feature extraction according to the user features and the video features to obtain a first feature vector; a second feature extraction unit 903, configured to perform feature importance learning based on the first feature vector to obtain a second feature vector; a prediction unit 904, configured to input the second feature vector into the tree-shaped full connection network for playback duration to perform progressive playback duration learning to obtain a predicted playback duration corresponding to the historical videos watched, and the tree-shaped full connection network for playback duration is obtained by modeling the playback duration using a tree structure; an adjustment unit 905, configured to adjust the parameters of the playback duration prediction model based on the predicted playback duration and the reference playback duration.
[0114] In some embodiments, the tree structure of the tree-shaped full connection network for playback duration includes multiple nodes at a set level, and each node in the multiple nodes includes a playback duration range of the historical videos watched passing through the node; the non-leaf nodes in the multiple nodes are binary classifiers, and the non-leaf nodes perform binary classification processing on the historical videos watched according to the playback duration range of the non-leaf nodes to determine the child nodes of the non-leaf nodes corresponding to the historical videos watched; the leaf nodes in the multiple nodes correspond to a set playback duration.
[0115] In some embodiments, the prediction unit 904 includes: a determination module (not shown in the figure), configured to input the second feature vector into the tree-shaped full connection network for playback duration, and the tree-shaped full connection network for playback duration determines a target node path and a target leaf node from the root node to the leaf node in the tree structure corresponding to the historical videos watched according to the second feature vector; a first acquisition module (not shown in the figure), configured to acquire a target probability value corresponding to the target leaf node according to the probability value corresponding to the target node path; a second acquisition module (not shown in the figure), configured to obtain a predicted playback duration of the historical videos watched based on the set playback duration corresponding to the target leaf node and the target probability value.
[0116] In some embodiments, the playback duration prediction model further includes a playback duration expert network, and the first feature extraction unit 902 includes: a first feature extraction module (not shown in the figure), configured to input the user features and the video features into the playback duration expert network to perform historical sequence feature extraction to obtain a first feature vector output by the playback duration expert network.
[0117] In some embodiments, the playing duration prediction model further includes a playing duration gating network. The second feature extraction unit 903 includes: a second feature extraction module (not shown in the figure), configured to input the first feature vector into the playing duration gating network for feature importance learning, and obtain a second feature vector output by the playing duration gating network.
[0118] In some embodiments, the adjustment unit 905 includes: a third acquisition module (not shown in the figure), configured to acquire a first loss value corresponding to each node in the target node path when the historical watched video passes through the nodes; a fourth acquisition module (not shown in the figure), configured to acquire a second loss value according to the predicted playing duration and the reference playing duration; a fifth acquisition module (not shown in the figure), configured to perform weighted fusion processing on the first loss value and the second loss value to obtain a target loss value; an adjustment module (not shown in the figure), configured to adjust the parameters of the playing duration prediction model according to the target loss value.
[0119] Figure 9 The provided model training device can execute the steps in the corresponding method embodiments of the above model training method, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0120] Figure 10 is a schematic diagram according to the eighth embodiment of the present disclosure. As Figure 10 shown, the video recommendation device 1000 provided by the eighth embodiment of the present disclosure includes: an acquisition unit 1001, configured to acquire user features of a target user, video features of historical watched videos of the target user, and candidate video features corresponding to multiple candidate videos respectively; a prediction unit 1002, configured to, for each candidate video among the multiple candidate videos, predict the playing duration of the candidate video through a playing duration prediction model according to the user features, video features, and candidate video features, and obtain a target playing duration corresponding to the candidate video, where the playing duration prediction model is obtained by using the model training method in any of the above method embodiments; a determination unit 1003, configured to determine a target candidate video recommended for the target user based on the target playing duration.
[0121] Figure 10 The provided video recommendation device can execute the steps in the corresponding method embodiments of the above video recommendation method, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0122] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device. The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the solution provided in any of the above embodiments.
[0123] According to an embodiment of the present disclosure, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the solution provided in any of the above embodiments.
[0124] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product, which includes: a computer program stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to execute the solution provided in any of the above embodiments.
[0125] Figure 11 FIG. is a schematic block diagram of an exemplary electronic device 1100 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0126] As Figure 11 shown, the electronic device 1100 includes a computing unit 1101, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) ( Figure 11 taking ROM 1102 as an example) or a computer program loaded from a storage unit 1108 into a random access memory (RAM) ( Figure 11 taking RAM 1103 as an example). In the RAM 1103, various programs and data required for the operation of the electronic device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. An input / output (I / O) interface ( Figure 11 taking the I / O interface 1105 as an example) is also connected to the bus 1104.
[0127] Multiple components in the electronic device 1100 are connected to the I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disc, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the electronic device 1100 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0128] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include but are not limited to a central processing unit (CPU), a graphic processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as the model training method. For example, in some embodiments, the model training method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the model training method described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the model training method in any other suitable manner (e.g., by means of firmware).
[0129] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on a chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0130] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0131] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0132] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0133] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0134] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0135] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0136] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, fusions, sub-fusions, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A model training method for training a playback time prediction model, wherein the playback time prediction model comprises a tree-shaped playback time fully connected network, and the model training method comprises: Acquire a training sample, wherein the training sample includes user features of a sample user, video features of historically watched videos of the sample user, and reference play durations corresponding to the historically watched videos; Perform historical sequence feature extraction according to the user feature and the video feature to obtain a first feature vector; Performing feature importance learning based on the first feature vector to obtain a second feature vector; Inputting the second feature vector into the tree-shaped playback time fully connected network to perform progressive playback time learning to obtain the predicted playback time corresponding to the historically watched video, wherein the tree-shaped playback time fully connected network is obtained by modeling the playback time using a tree structure; Based on the predicted playback duration and the reference playback duration, parameters of a playback duration prediction model are adjusted.
2. The model training method according to claim 1, wherein: The tree structure of the tree-shaped playback time fully connected network includes multiple nodes of set levels, and each of the multiple nodes contains the playback time range of the historical viewing video passing through the node; the non-leaf nodes among the multiple nodes are binary classifiers, and the non-leaf nodes perform binary classification processing on the playback time of the historical viewing video according to the playback time range of the non-leaf nodes, and determine the child nodes of the non-leaf nodes corresponding to the historical viewing video; the leaf nodes among the multiple nodes correspond to the set playback time.
3. The model training method according to claim 2, wherein: The step of inputting the second feature vector into the tree-shaped playback duration fully connected network to perform progressive playback duration learning to obtain the predicted playback duration corresponding to the historically viewed video includes: Inputting the second feature vector into the tree-shaped playback duration fully connected network, the tree-shaped playback duration fully connected network determining the target node path and the target leaf node from the root node to the leaf node in the tree structure corresponding to the historical viewing video according to the second feature vector; According to the probability value corresponding to the target node path, obtain the target probability value corresponding to the target leaf node; Based on the set playback duration and the target probability value corresponding to the target leaf node, the predicted playback duration corresponding to the historically viewed video is obtained.
4. The model training method according to any one of claims 1 to 3, wherein: The playback time prediction model also includes a playback time expert network, and the historical sequence feature extraction is performed according to the user feature and the video feature to obtain a first feature vector, including: The user features and video features are input into the playback time expert network to extract historical sequence features, and a first feature vector output by the playback time expert network is obtained.
5. The model training method according to any one of claims 1 to 3, wherein: The playback time prediction model further includes a playback time gating network, and the feature importance learning based on the first feature vector to obtain a second feature vector includes: The first feature vector is input into the playback duration gating network to perform feature importance learning, so as to obtain a second feature vector output by the playback duration gating network.
6. The model training method according to claim 3, wherein: The adjusting the parameters of the playback duration prediction model based on the predicted playback duration and the reference playback duration includes: Obtaining a first loss value corresponding to each node in the target node path when the historical viewing video passes through the target node path; Acquire a second loss value according to the predicted playback duration and the reference playback duration; Performing weighted fusion processing on the first loss value and the second loss value to obtain a target loss value; According to the target loss value, the parameters of the playback duration prediction model are adjusted.
7. A video recommendation method, comprising: Acquire user features of a target user, video features of historically watched videos of the target user, and candidate video features corresponding to a plurality of candidate videos; For each candidate video among the multiple candidate videos, predict the playback time of the candidate video by using a playback time prediction model according to the user features, the video features and the candidate video features, so as to obtain a target playback time corresponding to the candidate video, wherein the playback time prediction model is obtained by using the model training method according to any one of claims 1 to 6; Based on the target playback duration, a target candidate video recommended for the target user is determined.
8. A model training device for training a play time prediction model, wherein the play time prediction model comprises a tree-shaped play time fully connected network, and the model training device comprises: An acquisition unit, configured to acquire a training sample, wherein the training sample includes user features of a sample user, video features of a historically watched video of the sample user, and a reference playback duration corresponding to the historically watched video; A first feature extraction unit, configured to extract historical sequence features according to the user features and the video features to obtain a first feature vector; A second feature extraction unit, used for performing feature importance learning based on the first feature vector to obtain a second feature vector; A prediction unit, configured to input the second feature vector into the tree-shaped playback time fully connected network for progressive playback time learning, and obtain a predicted playback time corresponding to the historically watched video, wherein the tree-shaped playback time fully connected network is obtained by modeling the playback time using a tree structure; An adjustment unit is used to adjust parameters of a playback duration prediction model based on the predicted playback duration and the reference playback duration.
9. The model training device according to claim 8, wherein: The tree structure of the tree-shaped playback time fully connected network includes multiple nodes of set levels, and each of the multiple nodes contains the playback time range of the historical viewing video passing through the node; the non-leaf nodes among the multiple nodes are binary classifiers, and the non-leaf nodes perform binary classification processing on the playback time of the historical viewing video according to the playback time range of the non-leaf nodes, and determine the child nodes of the non-leaf nodes corresponding to the historical viewing video; the leaf nodes among the multiple nodes correspond to the set playback time.
10. The model training device according to claim 9, wherein: The prediction unit comprises: A determination module, configured to input the second feature vector into the tree-shaped playback duration fully connected network, wherein the tree-shaped playback duration fully connected network determines, based on the second feature vector, a target node path and a target leaf node from a root node to a leaf node in the tree structure corresponding to the historically viewed video; A first acquisition module, used to acquire a target probability value corresponding to the target leaf node according to the probability value corresponding to the target node path; The second acquisition module is used to obtain the predicted playback time corresponding to the historical viewing video based on the set playback time and the target probability value corresponding to the target leaf node.
11. The model training device according to any one of claims 8 to 10, wherein: The playback duration prediction model further includes a playback duration expert network, and the first feature extraction unit includes: The first feature extraction module is used to input the user features and video features into the play time expert network to perform historical sequence feature extraction, and obtain a first feature vector output by the play time expert network.
12. The model training device according to any one of claims 8 to 10, wherein: The playback duration prediction model further includes a playback duration gating network, and the second feature extraction unit includes: The second feature extraction module is used to input the first feature vector into the playback time gating network to perform feature importance learning, and obtain a second feature vector output by the playback time gating network.
13. The model training device according to claim 10, wherein: The adjustment unit comprises: A third acquisition module is used to obtain a first loss value corresponding to each node in the target node path when the historical viewing video passes through the target node path; A fourth acquisition module, used for acquiring a second loss value according to the predicted playback time and the reference playback time; a fifth acquisition module, configured to perform weighted fusion processing on the first loss value and the second loss value to obtain a target loss value; An adjustment module is used to adjust the parameters of the playback time prediction model according to the target loss value.
14. A video recommendation device, comprising: An acquisition unit, configured to acquire user features of a target user, video features of videos that the target user has watched in history, and candidate video features corresponding to a plurality of candidate videos; A prediction unit, configured to predict the playback time of each candidate video among a plurality of candidate videos according to the user features, the video features and the candidate video features by using a playback time prediction model to obtain a target playback time corresponding to the candidate video, wherein the playback time prediction model is obtained by using the model training method according to any one of claims 1 to 6; A determination unit is used to determine a target candidate video to be recommended to the target user based on the target playback duration.
15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the model training method according to any one of claims 1 to 6 or the video recommendation method according to claim 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the model training method according to any one of claims 1 to 6 or the video recommendation method according to claim 7.
17. A computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the model training method according to any one of claims 1 to 6 or the steps of the video recommendation method according to claim 7.
Citation Information
Cited By
Recommendation method and device, equipment and medium
CN120745824A