Model training method and apparatus, video recommendation method and apparatus, device, medium, and product
By extracting the user's short video behavior sequence and item behavior sequence in the e-commerce scenario, combined with the interest decoupling network, the problem of poor short video recommendation effect in the e-commerce scenario is solved, and higher recommendation accuracy and interpretability are achieved.
Patent Information
- Application Number
- PCT/CN2024/128241
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-08
AI Technical Summary
The recommended short videos in e-commerce scenarios are poor, mainly due to the sparse video behavior of users and the differences in users' interest in items and short videos that have not been effectively considered.
A model training method is proposed. By obtaining the user's short video behavior sequence, item behavior sequence and target short video, the user's video interest vector and item interest vector are extracted using the initial video recommendation model, and the video interest vector is decoupled into content interest and item interest through the interest decoupling network. The total loss function is determined based on these vectors, and the model training is performed to obtain the video recommendation model.
It improves the interpretability of the model and the accuracy of video recommendations, can more effectively identify users' short video interest preferences, and improves the short video recommendation effect in e-commerce scenarios.
Smart Images

Figure CN2024128241_08052025_PF_FP_ABST
Abstract
Description
Model training methods, video recommendation methods, devices, equipment, media, and products
[0001] This disclosure claims priority to Chinese patent application number 202311422536.8, filed on October 30, 2023, entitled “Model training method, video recommendation method, device, apparatus and medium”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present disclosure relates to the field of computer technology, and in particular to a model training method, a video recommendation method, a device, an electronic device, and a computer storage medium. Background Art
[0003] Short videos are an emerging content type for e-commerce platforms, focusing on introducing products through various video presentations, such as user experiences shared by regular users and sitcom commercials produced by suppliers or merchants. Short videos offer the following key advantages: 1) They provide users with high-quality short videos, enabling diverse pan-entertainment experiences and significantly increasing user engagement with e-commerce platforms. 2) Well-produced short videos attract user interest in related products and encourage purchases. 3) Short videos provide merchants with a high-quality channel for brand promotion and online marketing.
[0004] In related technologies, in order to improve the recommendation effect of short videos in e-commerce scenarios, short video recommendation methods for entertainment scenarios are usually applied. However, there are significant differences between short videos in e-commerce scenarios and short videos in entertainment scenarios. If the short video recommendation method for entertainment scenarios is directly applied, the recommendation effect of short videos in e-commerce scenarios will be reduced, and the recommendation accuracy of short videos will be reduced.
[0005] Summary of the Invention
[0006] The embodiments of the present disclosure provide a model training method, a video recommendation method, an apparatus, a device, a medium, and a product.
[0007] The present disclosure provides a model training method, the method comprising:
[0008] Obtain multiple training samples, each training sample includes a user's short video behavior sequence, an item behavior sequence, and one of the exposed target short videos;
[0009] Extracting and processing each training sample using an initial video recommendation model to obtain a first video interest vector of the user for the target short video and a first item interest vector for the target item; the target item is an item in the target short video;
[0010] Decoupling the first video interest vector using an initial video recommendation model to obtain a first video content interest vector and a first video item interest vector;
[0011] Based on the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector, a total loss function is determined; based on the total loss function, the initial video recommendation model is trained to obtain a video recommendation model.
[0012] The present disclosure provides a video recommendation method, the method comprising:
[0013] Acquire characteristic data of a target user; the characteristic data includes a short video behavior sequence, an item behavior sequence, and a candidate short video set of the target user; the candidate short video set includes a plurality of candidate short videos to be recommended;
[0014] Processing the feature data using a video recommendation model to obtain a second video interest vector, a second video content interest vector, a second video item interest vector, and a second item interest vector for each candidate short video of the target user; the candidate items are items in the candidate short videos;
[0015] Determining a prediction result based on the second video interest vector, the second item interest vector, the second video content interest vector, and the second video item interest vector; the prediction result includes a click probability of the target user on the multiple candidate short videos;
[0016] Based on the prediction result, short videos are recommended to the target user.
[0017] The embodiment of the present disclosure provides a model training device, which includes a first acquisition module, a processing module and a training module, wherein:
[0018] A first acquisition module is used to acquire multiple training samples, each training sample includes a user's short video behavior sequence, an item behavior sequence, and one of the exposed target short videos;
[0019] a processing module configured to extract and process each training sample using an initial video recommendation model to obtain a first video interest vector of the user for the target short video and a first item interest vector for a target item; the target item being an item in the target short video; and decoupling the first video interest vector using the initial video recommendation model to obtain a first video content interest vector and a first video item interest vector;
[0020] A training module is used to determine a total loss function based on the first video interest vector, the first item interest vector, the first video content interest vector and the first video item interest vector; and to train the initial video recommendation model based on the total loss function to obtain a video recommendation model.
[0021] The embodiment of the present disclosure provides a video recommendation device, which includes a second acquisition module, a prediction module and a recommendation module, wherein:
[0022] The second acquisition module is used to acquire characteristic data of the target user; the characteristic data includes the short video behavior sequence, the item behavior sequence and the candidate short video set of the target user; the candidate short video set includes multiple candidate short videos to be recommended;
[0023] a prediction module, configured to process the feature data using a video recommendation model to obtain a second video interest vector, a second video content interest vector, a second video item interest vector, and a second item interest vector for each candidate short video of the target user; the candidate item being an item in the candidate short video; and determining a prediction result based on the second video interest vector, the second item interest vector, the second video content interest vector, and the second video item interest vector; the prediction result including a click probability of the target user on the multiple candidate short videos;
[0024] A recommendation module is used to recommend short videos to the target user based on the prediction results.
[0025] An embodiment of the present disclosure provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the aforementioned model training method or video recommendation method is implemented.
[0026] An embodiment of the present disclosure provides a computer storage medium storing a computer program; after the computer program is executed, it can implement the aforementioned model training method or video recommendation method.
[0027] The embodiments of the present disclosure propose a model training method, a video recommendation method, an apparatus, an electronic device, and a computer storage medium, the method comprising: obtaining a plurality of training samples, each training sample comprising a user's short video behavior sequence, an item behavior sequence, and one of the exposed target short videos; using an initial video recommendation model to extract and process each of the training samples to obtain the user's first video interest vector for the target short video and a first item interest vector for the target item; the target item is an item in the target short video; using the initial video recommendation model to decouple the first video interest vector to obtain a first video content interest vector and a first video item interest vector; determining a total loss function based on the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector; and training the initial video recommendation model based on the total loss function to obtain a video recommendation model.
[0028] It can be seen that the embodiment of the present disclosure determines the video interest vector and item interest vector of each user according to the short video behavior sequence, item behavior sequence and exposed short video of each training sample. These interests can be used to determine the model loss later, that is, the model training process not only takes into account the interaction between the user and the short video, but also the interaction between the user and the item. In other words, the embodiment of the present disclosure can use the behavior data of the user and the item to help identify the user's short video interest preference, which can solve the problem of poor recommendation effect caused by only considering the interaction between the user and the video in the related technology; in addition, by decoupling the video interest vector, the user's interest in the item itself and the interest in the video content can be determined, that is, the difference in the user's interest in the item itself and the video content is also considered in the model training process. In this way, the interpretability of the model can be improved to ensure the video recommendation effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] FIG1 is a flow chart of a model training method according to an embodiment of the present disclosure;
[0030] FIG2 is a schematic diagram of the structure of a Transformer encoder according to an embodiment of the present disclosure;
[0031] FIG3 is a schematic diagram of the structure of an interest decoupling network according to an embodiment of the present disclosure;
[0032] FIG4 is a flow chart of a video recommendation method according to an embodiment of the present disclosure;
[0033] FIG5 is a schematic structural diagram of an initial video recommendation model according to an embodiment of the present disclosure;
[0034] FIG6 is a schematic diagram of the structure of a model training device according to an embodiment of the present disclosure;
[0035] FIG7 is a schematic diagram of the composition structure of a video recommendation device according to an embodiment of the present disclosure;
[0036] FIG8 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0037] The present disclosure will be further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the examples provided herein are merely for explaining the present disclosure and are not intended to limit the present disclosure. In addition, the examples provided below are partial examples for implementing the present disclosure, rather than providing all examples for implementing the present disclosure. In the absence of conflict, the technical solutions described in the examples of the present disclosure may be implemented in any combination.
[0038] It should be noted that, in the embodiments of the present disclosure, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a method or apparatus comprising a series of elements includes not only the elements explicitly stated, but also other elements not explicitly listed, or also includes elements inherent to the implementation of the method or apparatus. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other related elements (such as steps in the method or units in the apparatus, for example, a unit may be a portion of a circuit, a portion of a processor, a portion of a program or software, etc.) in the method or apparatus comprising the element.
[0039] The term "and / or" herein is merely a description of an association relationship between associated objects, indicating that three relationships may exist. For example, "I and / or J" may represent the existence of I alone, the existence of both I and J, and the existence of J alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of I, J, and R" may represent any one or more elements selected from the set consisting of I, J, and R.
[0040] For example, the model training method provided in the embodiment of the present disclosure includes a series of steps, but the model training method provided in the embodiment of the present disclosure is not limited to the recorded steps. Similarly, the model training device provided in the embodiment of the present disclosure includes a series of modules, but the model training device provided in the embodiment of the present disclosure is not limited to including the modules explicitly recorded, and may also include modules required to obtain relevant time series data or perform processing based on time series data.
[0041] Currently, short video recommendation in e-commerce contexts faces two major challenges. First, user video behavior is sparse. Existing short video recommendation methods typically focus solely on user-video interaction. However, in e-commerce contexts, user-video interaction is relatively sparse, so methods that directly apply short video recommendation methods to entertainment scenarios do not achieve effective results. Second, users differ in their interests between items and short videos. Although short videos in e-commerce contexts are about items, the way they are presented differs significantly, which is a key factor influencing user behavior in these two domains. Item pages feature promotional information, concise cover images, and keyword-based titles, making users more inclined to purchase products directly and quickly. In contrast, short videos offer more diverse multimedia content. Clicking on a video leads to a play page displaying more item details. Furthermore, the different multimodal information presentation formats of short videos can also lead to different popularity levels. However, most video recommendation algorithms fail to account for the differences in user interest in items and video content, which impacts the accuracy of video recommendations.
[0042] In order to solve the above technical problems, the following embodiments are proposed.
[0043] In some embodiments of the present disclosure, the model training method can be implemented using a processor in a model training device, and the above-mentioned processor can be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor.
[0044] It should be noted that in the technical solution of the present disclosure, the collection, use, storage, sharing and transfer of user personal information involved are in compliance with the provisions of relevant laws and regulations, and require notification to users and obtaining their consent or authorization. When applicable, user personal information is de-identified and / or anonymized and / or encrypted.
[0045] In the disclosed embodiments, the model training method can be applied to e-commerce scenarios of short video recommendations; here, short videos refer to videos with a playback time of less than a short time (such as 1 minute or 2 minutes, etc.); the purpose of short video recommendations is to divert traffic to related items in e-commerce scenarios and improve the conversion rate of items in short videos.
[0046] FIG1 is a flow chart of a model training method according to an embodiment of the present disclosure. As shown in FIG1 , the method includes the following steps:
[0047] Step 100: Obtain multiple training samples.
[0048] In an embodiment of the present disclosure, each of the multiple training samples may include a user's short video behavior sequence, an item behavior sequence, and one of the exposed target short videos; for each training sample, the user's short video behavior sequence may include video features of multiple short videos clicked by the user in a set time window; the user's item behavior sequence may include item features of multiple items clicked by the user in a set time window; the target short video may be one of the short videos exposed to the user during the item browsing process, and the target short video may be a short video that has been exposed and clicked by the user, or a short video that has been exposed but not clicked by the user.
[0049] Here, the value of the set time window can be set according to actual conditions, and the embodiment of the present disclosure does not specifically limit this. For example, it can be the past two months, or the past three months, etc.
[0050] For example, video features are also called video attribute features, which may include feature data such as video identification number (Identity Document, ID), video type, video length and cover image vector; item features are also called item attribute features, which may include item ID, item category and item brand ID, etc.
[0051] In the disclosed embodiment, each training sample may include not only the user's short video behavior sequence, the item behavior sequence, and the target short video, but also the user's basic features, context features, and label information; after obtaining the target short video, the target short video features and target item features corresponding to the target short video may be determined; the target item is the item in the target short video.
[0052] Exemplarily, the basic characteristics of a user may include the attribute characteristics and statistical characteristics of the user; wherein the attribute characteristics may include characteristic data such as the user's gender, age and author type, and the statistical characteristics may include characteristic data such as the number of clicks and click-through rate of each short video by the user in a set time window; the target short video characteristics represent the video characteristics of the target short video, and the target item characteristics represent the item characteristics of the target item; the characteristic data included in the target short video characteristics and target item characteristics can refer to the above-mentioned video characteristics and item characteristics, and will not be elaborated here; the contextual characteristics may include characteristic data such as weather, location, and operation time; the tag information is also called a click tag, and each target short video corresponds to a tag information, which is used to indicate whether the user has clicked on the target short video.
[0053] Step 101: Utilize the initial video recommendation model to extract and process each training sample to obtain the user's first video interest vector for the target short video and the first item interest vector for the target item.
[0054] In the embodiment of the present disclosure, before executing step 101, an initial video recommendation model may be constructed first; illustratively, the main component of the initial video recommendation model may include a click-through rate (CTR) prediction model.
[0055] For example, for the CTR prediction model, its input can be defined as (x, y) ~ (X, Y), where (x, y) corresponds to a training sample, x represents the input feature, y∈{0,1} represents the click label, and X, Y represents the value space of x, y. The input features corresponding to each training sample can include the following: item behavior sequence, which can be represented as The short video behavior sequence can be represented as Among them, l and n are the lengths of the item behavior sequence and the short video behavior sequence, that is, the number of items and short videos clicked by the user in the set time window, the user's basic characteristics x u , target short video features x v 、Target item feature x i and contextual features x c Assume that the target short video is defined as v t , then the target short video estimation problem can be defined as follows:
[0056] Among them, prob(v t |x) indicates that the user clicks on the target short video v t The probability of can be simplified to p(x), where f represents the function of estimated probability.
[0057] Understandably, in order to address the technical problem of poor recommendation results due to sparse user video behavior and differences in user interest in items and video content when recommending short videos in e-commerce scenarios, the disclosed embodiment constructs an interest extraction network, an interest decoupling network, and a multi-scale interest fusion network based on the CTR prediction model. That is, the constructed initial video recommendation model can include the CTR prediction model, the interest extraction network, the interest decoupling network, and the multi-scale interest fusion network.
[0058] For example, the interest extraction network is used to extract interest from the input features corresponding to each training sample, and obtain the user's first video interest vector I for the target short video. v and the first item interest vector I for the target item p ; Next, the first video interest vector I v and the first item interest vector I p The extraction process is illustrated in FIG.
[0059] For example, the initial video recommendation model may include a shared embedding layer. The input features of the initial video recommendation model are the input features of the CTR estimation model. x u , x v , x i , x c After that, because these features are all ID type features, each input feature can be one-hot encoded to obtain the corresponding one-hot vector; because the one-hot vector is highly sparse, the shared embedding layer can be used to further transform the one-hot vector corresponding to each input feature into a low-dimensional embedding vector, which can be expressed as e u , e v , e i , e c ;in, Represents the embedding vector corresponding to the item in the item behavior sequence, Represents the embedding vector corresponding to the nth short video in the short video behavior sequence.
[0060] Exemplarily, the interest extraction network includes a Transformer encoder corresponding to each of the short video behavior sequence and the item behavior sequence; here, the Transformer encoder is the encoder of the Transformer layer, and the encoder of the Transformer layer can be used to model the embedding vector corresponding to the short video behavior sequence and the embedding vector corresponding to the item behavior sequence, respectively. Figure 2 is a structural diagram of a Transformer encoder in an embodiment of the present disclosure. As shown in Figure 2, the Transformer encoder includes a two-layer network. The first layer of the network includes a multi-head self-attention layer (Multi-Head Self Attention) and a normalization layer and a residual (Add&Norm). The second layer of the network includes a feedforward neural network layer (Feed Forward) and a normalization layer and a residual. Norm is a standard normalization layer. The Transformer encoder can obtain high-level representations by stacking multiple self-attention layers.
[0061] For example, the Transformer encoder shown in Figure 2 is first used to learn the relationship between a short video and other short videos in the short video behavior sequence, as well as the relationship between an item and other items in the item behavior sequence, to achieve a deeper representation of each short video and item. The multi-head self-attention layer included in the Transformer encoder contains multiple self-attention layers. Assume that the number of self-attention layers is n h , then use n h Each head transforms the input vector n through different linear projections h Secondary mapping to d h dimensional latent vector, where the input vector represents the embedding vector of a short video or item in the behavior sequence. Next, a scaled dot product attention mechanism is performed on each head, and the output vectors of all scaled dot product attentions are merged. The specific definition is shown in formula (2):
[0062] Where Q∈R d ,K∈R d ,V∈R d Represents Query, Key and Value respectively, represents the weight matrix, and the dot product attention function with scaling is defined as shown in formula (3):
[0063] Among them, d k represents the dimension of K, d k with d h Equal; after getting the output of the multi-head self-attention layer,
[0064] A feedforward neural network layer is used to perform further nonlinear processing on the output to enhance the nonlinear capability of the model. The feedforward neural network layer includes two fully connected layers for performing nonlinear processing on each input vector.
[0065] For example, the embedding vector corresponding to the item behavior sequence Embedding vector corresponding to the short video behavior sequence After passing through different Transformer encoders, we can get the item representation sequence after internal interaction and short video representation sequences
[0066] Furthermore, after obtaining the item representation sequence and the short video representation sequence, the target short video v is obtained from the user's historical behavior sequence (corresponding to the short video behavior sequence and the item behavior sequence) through the target short video attention mechanism (TVA) and the target product attention mechanism (TPA). t Related first video interest vector I v and the target item p t Related first item interest vector I p At this time, we can get the user's response to the target short video v t The first video interest vector I v and the target item p t The first item interest vector I p , the specific definition is shown in formula (4):
[0067] Among them, a v (·) and a p (·) are two feedforward neural networks that calculate attention weights, and the last layer activation function is softmax, that is
[0068] It can be seen that in the embodiment of the present disclosure, interest extraction is performed on the input features corresponding to each training sample through the interest extraction network, and the user's first video interest vector for the target short video and the first item interest vector for the target item can be obtained; these interests can be used to determine the model loss later. It can be seen that the model training process not only takes into account the interaction between the user and the short video, but also the interaction between the user and the item, that is, the behavioral data of the user and the item can be used to help identify the user's short video interest preference, which can solve the problem of poor recommendation effect in related technologies due to only considering the interaction between the user and the video.
[0069] Step 102: Decoupling the first video interest vector using the initial video recommendation model to obtain a first video content interest vector and a first video item interest vector.
[0070] It can be understood that the first video interest vector I obtained from the short video behavior sequence v It couples the user's interest in video content and interest in video items. For example, a user may not be interested in an item in a short video, but its unique presentation makes the user feel very interested. The item behavior sequence can only represent the user's interest in the item. Therefore, for the short video behavior sequence, two types of interests are explicitly modeled to obtain two interest representations, namely the user's first video content interest vector and the first video item interest vector for the target short video, so that the model can be used to extract the user's interest from the first video interest vector I. v Extract the user's first video content interest vector for the target short video It not only increases the interpretability of the model, but also prevents the introduction of the first video interest vector I v The first video item interest vector in It may be related to the first item interest vector I extracted from the item behavior sequence p There is a large degree of duplication, which leads to a decrease in model performance.
[0071] In order to effectively extract the fine-grained interest of non-items, that is, the first video content interest vector Extract the first video content interest vector by creating an interest decoupling network The interest decoupling network includes a domain adaptation network, a domain decoupling network and a reconstruction network. The domain adaptation network is used to transform the first video item interest vector and the first item interest vector I p Align in latent space to ensure It represents the user's interest in the video item. The domain decoupling network is used to analyze the first video content interest vector and the first video item interest vector Decoupling to ensure the first video content interest vector It is the user’s interest in the short video content. The reconstruction network is used to reconstruct information to ensure that the information is not lost.
[0072] FIG3 is a schematic diagram of the structure of an interest decoupling network in an embodiment of the present disclosure. As shown in FIG3 , the first video interest vector I is extracted from the short video behavior sequence. v After that, the user’s first video content interest vector for the target short video can be obtained through the video encoder and the item encoder (SKU Encoder). and the first video item interest vector Here, the video encoder and the object encoder can be a two-layer fully connected network. v Input the video interest decoupling network for decoupling processing, and the user's first video content interest vector for the target short video can be obtained. and the first video item interest vector
[0073] Step 103: Determine a total loss function based on the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector; and train the initial video recommendation model based on the total loss function to obtain a video recommendation model.
[0074] In the embodiment of the present disclosure, after obtaining the first video interest vector I according to the above steps, v , the first item interest vector I p , first video content interest vector and the first video item interest vector Finally, the total loss function of the model can be determined based on these vectors.
[0075] In some embodiments, determining a total loss function based on the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector may include: performing feature alignment on the first video item interest vector and the first item interest vector using Maximum Mean Discrepancy (MMD), and determining the MMD loss based on the feature alignment result; and determining a total loss function based on the MMD loss, the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector.
[0076] It is understandable that when the item encoder is a two-layer fully connected network, it is not guaranteed that the first video item interest vector extracted by the item encoder alone To represent the user's interest in the video item, the first item interest vector I can be selected p For the agent, if the first video item interest vector and the first item interest vector I p The interest vector of the item encoder is close to that of the first video item extracted from the short video behavior sequence.
[0077] For example, the interest vector of the first video item can be calculated using MMD. and the first item interest vector I pPerform feature alignment and determine the MMD loss, and realize the first video item interest vector based on the MMD loss and the first item interest vector I p The interest vector of is constantly approaching, where the MMD loss The definition of is shown in formula (5):
[0078] Among them, |X v | and |X p | represents the number of samples in a Batch, and the two are equal; stop_grad is gradient blocking, which prevents the MMD loss from affecting the first item interest vector I during gradient back propagation. p interference.
[0079] For example, the first video item interest vector can be minimized by minimizing the MMD loss. and the first item interest vector I p Keep getting closer to ensure that the output of the item encoder is the first video item interest vector extracted from the short video behavior sequence
[0080] In some embodiments, the video interest decoupling network may further include a gradient reversal layer, and the above method may further include: determining the adversarial loss of the first item interest vector and the first video content interest vector through the gradient reversal layer.
[0081] For example, after obtaining the first video item interest vector After that, it is necessary to ensure that the video encoder outputs the user’s interest in the short video content. Before determining the MMD loss, a gradient reversal layer can be introduced (corresponding to the one shown in Figure 3). ), the adversarial loss between the first item interest vector and the first video content interest vector is determined by the gradient reversal layer, and the first video content interest vector can be maximized by maximizing the adversarial loss. With the first item interest vector I p Separate; that is, use the adversarial training method of gradient reversal to separate the first video content interest vector obtained from the short video behavior sequence The first item interest vector I obtained from the item behavior sequence p Separate to achieve the first video content interest vector of the target short video and the first video item interest vector decoupling.
[0082] In some embodiments, determining a total loss function based on the MMD loss, the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector may include: determining the mutual information loss between the first video content interest vector and the first video item interest vector; determining a total loss function based on the MMD loss, the mutual information loss, the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector.
[0083] Here, mutual information loss is also called mutual information. It can be understood that in order to enhance the first video content interest vector and the first video item interest vector The decoupling capability can further determine the first video content interest vector and the first video item interest vector The mutual information loss between them is minimized so that the first video content interest vector and the first video item interest vector Mutual information is a basic indicator for measuring the correlation between two random variables. Mathematically, the mutual information between two random variables x and y is as shown in formula (6):
[0084] Here, the joint distribution of random variables x and y is p(x,y), the marginal distributions are p(x)p(y) respectively, and the mutual information I(x;y) is the relative entropy between the joint distribution p(x,y) and the marginal distributions p(x)p(y).
[0085] Since the mutual information loss needs to be minimized and the mutual information cannot be directly calculated, here, the upper bound of the mutual information is used to estimate the CLUB approximate mutual information loss The definition is shown in formula (7):
[0086] in, is a batch of sample pairs from the joint distribution p(x,y), q θ It is an approximation of the joint distribution p(x,y). In actual processing, a neural network q parameterized by θ can be used. θ To learn the distribution, x, y represent the first video content interest vector and the first video item interest vector y i ′ is defined as the embedding vector of the i-th sample y′ in the Batch, and y′ is the randomly shuffled data within the Batch of y.
[0087] In some embodiments, determining a total loss function based on MMD loss, mutual information loss, the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector may include: reconstructing the first video content interest vector and the first video item interest vector to obtain a reconstructed first video interest vector; determining a reconstruction loss based on the first video interest vectors before and after reconstruction; and determining a total loss function based on MMD loss, mutual information loss, reconstruction loss, the first item interest vector, and the first video content interest vector.
[0088] For example, the first video content interest vector can be reconstructed by a reconstruction network (corresponding to the Reconstructor shown in FIG3 ). and the first video item interest vector Reconstruct and obtain the reconstructed first video interest vector I v , and then according to the first video interest vector I before reconstruction v and the reconstructed first video interest vector I′ v , determine the reconstruction loss L Reconstructor , defined as shown in formula (8): L Reconstructor = ||I v -I′ v || (8)
[0089] For example, the reconstruction loss can be minimized so that the interest vectors of the first video items before and after reconstruction are constantly close, which greatly reduces the time required to extract the interest vectors of the first video content. and the first video item interest vector There is a risk of information loss during the process.
[0090] In some embodiments, determining a total loss function based on MMD loss, mutual information loss, reconstruction loss, the first item interest vector, and the first video content interest vector may include: inputting the first item interest vector and the first video content interest vector into the multi-scale interest fusion network in the initial video recommendation model for processing to obtain the user's overall interest vector; determining the target short video features and target item features corresponding to the target short video, and using the user's basic features, the target short video features, and the target item features to determine the user's overall preference vector, the video preference vector for the target short video, and the item preference vector for the target item from the multi-scale interest set respectively; the multi-scale interest set includes the overall interest vector, the first item interest vector, and the first video content interest vector; determining the cross entropy loss based on the overall preference vector, the video preference vector, the item preference vector, the user's basic features, the target short video features, and the target item features; and determining the total loss function based on MMD loss, mutual information loss, reconstruction loss, and cross entropy loss.
[0091] Exemplarily, the multi-scale interest fusion network is used to derive the first item interest vector I p and the first video content interest vector Perform adaptive fusion; because the first video item interest vector and the first item interest vector I p The vector representation is basically the same. In order to avoid introducing duplicate information, the first video item interest vector It will not be input into the multi-scale interest fusion network. The multi-scale interest fusion network includes the Gate network, also known as the adaptive fusion network. The Gate network first adaptively fuses the first item interest vector I through the Gating method. p and the first video content interest vector Get the user's overall interest vector I g , the specific calculation is shown in formula (9):
[0092] Among them, the Gate network adapts to the first item interest vector I p and the first video content interest vector Assign fusion weights and adopt a method similar to the self-attention mechanism to fuse, w u is the adaptive fusion weight, also known as weight score, which is defined as shown in formula (10):
[0093] Among them, h1(u), h2(u) and h3(u) are three linear mapping functions, d represents the dimension,
[0094] Furthermore, after obtaining the user's overall interest vector I g Then, combined with the first item interest vector I p and the first video content interest vector A multi-scale interest set I can be determined; wherein the multi-scale interest set I includes the overall interest vector I g , the first item interest vector I p and the first video content interest vector Here, I g ∈R d ,I p ∈R d , That is, the multi-scale interest set I is represented as I∈R 3×d After determining the multi-scale interest set I, the target-aware multi-head attention mechanism (MHTA) can be used to determine the user's interest preference vectors at different scales.
[0095] For example, the embedding vector e corresponding to the basic features of the user can be used u And MHTA determines the user's overall preference vector a from the multi-scale interest set I u , using the embedding vector e corresponding to the target short video feature v And MHTA determines the user's video preference vector a for the target short video from the multi-scale interest set I v , using the embedding vector e corresponding to the target item feature i And MHTA determines the user's item preference vector a for the target item from the multi-scale interest set I i , the corresponding definition is shown in formula (11): a u =MHTA(e u ,I,I)=MultiHead(e u ,I,I) a v =MHTA(e v ,I,I)=MultiHead(e v ,I,I) a i =MHTA(e i ,I,I)=MultiHead(e i ,I,I) (11)
[0096] It can be seen that in the embodiment of the present disclosure, the multi-scale interest fusion network can be used to obtain the user's multi-scale interest vectors, and by splicing these interest vectors, the recommendation effect of the model can be improved.
[0097] In some embodiments, determining the cross entropy loss based on the overall preference vector, video preference vector, item preference vector, user basic features, target short video features, and target item features can include: splicing the overall preference vector, video preference vector, item preference vector, user basic features, target short video features, target item features, and context features to obtain a spliced feature vector; inputting the spliced feature vector into the multi-layer perceptron in the initial video recommendation model for processing to obtain a model output result; and determining the cross entropy loss based on the model output result and label information.
[0098] For example, after obtaining the overall preference vector a u , video preference vector a v and the item preference vector a i After that, the overall preference vector a can be u , video preference vector a v 、Item preference vector a i , the embedding vector e corresponding to the user's basic features u , the embedding vector e corresponding to the target short video featurev , the embedding vector e corresponding to the target item feature i The embedding vector e corresponding to the context feature c The spliced feature vector is obtained; then, the spliced feature vector is input into the multi-layer perceptron in the initial video recommendation model for processing to obtain the model output result; it should be noted that, in order to predict whether a user clicks on the target short video, the output layer of the initial video recommendation model can use the sigmoid function for binary classification, and the sigmoid function will map the model output result to a value between 0 and 1; at this time, the model output result is used to characterize the user's click probability on the target short video; based on the model output result and label information, the cross entropy loss L can be determined ce , as shown in formula (12):
[0099] In some embodiments, determining a total loss function based on MMD loss, mutual information loss, reconstruction loss, and cross entropy loss may include: determining a first product of MMD loss and a first hyperparameter, a second product of mutual information loss and a second hyperparameter, and a third product of reconstruction loss and a third hyperparameter; and accumulating the first product, the second product, the third product, and the cross entropy loss to determine the total loss function.
[0100] Exemplarily, the first hyperparameter, the second hyperparameter and the third hyperparameter can be represented as α, β, and γ respectively. Here, the values of α, β, and γ can be set according to actual conditions, and the embodiments of the present disclosure do not limit this. For example, the values can be 0.1, 0.1, and 0.01 respectively.
[0101] In the embodiment of the present disclosure, after obtaining the MMD loss Mutual Information Loss Reconstruction loss L Reconstructor and cross entropy loss L ce After that, the total loss function L of the initial video recommendation model can be determined all , as shown in formula (13):
[0102] For example, based on the total loss function L all The loss value is calculated, and the parameters of the initial video recommendation model are adjusted. Multiple rounds of training are repeated to obtain a trained video recommendation model. Because the video recommendation model can take into account the differences and coupling between video item interest and video content interest in short video behavior sequences in e-commerce scenarios, it can improve the model's interpretability.
[0103] An embodiment of the present disclosure proposes a model training method, which includes: obtaining multiple training samples, each training sample including a user's short video behavior sequence, an item behavior sequence, and one of the exposed target short videos; using an initial video recommendation model to extract and process each training sample to obtain the user's first video interest vector for the target short video and a first item interest vector for the target item; the target item is an item in the target short video; using the initial video recommendation model to decouple the first video interest vector to obtain a first video content interest vector and a first video item interest vector; based on the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector, determining a total loss function; based on the total loss function, training the initial video recommendation model to obtain a video recommendation model. It can be seen that the embodiment of the present disclosure determines the video interest vector and item interest vector of each user according to the short video behavior sequence, item behavior sequence and exposed short video of each training sample. These interests can be used to determine the model loss later, that is, the model training process not only takes into account the interaction between the user and the short video, but also the interaction between the user and the item. In other words, the embodiment of the present disclosure can use the behavior data of the user and the item to help identify the user's short video interest preference, which can solve the problem of poor recommendation effect caused by only considering the interaction between the user and the video in the related technology; in addition, by decoupling the video interest vector, the user's interest in the item itself and the interest in the video content can be determined, that is, the difference in the user's interest in the item itself and the video content is also considered in the model training process. In this way, the interpretability of the model can be improved to ensure the video recommendation effect of the model.
[0104] In order to better reflect the purpose of the present disclosure, further examples are given based on the above embodiments of the present disclosure.
[0105] FIG4 is a flow chart of a video recommendation method according to an embodiment of the present disclosure. As shown in FIG4 , the method includes the following steps:
[0106] Step 200: Obtain characteristic data of the target user;
[0107] Step 201: Process the feature data using a video recommendation model to obtain the target user's second video interest vector, second video content interest vector, second video item interest vector, and second item interest vector for each candidate short video;
[0108] Step 202: Determine a prediction result based on the second video interest vector, the second item interest vector, the second video content interest vector, and the second video item interest vector;
[0109] Step 203: Recommend short videos to target users based on the prediction results.
[0110] In the disclosed embodiment, the candidate short video set may include multiple candidate short videos to be recommended; the feature data of the target user may include the target user's short video behavior sequence, item behavior sequence, candidate short video set, basic features of the target user, and contextual features; the composition of each item of information in the feature data may refer to the composition of each item of information in the above-mentioned training samples. To avoid repeated description, it will not be repeated here.
[0111] For example, after obtaining a set of candidate short videos, the candidate short video features and candidate item features corresponding to each candidate short video can be determined. When performing short video tag click prediction, the target user's short video behavior sequence, item behavior sequence, basic features, candidate short video features, candidate item features, and context features can be input into the trained video recommendation model to predict the click probability and obtain the prediction result.
[0112] In the disclosed embodiment, the video recommendation model is derived from steps 100 to 103 described above; the prediction results include the click probability of the target user for multiple candidate short videos. The click probability prediction process for each candidate short video is similar to the training process for the target short video described above and will not be further described here.
[0113] FIG5 is a schematic diagram of the structure of an initial video recommendation model in an embodiment of the present disclosure. As shown in FIG5 , the input features of the model include basic user features, short video behavior sequences, item behavior sequences, short video features, and context features. Here, short video features include the target short video features and target item features described above. Transformer corresponds to the Transformer encoder described above, Domain adaption corresponds to the domain adaptation network described above, Domain Disentangle corresponds to the domain decoupling network described above, Side Info Merger corresponds to the adaptive fusion network described above, and is part of the multiscale interest fusion network. Multiscale Interest Adaptive Merger corresponds to the multiscale interest fusion network described above, and MLP corresponds to the multi-layer perceptron described above. The processing procedures corresponding to each component of the initial video recommendation model have been described in the above embodiments and will not be repeated here.
[0114] It can be seen that in the embodiment of the present disclosure, the differences and coupling issues between video item interests and video content interests in short video behavior sequences in e-commerce scenarios are taken into account during the model training process, and the user's interest in the short video content itself is extracted through the interest decoupling network to improve the interpretability of the model; the item interest obtained from the item behavior sequence is supplemented with the user's short video interest, which can improve the recommendation effect of short videos; the user's multi-scale interests are obtained through the multi-scale interest fusion network, including the user's overall interest, short video content interest and item interest, and the user's interest preference vector is determined through the target-aware multi-head attention mechanism, which can further improve the model effect.
[0115] FIG6 is a schematic diagram of the structure of a model training device according to an embodiment of the present disclosure. As shown in FIG6 , the device includes: a first acquisition module 300, a processing module 301, and a training module 302, wherein:
[0116] A first acquisition module 300 is configured to acquire a plurality of training samples, each of which includes a user's short video behavior sequence, an item behavior sequence, and one of the exposed target short videos;
[0117] Processing module 301 is configured to extract and process each training sample using an initial video recommendation model to obtain a first video interest vector of the user for the target short video and a first item interest vector for the target item; the target item is an item in the target short video; and decouple the first video interest vector using the initial video recommendation model to obtain a first video content interest vector and a first video item interest vector.
[0118] The training module 302 is used to determine a total loss function based on the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector; and train the initial video recommendation model based on the total loss function to obtain a video recommendation model.
[0119] In some embodiments, the training module 302 is configured to:
[0120] Performing feature alignment on the first video item interest vector and the first item interest vector using maximum mean difference (MMD), and determining an MMD loss based on the feature alignment result;
[0121] The total loss function is determined based on the MMD loss, the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector.
[0122] In some embodiments, the training module 302 is configured to:
[0123] determining a mutual information loss between the first video content interest vector and the first video item interest vector;
[0124] The total loss function is determined based on the MMD loss, the mutual information loss, the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector.
[0125] In some embodiments, the training module 302 is configured to:
[0126] Reconstructing the first video content interest vector and the first video object interest vector to obtain a reconstructed video interest vector;
[0127] Determine the reconstruction loss based on the video interest vectors before and after reconstruction;
[0128] The total loss function is determined based on the MMD loss, the mutual information loss, the reconstruction loss, the first item interest vector, and the first video content interest vector.
[0129] In some embodiments, each training sample further includes basic user characteristics, and the training module 302 is configured to:
[0130] Inputting the first item interest vector and the first video content interest vector into the multi-scale interest fusion network in the initial video recommendation model for processing to obtain the overall interest vector of the user;
[0131] Determining target short video features and target item features corresponding to the target short video, and determining the user's overall preference vector, the video preference vector for the target short video, and the item preference vector for the target item from a multi-scale interest set using the user's basic features, the target short video features, and the target item features, respectively; the multi-scale interest set includes the overall interest vector, the first item interest vector, and the first video content interest vector;
[0132] Determining a cross entropy loss based on the overall preference vector, the video preference vector, the item preference vector, the user basic features, the target short video features, and the target item features;
[0133] The total loss function is determined based on the MMD loss, the mutual information loss, the reconstruction loss, and the cross entropy loss.
[0134] In some embodiments, each training sample further includes context features and label information, where the label information is used to indicate whether the user has clicked on the target short video. The training module 302 is configured to:
[0135] splicing the overall preference vector, the video preference vector, the item preference vector, the user basic features, the target short video features, the target item features, and the context features to obtain a spliced feature vector;
[0136] Inputting the spliced feature vector into the multi-layer perceptron in the initial video recommendation model for processing to obtain a model output result;
[0137] The cross entropy loss is determined based on the model output result and the label information.
[0138] In some embodiments, the training module 302 is configured to:
[0139] Determining a first product of the MMD loss and a first hyperparameter, a second product of the mutual information loss and a second hyperparameter, and a third product of the reconstruction loss and a third hyperparameter;
[0140] The first product, the second product, the third product, and the cross entropy loss are accumulated to determine the total loss function.
[0141] FIG7 is a schematic diagram of the composition structure of a video recommendation device according to an embodiment of the present disclosure. As shown in FIG7 , the device includes: a second acquisition module 400, a prediction module 401, and a recommendation module 402, wherein:
[0142] The second acquisition module 400 is used to acquire characteristic data of a target user; the characteristic data includes a short video behavior sequence, an item behavior sequence, and a candidate short video set of the target user; the candidate short video set includes a plurality of candidate short videos to be recommended;
[0143] Prediction module 401 is configured to process the feature data using a video recommendation model to obtain a second video interest vector, a second video content interest vector, a second video item interest vector, and a second item interest vector for each candidate short video of the target user; the candidate item being an item in the candidate short video; and determine a prediction result based on the second video interest vector, the second item interest vector, the second video content interest vector, and the second video item interest vector; the prediction result includes a click probability of the target user on the multiple candidate short videos.
[0144] The recommendation module 402 is configured to recommend short videos to the target user based on the prediction result.
[0145] The video recommendation model is obtained according to the model training method provided by one or more of the aforementioned technical solutions.
[0146] In actual applications, the above-mentioned first acquisition module 300, processing module 301 and training module 302, second acquisition module 400, prediction module 401 and recommendation module 402 can all be implemented by a processor located in an electronic device, and the processor can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.
[0147] In addition, the functional modules in this embodiment may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or software functional modules.
[0148] If the integrated unit is implemented as a software functional module and is not sold or used as an independent item, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, or the portion that contributes to the relevant technology, or all or part of the technical solution can be embodied in the form of a software item. This computer software item is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0149] Specifically, the computer program instructions corresponding to a model training method or video recommendation method in this embodiment can be stored on a storage medium such as a CD, a hard disk, or a USB flash drive. When the computer program instructions corresponding to a model training method or a video recommendation method in the storage medium are read or executed by an electronic device, any model training method or video recommendation method in the aforementioned embodiments is implemented.
[0150] Based on the same technical concept as the above embodiments, referring to FIG8 , an electronic device 500 provided by the present disclosure may include: a memory 501 and a processor 502 ; wherein,
[0151] Memory 501, used to store computer programs and data;
[0152] The processor 502 is configured to execute a computer program stored in the memory to implement any one of the model training methods or video recommendation methods of the aforementioned embodiments.
[0153] In practical applications, the memory 501 may be a volatile memory, such as RAM; or a non-volatile memory, such as ROM, flash memory, hard disk drive (HDD) or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 502.
[0154] The processor 502 may be at least one of an ASIC, a DSP, a DSPD, a PLD, an FPGA, a CPU, a controller, a microcontroller, and a microprocessor. It is understood that the electronic device used to implement the functions of the processor may also be other electronic devices, which are not specifically limited in the embodiments of the present disclosure.
[0155] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0156] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.
[0157] The methods disclosed in the various method embodiments provided in this disclosure can be arbitrarily combined without conflict to obtain new method embodiments.
[0158] The features disclosed in the various article embodiments provided in this disclosure can be arbitrarily combined to obtain new article embodiments without conflict.
[0159] The features disclosed in the various method or device embodiments provided in this disclosure may be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0160] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program articles. Therefore, the present disclosure may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present disclosure may take the form of a computer program article implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage) containing computer-usable program code.
[0161] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program articles according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, and the combination of the processes and / or boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable model training device to produce a machine, so that the instructions executed by the processor of the computer or other programmable model training device generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0162] These computer program instructions may also be loaded onto a computer or other programmable model training device so that a series of operating steps are executed on the computer or other programmable device to produce computer-implemented processing, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0163] The above are only preferred embodiments of the present disclosure and are not intended to limit the scope of protection of the present disclosure. Industrial Applicability
[0164] The present disclosure determines the video interest vector and item interest vector of each user according to the short video behavior sequence, item behavior sequence and exposed short videos of each training sample. These interests can be used to determine the model loss later, that is, the model training process not only considers the interaction between the user and the short video, but also considers the interaction between the user and the item. In other words, the embodiment of the present disclosure can use the behavior data of the user and the item to help identify the user's short video interest preference, which can solve the problem of poor recommendation effect caused by only considering the interaction between the user and the video in the related technology; in addition, by decoupling the video interest vector, the user's interest in the item itself and the interest in the video content can be determined, that is, the difference in the user's interest in the item itself and the video content is also considered in the model training process. In this way, the interpretability of the model can be improved to ensure the video recommendation effect of the model.
Claims
1. A model training method, characterized in that: The method comprises: Acquire multiple training samples, each training sample includes a user's short video behavior sequence, an item behavior sequence, and one of the exposed target short videos; Using the initial video recommendation model to extract and process each training sample, obtaining the user's first video interest vector for the target short video and a first item interest vector for the target item; the target item is an item in the target short video; Decoupling the first video interest vector using an initial video recommendation model to obtain a first video content interest vector and a first video item interest vector; Based on the first video interest vector, the first item interest vector, the first video content interest vector and the first video item interest vector, a total loss function is determined; based on the total loss function, the initial video recommendation model is trained to obtain a video recommendation model.
2. The method according to claim 1, characterized in that The determining of a total loss function based on the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector includes: Performing feature alignment on the first video item interest vector and the first item interest vector using a maximum mean difference (MMD), and determining an MMD loss based on a result of the feature alignment; The total loss function is determined based on the MMD loss, the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector.
3. The method according to claim 2, characterized in that The determining the total loss function based on the MMD loss, the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector comprises: determining a mutual information loss between the first video content interest vector and the first video item interest vector; The total loss function is determined based on the MMD loss, the mutual information loss, the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector.
4. The method according to claim 3, characterized in that The determining the total loss function based on the MMD loss, the mutual information loss, the first video interest vector, the first item interest vector, the first video content interest vector, and the first video item interest vector comprises: Reconstructing the first video content interest vector and the first video object interest vector to obtain a reconstructed first video interest vector; Determining a reconstruction loss according to the first video interest vector before and after reconstruction; The total loss function is determined based on the MMD loss, the mutual information loss, the reconstruction loss, the first item interest vector and the first video content interest vector.
5. The method according to claim 4, characterized in that Each training sample further includes basic features of a user, and determining the total loss function based on the MMD loss, the mutual information loss, the reconstruction loss, the first item interest vector, and the first video content interest vector includes: Inputting the first item interest vector and the first video content interest vector into the multi-scale interest fusion network in the initial video recommendation model for processing to obtain the overall interest vector of the user; Determine the target short video features and target item features corresponding to the target short video, and respectively determine the user's overall preference vector, the video preference vector for the target short video, and the item preference vector for the target item from a multi-scale interest set using the user's basic features, the target short video features, and the target item features; the multi-scale interest set includes the overall interest vector, the first item interest vector, and the first video content interest vector; According to the overall preference vector, the video preference vector, the item preference vector, the user basic characteristics, Determine the cross entropy loss by using the target short video features and the target object features; The total loss function is determined based on the MMD loss, the mutual information loss, the reconstruction loss and the cross entropy loss.
6. The method according to claim 5, characterized in that Each training sample further includes context features and label information, wherein the label information is used to indicate whether the user has clicked on the target short video. The cross entropy loss is determined according to the overall preference vector, the video preference vector, the item preference vector, the user basic features, the target short video features, and the target item features, including: Concatenate the overall preference vector, the video preference vector, the item preference vector, the user basic features, the target short video features, the target item features, and the context features to obtain a concatenated feature vector; Inputting the concatenated feature vector into the multi-layer perceptron in the initial video recommendation model for processing to obtain a model output result; The cross entropy loss is determined according to the model output result and the label information.
7. The method according to claim 5 or 6, characterized in that: The determining the total loss function based on the MMD loss, the mutual information loss, the reconstruction loss and the cross entropy loss comprises: Determining a first product of the MMD loss and a first hyperparameter, a second product of the mutual information loss and a second hyperparameter, and a third product of the reconstruction loss and a third hyperparameter; The first product, the second product, the third product, and the cross entropy loss are accumulated to determine the total loss function.
8. A video recommendation method, characterized in that: The method comprises: Acquire feature data of a target user; the feature data includes a short video behavior sequence, an item behavior sequence, and a candidate short video set of the target user; the candidate short video set includes a plurality of candidate short videos to be recommended; The feature data is processed using a video recommendation model to obtain a second video interest vector, a second video content interest vector, a second video item interest vector, and a second item interest vector for each candidate short video of the target user; the candidate item is an item in the candidate short video; Determine a prediction result based on the second video interest vector, the second item interest vector, the second video content interest vector, and the second video item interest vector; the prediction result includes the click probability of the target user on the multiple candidate short videos; According to the prediction result, a short video is recommended to the target user.
9. A model training device, characterized in that: The device comprises: A first acquisition module is used to acquire multiple training samples, each training sample includes a user's short video behavior sequence, an item behavior sequence and one of the exposed target short videos; A processing module is used to extract and process each training sample using an initial video recommendation model to obtain a first video interest vector of the user for the target short video and a first item interest vector for the target item; the target item is an item in the target short video; and decouple the first video interest vector using the initial video recommendation model to obtain a first video content interest vector and a first video item interest vector; A training module is used to determine a total loss function based on the first video interest vector, the first item interest vector, the first video content interest vector and the first video item interest vector; based on the total loss function, the initial video recommendation model is trained to obtain a video recommendation model.
10. A video recommendation device, characterized in that: The device comprises: The second acquisition module is used to acquire feature data of a target user; the feature data includes a short video behavior sequence, an item behavior sequence and a candidate short video set of the target user; the candidate short video set includes a plurality of candidate short videos to be recommended; The prediction module is used to process the feature data using the video recommendation model to obtain the second video interest vector, the second video content interest vector, the second video item interest vector and the second video item interest vector of the target user for each candidate short video. a second item interest vector of a selected item; the candidate item is an item in the candidate short video; based on the second video interest vector, the second item interest vector, the second video content interest vector and the second video item interest vector, determining a prediction result; the prediction result includes the click probability of the target user on the multiple candidate short videos; A recommendation module is used to recommend short videos to the target user based on the prediction results.
11. An electronic device, characterized in that: The device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the method according to any one of claims 1 to 8 is implemented when the processor executes the program.
12. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the service control method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Short video recommendation method
CN113268633A
Multi-task model training method and device and content recommendation method and device
CN115545114A
Video recommendation method, model training method, electronic equipment and storage medium
CN116186326A
Short video recommendation method for decoupling user interest and short video duration preference
CN116723346A
Model training method and device, video recommendation method and device, equipment and medium
CN117390222A
Cited By
Cross-modal content recommendation method and device, equipment, medium and product
CN120407952A
Wafer surface treatment process recommendation method and device, equipment and storage medium
CN120611277A
Intelligent recommendation method for video content of IPTV set top box based on deep learning
CN120769118A