Training method and device of video recommendation model, electronic equipment and storage medium

By generating second interactive features to simulate the interactive information of cold-start videos and adjusting the model parameters, the accuracy problem in cold-start video recommendation is solved, and the overall accuracy of video recommendation is improved.

CN117688204BActive Publication Date: 2026-08-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211020873.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2026-08-25
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

In video recommendation scenarios, cold-start videos are difficult to recommend accurately due to a lack of interaction records, leading to a decline in the quality of the recommendation system.

Method used

By generating a second interactive feature based on content features to simulate the interactive information of cold start videos, and combining the prediction results with the label differences to adjust the model parameters, the representation differences between non-cold start and cold start videos are reduced.

Benefits of technology

It improves the prediction accuracy of cold-start videos during the testing phase, alleviates the disadvantage of cold-start videos in recommendation, and enhances the overall accuracy of video recommendation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117688204B_ABST
    Figure CN117688204B_ABST
Patent Text Reader

Abstract

The application relates to the computer technical field, in particular to the artificial intelligence technical field, and provides a video recommendation model training method and device, an electronic device and a storage medium. The method comprises the following steps: inputting each selected training sample into a video recommendation model to be trained; performing feature extraction on each selected training sample based on the video recommendation model to obtain a corresponding extraction result; performing feature mapping on the content features corresponding to the sample videos contained in each training sample to obtain a corresponding second interaction feature; performing recommendation prediction based on the extraction result and the second interaction feature corresponding to each training sample to obtain a corresponding prediction result; and adjusting the parameters of the video recommendation model based on the difference between each prediction result and a corresponding sample label and the difference between the prediction results. Since the second interaction feature is used to simulate interaction information during model training, the performance disadvantage of a cold start video is weakened, and the accuracy of video recommendation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more particularly to the field of artificial intelligence technology, providing a training method, apparatus, electronic device, and storage medium for a video recommendation model. Background Technology

[0002] In video recommendation scenarios, videos where objects have interacted (hereinafter referred to as non-cold start videos) typically have some interaction information captured, thus being recommended to the recommendation list of objects matching that video's interaction information. However, video platforms also add newly published videos (hereinafter referred to as cold start videos). These newly added videos have no interaction records, making it more difficult to extract their interaction characteristics compared to regular videos. Therefore, how to more accurately add cold start videos to the recommendation list of objects has become an important issue in video recommendation.

[0003] In related technologies, since cold start videos do not have corresponding interaction records, all-zero vectors are often used to replace the interaction representation of cold start videos. For non-cold start videos, the normal representation is based on the corresponding interaction records. This difference means that even if the content of the cold start video is more in line with the object's preferences, the cold start video may not be recommended because of this difference, thus affecting the recommendation quality of the recommendation system.

[0004] In conclusion, improving the accuracy of video recommendations during cold starts is an urgent issue that needs to be addressed. Summary of the Invention

[0005] This application provides a training method, apparatus, electronic device, and storage medium for a video recommendation model to improve the accuracy of video recommendations.

[0006] This application provides a method for training a video recommendation model, comprising:

[0007] Each selected training sample is input into the video recommendation model to be trained. Each training sample contains a sample object, a sample video, and a sample label. The sample label represents whether the corresponding sample object interacts with the corresponding sample video.

[0008] Based on the video recommendation model, feature extraction is performed on each selected training sample to obtain corresponding extraction results. Each extraction result includes: the content features and first interaction features of the corresponding sample video, and the object features of the corresponding sample object; and the content features of the sample videos contained in each training sample are respectively mapped to obtain corresponding second interaction features.

[0009] Recommendation predictions are made based on the extraction results and second interaction features corresponding to each training sample to obtain the corresponding prediction results.

[0010] Based on the differences between each prediction result and the corresponding sample label, as well as the differences between the prediction results, the parameters of the video recommendation model are adjusted.

[0011] This application provides a training device for a video recommendation model, comprising:

[0012] The input unit is used to input the selected training samples into the video recommendation model to be trained. Each training sample contains a sample object, a sample video, and a sample label. The sample label represents whether the corresponding sample object interacts with the corresponding sample video.

[0013] The first acquisition unit is used to extract features from each selected training sample based on the video recommendation model to obtain corresponding extraction results. Each extraction result includes: content features and first interaction features of the corresponding sample video, and object features of the corresponding sample object; and to perform feature mapping on the content features corresponding to the sample videos contained in each training sample to obtain corresponding second interaction features.

[0014] The second acquisition unit is used to perform recommendation prediction based on the extraction results and second interaction features corresponding to each training sample, and obtain the corresponding prediction results.

[0015] The adjustment unit is used to adjust the parameters of the video recommendation model based on the differences between each prediction result and the corresponding sample label, as well as the differences between the prediction results.

[0016] Optionally, for a sample video, if the sample video is not a cold start video, then the first interaction feature of the sample video is extracted based on the interaction information with the corresponding sample object; if the sample video is a cold start video, then the first interaction feature of the sample video is obtained by feature mapping based on the corresponding content features.

[0017] Optionally, the second acquisition unit is specifically used for:

[0018] The content features and first interaction features corresponding to each training sample are combined to obtain the corresponding first video features, and the content features and second interaction features corresponding to each training sample are combined to obtain the corresponding second video features.

[0019] Recommendation predictions are made based on the first video features and object features corresponding to each training sample, respectively, to obtain the corresponding first prediction sub-results;

[0020] Recommendation predictions are made based on the second video features and object features corresponding to each training sample, respectively, to obtain the corresponding second prediction sub-results;

[0021] Recommendation predictions are made based on the second interaction features and object features corresponding to each training sample, respectively, to obtain the corresponding third prediction sub-results;

[0022] Based on the first, second, and third prediction sub-results corresponding to each training sample, the corresponding prediction results are obtained.

[0023] Optionally, the adjustment unit is specifically used for:

[0024] Based on the difference between the third prediction sub-result of each training sample and the corresponding sample label, a meta-loss function is constructed;

[0025] Based on the difference between the first and second prediction results of each training sample, a pairwise loss function is constructed.

[0026] A target loss function is constructed based on the meta-loss function and the pairwise loss function, and the parameters of the video recommendation model are adjusted based on the target loss function.

[0027] Optionally, the device further includes:

[0028] The mapping unit is used to map the first video feature and the second video feature corresponding to each training sample based on a pre-configured video mapping matrix before the second acquisition unit makes recommendation predictions based on the extraction results and the second interaction features corresponding to each training sample and obtains the corresponding prediction results.

[0029] Based on the pre-configured object mapping matrix, the object features corresponding to each training sample are mapped respectively;

[0030] The adjustment unit is specifically used for:

[0031] The target loss function is constructed based on the meta-loss function and the pairwise loss function, as well as the video mapping matrix, the object mapping matrix, and the embedding factor for extracting content features.

[0032] Optionally, the training sample set consists of training sample subsets containing different sample objects, with each training sample subset corresponding to the same sample object;

[0033] The input unit is also used to select each training sample in the following manner:

[0034] A subset of training samples is selected from the training sample set, and a main training sample group and an auxiliary training sample group are selected from the training sample subset, wherein the main training sample group and the auxiliary training sample group contain the same number of training samples;

[0035] The input unit is specifically used for:

[0036] The training samples from the main training sample group and the auxiliary training sample group are respectively input into the video recommendation model.

[0037] Optionally, the second acquisition unit is specifically used for:

[0038] Based on the second interaction features and object features corresponding to each main training sample in the main training sample group, recommendation prediction is performed to obtain the third prediction sub-result corresponding to each main training sample.

[0039] Based on the differences between the obtained third prediction sub-results and the corresponding sample labels, the gradient of the interaction representation network in the video recommendation model is updated. The interaction representation network is used to perform feature mapping on the content features corresponding to the sample video to obtain the corresponding second interaction features.

[0040] Based on the second interaction features and object features corresponding to each auxiliary training sample in the auxiliary training sample group, recommendation prediction is performed to obtain the third prediction sub-result corresponding to each auxiliary training sample.

[0041] Optionally, the adjustment unit is specifically used for:

[0042] Based on the differences between the third prediction sub-result and the corresponding sample label of each main training sample in the main training sample group, the main loss function is constructed.

[0043] Based on the differences between the third prediction sub-result and the corresponding sample label of each auxiliary training sample in the auxiliary training sample group, an auxiliary loss function is constructed.

[0044] The primary loss function and the secondary loss function are weighted and summed to obtain the meta-loss function.

[0045] Optionally, the pairwise loss function includes: the pairwise loss function corresponding to the main training sample group and the auxiliary training sample group respectively; each training sample group includes: at least one sample label representing positive samples that interact, and at least one sample label representing negative samples that do not interact;

[0046] The adjustment unit is specifically used to construct the pairwise loss function corresponding to each training sample group in the following manner:

[0047] For a training sample group, a first loss function is constructed based on the difference between the first prediction result of the positive sample and the first prediction result of the negative sample in the training sample group;

[0048] Based on the difference between the second prediction result of positive samples and the first prediction result of negative samples in the training sample group, a second loss function is constructed;

[0049] The pairwise loss function is determined based on the first loss function and the second loss function.

[0050] Optionally, the adjustment unit is specifically used for:

[0051] Based on the difference between the first prediction result of positive samples and the second prediction result of negative samples in the training sample group, and the difference between the second prediction result of positive samples and the second prediction result of negative samples, a third loss function is constructed.

[0052] The pairwise loss function is determined by weighted summation of the first loss function, the second loss function, and the third loss function.

[0053] Optionally, the device further includes:

[0054] The prediction unit is used to input the video to be detected and the object to be detected into the trained video recommendation model;

[0055] Based on the trained video recommendation model, feature extraction is performed on the video to be detected and the object to be detected, respectively, to obtain the target content features and target interaction features corresponding to the video to be detected, and the target object features corresponding to the object to be detected.

[0056] The target content features and target interaction features of the video to be detected are combined to obtain the corresponding target video features;

[0057] Recommendation prediction is performed based on the target video features and the target object features to obtain the corresponding target prediction result. The target prediction result is used to characterize the probability of recommending the video to be detected to the target object.

[0058] An electronic device provided in this application includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of any of the above-described video recommendation model training methods.

[0059] This application provides a computer-readable storage medium including a computer program. When the computer program is run on an electronic device, the computer program is used to cause the electronic device to perform the steps of any of the above-described video recommendation model training methods.

[0060] This application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. When the processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of any of the above-described video recommendation model training methods.

[0061] The beneficial effects of this application are as follows:

[0062] This application provides a training method, apparatus, electronic device, and storage medium for a video recommendation model. Because this application proposes learning second interaction features based on video content features to simulate the interaction information of cold-start videos, it ensures that the interaction features of cold-start videos are not all zeros during video recommendation model training, mitigating the disadvantage caused by the lack of interaction information during training. Simultaneously, this application provides a method to predict whether a style object will interact with a sample video using the first interaction feature, second interaction feature, content feature, and object feature corresponding to each training sample. Furthermore, it adjusts the parameters of the video recommendation model by combining the differences between each prediction result and the corresponding sample label, as well as the differences between the prediction results themselves. Adjusting the model parameters based on the differences between each prediction result and the corresponding sample label effectively improves the accuracy of the second interaction feature learning. Furthermore, adjusting the model parameters based on the differences between each prediction result further reduces the representational differences between non-cold-start videos and cold-start videos. Ultimately, it can improve the disadvantage of cold-start videos in the prediction stage based on the prediction results, effectively improving the accuracy of video recommendations.

[0063] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0064] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0065] Figure 1 This is a schematic diagram illustrating an application scenario of a video recommendation model training method in this application embodiment;

[0066] Figure 2 This is an overall flowchart of a video recommendation model training method in an embodiment of this application;

[0067] Figure 3 This is a schematic diagram of the partitioning structure of a training sample set in an embodiment of this application;

[0068] Figure 4 This is a logical diagram illustrating one method for determining a third prediction result in an embodiment of this application.

[0069] Figure 5A This is a logical diagram illustrating how to determine the first and second prediction sub-results of the main training sample group according to an embodiment of this application.

[0070] Figure 5B This is a logical diagram illustrating how to determine the first and second prediction sub-results of an auxiliary training sample group according to an embodiment of this application.

[0071] Figure 6 This is a logical schematic diagram illustrating the determination of a meta-loss function according to an embodiment of this application;

[0072] Figure 7 This is a logical schematic diagram illustrating one embodiment of the present application for determining the pairwise loss function of the main training sample group;

[0073] Figure 8 This is a module representation diagram in an embodiment of this application;

[0074] Figure 9 This is an overall flowchart of another video recommendation model training method in this application embodiment;

[0075] Figure 10 This is a flowchart illustrating the application process of a detection and segmentation model in an embodiment of this application.

[0076] Figure 11 A flowchart illustrating the practical application of a video recommendation model from one of the embodiments of this application;

[0077] Figure 12 This is a schematic diagram of the composition structure of a video recommendation model training device according to an embodiment of this application;

[0078] Figure 13 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application;

[0079] Figure 14 This is a schematic diagram of the hardware structure of another electronic device using an embodiment of this application. Detailed Implementation

[0080] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0081] The following describes some of the concepts involved in the embodiments of this application.

[0082] Training sample set: This is the set of samples used to train the model. The content of the sample set varies depending on the training objective. In this application, the training objective is to predict whether an object will interact with a video when it is recommended to them. Therefore, the training samples in the training sample set include: sample objects, sample videos, and sample labels representing whether the corresponding sample objects interact with the corresponding sample videos. Furthermore, the training sample set can be divided into multiple training sample subsets based on the different sample objects. Each training sample subset contains the same sample objects. Therefore, during model training, small batches of training samples from the training sample subsets can be used to form training sample groups, such as the main training sample group and auxiliary training sample group in this application.

[0083] Cold start videos: These are videos that have not yet interacted with the target audience due to reasons such as being recently released. These videos do not have corresponding interactive information. In contrast, non-cold start videos are videos that interact with one or more targets and have some interactive information.

[0084] Content features: These describe the content contained in the video. They can include the type of video content, the video's style, editing style, the video author, or specific objects contained in the video. They are obtained by extracting features from the video content using a model. In this embodiment, the content features can be in vector form, and therefore can also be called content feature vectors.

[0085] First interaction feature: If the video is not a cold start video, the first interaction feature of the video is extracted by the model based on the interaction information with the corresponding object; if the video is a cold start video, the first interaction feature of the video is obtained by feature mapping based on the corresponding content features; it reflects the interaction between the video and the object. The interaction information may include one or more of the following: the time the object stays in the video, whether it sends interactive information (comments, bullet comments, etc.) in the video, whether it likes the video (likes, favorites, etc.).

[0086] The second interaction feature is obtained by feature mapping based on the content features corresponding to the video, and is mainly used to simulate the interaction information between the cold start video and the object. The first and second interaction features of the cold start video are the same. However, for non-cold start videos, these videos have already interacted with the object and have interaction information. Therefore, the determination methods of the first and second interaction features of non-cold start videos are different.

[0087] Video features: These are used to represent the video. In this application, video features are divided into first video features and second video features. The first video feature corresponds to the normal representation and is obtained by combining the content features and the first interaction feature corresponding to each training sample. The second video feature corresponds to the cold start representation and is obtained by combining the content features and the second interaction feature corresponding to each training sample. Since the first and second interaction features are the same for cold start videos, the first video features and the second video features are also the same for cold start videos.

[0088] Prediction result: Represents the preference score of the object to the video predicted by the video recommendation model, and can also characterize the probability of the video being recommended to the object, etc.; In this application, the prediction result includes a first prediction sub-result, a second prediction sub-result, and a third prediction sub-result, wherein the first prediction sub-result is obtained by recommendation prediction using the first video feature and object feature corresponding to each training sample; the second prediction sub-result is obtained by recommendation prediction using the second video feature and object feature corresponding to each training sample; and the third prediction result is obtained by recommendation prediction using the second interaction feature and object feature corresponding to each training sample.

[0089] The target loss function, constructed based on the meta-loss function and the pairwise loss function, represents the difference between the model's prediction results and the actual situation. The meta-loss function is derived from the difference between the third prediction sub-result of each sample video and its corresponding sample label. This process simulates the interaction information of cold-start videos using second interaction features, ensuring that the interaction features of cold-start videos are not all zeros. This meta-loss function is used to adjust the parameters of the video recommendation model, mitigating the disadvantage caused by all-zero interaction features in cold-start videos. The pairwise loss function is derived from the difference between the first and second prediction sub-results of each sample video. Adjusting the model parameters using the pairwise loss function combines the differences between prediction results, further reducing the representational differences between non-cold-start videos and cold-start videos, effectively improving the accuracy of video recommendations.

[0090] The embodiments of this application relate to artificial intelligence (AI) and machine learning (ML) technologies, and are designed based on deep learning in artificial intelligence.

[0091] Artificial intelligence (AI) technology mainly includes computer vision, natural language processing, machine learning / deep learning, autonomous driving, and intelligent transportation. With the research and advancement of AI technology, it is being studied and applied in various fields, such as smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, autonomous driving, robotics, and smart healthcare. It is believed that with technological development, AI will be applied in more fields and play an increasingly important role. The training method of the video recommendation model in this application embodiment can be applied to fields such as intelligent customer service, virtual assistants, smart speakers, and robots. By combining AI with recommendation services and preference prediction in these fields, suitable targets can be recommended to different users efficiently and accurately.

[0092] Furthermore, the video recommendation model in this embodiment is trained using machine learning or deep learning techniques. After training the video recommendation model using these techniques, it can be applied to recommend suitable video content to individuals with different preferences.

[0093] The design concept of the embodiments of this application will be briefly described below:

[0094] As people's leisure time becomes increasingly rich, various entertainment applications (APPs) are proliferating; among them, video apps have emerged and become an important part of people's lives. To encourage users to use video apps for an extended period, these apps typically include systems that incorporate video recommendation technology.

[0095] In related technologies, for videos with interactivity (i.e., non-cold start videos), video recommendation systems can capture some interaction information, further determine potential audiences interested in the video based on this information, and recommend the video to a recommendation list of audiences matching the video's interaction information. However, for newly published videos without any interaction records (i.e., cold start videos), it is more difficult to extract interactive features compared to regular videos. Therefore, cold start videos are usually treated by capturing their content features to alleviate the feature representation problem. However, even with cold start videos, there is still only content information, not interaction information. In contrast, non-cold start videos contain both interaction and content information. This difference can put cold start videos at a disadvantage when being recommended, even if their content better matches audience preferences.

[0096] Similarly, when training a video recommendation model, since there is no cold-start interaction data in the training set, the interaction representation of cold-start videos remains in an initial state of all-zero vectors during training. Therefore, cold-start videos can only be represented using content features alone. Ultimately, this results in cold-start videos performing at a disadvantage in prediction scores compared to non-cold-start videos during the testing phase, further putting them at a disadvantage in recommendation.

[0097] In addition, some technologies incorporate visual information into the process of predicting an object's preference for a video, while also considering the impact of the interaction between the object and the video on the prediction of preference, which can alleviate the recommendation effect of cold-start videos to some extent. However, the problem of the interaction representation of cold-start videos remaining in an initial state still exists, causing cold-start videos to be at a disadvantage compared to non-cold-start videos in the prediction during the testing phase.

[0098] In view of this, embodiments of this application provide a training method, apparatus, electronic device, and storage medium for a video recommendation model. This application learns and generates second interaction features based on the content features of the video, using these second interaction features to simulate the interaction information of a cold-start video. This ensures that the interaction features of the cold-start video are not all zeros during video recommendation model training, mitigating the disadvantage caused by all-zero interaction features in the cold-start video. Simultaneously, this application provides a method to predict whether a style object will interact with a sample video using the first interaction features, second interaction features, content features, and object features corresponding to each training sample. Furthermore, by combining the differences between each prediction result and the corresponding sample label, as well as the differences between the prediction results, the parameters of the video recommendation model are adjusted. Adjusting the model parameters based on the differences between each prediction result and the corresponding sample label effectively improves the accuracy of the second interaction feature learning. Furthermore, adjusting the model parameters based on the differences between the prediction results further reduces the representational differences between non-cold-start videos and cold-start videos. Ultimately, this improves the disadvantage of cold-start videos in the prediction stage based on the prediction results, effectively enhancing the accuracy of video recommendations. The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0099] like Figure 1 The diagram shown is an application scenario illustration of an embodiment of this application. The application scenario diagram includes two terminal devices 110 and one server 120.

[0100] In this embodiment, the terminal device 110 includes, but is not limited to, mobile phones, tablets, laptops, desktop computers, e-book readers, smart voice interaction devices, smart home appliances, and in-vehicle terminals. The terminal device may have a client installed related to video recommendation model training. This client can be software (e.g., a browser, video recommendation software), a webpage, or a mini-program. The server 120 is the backend server corresponding to the software, webpage, or mini-program, or a server specifically used for training video recommendation models; this application does not impose specific limitations. The server 120 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0101] It should be noted that the training method of the video recommendation model in each embodiment of this application can be executed by an electronic device, which can be a terminal device 110 or a server 120. That is, the method can be executed by the terminal device 110 or the server 120 alone, or by both the terminal device 110 and the server 120. For example, when executed by the server 120 alone, suppose a short video APP needs to train a video recommendation model. The server 120 inputs the training samples of a pre-set training sample set into the video recommendation model to obtain the content features, object features, first interaction features, and second interaction features of the sample video.

[0102] Subsequently, server 120 combines the content features and first interaction features corresponding to each training sample to obtain the corresponding first video features; it also combines the content features and second interaction features corresponding to each training sample to obtain the corresponding second video features. Further, server 120 obtains the corresponding first prediction sub-result based on the first video features and object features corresponding to each training sample, and obtains the corresponding second prediction sub-result based on the second video features and object features corresponding to each training sample; and obtains a pairwise loss function based on the first and second prediction sub-results.

[0103] In addition, server 120 obtains the third prediction sub-result corresponding to each training sample based on the second interaction feature and object feature corresponding to each training sample; and constructs a meta-loss function based on the difference between the obtained third prediction sub-result and the corresponding sample label.

[0104] Ultimately, server 120 constructs a target loss function based on the meta loss function and the pairwise loss function, in order to adjust the model parameters through the target loss function and improve the accuracy of the video recommendation model.

[0105] In the model application phase, the terminal device 110 acquires the object features of the target object and sends these features to the server 120. Upon receiving the object features, the server acquires target content features and target interaction features from the video library, and combines these features to obtain target video features. Then, the server 120 performs recommendation prediction based on the target video features and target object features to obtain the target prediction result for the target video. This prediction result can be a number, representing the model's prediction of the target object's level of interest in the target video. Finally, based on the prediction result, the server 120 selects videos that the target object may be interested in. Specifically, an interest threshold can be preset. For videos exceeding this threshold, the server 120 sends one or more to the client 110 for display, or sorts the prediction results for these videos, recommends a certain number of videos in a specific order, and sends the filtered video results back to the client 110, which then displays the recommendations to the target object. This paper does not impose specific limitations on this approach.

[0106] In one alternative implementation, the terminal device 110 and the server 120 can communicate via a communication network.

[0107] In one alternative implementation, the communication network is a wired network or a wireless network.

[0108] It should be noted that, Figure 1 The examples shown are merely illustrative; in reality, the number of terminal devices and servers is unlimited and is not specifically limited in the embodiments of this application.

[0109] In this embodiment of the application, when there are multiple servers, the multiple servers can form a blockchain, and the servers are nodes on the blockchain; as in the video recommendation model training method disclosed in this embodiment of the application, the model training data involved can be stored on the blockchain, such as the content features, object features, first interaction features, second interaction features of the video to be detected, and embedding factors used to extract content features, etc.

[0110] Furthermore, the embodiments of this application can be applied to various scenarios, including not only video recommendation scenarios, but also scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0111] The following describes the training method of the video recommendation model provided by the exemplary embodiments of this application, in conjunction with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way in this respect.

[0112] See Figure 2 The diagram shows a flowchart of a training method for a video recommendation model provided in this application. The specific implementation process of this method is as follows: S201-S204:

[0113] S201: Input the selected training samples into the video recommendation model to be trained.

[0114] Each training sample contains a sample object, a sample video, and a sample label. The sample label indicates whether the corresponding sample object interacts with the corresponding sample video.

[0115] In the embodiments of this application, the interaction of a sample object (such as a user) with the sample video includes, but is not limited to, some or all of the following: the sample object clicks on the sample video to watch it, the time spent in the sample video, whether to send interactive information in the sample video, whether to click "like" on the sample video, etc. This article uses the example of a sample object clicking on the sample video to watch it for illustration.

[0116] In the embodiments of this application, when training the video recommendation model, it is necessary to perform multiple rounds of iterative training on the video recommendation model to be trained based on the training sample set. In each round of iterative training, the following steps S201-S204 can be executed to adjust the parameters of the video recommendation model.

[0117] The training sample set includes both cold start videos and non-cold start videos. During each iteration, training samples need to be selected in batches from the training sample set.

[0118] An optional implementation involves the training sample set being composed of subsets of training samples containing different sample objects. Each subset of training samples corresponds to the same sample object. Each training sample is selected as follows: a subset of training samples is selected from the training sample set, and two training sample groups, i.e., two small batches of training samples, are selected from the subset. These groups are named the main training sample group and the auxiliary training sample group according to their input order. The main training sample group and the auxiliary training sample group contain the same number of training samples.

[0119] In summary, inputting the selected training samples into the video recommendation model to be trained involves inputting the training samples from the main training sample group and the auxiliary training sample group into the video recommendation model, respectively. Specifically, first, the training samples from the main training sample group are input into the video recommendation model; then, the training samples from the auxiliary training sample group are input into the video recommendation model.

[0120] Taking a specific scenario as an example, suppose we are training a video recommendation model, such as... Figure 3 The diagram shown illustrates a training sample set partitioning structure in an embodiment of this application. The server divides the training sample set D into several tasks, i.e., multiple training sample subsets D1, D2, D3, D4, D5, D6, D7, D8, D9, D1, D1, D2 ... i ...with D i For example, the training sample subset D i The task is to train a subset D of samples based on sample objects i. i There is N i One training sample; further, a subset D of the training samples. i It contains multiple training sample groups, from which two disjoint training sample groups can be selected, namely the main training sample group. and auxiliary training sample group Similarly, the training sample subset D1 also contains multiple training sample groups, from which two disjoint training sample groups can be selected, namely the main training sample group. and auxiliary training sample group The training sample subset D2 also contains multiple training sample groups, from which two disjoint training sample groups can be selected, namely the main training sample group. and auxiliary training sample group

[0121] In this embodiment of the application, each training sample subset may contain multiple training sample groups, and each training sample group contains K samples, where K << N. i / 2, during one iteration, one of these multiple training sample groups can be selected as the main training sample group. Select one as the auxiliary training sample group Then, the server will distribute the main training sample groups respectively. and auxiliary training sample group The training samples in the video recommendation model are used to train the model.

[0122] It should be noted that the above-mentioned methods for selecting training samples are merely illustrative examples, and any selection method is applicable to the embodiments of this application, without any specific limitation.

[0123] S202: Based on the video recommendation model, feature extraction is performed on each selected training sample to obtain the corresponding extraction results, and feature mapping is performed on the content features corresponding to the sample videos contained in each training sample to obtain the corresponding second interaction features.

[0124] Each extraction result includes: the content features and first interaction features of the corresponding sample video, and the object features of the corresponding sample object. For a sample video, if it is not a cold start video, the first interaction feature is extracted based on the interaction information with the corresponding sample object; if it is a cold start video, the first interaction feature is obtained by feature mapping based on the corresponding content features, meaning the first interaction feature and the second interaction feature of a cold start video are the same.

[0125] Taking the hypothetical scenario in S201 as an example, for a sample video j, the content of the sample video is first represented as a one-hot vector t. j This leads to the one-hot vector t j Convert to an embedding vector c j That is, the content feature is represented as c j There is c j =W×t j Where W is the mapping matrix, and the first interactive feature is represented as e j For a sample object i, its object features are represented as u. i In the video recommendation model, the model parameters used to obtain content features and object features are pre-trained with a large amount of data and do not need to be updated.

[0126] The second interaction feature is generated by the interaction representation network in the video recommendation model of this application, through the content features c corresponding to the sample video. j The feature mapping is used to obtain the following representation: and Here, w represents the parameters of the interactive representation network, which can be composed of fully connected layers, i.e., w represents the parameters of the fully connected layers.

[0127] S203: Make recommendation predictions based on the extraction results and second interaction features of each training sample to obtain the corresponding prediction results.

[0128] The prediction result represents the object's preference score for the video as predicted by the video recommendation model. It can also represent the probability that the video will be recommended to the object. The following text uses the preference score of each object for the video as an example.

[0129] In this embodiment of the application, the server can first perform recommendation prediction based on the second interaction features and object features corresponding to each training sample to obtain the corresponding third prediction sub-result.

[0130] Specifically, such as Figure 4 The diagram illustrates a logical representation of determining a third prediction sub-result according to an embodiment of this application. For each subset of training samples, the server first trains a video recommendation model based on a main training sample group (containing K training samples, such as training samples 1-10 when K=10). Based on the second interaction features of each sample video and the object features of the sample objects in the main training sample group, the server obtains the corresponding third prediction sub-result. Based on the difference between the obtained third prediction sub-results and the corresponding sample labels, the server performs gradient updates on the interaction representation network responsible for obtaining the second interaction features in the video recommendation model. On this basis, the server then trains the video recommendation model that has undergone gradient updates again based on an auxiliary training sample group (containing K training samples, such as training samples 11-20 when K=10). Based on the object features of the sample objects in an auxiliary training sample group and the second interaction features of each sample video obtained after gradient updates, the server obtains the third prediction sub-result corresponding to each auxiliary training sample.

[0131] Then, the server combines the content features and the first interaction features corresponding to each training sample to obtain the corresponding first video features, and combines the content features and the second interaction features corresponding to each training sample to obtain the corresponding second video features.

[0132] In the above, the server provides two feature representations for each video: a first video feature and a second video feature. This addresses the problem caused by the significant difference in feature representations between the training and testing phases. Specifically, non-cold-start videos are represented using the first video feature, while cold-start videos are represented using the second video feature. The second video feature can simulate cold-start scenarios in the test set.

[0133] Optionally, the server maps the first video feature and the second video feature corresponding to each training sample based on a pre-configured video mapping matrix, and maps the object feature corresponding to each training sample based on a pre-configured object mapping matrix, so that the first video feature, the second video feature, and the object feature have the same dimension. The dimensions of the first video feature, the second video feature, and the object feature can all be adjusted to one of the three dimensions, or they can be adjusted to a specific dimension, etc. Specific adjustment methods can be obtained through experiments, and this paper does not impose specific limitations.

[0134] The parameters in each of the mapping matrices mentioned above are also adjustable.

[0135] Furthermore, for each training sample subset, each training sample group includes at least one sample label representing positive samples with interaction and at least one sample label representing negative samples without interaction. Based on the positive and negative samples, and the two video representation methods proposed in this paper (normal representation and cold start representation), pairwise learning can be performed. Cold start videos use the cold start representation, and non-cold start videos use the normal representation. The specific process is as follows:

[0136] Based on the above, the training videos can be divided into the following four categories according to whether each training sample is a positive or negative sample, and whether the video samples contained in each training sample are represented in normal or cold start mode:

[0137] Positive samples represented by cold start, positive samples represented by normal, negative samples represented by cold start, and negative samples represented by normal.

[0138] like Figure 5A The diagram illustrates a logical representation of an embodiment of this application, showing how to determine the first and second prediction sub-results of a main training sample group. The server first performs recommendation prediction based on the first video features and object features corresponding to each positive sample video using normal representation in the main training sample group, obtaining the first prediction sub-result for each positive sample video that is not a cold start video. Then, based on the second video features and object features corresponding to each positive sample video using cold start representation in the main training sample group, it performs recommendation prediction to obtain the second prediction sub-result for each positive sample video that is a cold start video. Finally, based on the first video features and object features corresponding to each negative sample video using normal representation in the main training sample group, it performs recommendation prediction to obtain the first prediction sub-result for each negative sample video that is not a cold start video. And finally, based on the second video features and object features corresponding to each negative sample video using cold start representation in the main training sample group, it performs recommendation prediction to obtain the second prediction sub-result for each negative sample video that is a cold start video.

[0139] After that, as Figure 5B The diagram illustrates a logical representation of an embodiment of this application, showing how to determine the first and second prediction sub-results of an auxiliary training sample group. The server, based on a video recommendation model that has undergone gradient updates to the interactive representation network, performs the aforementioned process on the auxiliary training sample group within the training sample subset to obtain the first prediction sub-results for each positive sample video using normal representation, the second prediction sub-results for each positive sample video using cold-start representation, and the first prediction sub-results for each negative sample video using normal representation, and the second prediction sub-results for each negative sample video using cold-start representation.

[0140] That is, in the above, the second interaction feature of the training sample in the auxiliary training sample group is obtained based on the interaction representation network that has undergone gradient update. The content features corresponding to each training sample in the auxiliary training sample group are combined with the second interaction feature obtained based on the interaction representation network that has undergone gradient update to obtain the second video feature corresponding to the training sample in the auxiliary training sample group.

[0141] Finally, the server will combine the first, second, and third prediction sub-results corresponding to each training sample to form the prediction result in S203.

[0142] It should be noted that the processes of obtaining the first, second, and third prediction sub-results can be performed in parallel or sequentially, and this application does not impose any specific limitations on them.

[0143] Continuing with the hypothetical scenario in S202, for each training sample subset D i , N i D represents the training sample subset i There is N i There are training samples, and the vector v is... j Based on content feature c j and the first interaction feature e j Composition, v j =δ(e j ,c j ), δ(·) represents a combinatorial function, which is generally implemented by connecting δ(e) i ,c i ) = e i ||c i ; This represents the preference score predicted by the video recommendation model for an object in relation to a video. Where F(·) represents the interaction function.

[0144] The server processes each training sample subset D i The main training sample group In this process, for each sample video, the following procedure is performed: For sample video j1, based on the second interaction features of j1... Object features u of sample object i i The third prediction result for the corresponding sample video j1 is obtained. aj1 means The j1-th sample video is used, where θ represents the prediction model parameters, which are pre-trained and do not need further updating. Then, based on the differences between the obtained third prediction sub-results and the corresponding sample labels, the server performs gradient updates on the interaction representation network responsible for obtaining the second interaction features in the video recommendation model. On this basis, the server further updates the gradients based on the auxiliary training sample group in the training sample subset. The video recommendation model that has undergone gradient updates is retrained, and the auxiliary training sample group is used. In this process, each sample video performs the following procedure: based on the object features u of sample object i... i and the second interaction features of sample video j2 obtained after gradient update. Obtain the third prediction result corresponding to the auxiliary training sample j2 bj2 means The j2nd sample video in the dataset.

[0145] Then, the server processes each training sample and assigns its corresponding content features c. j and the first interaction feature e j Combine them to obtain the corresponding first video features. δ(·) represents the combination function; and the content features c corresponding to each training sample are respectively... j Second interaction features By combining the features, the corresponding second video features can be obtained.

[0146] Optionally, the server can also be based on a pre-configured video mapping matrix W v Map the first video features and second video features corresponding to each training sample respectively, and then... and Mapped to a common space, there are and And based on the pre-configured object mapping matrix W u Each training sample's corresponding object features are mapped to a common space. This ensures that the first video feature, the second video feature, and the object feature have the same dimension.

[0147] In the above, Video feature representation for non-cold start videos. The video feature representation used for cold start videos is that non-cold start videos use the first video feature representation, while cold start videos use the second video feature representation. This allows for optimization during training to compensate for the representational disadvantage of cold start videos, further reducing the disadvantage and difference between cold start videos and non-cold start videos in prediction on the test set.

[0148] Furthermore, for each training sample subset D i In this context, let g represent a positive sample and k represent a negative sample. This represents the first video feature corresponding to a positive sample video using the normal representation. This represents the second video feature corresponding to the positive sample video using cold start representation. This represents the first video feature corresponding to the negative sample video represented using the normal representation. This represents the second video feature corresponding to the negative sample video using cold start representation. These are the mapped object features. The server first processes the main training sample group... In the middle, the positive samples in each video using normal representation correspond to the first video feature. and object characteristics Perform recommendation predictions to obtain the first prediction sub-result for positive samples in each video using normal representation. Based on the main training sample group In the middle, the positive samples in each video using cold start representation correspond to the second video features. and object characteristics Perform recommendation predictions to obtain the second prediction sub-results for positive samples in each video using the cold start representation. Based on the main training sample group In the middle, the negative samples in each video using normal representation correspond to the first video features. and object characteristics Perform recommendation predictions to obtain the first prediction sub-result for negative samples in each video using normal representation. and based on the main training sample group In the middle, the negative samples in each video using cold start representation correspond to the second video features. and object characteristics Perform recommendation predictions to obtain the second prediction sub-results for negative samples in each video represented by cold start.

[0149] Subsequently, the server, based on the video recommendation model that has undergone gradient updates to the interaction representation network, performs the same process on the auxiliary training sample group in the training sample subset to obtain the first prediction sub-result for positive samples in each video using normal representation. Second prediction sub-results for positive samples in each video using cold start representation And the first prediction sub-results of negative samples in each video using normal representation. Second prediction sub-results of negative samples in each video using cold start representation

[0150] Furthermore, for all training samples in the training sample set, the click records of sample objects on them can be represented as S = {(v j1 u i1 ), (v j2 u i2 ), ..., (v jNs u iNs )}, where Ns is the total number of click records. A matrix can be used. This represents the interaction between objects and videos, where M represents the number of objects, N represents the number of videos, and r is the value in the i-th row and j-th column of R. ij When r ij When r = 1, it indicates that object i interacts with video j; otherwise, there is no interaction. Using R = {(i, j) | r ij =1} represents the interactive dataset. Bayesian Personalized Ranking (BPR) can be used to optimize the prediction function, trained using a triplet dataset, with the training sample set... Where i represents object i, which has positive sample video g with interaction and negative sample video k without interaction.

[0151] It is understood that in the specific implementation of this application, data such as the click record information of the sample object on the sample video are involved. When the above embodiments of this application are applied to specific products or technologies, permission or consent from the object is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0152] Optionally, the above mapping process can be implemented based on a system containing a simulation module.

[0153] S204: Adjust the parameters of the video recommendation model based on the differences between each prediction result and the corresponding sample label, as well as the differences between each prediction result.

[0154] Specifically, taking a pair of main training sample groups and auxiliary training sample groups as an example, such as Figure 6The diagram illustrates a logical representation of determining the meta-loss function according to an embodiment of this application. The server constructs the main loss function for the main training sample group (containing K training samples, e.g., when K=10, including training samples 1-10), based on the differences between the third prediction sub-results and corresponding sample labels for each main training sample. It also constructs an auxiliary loss function for the auxiliary training sample group (containing K training samples, e.g., when K=10, including training samples 11-20), based on the differences between the third prediction sub-results and corresponding sample labels for each auxiliary training sample. The main loss function and the auxiliary loss function are then weighted and summed to obtain the meta-loss function.

[0155] Subsequently, the server constructs a pairwise loss function based on the difference between the first and second prediction sub-results of each training sample. The pairwise loss function can be further divided into pairwise loss functions corresponding to the main training sample group and the auxiliary training sample group. Specifically, the pairwise loss function corresponding to each training sample group in the main training sample group and the auxiliary training sample group is constructed in the following way:

[0156] like Figure 7 The diagram illustrates a logical representation of determining the pairwise loss function for the main training sample group according to an embodiment of this application. For the main training sample group, the server constructs a first loss function based on the difference between the first predicted sub-result of a positive sample and the first predicted sub-result of a negative sample; it constructs a second loss function based on the difference between the second predicted sub-result of a positive sample and the first predicted sub-result of a negative sample; then, based on the first and second loss functions, the pairwise loss function is determined. Furthermore, a third loss function can be constructed based on the differences between the first predicted sub-result of a positive sample and the second predicted sub-result of a negative sample, as well as the differences between the second predicted sub-result of a positive sample and the second predicted sub-result of a negative sample. Then, a weighted sum is performed based on the first, second, and third loss functions to determine the pairwise loss function for each training sample group in the main training sample group.

[0157] The method for determining the pairwise loss function corresponding to the auxiliary training sample group is the same as that for the main training sample group. However, it should be noted that in the auxiliary training sample group, the second interaction feature of the training sample is obtained based on the video recommendation model that has undergone gradient updates to the interaction representation network; the second video feature corresponding to the training sample is obtained by combining its content feature and the second interaction feature obtained based on the video recommendation model that has undergone gradient updates to the interaction representation network; therefore, its second prediction sub-result is also obtained by performing recommendation prediction on the second video feature obtained after gradient updates and the object feature mentioned above.

[0158] Finally, the server constructs a target loss function based on the meta-loss function, pairwise loss function, video mapping matrix, object mapping matrix, and embedding factors for extracting content features. The parameters of the video recommendation model are then adjusted based on this target loss function.

[0159] Continuing with the hypothetical scenario in S203, for a main training sample group, its main loss function is:

[0160]

[0161] in, The third prediction result corresponding to the main training sample j1 has a value between 0 and 1, y aj1 The label represents the sample label, indicating whether the sample object actually interacts with the sample video; if it does, it is 1, otherwise it is 0.

[0162] Furthermore, the gradient update in S202 and S203 specifically involves: calculating l a right The gradient of the interaction representation network is used to update the gradient. After the update, for an auxiliary training sample j2, its second interaction feature is updated as follows: Where h is the learning rate.

[0163] Similarly, for an auxiliary training sample set, its auxiliary loss function is:

[0164]

[0165] The main loss function and the auxiliary loss function are weighted and summed to obtain the meta-loss function l. meta =αl a +(1-α)l b .

[0166] Where α is the correlation coefficient between the main loss function and the auxiliary loss function. l is calculated using the gradient descent method. meta The gradient of the fully connected layer parameter w is used to optimize the parameter w.

[0167] The learning of the second interactive feature and the acquisition of the meta-loss function mentioned above can be implemented based on a system that includes a meta-embedding module.

[0168] Then, the server constructs a pairwise loss function based on the difference between the first and second prediction results for each training sample:

[0169] For a main training sample group The first loss function is constructed by the difference between the first prediction result of the positive sample g and the first prediction result of the negative sample k. Where σ(·) is the sigmoid function, and L1(i,g,k) is used for learning the non-cold-start video ranking task.

[0170] A second loss function is constructed based on the difference between the second predictor result of the positive sample g and the first predictor result of the negative sample k. The role of L2(i,j,k) is to make the representation of positive sample cold start videos more advantageous than the representation of negative sample non-cold start videos, which is the key to reducing the disadvantage of cold start videos.

[0171] Then, based on the first loss function and the second loss function, the pairwise loss function is determined.

[0172] For example, the pairwise loss function between a positive sample g and a negative sample k corresponding to sample object i is: L(i,g,k)=L1(i,g,k)+L2(i,g,k).

[0173] Furthermore, a third loss function can be constructed by using the differences between the first prediction result of positive samples and the second prediction result of negative samples, as well as the differences between the second prediction result of positive samples and the second prediction result of negative samples. L3(i,g,k) represents the positive sample representation of the cold start video and the positive sample representation of the non-cold start video. Overall, it should be better than the negative sample representation of the cold start video. However, since the reason why the cold start video becomes a negative sample may be due to the disadvantage caused by the lack of interactive data, the negative sample representation of the cold start video should not be too affected. Therefore, the influence of L3 on the overall loss function is reduced by setting weights. Then, the pairwise loss function between a positive sample g and a negative sample k corresponding to sample object i is:

[0174] L(i,g,k)=β(L1(i,g,k)+L2(i,g,k))+(1-β)L3(i,g,k); where β is the correlation coefficient, and β is set to >0.5.

[0175] Similarly, for an auxiliary training sample group Its first loss function Second loss function Third loss function Furthermore, pairwise loss function

[0176] In the above, the pairwise loss function can be obtained based on a system containing a pairwise learning module. The objective function of the pairwise learning module is: It is the sum of the pairwise loss functions of all training samples in the training sample set D.

[0177] Specifically, in each iteration, training is not based on the entire training sample set D, but rather on a main training sample set and an auxiliary training sample set selected from D. Therefore, during each iteration, for each pair of main and auxiliary training sample sets, the server uses the meta-loss function, pairwise loss function, and video mapping matrix Θ. q Object mapping matrix Θ w and embedding factor Θ used to extract content features e We construct a target loss function and adjust the parameters of the video recommendation model based on the target loss function.

[0178] With the main training sample group and auxiliary training sample group For example, its objective loss function is:

[0179]

[0180]

[0181] in, This indicates the main training sample group. The average of the pairwise loss functions of all training samples; Represented as auxiliary training sample group The average of the pairwise loss functions of all training samples.

[0182] Optionally, the above process can be implemented based on a system containing a training module. The training module is responsible for training and updating the parameters in the meta-embedding module and the pairwise learning module framework together, and training is performed iteratively with a pair of main training sample groups and auxiliary training sample groups each time.

[0183] It is important to note that during the testing phase of the video recommendation model, when obtaining prediction results for samples, whether it is a cold start video or a non-cold start video, the first video feature is used for prediction. This is because the first and second video features of a cold start video are represented by the same feature. For non-cold start items, the first video feature obtained from their actual interaction data can be used directly for prediction.

[0184] In summary, the method described in this application can be implemented by a system comprising a meta-embedding module, a simulation module, a pairwise learning module, and a training module, such as... Figure 8The diagram shown is a module representation diagram in an embodiment of this application. After obtaining the content features, first interaction features, object features, and second interaction features of the training samples, the meta-embedding module is responsible for calculating the meta loss function of the training samples, the simulation module is responsible for mapping, the pairwise learning module is responsible for calculating the pairwise loss function of the training samples, and the training module is responsible for combining the results of the meta-embedding module and the pairwise loss module to obtain the final target loss function. Based on the target loss function, the parameters of the video recommendation model are updated, that is, the parameters in the meta-embedding module and the pairwise learning module are trained and updated together.

[0185] Furthermore, for the specific scenarios mentioned above, taking a pair of main training sample groups and auxiliary training sample groups as an example, one possible implementation method is as follows: Figure 9 The flowchart shown implements S201-S204, including the following steps:

[0186] Step 901: In the training sample subset D i In this process, a main training sample group and an auxiliary training sample group are selected.

[0187] Step 902: Input the training samples from the main training sample group and the auxiliary training sample group into the video recommendation model.

[0188] Step 903: Obtain the content features, object features, first interaction features, and second interaction features of the training samples.

[0189] Step 904: Based on the second interaction features and object features corresponding to each main training sample in the main training sample group, obtain the third prediction sub-result corresponding to each main training sample.

[0190] Step 905: Construct the main loss function based on the differences between the obtained third prediction sub-results and the corresponding sample labels.

[0191] Step 906: Based on the main loss function, perform gradient updates on the interaction representation network in the video recommendation model.

[0192] Step 907: Based on the object features and second interaction features corresponding to each auxiliary training sample in the auxiliary training sample group, obtain the third prediction sub-result corresponding to each auxiliary training sample.

[0193] The second interactive feature of the auxiliary training sample is obtained after gradient update in step 906.

[0194] Step 908: Construct an auxiliary loss function based on the difference between the third prediction result and the corresponding sample label.

[0195] Step 909: Obtain the meta-loss function based on the main loss function and the auxiliary loss function.

[0196] Step 910: Combine the content features and first interaction features corresponding to each training sample to obtain the corresponding first video features, and combine the content features and second interaction features corresponding to each training sample to obtain the corresponding second video features.

[0197] Step 911: Map the first video features and second video features of each training sample and map the object features.

[0198] Step 912: Based on the first video features and object features corresponding to each training sample, obtain the corresponding first prediction sub-result, and based on the second video features and object features corresponding to each training sample, obtain the corresponding second prediction sub-result.

[0199] Step 913: Based on the difference between the first prediction result of the positive sample and the first prediction result of the negative sample in the main training sample group, construct the first loss function of the main training sample group; based on the difference between the second prediction result of the positive sample and the first prediction result of the negative sample, construct the second loss function of the main training sample group; based on the difference between the first prediction result of the positive sample and the second prediction result of the negative sample, and the difference between the second prediction result of the positive sample and the second prediction result of the negative sample, construct the third loss function of the main training sample group.

[0200] Step 914: Based on the difference between the first prediction result of the positive sample and the first prediction result of the negative sample in the auxiliary training sample group, construct the first loss function of the auxiliary training sample group; based on the difference between the second prediction result of the positive sample and the first prediction result of the negative sample, construct the second loss function of the auxiliary training sample group; based on the difference between the first prediction result of the positive sample and the second prediction result of the negative sample, and the difference between the second prediction result of the positive sample and the second prediction result of the negative sample, construct the third loss function of the auxiliary training sample group.

[0201] It should be noted that the second interaction feature, the second video feature, and the second prediction sub-result of the auxiliary training sample are all obtained after gradient update in step 906.

[0202] Step 915: Obtain the pairwise loss function from the first loss function, the second loss function, and the third loss function.

[0203] Step 916: Construct the target loss function based on the meta-loss function, pairwise loss function, video mapping matrix, object mapping matrix, and embedding factor for extracting content features.

[0204] The training method of the detection and segmentation model in this application embodiment is further described below from the perspective of model application:

[0205] See Figure 10 The diagram shown is a flowchart illustrating the application process of a detection and segmentation model in this embodiment of the application. Taking the server as the execution entity as an example, the specific implementation process of this method is as follows:

[0206] S1001: Input the video to be detected and the object to be detected into the trained video recommendation model.

[0207] S1002: Based on the trained video recommendation model, feature extraction is performed on the video to be detected and the object to be detected, respectively, to obtain the target content features and target interaction features corresponding to the video to be detected, as well as the target object features corresponding to the object to be detected.

[0208] Among them, the target content features are used to describe the content contained in the video to be detected. These can be the type of video content, the content style, editing style, video author, or specific objects contained in the video, etc., and are obtained by the model through feature extraction of the video content. The target interaction features are divided into first interaction features and second interaction features. For the first interaction features, if the video is not a cold start video, the first interaction features of the video are extracted by the model based on the interaction information with the corresponding object. If the video is a cold start video, the first interaction features of the video are obtained by feature mapping based on the corresponding content features. The second interaction features are obtained by feature mapping based on the corresponding content features of the video, and are mainly used to simulate the interaction information between the cold start video and the object. The first interaction features and the second interaction features of the cold start video are the same.

[0209] S1003: Combine the target content features and target interaction features of the video to be detected to obtain the corresponding target video features.

[0210] Here, the target video feature is the first video feature of the video to be detected. Because the first and second video features of a cold-start video are represented identically during the testing or application phase, non-cold-start items can directly use the first video feature obtained from their actual interaction data for prediction. The prediction score of object i for video j is:

[0211] S1004: Make recommendations and predictions based on the features of the target video and the features of the target object to obtain the corresponding target prediction results.

[0212] The target prediction result is used to characterize the probability of recommending the video to be detected to the target object.

[0213] After training and deployment, a well-trained video recommendation model can more accurately recommend videos that users might be interested in, based on different user groups. Figure 11 The diagram shows a flowchart of a video recommendation model in practical application according to an embodiment of this application. The server obtains the object feature information of object i and predicts the videos in the video library based on the video recommendation model. The server obtains the content features, first interaction features, and second interaction features of the video, as well as the object features, and performs preference prediction on the video to obtain the prediction result. Among them, the prediction result of A is 0.1, the prediction result of B is 0.2, the prediction result of C is 0.1, D is 0.9, E is 0.87, and so on. The numbers in the prediction result represent the object's interest in the corresponding video, with values ​​between 0 and 1, where 0 represents no interest. An interest threshold can also be preset. For videos that exceed the interest threshold, one or more are recommended to object i for display on the client, or the prediction results of these videos are sorted and a certain number of videos are recommended in a certain order. This paper does not make specific limitations. Assuming the interest threshold is 0.6, the server sends video D and video E to the client, and the client displays the recommendations to object i.

[0214] Based on the same inventive concept, embodiments of this application also provide a training device for a video recommendation model. For example... Figure 12 As shown, this is a schematic diagram of the structure of the training device 1200 for the video recommendation model, which may include:

[0215] The input unit 1201 is used to input the selected training samples into the video recommendation model to be trained. Each training sample contains a sample object, a sample video, and a sample label. The sample label represents whether the corresponding sample object interacts with the corresponding sample video.

[0216] The first acquisition unit 1202 is used to extract features from each selected training sample based on the video recommendation model to obtain corresponding extraction results. Each extraction result includes: the content features and first interaction features of the corresponding sample video, and the object features of the corresponding sample object; and to perform feature mapping on the content features corresponding to the sample video contained in each training sample to obtain the corresponding second interaction features.

[0217] The second acquisition unit 1203 is used to make recommendation predictions based on the extraction results and second interaction features of each training sample respectively, and obtain the corresponding prediction results.

[0218] The adjustment unit 1204 is used to adjust the parameters of the video recommendation model based on the differences between each prediction result and the corresponding sample label, as well as the differences between each prediction result.

[0219] Optionally, for a sample video, if the sample video is not a cold start video, the first interaction feature of the sample video is extracted based on the interaction information with the corresponding sample object; if the sample video is a cold start video, the first interaction feature of the sample video is obtained by feature mapping based on the corresponding content features.

[0220] Optionally, the second acquisition unit 1203 is specifically used for:

[0221] The content features and first interaction features corresponding to each training sample are combined to obtain the corresponding first video features, and the content features and second interaction features corresponding to each training sample are combined to obtain the corresponding second video features.

[0222] Recommendation predictions are made based on the first video features and object features corresponding to each training sample, respectively, to obtain the corresponding first prediction sub-results;

[0223] Recommendation predictions are made based on the second video features and object features corresponding to each training sample, and the corresponding second prediction sub-results are obtained.

[0224] Recommendation predictions are made based on the second interaction features and object features corresponding to each training sample, and the corresponding third prediction sub-results are obtained.

[0225] Based on the first, second, and third prediction sub-results corresponding to each training sample, the corresponding prediction results are obtained.

[0226] Optionally, the adjustment unit 1204 is specifically used for:

[0227] Based on the difference between the third prediction result and the corresponding sample label of each training sample, a meta-loss function is constructed;

[0228] Based on the difference between the first and second prediction results of each training sample, a pairwise loss function is constructed.

[0229] A target loss function is constructed based on the meta-loss function and the pairwise loss function, and the parameters of the video recommendation model are adjusted based on the target loss function.

[0230] Optionally, the device also includes:

[0231] The mapping unit 1205 is used to map the first video feature and the second video feature corresponding to each training sample based on a pre-configured video mapping matrix before making recommendation predictions based on the extraction results and second interaction features corresponding to each training sample and obtaining the corresponding prediction results.

[0232] Based on the pre-configured object mapping matrix, the object features corresponding to each training sample are mapped respectively;

[0233] Adjustment unit 1204 is specifically used for:

[0234] A target loss function is constructed based on the meta-loss function and pairwise loss function, as well as the video mapping matrix, object mapping matrix, and embedding factors for extracting content features.

[0235] Optionally, the training sample set consists of training sample subsets containing different sample objects, with each training sample subset corresponding to the same sample object;

[0236] The input unit is also used to select training samples in the following manner:

[0237] Select a subset of training samples from the training sample set, and then select a main training sample group and an auxiliary training sample group from the training sample subset. The main training sample group and the auxiliary training sample group contain the same number of training samples.

[0238] Input unit 1201 is specifically used for:

[0239] The training samples from the main training sample group and the auxiliary training sample group are respectively input into the video recommendation model.

[0240] Optionally, the second acquisition unit 1203 is specifically used for:

[0241] Based on the second interaction features and object features corresponding to each main training sample in the main training sample group, recommendation prediction is performed to obtain the third prediction sub-result corresponding to each main training sample.

[0242] Based on the differences between the obtained third prediction sub-results and the corresponding sample labels, the gradient of the interaction representation network in the video recommendation model is updated. The interaction representation network is used to perform feature mapping on the content features corresponding to the sample video to obtain the corresponding second interaction features.

[0243] Recommendation prediction is performed based on the second interaction features and object features corresponding to each auxiliary training sample in the auxiliary training sample group, and the third prediction sub-result corresponding to each auxiliary training sample is obtained.

[0244] Optionally, the adjustment unit 1204 is specifically used for:

[0245] Based on the differences between the third prediction sub-result and the corresponding sample label of each main training sample in the main training sample group, the main loss function is constructed.

[0246] Based on the differences between the third prediction sub-result and the corresponding sample label of each auxiliary training sample in the auxiliary training sample group, an auxiliary loss function is constructed.

[0247] The meta-loss function is obtained by weighted summation of the main loss function and the auxiliary loss function.

[0248] Optionally, the pairwise loss function includes: the pairwise loss function corresponding to the main training sample group and the auxiliary training sample group respectively; each training sample group includes: at least one sample label representing the positive sample with interaction, and at least one sample label representing the negative sample without interaction;

[0249] The adjustment unit 1204 is specifically used to construct the pairwise loss function corresponding to each training sample group in the following manner:

[0250] For a training sample group, a first loss function is constructed based on the difference between the first prediction result of the positive sample and the first prediction result of the negative sample in the training sample group;

[0251] A second loss function is constructed based on the difference between the second prediction result of positive samples and the first prediction result of negative samples in a training sample set.

[0252] Based on the first loss function and the second loss function, the pairwise loss function is determined.

[0253] Optionally, the adjustment unit 1204 is specifically used for:

[0254] A third loss function is constructed based on the difference between the first prediction result of positive samples and the second prediction result of negative samples in a training sample group, as well as the difference between the second prediction result of positive samples and the second prediction result of negative samples.

[0255] The pairwise loss function is determined by weighted summation of the first loss function, the second loss function, and the third loss function.

[0256] Optionally, the device also includes:

[0257] Prediction unit 1206 is used to input the video to be detected and the object to be detected into the trained video recommendation model;

[0258] Based on the trained video recommendation model, feature extraction is performed on the video to be detected and the object to be detected, respectively, to obtain the target content features and target interaction features corresponding to the video to be detected, and the target object features corresponding to the object to be detected.

[0259] The target content features and target interaction features of the video to be detected are combined to obtain the corresponding target video features;

[0260] Recommendation prediction is performed based on the features of the target video and the target object to obtain the corresponding target prediction results. The target prediction results are used to characterize the probability of recommending the video to be detected to the target object.

[0261] For ease of description, the above sections are divided into modules (or units) according to their functions and described separately. Of course, in implementing this application, the functions of each module (or unit) can be implemented in one or more software or hardware components.

[0262] Having introduced the training method and apparatus for the video recommendation model according to exemplary embodiments of this application, the electronic device according to another exemplary embodiment of this application will now be described.

[0263] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."

[0264] Based on the same inventive concept as the above-described method embodiments, this application also provides an electronic device. In one embodiment, the electronic device may be a server, such as... Figure 1 The server 120 is shown. In this embodiment, the structure of the electronic device can be as follows: Figure 13 As shown, it includes a memory 1301, a communication module 1303, and one or more processors 1302.

[0265] The memory 1301 is used to store computer programs executed by the processor 1302. The memory 1301 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0266] Memory 1301 may be volatile memory, such as random-access memory (RAM); memory 1301 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 1301 may be any other medium capable of carrying or storing a desired computer program having the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1301 may be a combination of the above-described memories.

[0267] The processor 1302 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 1302 is used to implement the training method of the video recommendation model described above when it calls the computer program stored in the memory 1301.

[0268] The communication module 1303 is used to communicate with terminal devices and other servers.

[0269] This application embodiment does not limit the specific connection medium between the memory 1301, communication module 1303, and processor 1302. This application embodiment... Figure 13 The memory 1301 and the processor 1302 are connected via a bus 1304, and the bus 1304 is in Figure 13 The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. The 1304 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 13 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.

[0270] The memory 1301 stores a computer storage medium, which stores computer-executable instructions. These instructions are used to implement the training method of the video recommendation model according to embodiments of this application. The processor 1302 is used to execute the aforementioned training method of the video recommendation model, such as... Figure 2 As shown.

[0271] In another embodiment, the electronic device may also be other electronic devices, such as... Figure 1 The terminal device 110 is shown. In this embodiment, the electronic device can be structured as follows: Figure 14As shown, it includes components such as: communication component 1410, memory 1420, display unit 1430, camera 1440, sensor 1450, audio circuit 1460, Bluetooth module 1470, processor 1480, etc.

[0272] The communication component 1410 is used to communicate with the server. In some embodiments, it may include a Circuit-Based Wireless Fidelity (WiFi) module, which is a short-range wireless transmission technology. Electronic devices can use the WiFi module to help users send and receive information.

[0273] The memory 1420 can be used to store software programs and data. The processor 1480 executes various functions of the terminal device 110 and performs data processing by running the software programs or data stored in the memory 1420. The memory 1420 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The memory 1420 stores an operating system that enables the terminal device 110 to run. In this application, the memory 1420 may store the operating system and various application programs, and may also store a computer program that executes the training method of the video recommendation model of the embodiments of this application.

[0274] The display unit 1430 can also be used to display information input by the user or information provided to the user, as well as various menus of the terminal device 110, in a graphical user interface (GUI). Specifically, the display unit 1430 may include a display screen 1432 disposed on the front of the terminal device 110. The display screen 1432 may be configured as a liquid crystal display, a light-emitting diode, or the like. The display unit 1430 can be used to display the video recommendation interface, etc., as described in this application embodiment.

[0275] The display unit 1430 can also be used to receive input digital or character information and generate signal inputs related to user settings and function control of the terminal device 110. Specifically, the display unit 1430 may include a touch screen 1431 disposed on the front of the terminal device 110, which can collect touch operations of the user on or near it, such as clicking buttons, dragging scroll boxes, etc.

[0276] The touchscreen 1431 can be placed on top of the display screen 1432, or the touchscreen 1431 and the display screen 1432 can be integrated to realize the input and output functions of the terminal device 110. After integration, it can be referred to as a touch display screen. In this application, the display unit 1430 can display the application and the corresponding operation steps.

[0277] Camera 1440 can be used to capture still images, which users can then share via an application. There can be one or multiple cameras 1440. An object is projected onto a photosensitive element through a lens, generating an optical image. This photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the processor 1480 for conversion into a digital image signal.

[0278] The terminal device may also include at least one sensor 1450, such as an accelerometer 1451, a proximity sensor 1452, a fingerprint sensor 1453, and a temperature sensor 1454. The terminal device may also be equipped with other sensors such as a gyroscope, barometer, hygrometer, thermometer, infrared sensor, light sensor, and motion sensor.

[0279] Audio circuitry 1460, speaker 1461, and microphone 1462 provide an audio interface between the user and terminal device 110. Audio circuitry 1460 converts received audio data into electrical signals, which are then transmitted to speaker 1461, where they are converted into sound signals for output. Terminal device 110 may also be equipped with volume buttons for adjusting the volume of the sound signal. Conversely, microphone 1462 converts collected sound signals into electrical signals, which are then received by audio circuitry 1460, converted back into audio data, and output to communication component 1410 for transmission to, for example, another terminal device 110, or to memory 1420 for further processing.

[0280] The Bluetooth module 1470 is used to interact with other Bluetooth devices that also have a Bluetooth module via the Bluetooth protocol. For example, a terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smartwatch) that also has a Bluetooth module through the Bluetooth module 1470, thereby exchanging data.

[0281] The processor 1480 is the control center of the terminal device, connecting various parts of the terminal through various interfaces and lines. It executes various functions and processes data by running or executing software programs stored in the memory 1420 and calling data stored in the memory 1420. In some embodiments, the processor 1480 may include one or more processing units; the processor 1480 may also integrate an application processor and a baseband processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the baseband processor mainly handles wireless communication. It is understood that the baseband processor may not be integrated into the processor 1480. In this application, the processor 1480 can run the operating system, applications, user interface display and touch response, and the training method of the video recommendation model in this embodiment. Furthermore, the processor 1480 is coupled to the display unit 1430.

[0282] In some possible implementations, various aspects of the training method for the video recommendation model provided in this application can also be implemented in the form of a program product, which includes a computer program. When the program product is run on an electronic device, the computer program is used to cause the electronic device to perform the steps in the training method for the video recommendation model according to the various exemplary embodiments of this application described above. For example, the electronic device can perform actions such as... Figure 2 The steps are shown in the figure.

[0283] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0284] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on an electronic device. However, the program product of this application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with a command execution system, apparatus, or device.

[0285] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.

[0286] Computer programs contained on readable media can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0287] Computer programs for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The computer program can execute entirely on the user's electronic device, partially on the user's electronic device, as a standalone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In cases involving remote electronic devices, the remote electronic device can be connected to the user's electronic device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external electronic device (e.g., via the Internet using an Internet service provider).

[0288] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0289] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0290] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing a computer-usable computer program.

[0291] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0292] These computer program commands may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the commands stored in the computer-readable storage medium produce an article of manufacture including command means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0293] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0294] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A training method for a video recommendation model, characterized in that, The method includes: Each selected training sample is input into the video recommendation model to be trained. Each training sample contains a sample object, a sample video, and a sample label. The sample label represents whether the corresponding sample object interacts with the corresponding sample video. Based on the video recommendation model, feature extraction is performed on each selected training sample to obtain corresponding extraction results. Each extraction result includes: the content features and first interaction features of the corresponding sample video, and the object features of the corresponding sample object; and the content features of the sample videos contained in each training sample are respectively mapped to obtain corresponding second interaction features. The content features and first interaction features corresponding to each training sample are combined to obtain the corresponding first video features. The content features and second interaction features corresponding to each training sample are combined to obtain the corresponding second video features. Based on a pre-configured video mapping matrix, the first video features and second video features corresponding to each training sample are mapped. Based on a pre-configured object mapping matrix, the object features corresponding to each training sample are mapped. Recommendation predictions are made based on the extraction results and second interaction features corresponding to each training sample, respectively, to obtain corresponding prediction results; each prediction result includes a first prediction sub-result, a second prediction sub-result, and a third prediction sub-result. Based on the difference between the third prediction sub-result and the corresponding sample label of each training sample, a meta-loss function is constructed; based on the difference between the first prediction sub-result and the second prediction sub-result of each training sample, a pairwise loss function is constructed; based on the meta-loss function and the pairwise loss function, as well as the video mapping matrix, the object mapping matrix and the embedding factor used to extract content features, a target loss function is constructed, and the parameters of the video recommendation model are adjusted based on the target loss function.

2. The method as described in claim 1, characterized in that, For a sample video, if the sample video is not a cold start video, then the first interaction feature of the sample video is extracted based on the interaction information with the corresponding sample object; if the sample video is a cold start video, then the first interaction feature of the sample video is obtained by feature mapping based on the corresponding content features.

3. The method as described in claim 1, characterized in that, The recommendation prediction is performed based on the extraction results and second interaction features corresponding to each training sample to obtain the corresponding prediction results, including: Recommendation predictions are made based on the first video features and object features corresponding to each training sample, respectively, to obtain the corresponding first prediction sub-results; Recommendation predictions are made based on the second video features and object features corresponding to each training sample, respectively, to obtain the corresponding second prediction sub-results; Recommendation predictions are made based on the second interaction features and object features corresponding to each training sample, respectively, to obtain the corresponding third prediction sub-results; Based on the first, second, and third prediction sub-results corresponding to each training sample, the corresponding prediction results are obtained.

4. The method as described in claim 1, characterized in that, The training sample set consists of training sample subsets containing different sample objects, with each training sample subset corresponding to the same sample object; The training samples were selected in the following manner: A subset of training samples is selected from the training sample set, and a main training sample group and an auxiliary training sample group are selected from the training sample subset, wherein the main training sample group and the auxiliary training sample group contain the same number of training samples; The step of inputting the selected training samples into the video recommendation model to be trained includes: The training samples from the main training sample group and the auxiliary training sample group are respectively input into the video recommendation model.

5. The method as described in claim 4, characterized in that, The recommendation prediction is performed based on the second interaction features and object features corresponding to each of the training samples, respectively, to obtain the corresponding third prediction sub-result, including: Based on the second interaction features and object features corresponding to each main training sample in the main training sample group, recommendation prediction is performed to obtain the third prediction sub-result corresponding to each main training sample. Based on the differences between the obtained third prediction sub-results and the corresponding sample labels, the gradient of the interaction representation network in the video recommendation model is updated. The interaction representation network is used to perform feature mapping on the content features corresponding to the sample video to obtain the corresponding second interaction features. Based on the second interaction features and object features corresponding to each auxiliary training sample in the auxiliary training sample group, recommendation prediction is performed to obtain the third prediction sub-result corresponding to each auxiliary training sample.

6. The method as described in claim 4, characterized in that, The meta-loss function is constructed based on the difference between the third prediction result and the corresponding sample label, including: Based on the differences between the third prediction sub-result and the corresponding sample label of each main training sample in the main training sample group, the main loss function is constructed. Based on the differences between the third prediction sub-result and the corresponding sample label of each auxiliary training sample in the auxiliary training sample group, an auxiliary loss function is constructed. The primary loss function and the secondary loss function are weighted and summed to obtain the meta-loss function.

7. The method as described in claim 4, characterized in that, The pairwise loss function includes: pairwise loss functions corresponding to the main training sample group and the auxiliary training sample group respectively; each training sample group includes: at least one sample label representing positive samples with interaction, and at least one sample label representing negative samples without interaction; the pairwise loss function corresponding to each training sample group is constructed in the following way: For a training sample group, a first loss function is constructed based on the difference between the first prediction result of the positive sample and the first prediction result of the negative sample in the training sample group; Based on the difference between the second prediction result of positive samples and the first prediction result of negative samples in the training sample group, a second loss function is constructed; The pairwise loss function is determined based on the first loss function and the second loss function.

8. The method as described in claim 7, characterized in that, Determining the pairwise loss function based on the first loss function and the second loss function includes: Based on the difference between the first prediction result of positive samples and the second prediction result of negative samples in the training sample group, and the difference between the second prediction result of positive samples and the second prediction result of negative samples, a third loss function is constructed. The pairwise loss function is determined by weighted summation of the first loss function, the second loss function, and the third loss function.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Input the video to be detected and the object to be detected into the trained video recommendation model; Based on the trained video recommendation model, feature extraction is performed on the video to be detected and the object to be detected, respectively, to obtain the target content features and target interaction features corresponding to the video to be detected, and the target object features corresponding to the object to be detected. The target content features and target interaction features of the video to be detected are combined to obtain the corresponding target video features; Recommendation prediction is performed based on the target video features and the target object features to obtain the corresponding target prediction result. The target prediction result is used to characterize the probability of recommending the video to be detected to the target object.

10. A training device for a video recommendation model, characterized in that, include: The input unit is used to input the selected training samples into the video recommendation model to be trained. Each training sample contains a sample object, a sample video, and a sample label. The sample label represents whether the corresponding sample object interacts with the corresponding sample video. The first acquisition unit is used to extract features from each selected training sample based on the video recommendation model to obtain corresponding extraction results. Each extraction result includes: content features and first interaction features of the corresponding sample video, and object features of the corresponding sample object; and to perform feature mapping on the content features corresponding to the sample videos contained in each training sample to obtain corresponding second interaction features. The second acquisition unit is used to combine the content features and the first interaction features corresponding to each training sample to obtain the corresponding first video features, and to combine the content features and the second interaction features corresponding to each training sample to obtain the corresponding second video features. The mapping unit is used to map the first video feature and the second video feature corresponding to each training sample based on a pre-configured video mapping matrix; and to map the object feature corresponding to each training sample based on a pre-configured object mapping matrix. The second acquisition unit is also used to perform recommendation prediction based on the extraction results and second interaction features corresponding to each training sample, and obtain corresponding prediction results; each prediction result includes a first prediction sub-result, a second prediction sub-result and a third prediction sub-result. The adjustment unit is used to construct a meta-loss function based on the difference between the third prediction sub-result and the corresponding sample label of each training sample; construct a pairwise loss function based on the difference between the first prediction sub-result and the second prediction sub-result of each training sample; construct a target loss function based on the meta-loss function and the pairwise loss function, as well as the video mapping matrix, the object mapping matrix and the embedding factor for extracting content features, and adjust the parameters of the video recommendation model based on the target loss function.

11. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of any of the methods described in claims 1 to 9.

12. A computer-readable storage medium, characterized in that, It includes a computer program that, when run on an electronic device, causes the electronic device to perform the steps of any of the methods described in claims 1 to 9.

13. A computer program product, characterized in that, The method includes a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Migration model training method and device based on artificial intelligence and storage medium

    CN111026970A