Multi-target model training method and device and related equipment

By generating and processing video information feature vectors, using the target binomial distribution probability and multi-objective learning algorithm to train multi-objective models, the problem of data sparsity in video recommendation scenarios is solved, and the modeling effect of the model's transformation behavior is improved.

CN120236179APending Publication Date: 2025-07-01BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510301854.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, the multi-objective model has data sparsity problems in video recommendation scenarios, resulting in poor modeling of user conversion behaviors.

Method used

By generating feature vectors based on video information on the server and user side, using the target binomial distribution probability for data expansion and dimensionality reduction, generating anchor vectors, and combining multi-objective unsupervised and supervised learning algorithms to calculate the loss value for multi-objective models.

Benefits of technology

The training samples of multi-objective models have been expanded, the modeling accuracy of the model's modeling of user conversion behavior has been improved, and the data sparseness problem has been alleviated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236179A_ABST
    Figure CN120236179A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-target model training method. The multi-target model training method comprises the steps of generating M first feature vectors based on video information of a server side and a user side; obtaining M first anchor point vectors and M second anchor point vectors associated with the target binomial distribution probability; and updating the multi-target model according to the first loss value and the second loss value to obtain an updated multi-target model. According to the method and the device, M first feature vectors are obtained through conversion according to video information of a server side and a user side, data expansion and dimension reduction processing are performed on the M first feature vectors according to a target binomial distribution probability to obtain M first anchor point vectors and M second anchor point vectors, and a first loss value of a multi-target unsupervised learning algorithm is generated according to the M first anchor point vectors and the M second anchor point vectors; and the multi-target model is updated in cooperation with the second loss value of the multi-target supervised learning algorithm, so that samples for training the multi-target model are expanded, and the problem of data sparsity during training of the multi-target model is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and particularly to a training method, device and related equipment for a multi-objective model. Background Art

[0002] A multi-objective model is a model that simultaneously considers multiple objectives in an optimization problem. Such models are usually used in situations where trade-offs need to be made between multiple conflicting objectives. In the video recommendation scenario, the conversion behaviors (such as behaviors like user likes, comments, consumption, etc. after clicking) in multi-objective modeling are often relatively sparse, which to a certain extent restricts the modeling effect of the model on user conversion behaviors. That is, in real-world data, in the video recommendation scenario, from exposure to click and then to conversion, the data continuously decreases, and the conversion part is often a small part of the video clicks, resulting in the problem of data sparsity in the multi-objective models for recommending videos in the prior art. Summary of the Invention

[0003] The purpose of the embodiments of the present invention is to provide a training method, device and related equipment for a multi-objective model to solve the problem of data sparsity in the multi-objective model in the prior art. The specific technical solutions are as follows:

[0004] In the first aspect of the implementation of the present invention, first, a training method for a multi-objective model is provided. The method includes:

[0005] Generating M first feature vectors based on the video information of the server side and the user side. The first feature vector is an N-dimensional feature vector for representing the video information, where M is a positive integer and N is a positive integer;

[0006] Performing target processing on the M first feature vectors based on the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors associated with the target binomial distribution probability, where the target processing includes data expansion and dimensionality reduction processing;

[0007] Calculating the loss value between the M first anchor vectors and the M second anchor vectors according to the first loss function to obtain the first loss value;

[0008] Training the multi-objective model according to the first loss value and the second loss value to obtain the trained multi-objective model. The multi-objective model is a model that simultaneously considers multiple input objectives in an optimization problem, and the second loss value is the loss value obtained by calculating the loss of the M first feature vectors according to the multi-objective supervised learning algorithm associated with the multi-objective model.

[0009] Optionally, the target processing includes data expansion, dimensionality reduction processing, information interaction processing, and compression processing. The target processing of the M first feature vectors based on the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors associated with the target binomial distribution probability includes:

[0010] Generating a target mask based on the target binomial distribution probability to perform data expansion on the M first feature vectors, and performing dimensionality reduction processing on the M first feature vectors to obtain M second feature vectors and M third feature vectors associated with the target binomial distribution probability. The M second feature vectors and the M third feature vectors correspond one by one. The target mask is used to determine the number of vectors for data expansion of the M first feature vectors;

[0011] Performing information interaction processing and compression processing on the M second feature vectors and the M third feature vectors respectively through a multi-layer perceptron to obtain the M first anchor vectors and the M second anchor vectors. The interaction processing includes embedding the first vector information of the second feature vector into the third feature vector, and / or embedding the second vector information of the third feature vector into the second feature vector. The first vector information is used to indicate the vector characteristics of the second feature vector, and the second vector information is used to indicate the vector characteristics of the third feature vector.

[0012] Optionally, the generation of the M first feature vectors based on the video information of the server side and the user side includes:

[0013] Obtaining log data, where the log data includes multiple video information of the server side and the user side. Among them, the video information includes at least one of the following: video content information, video statistical information, video interaction information, and video click information;

[0014] Generating M video features according to the multiple video information;

[0015] Mapping the M video features based on an information retrieval method to generate the M first feature vectors.

[0016] Optionally, before calculating the loss value between the M first anchor vectors and the M second anchor vectors according to the first loss function to obtain the first loss value, the method further includes:

[0017] Determining the first anchor vector and the second anchor vector generated from the same first feature vector as positive samples for each other;

[0018] Determine a plurality of target anchor vectors from the M first anchor vectors and the M second anchor vectors, where the target anchor vectors are anchor vectors related to the video interaction information and / or the video click information;

[0019] Determine the plurality of target anchor vectors as positive samples for each other, and determine the remaining first anchor vectors and second anchor vectors as negative samples;

[0020] Determine the M first anchor vectors, the M second anchor vectors, and the plurality of target anchor vectors as training samples for the multi-target model;

[0021] The training of the multi-target model according to the first loss value and the second loss value to obtain the trained multi-target model includes:

[0022] Train the multi-target model according to the first loss value, the second loss value, and the training samples to obtain the trained multi-target model.

[0023] Optionally, the calculating the loss value between the M first anchor vectors and the M second anchor vectors according to the first loss function to obtain the first loss value includes:

[0024] Generate a first loss function, where the first loss function is the sum of the ratio of a first exponential function and a second exponential function. The first exponential function is the exponential function of the M cosine similarities between the M first anchor vectors and the corresponding M second anchor vectors, and the second exponential function is the sum of the exponential functions of the M cosine similarities;

[0025] Substitute the M first anchor vectors and the M second anchor vectors into the first loss function for calculation to obtain the first loss value.

[0026] Optionally, the training of the multi-target model according to the first loss value and the second loss value to obtain the trained multi-target model includes:

[0027] Calculate a second loss value according to the multi-target supervised learning algorithm for the M first feature vectors;

[0028] Set a first weight coefficient and a second weight coefficient based on the multi-target model, where the first weight coefficient corresponds to the first loss value and the second weight coefficient corresponds to the second loss value;

[0029] Perform weighted calculation on the first loss value and the second loss value respectively based on the first weight coefficient and the second weight coefficient to obtain a target loss value;

[0030] Train the multi-objective model according to the target loss value to obtain a trained multi-objective model.

[0031] In a second aspect, a training device for a multi-objective model is provided. The device includes:

[0032] A generation module, configured to generate M first feature vectors based on video information of a server side and a user side, where the first feature vectors are N-dimensional feature vectors for representing the video information, M is a positive integer, and N is a positive integer;

[0033] A processing module, configured to perform target processing on the M first feature vectors based on a target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors associated with the target binomial distribution probability, where the target processing includes data expansion and dimensionality reduction processing;

[0034] A calculation module, configured to calculate a loss value between the M first anchor vectors and the M second anchor vectors according to a first loss function to obtain a first loss value;

[0035] A training module, configured to train a multi-objective model according to the first loss value and a second loss value to obtain a trained multi-objective model, where the multi-objective model is a model that simultaneously considers multiple input targets in an optimization problem, and the second loss value is a loss value obtained by calculating a loss for the M first feature vectors according to a multi-objective supervised learning algorithm associated with the multi-objective model.

[0036] In a third aspect of the embodiments of the present application, an electronic device is further provided, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the training method for the multi-objective model according to any one of the first aspects are implemented.

[0037] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is further provided. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps of the training method for the multi-objective model according to any one of the first aspects are implemented.

[0038] In the fifth aspect of the embodiments of the present application, the embodiments of the present application further provide a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the training method of the multi-objective model described in any item of the first aspect. The embodiments of the present invention provide a training method of a multi-objective model, and the method includes: generating M first feature vectors based on the video information of the server side and the user side, where the first feature vector is an N-dimensional feature vector for representing the video information, M is a positive integer, and N is a positive integer; performing target processing on the M first feature vectors based on the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors associated with the target binomial distribution probability, where the target processing includes data expansion and dimensionality reduction processing; calculating a loss value between the M first anchor vectors and the M second anchor vectors according to a first loss function to obtain a first loss value; training the multi-objective model according to the first loss value and a second loss value to obtain a trained multi-objective model, where the second loss value is a loss value calculated by performing loss calculation on the M first feature vectors according to a multi-objective supervised learning algorithm associated with the multi-objective model. The present application converts to obtain M first feature vectors according to the video information of the server side and the user side, performs data expansion and dimensionality reduction processing on the M first feature vectors according to the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors, thereby generating a first loss value of the multi-objective unsupervised learning algorithm, and updating the multi-objective model in cooperation with the second loss value of the multi-objective supervised learning algorithm, so as to expand the samples for training the multi-objective model and solve the problem of data sparsity when training the multi-objective model. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0040] Figure 1 It is a schematic flowchart of the training method of the multi-objective model in the embodiments of the present application;

[0041] Figure 2 It is a schematic diagram of the processing flow in the embodiments of the present application;

[0042] Figure 3 It is a schematic structural diagram of the training device of the multi-objective model in the embodiments of the present application;

[0043] Figure 4 It is a schematic structural diagram of the electronic device in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0045] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subprogram, and the like.

[0046] In addition, terms such as "first" and "second" may be used herein to describe various directions, actions, steps, or elements, etc., but these directions, actions, steps, or elements are not limited by these terms. These terms are only used to distinguish one direction, action, step, or element from another. For example, without departing from the scope of the present application, the first order request can be referred to as the second order request, and similarly, the second order request can be referred to as the first order request. Both the first order request and the second order request are order requests, but they are not the same order request. The terms "first", "second", etc. should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0047] The embodiments of the present application provide a training method for a multi-objective model, as Figure 1 shown, the method includes:

[0048] Step 101, generating M first feature vectors based on the video information of the server side and the user side, where the first feature vector is an N-dimensional feature vector for representing the video information, M is a positive integer, and N is a positive integer.

[0049] In this embodiment, the video information is the information content generated by the server side and the user side based on the video, such as the viewing data of the video, the comment data of the video, the consumption data of the video, etc., which is not specifically limited in this embodiment. It should be noted that in this embodiment, the video information of the server side is used in the training process of the multi-objective model. However, the conversion behaviors of the video (such as the behaviors of the user clicking to like, comment, consume, etc.) are often relatively sparse. Therefore, it is necessary to effectively convert the conversion behaviors into samples of the multi-objective model. Among them, the multi-objective model (Multi-Objective Model) is a model that simultaneously considers multiple objectives in an optimization problem and is usually used in situations where it is necessary to balance multiple conflicting objectives.

[0050] Among them, multiple video information is respectively vector-converted to generate M first feature vectors, where M is a positive integer. Among them, the first feature vector is an N-dimensional feature vector used to represent the video information, which can be expressed as the N-dimensional feature vector E n That is, each generated first feature vector includes N feature dimensions. Among them, N is a positive integer. In actual use, the specific value of N can be adjusted according to the actual situation. It should be noted that different feature dimensions are used to describe different attributes or variables of data samples, and each feature dimension represents a characteristic in the data. Exemplarily, the N feature dimensions can be numerical values, time, content, etc.

[0051] Step 102: Perform target processing on the M first feature vectors based on the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors associated with the target binomial distribution probability, where the target processing includes data expansion and dimensionality reduction processing.

[0052] In this embodiment, the binomial distribution is a discrete probability distribution used to describe the number of times a successful event occurs in n independent trials. Each trial has two possible outcomes (usually referred to as "success" and "failure"), and the probability of success in each trial is p, and the probability of failure is 1 - p. A target mask is generated by setting the target binomial distribution probability, where the target mask is used to perform mask processing on the M first feature vectors E n For mask processing, each element of the target mask independently determines whether to retain (usually 1) or discard (usually 0) according to the set success probability (p). For example, a certain element is retained with probability p and discarded with probability (1 - p), that is, each dimension feature in the N-dimensional vector has a probability p of being processed as 0. Being processed as 0 means discarding the data, and the probability p is a hyperparameter, and the value of p can be set according to the actual situation.

[0053] Specifically, the target processing in this embodiment is to use the target binomial distribution probability to process the M first feature vectors E nAfter data expansion, dimensionality reduction is performed on multiple feature vectors after data expansion, and finally M first anchor vectors h generated by the target binomial distribution probability are obtained. i and M second anchor vectors h j , thus completing the information expansion and compression process of the first feature vector.

[0054] It should be noted that the first anchor vector and the second anchor vector are used to define the reference vectors of a certain structure. For example, in Triplet Loss, the anchor vector, the positive sample vector, and the negative sample vector are used to train the model so that the distance between the anchor and the positive sample is less than the distance between the anchor and the negative sample.

[0055] Step 103: Calculate the loss value between the M first anchor vectors and the M second anchor vectors according to the first loss function to obtain the first loss value.

[0056] In this embodiment, the loss value between the M first anchor vectors and the M second anchor vectors is calculated through the set first loss function to obtain the first loss value. Among them, the first loss function can be a conventional loss function, such as cross-entropy loss function, binary cross-entropy loss function, mean squared error loss function, etc., which are not specifically limited in this embodiment. Among them, the M first anchor vectors and the M second anchor vectors in this embodiment can be directly obtained through operations such as an encoder, and then the first loss value can be calculated through the set first loss function, without calculating the loss value through a model trained with data. Therefore, the first loss value in this embodiment is obtained through an unsupervised learning algorithm, that is, without labeled data, and the loss value is calculated by directly inputting the first anchor vector and the second anchor vector, thereby updating the model. Among them, the unsupervised learning algorithm is a machine learning method in which the model learns on a dataset without labels. Different from supervised learning, unsupervised learning does not rely on known outputs or targets, but tries to discover potential structures or patterns from the input data, such as through clustering, dimensionality reduction, and association learning to train the model, thereby improving the recognition effect of the model.

[0057] Step 104: Train the multi-objective model according to the first loss value and the second loss value to obtain the trained multi-objective model. The multi-objective model is a model that simultaneously considers multiple input targets in an optimization problem. The second loss value is the loss value calculated according to the multi-objective supervised learning algorithm associated with the multi-objective model for the M first feature vectors.

[0058] In this embodiment, supervised learning is a machine learning method in which the model learns using a labeled dataset during training. Each training sample has an input feature and a corresponding output label, and the model makes predictions by learning the relationship between these inputs and outputs. The process of training a multi-objective model using training labels and training samples is a multi-objective supervised learning algorithm, thereby generating a second loss value. The multi-objective model is updated by comprehensively combining the first loss value and the second loss value, thereby obtaining an updated multi-objective model.

[0059] It should be noted that the inputs of the multi-objective model mainly include multiple objectives to be optimized such as objective functions, decision variables, constraint conditions, and relevant data. In this embodiment, the inputs are, for example, click data, user comment data, purchase data, share data, and so on. The output of the multi-objective model is usually the optimal solution or decision recommendation, as well as the objective function value related thereto. In this embodiment, the output is, for example, a video optimization strategy, thereby helping to determine the best solution.

[0060] This application generates M first feature vectors by converting the video information of the server side and the user side, performs data expansion and dimensionality reduction processing on the M first feature vectors according to the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors, thereby generating the first loss value of the multi-objective unsupervised learning algorithm, and updates the multi-objective model in cooperation with the second loss value of the multi-objective supervised learning algorithm, thereby expanding the samples for training the multi-objective model and solving the problem of data sparsity that occurs when training the multi-objective model.

[0061] In some feasible implementation manners, optionally, step 101, generating M first feature vectors based on the video information of the server side and the user side, includes:

[0062] Obtain log data, where the log data includes multiple video information of the server side and the user side, and where the video information includes at least one of the following: video content information, video statistical information, video interaction information, and video click information;

[0063] Generate M video features according to the multiple video information;

[0064] Map the M video features based on an information retrieval method to generate the M first feature vectors.

[0065] In this embodiment, the log data is the log data obtained from the server-side and user-side mobile phone user behavior logs, which includes video information of different users watching different videos. The video information includes at least one of the following: video content information, video statistics information, video interaction information, and video click information. The video content information refers to the type of the video and the content included, such as action movies and documentary movies, etc. The video statistics information is the viewing situation of the video, such as the number of historical viewers, etc. The video interaction information includes the interaction behaviors between the user and the video, such as the user liking, commenting on, and forwarding the video, etc. The video click information includes the click information of the user on the comments or relevant links of the video.

[0066] After obtaining multiple video information, convert the multiple video information into M video features. Specifically, the information retrieval method can be used to convert the multiple video information. Specifically, the multiple video information is mapped into a feature vector including N dimensions by using the information retrieval method, that is, M first feature vectors E n . The information retrieval method refers to the process of obtaining specific information from a large amount of information resources. In this embodiment, the conversion is performed by means of mapping. Specifically, the mapping method may include steps of converting multiple video information into feature representations, mapping the feature representations to corresponding feature spaces, and finally calculating similarities to generate M video features, thereby expanding the number of feature vectors. It should be noted that N in the N dimensions is related to the video information, so the specific value of N is related to the actual situation and is not limited here. By mapping the M video features through the information retrieval method, the number of generated feature vectors is increased, thus enriching the samples during the training process.

[0067] Optionally, the target processing includes data expansion, dimensionality reduction processing, information interaction processing, and compression processing. Step 102, perform target processing on the M first feature vectors based on the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors associated with the target binomial distribution probability, including:

[0068] Generate a target mask based on the target binomial distribution probability to perform data expansion on the M first feature vectors, and perform dimensionality reduction processing on the M first feature vectors to obtain M second feature vectors and M third feature vectors associated with the target binomial distribution probability. The M second feature vectors and the M third feature vectors correspond one by one. The target mask is used to determine the number of vectors for performing data expansion on the M first feature vectors;

[0069] The information interaction processing and compression processing are respectively performed on the M second feature vectors and the M third feature vectors through a multi-layer perceptron to obtain the M first anchor vectors and the M second anchor vectors. The interaction processing includes embedding the first vector information of the second feature vector into the third feature vector, and / or embedding the second vector information of the third feature vector into the second feature vector. The first vector information is used to indicate the vector feature of the second feature vector, and the second vector information is used to indicate the vector feature of the third feature vector.

[0070] In this embodiment, the target processing further includes information interaction processing and compression processing. Specifically, a target mask is first generated through a target binomial distribution probability to perform data expansion on the M first feature vectors E n and then dimensionality reduction processing is performed on the vectors obtained after data expansion, and finally M second feature vectors e i and M third feature vectors e j .

[0071] Specifically, a target mask is generated by setting a target binomial distribution probability. The target mask is used to perform masking processing on the M first feature vectors E n . Each element of the target mask independently determines whether to retain (usually 1) or discard (usually 0) according to the set success probability (p). For example, a certain element is retained with probability p and discarded with probability (1 - p), that is, each dimension feature in the N-dimensional vector has a probability p of being processed as 0. Being processed as 0 means discarding the data. The probability p is a hyperparameter, and the value of p can be set according to the actual situation. Through the above operations, the first feature vectors E n are processed based on the binomial distribution probability into two different e i and e j , thus completing the data expansion of the M first feature vectors E n . Subsequently, dimensionality reduction processing is further performed on the expanded multiple vectors to generate M second feature vectors e i and M third feature vectors e j .

[0072] After generating the M second feature vectors e i and the M third feature vectors e j , the information interaction processing and compression processing are respectively performed on the M second feature vectors e i and the M third feature vectors e j through a multi-layer perceptron (MLP) to generate M first anchor vectors h i and M second anchor vectors h jAmong them, the multi-layer perceptron is a feed-forward neural network composed of multiple layers, including an input layer, a hidden layer, and an output layer. Each neuron is connected to all neurons in the previous layer and is commonly used for classification and regression tasks. In this embodiment, the multi-layer perceptron can embed the first vector information in the second feature vector e i into the third feature vector e j , and can also embed the second vector information in the third feature vector e j into the second feature vector e i to achieve information interaction between the second feature vector e i and the third feature vector e j . Specifically, the first vector information can be certain dimensional features in the second feature vector e j , and the second vector information can be certain dimensional features in the third feature vector e j , which can be selected according to the actual situation, such as selecting time dimensional features, etc., and are not specifically limited in this embodiment. Thus, the features of the vector are enriched, that is, among a fixed number of feature vectors, the feature vectors are associated and expanded through information interaction. In addition, the multi-layer perceptron can also compress M second feature vectors e i and M third feature vectors e j , thereby reducing the sizes of the first anchor vector h i and the second anchor vector h j , and improving the processing efficiency when batch processing the anchor vectors.

[0073] Optionally, before step 103, calculating the loss value between the M first anchor vectors and the M second anchor vectors according to the first loss function to obtain the first loss value, the method further includes:

[0074] Determining the first anchor vector and the second anchor vector generated from the same first feature vector as positive samples for each other;

[0075] Determining multiple target anchor vectors among the M first anchor vectors and the M second anchor vectors, where the target anchor vectors are anchor vectors related to the video interaction information and / or the video click information;

[0076] Determining the multiple target anchor vectors as positive samples for each other, and determining the remaining first anchor vectors and second anchor vectors as negative samples;

[0077] Determining the M first anchor vectors, the M second anchor vectors, and the multiple target anchor vectors as the training samples of the multi-target model;

[0078] Step 104. Training the multi-objective model according to the first loss value and the second loss value to obtain a trained multi-objective model, including:

[0079] Training the multi-objective model according to the first loss value, the second loss value and the training samples to obtain the trained multi-objective model.

[0080] In this embodiment, the method for dividing positive and negative samples is set. Specifically, first, the first anchor vector and the second anchor vector obtained by processing the same first feature vector are determined to be positive samples of each other, that is, the second anchor vector corresponding one-to-one to the first anchor vector is determined to be a positive sample of each other. It should be noted that being positive samples of each other means that in a group of samples, there is a certain relationship between two samples, so that they are both regarded as positive samples of the same category.

[0081] Secondly, among the M first anchor vectors and the M second anchor vectors, the anchor vectors obtained by conversion through video interaction information and / or the video click information are determined, that is, the anchor vectors obtained from the conversion behavior of the user to the video. It should be noted that the conversion behavior refers to the user's interactive operation on the video after watching the video, such as clicking on the video comment, converting the video, liking, etc. Thus, the anchor vectors related to the video interaction information and / or the video click information are determined as target anchor vectors, and multiple anchor vectors are determined to be positive samples of each other in the contrast learning. The remaining first anchor vectors and second anchor vectors are determined as negative samples, thereby obtaining the training samples of the multi-objective model. The multi-objective model is trained and updated through the training samples, the first loss value, and the second loss value to obtain an updated multi-objective model.

[0082] In this embodiment, the conversion behavior samples are introduced in the modeling process of contrast learning, and the conversion behavior is also used as a positive sample in the contrast learning training process, alleviating the sparsity of the conversion behavior samples.

[0083] Optionally, step 103. Calculating the loss value between the M first anchor vectors and the M second anchor vectors according to the first loss function to obtain the first loss value, including:

[0084] Generating a first loss function, where the first loss function is the sum of the ratio of a first exponential function and a second exponential function. The first exponential function is the exponential function of the M cosine similarities between the M first anchor vectors and the corresponding M second anchor vectors, and the second exponential function is the sum of the exponential functions of the M cosine similarities;

[0085] Substituting the M first anchor vectors and the M second anchor vectors into the first loss function for calculation to obtain the first loss value.

[0086] For example, the first loss function is set as the contrastive learning loss function L0. Among them, the first loss function is respectively composed of the ratio of the first exponential function and the second exponential function. Specifically, the first loss function L0

[0087]

[0088] is as follows:

[0089] Among them, exp represents the exponential function, and the function s(·) represents the anchor vector h i and h j 's cosine similarity, h k is the number of negative samples, τ is the temperature hyperparameter, N is the number of positive samples, i is the number parameter in the positive samples, h i is the first anchor vector and h j is the second anchor vector.

[0090] In this embodiment, the contrastive learning loss function is used to optimize the model, so that similar samples (positive samples) are closer in the embedding space, while dissimilar samples (negative samples) are farther away. Based on the positive and negative samples set in this embodiment, the loss value can be accurately calculated, thereby improving the training accuracy of the multi-objective model.

[0091] Optionally, step 104, training the multi-objective model according to the first loss value and the second loss value to obtain a trained multi-objective model, includes:

[0092] Calculating the second loss value according to the multi-objective supervised learning algorithm for the M first feature vectors;

[0093] Based on the multi-objective model, set the first weight coefficient and the second weight coefficient. The first weight coefficient corresponds to the first loss value, and the second weight coefficient corresponds to the second loss value;

[0094] Based on the first weight coefficient and the second weight coefficient, perform weighted calculations on the first loss value and the second loss value respectively to obtain the target loss value;

[0095] Train the multi-objective model according to the target loss value to obtain a trained multi-objective model.

[0096] In this embodiment, the first weight coefficient α and the second weight coefficient β are set according to the multi-objective model to balance the first loss value and the second loss value. Specifically, the overall loss function L is set allIt is the weighted sum of the cross-entropy loss function (Supervised loss) L1 of the multi-object supervised learning algorithm ESMM and the unsupervised contrastive learning loss function (Contrastive loss) L0, that is:

[0097] L all = α * L1 + β * L0

[0098] Among them, L1 is the second loss value, L0 is the first loss value, α is the first weight coefficient, and β is the second weight coefficient.

[0099] By the overall loss function L all Calculate the loss value, realizing the integration of supervised learning and unsupervised contrastive learning. The two learning methods are integrated by means of loss function weighting, complementing each other and improving the calculation accuracy of the loss value.

[0100] As Figure 2 shown, Figure 2 This is the process schematic diagram in this application. The first feature vector respectively obtains the second loss value through the multi-object supervised learning algorithm, and after the first feature vector is converted into the second feature vector and the third feature vector, the first anchor vector and the second anchor vector are generated through the encoder, and the first loss value is calculated. Finally, the target loss value is generated by integrating the first loss value and the second loss value to train and update the multi-object model.

[0101] This application adopts an end-to-end training method. In the training, the gradient descent method is used to update the parameters of the model until the loss function reaches the preset threshold. By introducing the unsupervised contrastive learning loss function to assist in the learning of multi-object tasks, especially conversion behaviors, the training effect is improved.

[0102] This application converts M first feature vectors according to the video information of the server side and the user side, performs data expansion and dimensionality reduction processing on the M first feature vectors according to the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors, generates the first loss value of the multi-object unsupervised learning algorithm accordingly, and updates the multi-object model in cooperation with the second loss value of the multi-object supervised learning algorithm, thereby expanding the samples for training the multi-object model and solving the problem of data sparsity when training the multi-object model.

[0103] This application embodiment provides a training device for a multi-object model. As Figure 3 shown, the device includes:

[0104] A generation module 310, configured to generate M first feature vectors based on the video information of the server side and the user side. The first feature vector is an N-dimensional feature vector for representing the video information. M is a positive integer, and N is a positive integer;

[0105] A processing module 320, configured to perform target processing on the M first feature vectors based on a target binomial distribution probability, to obtain M first anchor vectors and M second anchor vectors associated with the target binomial distribution probability, where the target processing includes data expansion and dimensionality reduction processing;

[0106] A calculation module 330, configured to calculate a loss value between the M first anchor vectors and the M second anchor vectors according to a first loss function, to obtain a first loss value;

[0107] A training module 340, configured to train a multi-target model according to the first loss value and a second loss value, to obtain a trained multi-target model, where the multi-target model is a model that simultaneously considers multiple input targets in an optimization problem, and the second loss value is a loss value obtained by calculating a loss for the M first feature vectors according to a multi-target supervised learning algorithm associated with the multi-target model.

[0108] Optionally, the target processing includes data expansion, dimensionality reduction processing, information interaction processing, and compression processing, and the processing module 320 includes:

[0109] A first processing sub-module, configured to generate a target mask based on a target binomial distribution probability to perform data expansion on the M first feature vectors, and perform dimensionality reduction processing on the M first feature vectors, to obtain M second feature vectors and M third feature vectors associated with the target binomial distribution probability, where the M second feature vectors and the M third feature vectors correspond one by one, and the target mask is used to determine the number of vectors for performing data expansion on the M first feature vectors;

[0110] A second processing sub-module, configured to perform information interaction processing and compression processing on the M second feature vectors and the M third feature vectors respectively through a multi-layer perceptron, to obtain the M first anchor vectors and the M second anchor vectors, where the interaction processing includes embedding first vector information of the second feature vector into the third feature vector, and / or, embedding second vector information of the third feature vector into the second feature vector, the first vector information is used to indicate a vector feature of the second feature vector, and the second vector information is used to indicate a vector feature of the third feature vector.

[0111] Optionally, the generation module 310 includes:

[0112] An acquisition sub-module, configured to acquire log data, where the log data includes multiple video information of the server side and the user side, and the video information includes at least one of the following: video content information, video statistical information, video interaction information, and video click information;

[0113] The first generation sub-module is used to generate M video features according to the multiple video information;

[0114] The second generation sub-module is used to map the M video features based on an information retrieval method to generate the M first feature vectors.

[0115] Optionally, it further includes:

[0116] The first determination module is used to determine the first anchor vector and the second anchor vector generated by the same first feature vector as positive samples for each other;

[0117] The second determination module is used to determine multiple target anchor vectors among the M first anchor vectors and the M second anchor vectors, and the target anchor vectors are anchor vectors related to the video interaction information and / or the video click information;

[0118] The third determination module is used to determine the multiple target anchor vectors as positive samples for each other, and determine the remaining first anchor vectors and second anchor vectors as negative samples;

[0119] The fourth determination module is used to determine the M first anchor vectors, the M second anchor vectors and the multiple target anchor vectors as the training samples of the multi-target model;

[0120] The training module 340 includes:

[0121] The first update sub-module is used to train the multi-target model according to the first loss value, the second loss value and the training samples to obtain the trained multi-target model.

[0122] The calculation module 330 includes:

[0123] The first setting sub-module is used to generate a first loss function, and the first loss function is the sum of the ratio of a first exponential function and a second exponential function. The first exponential function is the exponential function of the M cosine similarities between the M first anchor vectors and the corresponding M second anchor vectors, and the second exponential function is the sum of the exponential functions of the M cosine similarities;

[0124] The first calculation sub-module is used to substitute the M first anchor vectors and the M second anchor vectors into the first loss function for calculation to obtain the first loss value.

[0125] The training module 340 includes:

[0126] The second calculation sub-module is used to calculate a second loss value according to a multi-target supervised learning algorithm for the M first feature vectors;

[0127] A second setting sub-module, configured to set a first weight coefficient and a second weight coefficient based on the multi-objective model, where the first weight coefficient corresponds to the first loss value, and the second weight coefficient corresponds to the second loss value;

[0128] A third calculation sub-module, configured to perform weighted calculations on the first loss value and the second loss value respectively based on the first weight coefficient and the second weight coefficient to obtain a target loss value;

[0129] A second update sub-module, configured to train the multi-objective model according to the target loss value to obtain a trained multi-objective model.

[0130] In this application, M first feature vectors are obtained by converting video information of a server side and a user side, and M first anchor vectors and M second anchor vectors are obtained by performing data expansion and dimensionality reduction processing on the M first feature vectors according to a target binomial distribution probability. Accordingly, a first loss value of a multi-objective unsupervised learning algorithm is generated, and the multi-objective model is updated in cooperation with a second loss value of a multi-objective supervised learning algorithm, thereby expanding samples for training the multi-objective model and solving the problem of data sparsity during training of the multi-objective model.

[0131] Figure 4 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of this application. As Figure 4 shown, the electronic device 400 includes a memory 410 and a processor 420. The number of processors 420 in the electronic device 400 may be one or more. Figure 4 Here, one processor 420 is taken as an example; the memory 410 and the processor 420 in the server may be connected through a bus or other means. Figure 4 Here, connection through a bus is taken as an example.

[0132] The memory 410, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the push method for members in the embodiments of this application. The processor 420 runs the software programs, instructions, and modules stored in the memory 410 to execute various functional applications and data processing of the server / terminal / server, that is, to implement the above-mentioned training method for the multi-objective model.

[0133] Among them, the processor 420 is configured to run a computer program stored in the memory 410 to implement the following steps:

[0134] Generate M first feature vectors based on video information of a server side and a user side, where the first feature vector is an N-dimensional feature vector for representing the video information, M is a positive integer, and N is a positive integer;

[0135] Perform target processing on the M first feature vectors based on the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors associated with the target binomial distribution probability, where the target processing includes data expansion and dimensionality reduction processing;

[0136] Calculate the loss value between the M first anchor vectors and the M second anchor vectors according to the first loss function to obtain the first loss value;

[0137] Train the multi-objective model according to the first loss value and the second loss value to obtain the trained multi-objective model. The multi-objective model is a model that simultaneously considers multiple input targets in an optimization problem. The second loss value is the loss value obtained by calculating the loss of the M first feature vectors according to the multi-objective supervised learning algorithm associated with the multi-objective model.

[0138] Optionally, the target processing includes data expansion, dimensionality reduction processing, information interaction processing, and compression processing. The performing target processing on the M first feature vectors based on the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors associated with the target binomial distribution probability includes:

[0139] Generate a target mask based on the target binomial distribution probability to perform data expansion on the M first feature vectors and perform dimensionality reduction processing on the M first feature vectors to obtain M second feature vectors and M third feature vectors associated with the target binomial distribution probability. The M second feature vectors and the M third feature vectors correspond one by one. The target mask is used to determine the number of vectors for performing data expansion on the M first feature vectors;

[0140] Perform information interaction processing and compression processing on the M second feature vectors and the M third feature vectors respectively through a multi-layer perceptron to obtain the M first anchor vectors and the M second anchor vectors. The interaction processing includes embedding the first vector information of the second feature vector into the third feature vector and / or embedding the second vector information of the third feature vector into the second feature vector. The first vector information is used to indicate the vector characteristics of the second feature vector, and the second vector information is used to indicate the vector characteristics of the third feature vector.

[0141] Optionally, the generating M first feature vectors based on the video information of the server side and the user side includes:

[0142] Obtain log data, where the log data includes multiple video information of the server side and the user side, and the video information includes at least one of the following: video content information, video statistical information, video interaction information, and video click information;

[0143] Generate M video features based on the multiple video information;

[0144] Map the M video features based on an information retrieval method to generate the M first feature vectors.

[0145] Optionally, before calculating the loss value between the M first anchor vectors and the M second anchor vectors according to the first loss function to obtain the first loss value, the method further includes:

[0146] Determine the first anchor vector and the second anchor vector generated from the same first feature vector as positive samples for each other;

[0147] Determine a plurality of target anchor vectors among the M first anchor vectors and the M second anchor vectors, where the target anchor vectors are anchor vectors related to the video interaction information and / or the video click information;

[0148] Determine the plurality of target anchor vectors as positive samples for each other, and determine the remaining first anchor vectors and second anchor vectors as negative samples;

[0149] Determine the M first anchor vectors, the M second anchor vectors, and the plurality of target anchor vectors as training samples of the multi-target model;

[0150] The training the multi-target model according to the first loss value and the second loss value to obtain the trained multi-target model includes:

[0151] Train the multi-target model according to the first loss value, the second loss value, and the training samples to obtain the trained multi-target model.

[0152] Optionally, the calculating the loss value between the M first anchor vectors and the M second anchor vectors according to the first loss function to obtain the first loss value includes:

[0153] Generate a first loss function, where the first loss function is the sum of the ratio of a first exponential function and a second exponential function, the first exponential function is the exponential function of the M cosine similarities between the M first anchor vectors and the corresponding M second anchor vectors, and the second exponential function is the sum of the exponential functions of the M cosine similarities;

[0154] Substitute the M first anchor vectors and the M second anchor vectors into the first loss function for calculation to obtain the first loss value.

[0155] Optionally, training the multi-objective model according to the first loss value and the second loss value to obtain a trained multi-objective model includes:

[0156] Calculating the second loss value according to the multi-objective supervised learning algorithm for the M first feature vectors;

[0157] Setting a first weight coefficient and a second weight coefficient based on the multi-objective model, where the first weight coefficient corresponds to the first loss value and the second weight coefficient corresponds to the second loss value;

[0158] Performing weighted calculations on the first loss value and the second loss value respectively based on the first weight coefficient and the second weight coefficient to obtain a target loss value;

[0159] Training the multi-objective model according to the target loss value to obtain a trained multi-objective model.

[0160] In this application, M first feature vectors are obtained by converting video information of the server side and the user side. The M first feature vectors are processed for data expansion and dimensionality reduction according to the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors. Based on this, the first loss value of the multi-objective unsupervised learning algorithm is generated, and the multi-objective model is updated in cooperation with the second loss value of the multi-objective supervised learning algorithm, thereby expanding the samples for training the multi-objective model and solving the problem of data sparsity during the training of the multi-objective model.

[0161] The computer-readable storage medium of the embodiments of this application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0162] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0163] The program code contained on the storage medium can be transmitted using any appropriate medium, including - but not limited to - wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0164] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++ - and also conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or terminal. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0165] Another embodiment of this application provides a computer program product. The computer program product is stored in a storage medium and is executed by at least one processor to implement the various processes of the foregoing embodiment of the training method for the multi-objective model, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0166] Note that the above is only a preferred embodiment of this application and the applied technical principles. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of this application. Therefore, although this application has been described in more detail through the above embodiments, this application is not limited to the above embodiments. Without departing from the concept of this application, it can also include more other equivalent embodiments, and the scope of this application is determined by the scope of the appended claims.

Claims

1. A training method for a multi-objective model, characterized in that: The method comprises: Generate M first feature vectors based on the video information of the server and the user, where the first feature vector is an N-dimensional feature vector used to represent the video information, where M is a positive integer and N is a positive integer; Performing target processing on the M first feature vectors based on the target binomial distribution probability to obtain M first anchor point vectors and M second anchor point vectors associated with the target binomial distribution probability, wherein the target processing includes data expansion and dimensionality reduction processing; Calculating a loss value between the M first anchor point vectors and the M second anchor point vectors according to a first loss function to obtain a first loss value; The multi-objective model is trained according to the first loss value and the second loss value to obtain a trained multi-objective model, wherein the multi-objective model is a model that simultaneously considers multiple input objectives in an optimization problem, and the second loss value is a loss value obtained by performing loss calculation on the M first feature vectors according to a multi-objective supervised learning algorithm associated with the multi-objective model.

2. The method according to claim 1, characterized in that The target processing includes data expansion, dimensionality reduction, information interaction and compression. The target processing is performed on the M first feature vectors based on the target binomial distribution probability to obtain M first anchor vectors and M second anchor vectors associated with the target binomial distribution probability, including: Generate a target mask based on the target binomial distribution probability to perform data expansion on the M first eigenvectors, and perform dimensionality reduction processing on the M first eigenvectors to obtain M second eigenvectors and M third eigenvectors associated with the target binomial distribution probability, the M second eigenvectors and the M third eigenvectors correspond one to one, and the target mask is used to determine the number of vectors for data expansion of the M first eigenvectors; The M second eigenvectors and the M third eigenvectors are respectively subjected to information interaction processing and compression processing through a multilayer perceptron to obtain the M first anchor vectors and the M second anchor vectors, wherein the interaction processing includes embedding the first vector information of the second eigenvector into the third eigenvector, and / or embedding the second vector information of the third eigenvector into the second eigenvector, wherein the first vector information is used to indicate the vector features of the second eigenvector, and the second vector information is used to indicate the vector features of the third eigenvector.

3. The method according to claim 1, characterized in that The generating of M first feature vectors based on the video information of the server and the user includes: Obtaining log data, wherein the log data includes multiple video information of the server and the user, wherein the video information includes at least one of the following: video content information, video statistics information, video interaction information, and video click information; Generate M video features according to the multiple video information; The M video features are mapped based on an information retrieval method to generate the M first feature vectors.

4. The method according to claim 3, characterized in that Before calculating the loss values ​​between the M first anchor point vectors and the M second anchor point vectors according to the first loss function to obtain the first loss value, the method further includes: Determine a first anchor point vector and a second anchor point vector generated by the same first feature vector as positive samples of each other; Determine a plurality of target anchor vectors among the M first anchor vectors and the M second anchor vectors, wherein the target anchor vectors are anchor vectors related to the video interaction information and / or the video click information; Determine the multiple target anchor point vectors as positive samples of each other, and determine the remaining first anchor point vectors and the second anchor point vector as negative samples; Determine the M first anchor point vectors, the M second anchor point vectors, and the plurality of target anchor point vectors as training samples of the multi-target model; The step of training the multi-objective model according to the first loss value and the second loss value to obtain the trained multi-objective model includes: The multi-objective model is trained according to the first loss value, the second loss value and the training sample to obtain the trained multi-objective model.

5. The method according to claim 1, characterized in that The calculating the loss value between the M first anchor point vectors and the M second anchor point vectors according to the first loss function to obtain the first loss value includes: Generate a first loss function, where the first loss function is the sum of the ratios of a first exponential function and a second exponential function, where the first exponential function is an exponential function of M cosine similarities between the M first anchor point vectors and the corresponding M second anchor point vectors, and the second exponential function is the sum of the exponential functions of the M cosine similarities; Substitute the M first anchor point vectors and the M second anchor point vectors into the first loss function for calculation to obtain the first loss value.

6. The method according to claim 1, characterized in that The step of training the multi-objective model according to the first loss value and the second loss value to obtain the trained multi-objective model includes: Calculating the M first feature vectors according to a multi-objective supervised learning algorithm to obtain a second loss value; Setting a first weight coefficient and a second weight coefficient based on the multi-objective model, wherein the first weight coefficient corresponds to the first loss value, and the second weight coefficient corresponds to the second loss value; Performing weighted calculation on the first loss value and the second loss value based on the first weight coefficient and the second weight coefficient respectively to obtain a target loss value; The multi-objective model is trained according to the target loss value to obtain a trained multi-objective model.

7. A training device for a multi-objective model, characterized in that: The device comprises: A generating module, used to generate M first feature vectors based on the video information of the server and the user, wherein the first feature vector is an N-dimensional feature vector used to represent the video information, wherein M is a positive integer and N is a positive integer; A processing module, configured to perform target processing on the M first feature vectors based on a target binomial distribution probability to obtain M first anchor point vectors and M second anchor point vectors associated with the target binomial distribution probability, wherein the target processing includes data expansion and dimensionality reduction processing; A calculation module, configured to calculate a loss value between the M first anchor point vectors and the M second anchor point vectors according to a first loss function to obtain a first loss value; A training module is used to train the multi-objective model according to the first loss value and the second loss value to obtain a trained multi-objective model, wherein the multi-objective model is a model that simultaneously considers multiple input objectives in the optimization problem, and the second loss value is a loss value obtained by performing loss calculation on the M first feature vectors according to a multi-objective supervised learning algorithm associated with the multi-objective model.

8. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the multi-objective model training method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the training method of the multi-objective model as described in any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product is stored in a storage medium, and the computer program product is executed by at least one processor to implement the steps in the training method of the multi-objective model according to any one of claims 1 to 6.