Information recommendation, multi-target recommendation model training method and device, and computer equipment
By performing biased prediction and interaction prediction on a multi-objective recommendation model, and utilizing a combination of object attribute features and features of the information to be recommended, the likelihood of interaction is corrected, thereby improving the accuracy of information recommendation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-05-13
- Publication Date
- 2026-05-19
AI Technical Summary
Existing multi-objective recommendation models cannot accurately determine the recommended information in information recommendation, resulting in low recommendation accuracy.
By obtaining the object attribute features of the object to be recommended, the underlying features are extracted, bias prediction and interaction prediction are performed, the degree of bias is used to correct the interaction probability, and the features are fused to obtain the fused recommendation degree. Recommendation is made when the preset conditions are met.
It improves the accuracy of information recommendation, enhances the ability to represent recommended information, and ensures the accuracy of recommendation results.
Smart Images

Figure CN117112880B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to an information recommendation, multi-objective recommendation model training method, apparatus, computer device, storage medium and computer program product. Background Technology
[0002] With the development of artificial intelligence technology, intelligent recommendation technology has emerged. This technology typically uses a multi-objective recommendation model for information recommendation. This model simultaneously predicts multiple business objectives and then fuses the prediction results to determine whether to recommend information. Currently, the common approach is to obtain the attributes of the recommended object and the information to be recommended, perform a global modeling process to obtain the multi-objective recommendation model, and then use this globally modeled model for information recommendation. However, global modeling means that the resulting multi-objective recommendation model can only learn global information. Using global information for recommendation makes it impossible to accurately determine the recommended information, resulting in low accuracy of the recommended information. Summary of the Invention
[0003] Therefore, it is necessary to provide an information recommendation, multi-objective recommendation model training method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the accuracy of information recommendation in response to the above-mentioned technical problems.
[0004] On the one hand, this application provides an information recommendation method. The method includes:
[0005] Obtain the object attribute features and recommendation information features corresponding to the object to be recommended;
[0006] Based on object attribute features, low-level features are extracted to obtain object extracted features. Based on the object extracted features, bias prediction is performed on each business objective to obtain the degree of bias of the object to be recommended to each business objective.
[0007] The object attribute features and the information features to be recommended are combined to obtain combined features. Based on the combined features, object extraction features and information features to be recommended, interaction prediction is performed on each business objective to obtain the interaction probability of the object to be recommended with each business objective.
[0008] By correcting the corresponding interaction possibilities based on the degree of bias of each business objective, the interaction possibilities of each objective are obtained, and the interaction possibilities of each objective are merged to obtain the degree of fusion recommendation corresponding to the information to be recommended.
[0009] When the degree of fusion recommendation meets the preset recommendation conditions, the information to be recommended will be recommended to the terminal corresponding to the object to be recommended.
[0010] On the other hand, this application also provides an information recommendation device. The device includes:
[0011] The feature acquisition module is used to acquire the object attribute features and recommendation information features corresponding to the object to be recommended.
[0012] The bias prediction module is used to extract low-level features based on object attribute features to obtain object extracted features. Based on the object extracted features, bias prediction is performed on each business objective to obtain the degree of bias of the object to be recommended towards each business objective.
[0013] The interaction prediction module is used to combine the object attribute features and the features of the information to be recommended to obtain combined features. Based on the combined features, object extraction features and features of the information to be recommended, the module performs interaction prediction on each business objective to obtain the interaction probability of the object to be recommended with each business objective.
[0014] The fusion module is used to correct the corresponding interaction possibilities of each business objective based on the degree of bias of each objective, obtain the interaction possibilities of each objective, and fuse the interaction possibilities of each objective to obtain the fusion recommendation degree of the information to be recommended.
[0015] The recommendation module is used to recommend information to the terminal corresponding to the recommended object when the degree of fusion recommendation meets the preset recommendation conditions.
[0016] On the other hand, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0017] Obtain the object attribute features and recommendation information features corresponding to the object to be recommended;
[0018] Based on object attribute features, low-level features are extracted to obtain object extracted features. Based on the object extracted features, bias prediction is performed on each business objective to obtain the degree of bias of the object to be recommended to each business objective.
[0019] The object attribute features and the information features to be recommended are combined to obtain combined features. Based on the combined features, object extraction features and information features to be recommended, interaction prediction is performed on each business objective to obtain the interaction probability of the object to be recommended with each business objective.
[0020] By correcting the corresponding interaction possibilities based on the degree of bias of each business objective, the interaction possibilities of each objective are obtained, and the interaction possibilities of each objective are merged to obtain the degree of fusion recommendation corresponding to the information to be recommended.
[0021] When the degree of fusion recommendation meets the preset recommendation conditions, the information to be recommended will be recommended to the terminal corresponding to the object to be recommended.
[0022] On the other hand, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0023] Obtain the object attribute features and recommendation information features corresponding to the object to be recommended;
[0024] Based on object attribute features, low-level features are extracted to obtain object extracted features. Based on the object extracted features, bias prediction is performed on each business objective to obtain the degree of bias of the object to be recommended to each business objective.
[0025] The object attribute features and the information features to be recommended are combined to obtain combined features. Based on the combined features, object extraction features and information features to be recommended, interaction prediction is performed on each business objective to obtain the interaction probability of the object to be recommended with each business objective.
[0026] By correcting the corresponding interaction possibilities based on the degree of bias of each business objective, the interaction possibilities of each objective are obtained, and the interaction possibilities of each objective are merged to obtain the degree of fusion recommendation corresponding to the information to be recommended.
[0027] When the degree of fusion recommendation meets the preset recommendation conditions, the information to be recommended will be recommended to the terminal corresponding to the object to be recommended.
[0028] On the other hand, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0029] Obtain the object attribute features and recommendation information features corresponding to the object to be recommended;
[0030] Based on object attribute features, low-level features are extracted to obtain object extracted features. Based on the object extracted features, bias prediction is performed on each business objective to obtain the degree of bias of the object to be recommended to each business objective.
[0031] The object attribute features and the information features to be recommended are combined to obtain combined features. Based on the combined features, object extraction features and information features to be recommended, interaction prediction is performed on each business objective to obtain the interaction probability of the object to be recommended with each business objective.
[0032] By correcting the corresponding interaction possibilities based on the degree of bias of each business objective, the interaction possibilities of each objective are obtained, and the interaction possibilities of each objective are merged to obtain the degree of fusion recommendation corresponding to the information to be recommended.
[0033] When the degree of fusion recommendation meets the preset recommendation conditions, the information to be recommended will be recommended to the terminal corresponding to the object to be recommended.
[0034] The aforementioned information recommendation method, apparatus, computer equipment, storage medium, and computer program product extract low-level features through object attribute features to obtain object-extracted features. Based on these object-extracted features, bias prediction is performed on each business objective to obtain the degree of bias of the object to be recommended towards each business objective. Then, based on combined features, object-extracted features, and features of the information to be recommended, interaction prediction is performed on each business objective to obtain the interaction probability of the object to be recommended towards each business objective. The bias degree of each business objective is used to correct the corresponding interaction probability to obtain the interaction probability of each objective. Finally, the interaction probabilities of each objective are fused to obtain the fused recommendation degree corresponding to the information to be recommended. That is, by separating the object attribute part during multi-objective prediction and then using the bias degree to correct the interaction probability to obtain the target interaction probability, the representation ability of the information to be recommended is enhanced. Furthermore, the fusion of the interaction probabilities of each objective makes the obtained fused recommendation degree more accurate. Finally, when the fused recommendation degree meets the preset recommendation conditions, the information to be recommended is recommended to the terminal corresponding to the object to be recommended, thus improving the accuracy of information recommendation.
[0035] On the one hand, this application provides a method for training a multi-objective recommendation model. The method includes:
[0036] Obtain the positive and negative samples corresponding to the training recommendation objects. Positive samples include the features of the training recommendation objects and the positive recommendation information features, while negative samples include the features of the training recommendation objects and the negative recommendation information features.
[0037] The training recommendation object features and positive recommendation information features are input into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, so as to obtain the training bias degree and positive interaction probability of the training recommendation object for each business objective. The positive interaction probability corresponding to each business objective is corrected by the training bias degree of each business objective, so as to obtain the positive interaction probability of each objective. The positive interaction probabilities of each objective are fused to obtain the positive fusion recommendation degree corresponding to the positive recommendation information.
[0038] The training recommendation object features and negative recommendation information features are input into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, thereby obtaining the training bias degree and negative interaction probability of the training recommendation object for each business objective. The negative interaction probability corresponding to each business objective is corrected by the training bias degree of each business objective, thereby obtaining the negative interaction probability of each objective. The negative interaction probabilities of each objective are fused to obtain the negative fused recommendation degree corresponding to the negative recommendation information.
[0039] The model loss is calculated based on the positive and negative fusion recommendation levels to obtain model loss information;
[0040] The first initial multi-objective recommendation model is updated in reverse based on the model loss information to obtain the first updated multi-objective recommendation model. The first updated multi-objective recommendation model is used as the first initial multi-objective recommendation model, and the steps of obtaining the positive and negative samples corresponding to the training recommendation objects are returned to be executed until the first training completion condition is met, thus obtaining the first multi-objective recommendation model.
[0041] On the other hand, this application also provides a multi-objective recommendation model training device. The device includes:
[0042] The sample pair acquisition module is used to acquire positive and negative samples corresponding to the training recommendation objects. Positive samples include the features of the training recommendation objects and positive recommendation information features, while negative samples include the features of the training recommendation objects and negative recommendation information features.
[0043] The positive prediction module is used to input the training recommendation object features and positive recommendation information features into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, obtain the training bias degree and positive interaction probability of the training recommendation object for each business objective, and correct the corresponding positive interaction probability by the training bias degree of each business objective to obtain the positive interaction probability of each objective. Finally, the positive interaction probabilities of each objective are fused to obtain the positive fusion recommendation degree corresponding to the positive recommendation information.
[0044] The negative prediction module is used to input the features of the training recommendation objects and the features of negative recommendation information into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, obtain the training bias degree and negative interaction probability of the training recommendation objects for each business objective, and correct the corresponding negative interaction probability by the training bias degree of each business objective to obtain the negative interaction probability of each objective. Finally, the negative interaction probabilities of each objective are fused to obtain the negative fused recommendation degree corresponding to the negative recommendation information.
[0045] The loss calculation module is used to calculate the model loss based on the positive fusion recommendation degree and the negative fusion recommendation degree, and obtain the model loss information;
[0046] The first iteration module is used to back-update the first initial multi-objective recommendation model based on the model loss information to obtain the first updated multi-objective recommendation model. The first updated multi-objective recommendation model is used as the first initial multi-objective recommendation model, and the step of obtaining the positive and negative samples corresponding to the training recommendation objects is returned to be executed until the first training completion condition is met, and the first multi-objective recommendation model is obtained.
[0047] On the other hand, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0048] Obtain the positive and negative samples corresponding to the training recommendation objects. Positive samples include the features of the training recommendation objects and the positive recommendation information features, while negative samples include the features of the training recommendation objects and the negative recommendation information features.
[0049] The training recommendation object features and positive recommendation information features are input into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, so as to obtain the training bias degree and positive interaction probability of the training recommendation object for each business objective. The positive interaction probability corresponding to each business objective is corrected by the training bias degree of each business objective, so as to obtain the positive interaction probability of each objective. The positive interaction probabilities of each objective are fused to obtain the positive fusion recommendation degree corresponding to the positive recommendation information.
[0050] The training recommendation object features and negative recommendation information features are input into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, thereby obtaining the training bias degree and negative interaction probability of the training recommendation object for each business objective. The negative interaction probability corresponding to each business objective is corrected by the training bias degree of each business objective, thereby obtaining the negative interaction probability of each objective. The negative interaction probabilities of each objective are fused to obtain the negative fused recommendation degree corresponding to the negative recommendation information.
[0051] The model loss is calculated based on the positive and negative fusion recommendation levels to obtain model loss information;
[0052] The first initial multi-objective recommendation model is updated in reverse based on the model loss information to obtain the first updated multi-objective recommendation model. The first updated multi-objective recommendation model is used as the first initial multi-objective recommendation model, and the steps of obtaining the positive and negative samples corresponding to the training recommendation objects are returned to be executed until the first training completion condition is met, thus obtaining the first multi-objective recommendation model.
[0053] On the other hand, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0054] Obtain the positive and negative samples corresponding to the training recommendation objects. Positive samples include the features of the training recommendation objects and the positive recommendation information features, while negative samples include the features of the training recommendation objects and the negative recommendation information features.
[0055] The training recommendation object features and positive recommendation information features are input into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, so as to obtain the training bias degree and positive interaction probability of the training recommendation object for each business objective. The positive interaction probability corresponding to each business objective is corrected by the training bias degree of each business objective, so as to obtain the positive interaction probability of each objective. The positive interaction probabilities of each objective are fused to obtain the positive fusion recommendation degree corresponding to the positive recommendation information.
[0056] The training recommendation object features and negative recommendation information features are input into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, thereby obtaining the training bias degree and negative interaction probability of the training recommendation object for each business objective. The negative interaction probability corresponding to each business objective is corrected by the training bias degree of each business objective, thereby obtaining the negative interaction probability of each objective. The negative interaction probabilities of each objective are fused to obtain the negative fused recommendation degree corresponding to the negative recommendation information.
[0057] The model loss is calculated based on the positive and negative fusion recommendation levels to obtain model loss information;
[0058] The first initial multi-objective recommendation model is updated in reverse based on the model loss information to obtain the first updated multi-objective recommendation model. The first updated multi-objective recommendation model is used as the first initial multi-objective recommendation model, and the steps of obtaining the positive and negative samples corresponding to the training recommendation objects are returned to be executed until the first training completion condition is met, thus obtaining the first multi-objective recommendation model.
[0059] On the other hand, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0060] Obtain the positive and negative samples corresponding to the training recommendation objects. Positive samples include the features of the training recommendation objects and the positive recommendation information features, while negative samples include the features of the training recommendation objects and the negative recommendation information features.
[0061] The training recommendation object features and positive recommendation information features are input into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, so as to obtain the training bias degree and positive interaction probability of the training recommendation object for each business objective. The positive interaction probability corresponding to each business objective is corrected by the training bias degree of each business objective, so as to obtain the positive interaction probability of each objective. The positive interaction probabilities of each objective are fused to obtain the positive fusion recommendation degree corresponding to the positive recommendation information.
[0062] The training recommendation object features and negative recommendation information features are input into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, thereby obtaining the training bias degree and negative interaction probability of the training recommendation object for each business objective. The negative interaction probability corresponding to each business objective is corrected by the training bias degree of each business objective, thereby obtaining the negative interaction probability of each objective. The negative interaction probabilities of each objective are fused to obtain the negative fused recommendation degree corresponding to the negative recommendation information.
[0063] The model loss is calculated based on the positive and negative fusion recommendation levels to obtain model loss information;
[0064] The first initial multi-objective recommendation model is updated in reverse based on the model loss information to obtain the first updated multi-objective recommendation model. The first updated multi-objective recommendation model is used as the first initial multi-objective recommendation model, and the steps of obtaining the positive and negative samples corresponding to the training recommendation objects are returned to be executed until the first training completion condition is met, thus obtaining the first multi-objective recommendation model.
[0065] The aforementioned multi-objective recommendation model training method, apparatus, computer equipment, storage medium, and computer program product improve the accuracy of obtaining the positive fusion recommendation degree by performing multi-objective prediction on positive samples, and then performing multi-objective prediction on negative samples to obtain the negative fusion recommendation degree, thus improving the accuracy of obtaining the negative fusion recommendation degree. The positive and negative fusion recommendation degrees are then used to calculate the model loss, obtaining model loss information, which improves the accuracy of the model loss information. Finally, the model loss information is used to train a first initial multi-objective recommendation model. When the first training completion condition is met, a first multi-objective recommendation model is obtained, thereby improving the accuracy of the first multi-objective recommendation model, and ultimately enhancing the accuracy of information recommendation.
[0066] On the one hand, this application provides a method for training a multi-objective recommendation model. The method includes:
[0067] Obtain training samples, which include training recommendation object information, training recommendation information, and various interaction labels;
[0068] The training recommendation object features and training recommendation information features are input into the second initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective. This yields the training bias degree and training interaction probability of the training recommendation object for each business objective. The training bias degree of each business objective is used to correct the corresponding training interaction probability, resulting in the training interaction probability of each objective. The training interaction probabilities of each objective are then fused to obtain the training fusion recommendation degree corresponding to the training recommendation information. The training bias degree of each business objective is then fused to obtain the training bias recommendation degree corresponding to the training recommendation information. Finally, the sum of the training bias recommendation degree and the training fusion recommendation degree is calculated to obtain the training objective recommendation degree.
[0069] The fusion loss is calculated based on the recommendation degree of each interaction label and the training target to obtain the fusion loss information. The fusion loss information is then adjusted by using the degree of each training bias to obtain the target loss information.
[0070] The second initial multi-objective recommendation model is updated in reverse based on the target loss information to obtain the second updated multi-objective recommendation model. The second updated multi-objective recommendation model is used as the second initial multi-objective recommendation model, and the step of obtaining training samples is returned to execute until the second training completion condition is met, thus obtaining the second multi-objective recommendation model.
[0071] On the other hand, this application also provides a multi-objective recommendation model training device. The device includes:
[0072] The sample acquisition module is used to acquire training samples, which include training recommendation object information, training recommendation information, and various interaction tags.
[0073] The training module is used to input the features of the training recommendation objects and the features of the training recommendation information into the second initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective. It obtains the training bias degree and training interaction probability of the training recommendation objects for each business objective, and corrects the corresponding training interaction probability by the training bias degree of each business objective to obtain the training interaction probability of each objective. It then merges the training interaction probabilities of each objective to obtain the training fusion recommendation degree corresponding to the training recommendation information. Finally, it merges the training bias degrees of each business objective to obtain the training bias recommendation degree corresponding to the training recommendation information, and calculates the sum of the training bias recommendation degree and the training fusion recommendation degree to obtain the training objective recommendation degree.
[0074] The target loss calculation module is used to calculate the fusion loss based on each interaction label and the degree of recommendation of the training target, obtain the fusion loss information, and adjust the fusion loss information with each training bias degree to obtain the target loss information;
[0075] The second iteration module is used to update the second initial multi-objective recommendation model in reverse based on the target loss information, to obtain the second updated multi-objective recommendation model. The second updated multi-objective recommendation model is used as the second initial multi-objective recommendation model, and the step of obtaining training samples is returned to execute until the second training completion condition is met, and the second multi-objective recommendation model is obtained.
[0076] On the other hand, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0077] Obtain training samples, which include training recommendation object information, training recommendation information, and various interaction labels;
[0078] The training recommendation object features and training recommendation information features are input into the second initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective. This yields the training bias degree and training interaction probability of the training recommendation object for each business objective. The training bias degree of each business objective is used to correct the corresponding training interaction probability, resulting in the training interaction probability of each objective. The training interaction probabilities of each objective are then fused to obtain the training fusion recommendation degree corresponding to the training recommendation information. The training bias degree of each business objective is then fused to obtain the training bias recommendation degree corresponding to the training recommendation information. Finally, the sum of the training bias recommendation degree and the training fusion recommendation degree is calculated to obtain the training objective recommendation degree.
[0079] The fusion loss is calculated based on the recommendation degree of each interaction label and the training target to obtain the fusion loss information. The fusion loss information is then adjusted by using the degree of each training bias to obtain the target loss information.
[0080] The second initial multi-objective recommendation model is updated in reverse based on the target loss information to obtain the second updated multi-objective recommendation model. The second updated multi-objective recommendation model is used as the second initial multi-objective recommendation model, and the step of obtaining training samples is returned to execute until the second training completion condition is met, thus obtaining the second multi-objective recommendation model.
[0081] On the other hand, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0082] Obtain training samples, which include training recommendation object information, training recommendation information, and various interaction labels;
[0083] The training recommendation object features and training recommendation information features are input into the second initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective. This yields the training bias degree and training interaction probability of the training recommendation object for each business objective. The training bias degree of each business objective is used to correct the corresponding training interaction probability, resulting in the training interaction probability of each objective. The training interaction probabilities of each objective are then fused to obtain the training fusion recommendation degree corresponding to the training recommendation information. The training bias degree of each business objective is then fused to obtain the training bias recommendation degree corresponding to the training recommendation information. Finally, the sum of the training bias recommendation degree and the training fusion recommendation degree is calculated to obtain the training objective recommendation degree.
[0084] The fusion loss is calculated based on the recommendation degree of each interaction label and the training target to obtain the fusion loss information. The fusion loss information is then adjusted by using the degree of each training bias to obtain the target loss information.
[0085] The second initial multi-objective recommendation model is updated in reverse based on the target loss information to obtain the second updated multi-objective recommendation model. The second updated multi-objective recommendation model is used as the second initial multi-objective recommendation model, and the step of obtaining training samples is returned to execute until the second training completion condition is met, thus obtaining the second multi-objective recommendation model.
[0086] On the other hand, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0087] Obtain training samples, which include training recommendation object information, training recommendation information, and various interaction labels;
[0088] The training recommendation object features and training recommendation information features are input into the second initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective. This yields the training bias degree and training interaction probability of the training recommendation object for each business objective. The training bias degree of each business objective is used to correct the corresponding training interaction probability, resulting in the training interaction probability of each objective. The training interaction probabilities of each objective are then fused to obtain the training fusion recommendation degree corresponding to the training recommendation information. The training bias degree of each business objective is then fused to obtain the training bias recommendation degree corresponding to the training recommendation information. Finally, the sum of the training bias recommendation degree and the training fusion recommendation degree is calculated to obtain the training objective recommendation degree.
[0089] The fusion loss is calculated based on the recommendation degree of each interaction label and the training target to obtain the fusion loss information. The fusion loss information is then adjusted by using the degree of each training bias to obtain the target loss information.
[0090] The second initial multi-objective recommendation model is updated in reverse based on the target loss information to obtain the second updated multi-objective recommendation model. The second updated multi-objective recommendation model is used as the second initial multi-objective recommendation model, and the step of obtaining training samples is returned to execute until the second training completion condition is met, thus obtaining the second multi-objective recommendation model.
[0091] The aforementioned multi-objective recommendation model training method, apparatus, computer equipment, storage medium, and computer program product input the training recommendation object features and training recommendation information features into a second initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective. This yields the training bias degree and training interaction probability of the training recommendation object towards each business objective. The training bias degree of each business objective is used to correct its corresponding training interaction probability, resulting in the training interaction probability of each objective. These training interaction probabilities are then fused to obtain the training fusion recommendation degree corresponding to the training recommendation information. The training bias degrees of each business objective are also fused to obtain the training bias recommendation degree corresponding to the training recommendation information. The sum of the training bias recommendation degree and the training fusion recommendation degree is calculated to obtain the training target recommendation degree, thus improving the accuracy of the training target recommendation degree. This training target recommendation degree is then used to calculate the fusion loss information, and the training bias degree is used to adjust the fusion loss information to obtain the target loss information, thereby improving the accuracy of the obtained target loss information. Finally, the target loss information is used to train the second initial multi-objective recommendation model. When the second training completion condition is met, the second multi-objective recommendation model is obtained, improving the accuracy of the obtained second multi-objective recommendation model. Attached Figure Description
[0092] Figure 1 This is a diagram illustrating the application environment of an information recommendation method in one embodiment.
[0093] Figure 2 This is a flowchart illustrating an information recommendation method in one embodiment;
[0094] Figure 3 This is a schematic diagram illustrating the process of obtaining the degree of fusion recommendation in one embodiment;
[0095] Figure 4 This is a schematic diagram of the network structure of the first multi-objective prediction network in a specific embodiment;
[0096] Figure 5 This is a schematic diagram of the network structure of an interactive fusion network in a specific embodiment;
[0097] Figure 6 This is a schematic diagram of the network structure for obtaining the degree of biased recommendation in a specific embodiment;
[0098] Figure 7This is a flowchart illustrating a multi-objective recommendation model training method in one embodiment;
[0099] Figure 8 This is a schematic diagram of the process for obtaining model loss information in one embodiment;
[0100] Figure 9 This is a flowchart illustrating a multi-objective recommendation model training method in another embodiment;
[0101] Figure 10 This is a schematic diagram of the process for obtaining the recommendation level of the training target in one embodiment;
[0102] Figure 11 This is a schematic diagram of the process for obtaining target loss information in one embodiment;
[0103] Figure 12 This is a schematic diagram of the process for obtaining fusion loss information in one embodiment;
[0104] Figure 13 This is a schematic diagram illustrating the process of obtaining sample weights in one embodiment;
[0105] Figure 14 This is a schematic diagram illustrating the correction process for a second multi-objective recommendation model in a specific embodiment.
[0106] Figure 15 for Figure 14 A schematic diagram of the news recommendation page in a specific embodiment;
[0107] Figure 16-A This is a schematic diagram showing a test comparison of click-through rates for images and text in a specific embodiment.
[0108] Figure 16-B for Figure 16-A A comparative diagram of test result indicators in a specific embodiment;
[0109] Figure 17-A This is a schematic diagram comparing the click-through rates of images and text in another specific embodiment;
[0110] Figure 17-B for Figure 17-A A comparative diagram of test result indicators in a specific embodiment;
[0111] Figure 18 This is a structural block diagram of an information recommendation device in one embodiment;
[0112] Figure 19 This is a structural block diagram of a multi-objective recommendation model training device in one embodiment;
[0113] Figure 20 This is a structural block diagram of a multi-objective recommendation model training device in another embodiment;
[0114] Figure 21 This is an internal structural diagram of a computer device in one embodiment;
[0115] Figure 22 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation
[0116] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0117] The information recommendation method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another server. Server 104 obtains the object attribute features and recommendation information features corresponding to the object to be recommended from the data storage system. Server 104 performs low-level feature extraction based on the object attribute features to obtain object-extracted features. Based on the object-extracted features, it performs bias prediction on each business objective to obtain the degree of bias of the object to be recommended towards each business objective. Server 104 combines the object attribute features and the recommendation information features to obtain combined features. Based on the combined features, object-extracted features, and recommendation information features, it performs interaction prediction on each business objective to obtain the interaction probability of the object to be recommended towards each business objective. Server 104 corrects the corresponding interaction probability based on the degree of bias of each business objective to obtain the interaction probability of each objective, and merges the interaction probabilities of each objective to obtain the fused recommendation degree corresponding to the information to be recommended. When the fused recommendation degree meets the preset recommendation conditions, server 104 recommends the information to be recommended to the terminal 102 corresponding to the object to be recommended. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. Server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0118] In one embodiment, such as Figure 2 As shown, an information recommendation method is provided, which can be applied to... Figure 1Taking a server as an example, it can be understood that this method can also be applied to servers, and also to systems including terminals and servers, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0119] Step 202: Obtain the object attribute features and recommendation information features corresponding to the object to be recommended.
[0120] The "object to be recommended" refers to the object for which information needs to be recommended. This object can be a real object or a virtual object. Real objects include anything from people to animals. Virtual objects can refer to objects using virtual avatars in online live streaming. "Object attribute features" refers to the characteristics corresponding to the attributes of the object to be recommended. These object attribute features can be used to characterize the object to be recommended. These object attribute features include, but are not limited to, the object's basic attribute features, social relationship features, and spending power features. The basic attribute features characterize the object's basic attributes, the social relationship features characterize the object's social relationships, and the spending power features characterize the object's spending power. "Information to be recommended features" refers to the features of the information to be recommended. This information is the information that needs to be determined whether to recommend it to the object to be recommended. This information can be multimedia data, including but not limited to video, images, text, and audio.
[0121] Specifically, the server can retrieve the object attribute features and recommendation information features corresponding to the object to be recommended from the database. The server can also retrieve these features from the data service provider. Furthermore, the server can retrieve the object attribute features and recommendation information features corresponding to the object to be recommended uploaded by the terminal. Finally, the server can retrieve these features from the Internet. In one embodiment, the server can retrieve the object attribute information and recommendation information corresponding to the object to be recommended, and then perform data preprocessing on the object data and recommendation information to obtain the object attribute features and recommendation information features.
[0122] Step 204: Extract low-level features based on object attribute features to obtain object extracted features. Based on the object extracted features, predict the bias of each business objective to obtain the degree of bias of the object to be recommended to each business objective.
[0123] Among these, object extraction features refer to the semantic feature extraction obtained from the object's attribute features, which is a low-dimensional representation of the object. Business objectives refer to the goals to be achieved after the recommendation process, which may include clicks, reads, interactions, and exposures, etc. Bias degree is used to characterize the tendency of the recommended object towards the business objective; the higher the bias degree, the stronger the tendency of the recommended object towards that business objective, indicating that the business objective is more likely to be achieved. Different recommended objects tend to favor different business objectives. For example, older recommended objects tend to read for extended periods, while younger recommended objects tend to interact.
[0124] Specifically, the server uses a deep neural network for low-level feature extraction, that is, it compresses the object attribute features through the deep neural network to obtain object-extracted features. Then, the object-extracted features are used to perform multi-objective bias prediction through the deep neural network, that is, to predict the bias of each business objective and obtain the degree of bias of the object to be recommended to each business objective.
[0125] Step 206: Combine the object attribute features and the information features to be recommended to obtain combined features. Based on the combined features, object extraction features and information features to be recommended, perform interaction prediction on each business objective to obtain the interaction probability of the object to be recommended with each business objective.
[0126] Feature combination refers to the synthesis of features formed by combining individual features (through multiplication or Cartesian product). Feature combination helps to represent nonlinear relationships. Combined features refer to features obtained by combining or crossing features.
[0127] Interaction probability characterizes the likelihood that a recommended object will interact with a corresponding business objective. The higher the interaction probability of a business objective, the more likely the recommended object is to interact with the recommended information corresponding to that business objective. For example, a higher click interaction probability means a higher probability that the recommended object will click on the corresponding recommended information.
[0128] Specifically, the server multiplies or calculates a Cartesian product between object attribute features and features to be recommended, obtaining combined features. Alternatively, the server can cross-interchange object attribute features and features to be recommended to obtain combined features. Then, the combined features, object extraction features, and features to be recommended are used to perform multi-objective interaction prediction through a deep neural network, i.e., predicting the interaction between various business objectives, and obtaining the output probability of the recommended object's interaction with each business objective. In one embodiment, the network structure for multi-objective bias prediction and multi-objective interaction prediction is the same, but the network parameters are different.
[0129] Step 208: Correct the corresponding interaction possibilities of each business objective by the degree of bias, obtain the interaction possibilities of each objective, and merge the interaction possibilities of each objective to obtain the fused recommendation degree of the information to be recommended.
[0130] Among them, the target interaction probability refers to the interaction probability adjusted using a bias degree. This target interaction probability represents the part unrelated to the object to be recommended, i.e., the interaction probability predicted from the information to be recommended. By correcting the corresponding interaction probabilities using the bias degree of the business objectives, the accuracy of the obtained target interaction probabilities is improved. Each business objective obtains a corresponding target interaction probability. The fusion recommendation degree is used to characterize the degree to which the information to be recommended can be recommended to the object to be recommended. The higher the fusion recommendation degree, the more likely the corresponding information to be recommended can be recommended to the object to be recommended.
[0131] Specifically, the server can correct the interaction probabilities of each business objective by using the degree of bias of each business objective. This can be done by calculating the difference between the interaction probabilities of each business objective and its corresponding degree of bias to obtain the target interaction probability. Then, the interaction probabilities of each target can be fused to obtain the fused recommendation degree corresponding to the information to be recommended. The server can use a deep neural network to fuse the interaction probabilities of each target to obtain the fused recommendation degree. Alternatively, the server can use pre-set weights of business objectives to weight the interaction probabilities of each target and calculate the weighted sum to obtain the fused recommendation degree.
[0132] Step 210: When the degree of fusion recommendation meets the preset recommendation conditions, the information to be recommended is recommended to the terminal corresponding to the object to be recommended.
[0133] Among them, the preset recommendation conditions refer to the pre-set conditions for making recommendations. These conditions can be either reaching a recommendation level threshold or reaching the target ranking position in the fusion recommendation level sequence corresponding to each piece of information to be recommended.
[0134] Specifically, the server determines whether the degree of fusion recommendation meets the preset recommendation conditions. This can be done by comparing the degree of fusion recommendation with a recommendation degree threshold; if it exceeds the threshold, it meets the preset recommendation conditions. Alternatively, it can determine the position of the degree of fusion recommendation in the fusion recommendation degree sequence; if the position is within the target ranking position, it meets the recommendation conditions. In this case, the information to be recommended is recommended to the terminal corresponding to the target object. This terminal includes, but is not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, in-vehicle terminals, and aircraft.
[0135] The aforementioned information recommendation method, apparatus, computer equipment, storage medium, and computer program product extract low-level features through object attribute features to obtain object-extracted features. Based on these object-extracted features, bias prediction is performed on each business objective to obtain the degree of bias of the object to be recommended towards each business objective. Then, based on combined features, object-extracted features, and features of the information to be recommended, interaction prediction is performed on each business objective to obtain the interaction probability of the object to be recommended towards each business objective. The bias degree of each business objective is used to correct the corresponding interaction probability to obtain the interaction probability of each objective. Finally, the interaction probabilities of each objective are fused to obtain the fused recommendation degree corresponding to the information to be recommended. That is, by separating the object attribute part during multi-objective prediction and then using the bias degree to correct the interaction probability to obtain the target interaction probability, the representation ability of the information to be recommended is enhanced. Furthermore, the fusion of the interaction probabilities of each objective makes the obtained fused recommendation degree more accurate. Finally, when the fused recommendation degree meets the preset recommendation conditions, the information to be recommended is recommended to the terminal corresponding to the object to be recommended, thus improving the accuracy of information recommendation.
[0136] In one embodiment, such as Figure 3 As shown, information recommendation methods also include:
[0137] Step 302: Input the object attribute features and the information features to be recommended into the first multi-objective recommendation model;
[0138] Step 304: Extract low-level features based on object attribute features using the first multi-objective recommendation model to obtain object extraction features. Based on the object extraction features, predict the bias of each business objective to obtain the degree of bias of the object to be recommended towards each business objective.
[0139] Step 306: Combine the object attribute features and the information features to be recommended using the first multi-objective recommendation model to obtain combined features. Based on the combined features, object extraction features and information features to be recommended, perform interaction prediction on each business objective to obtain the interaction probability of the object to be recommended with each business objective.
[0140] Step 308: The first multi-objective recommendation model is used to correct the corresponding interaction probability by using the bias degree of each business objective, so as to obtain the interaction probability of each objective, and the interaction probability of each objective is fused to obtain the fused recommendation degree corresponding to the information to be recommended.
[0141] The first multi-objective recommendation model is a multi-objective prediction model built using a deep neural network. This model is used for recommendations based on multiple business objectives. The first multi-objective recommendation model can be trained using training sample pairs, which include positive samples and negative samples. Positive samples are those with recommendation information already provided to the training object, while negative samples are those with recommendation information not provided to the training object. The training objective of the first multi-objective recommendation model is to maximize the difference between the fused recommendation degree corresponding to positive samples and the fused recommendation degree corresponding to negative samples.
[0142] Specifically, the server pre-trains a first multi-objective recommendation model using training samples and then deploys it. After obtaining the object attribute features and the features of the information to be recommended, the server calls the first multi-objective recommendation model, using these features as input. Upon receiving the input features, the first multi-objective recommendation model performs bias prediction on each business objective to obtain the degree of bias of the recommended object towards each objective. Simultaneously, it performs interaction prediction on each business objective to obtain the likelihood of interaction between the recommended object and each objective. After bias correction, the results are fused to obtain the fused recommendation degree corresponding to the information to be recommended. By calling the pre-trained first multi-objective recommendation model for recommendation prediction and obtaining the fused recommendation degree for the information to be recommended, recommendation efficiency can be improved.
[0143] In one embodiment, the first multi-objective recommendation model includes a first multi-objective prediction network and an interactive fusion network, wherein the first multi-objective prediction network includes a first biased prediction sub-network and a first interactive prediction sub-network.
[0144] Step 304: Extract low-level features based on object attribute features to obtain object-extracted features. Based on these features, predict the bias of each business objective to obtain the degree of bias of the object to be recommended towards each business objective, including:
[0145] The object attribute features are input into the first bias prediction subnetwork in the first multi-objective prediction network; the object attribute features are extracted at the low level through the first bias prediction subnetwork to obtain object extraction features, and bias prediction is performed on each business objective based on the object extraction features to obtain the degree of bias of the object to be recommended to each business objective.
[0146] The first multi-objective prediction network refers to a deep neural network used for multi-objective prediction. The first bias prediction sub-network refers to a deep neural network that predicts the degree of bias for each business objective. This first bias prediction sub-network is a sub-network within the first multi-objective prediction network. The deep neural network can be a DNN (Deep Neural Network), a CNN (Recurrent Neural Network), an RNN (Convolutional Neural Network), etc.
[0147] Specifically, the server uses the first bias prediction sub-network in the first multi-objective recommendation model to extract underlying features, then performs bias prediction on each business objective, and outputs the degree of bias of the object to be recommended to each business objective.
[0148] Step 306: Based on the combined features, object extraction features, and features of the information to be recommended, perform interaction prediction for each business objective to obtain the interaction probability of the object to be recommended with each business objective, including the following steps:
[0149] The combined features, object extraction features, and features of the information to be recommended are input into the first interaction prediction subnetwork in the first multi-objective prediction network; the first interaction prediction subnetwork uses the combined features, object extraction features, and features of the information to be recommended to perform interaction prediction for each business objective, thereby obtaining the interaction probability of the object to be recommended with each business objective.
[0150] Among them, the first interaction prediction subnetwork refers to a deep neural network that predicts the interaction probability of various business objectives.
[0151] Specifically, while predicting the degree of bias of each business objective, the server can also predict the interaction of business objectives. That is, the combined features, object extraction features and features of the information to be recommended are input into the first interaction prediction sub-network in the first multi-objective prediction network to obtain the output of the interaction probability of the object to be recommended with each business objective.
[0152] Step 308 involves fusing the interaction possibilities of each target to obtain the fused recommendation level corresponding to the information to be recommended, including the following steps:
[0153] The interaction possibilities of each target are concatenated to obtain a concatenated vector. This concatenated vector is then input into an interaction fusion network for fusion to obtain the fusion recommendation degree corresponding to the information to be recommended.
[0154] Here, the concatenated vector is the vector obtained by concatenating the first and last elements. The interaction fusion network is a deep neural network used to fuse interaction possibilities.
[0155] Specifically, the server uses each target interaction possibility as a vector element, and concatenates the first and last elements sequentially to obtain a spliced vector. The spliced vector is then input into the interaction fusion network for fusion to obtain the fusion recommendation degree corresponding to the output recommendation information.
[0156] In one specific embodiment, the first multi-objective prediction network can be a multi-objective prediction network using a PLE (a novel hierarchical extraction multi-task learning network architecture) structure, such as... Figure 4 The diagram shows the network structure of the first multi-objective prediction network, used for news recommendation. This news recommendation has multiple business objectives, including click objectives, duration objectives, and interaction objectives. The first multi-objective prediction network comprises a two-layer network structure, each layer being a multi-gated hybrid expert network. This multi-gated hybrid expert network is a commonly used network structure for multi-objective learning; the expert networks are mostly DNN network structures. Multiple expert networks are used to extract different features, and gating is used to assign weights to each expert. The features of the object to be recommended are input into the first multi-objective prediction network. First, the first-layer multi-gated hybrid expert network performs low-level feature extraction to obtain object-extracted features. Then, the object-extracted features are passed through a multi-task network to predict various interaction biases, resulting in the output interaction biases. Simultaneously, the combined features, object-extracted features, and the features of the information to be recommended are concatenated and input into the second-layer multi-gated hybrid expert network for low-level feature extraction, followed by interaction probability prediction, resulting in various interaction probabilities. Each prediction task predicts the corresponding interaction probability and bias. That is, the first multi-objective prediction network finally outputs the bias score and interaction score corresponding to the click business objective, the bias score and interaction score corresponding to the duration business objective, and the bias score and interaction score corresponding to the interaction business objective. Further, such as... Figure 5 As shown, a schematic diagram of an interactive fusion network is provided, in which the likelihood and bias of interaction can be represented by probability or score. The interaction score for the click target is obtained by subtracting the bias score from the interaction score for the click target; the interaction score for the duration target is obtained by subtracting the bias score from the interaction score for the duration target; and the interaction score for the interactive target is obtained by subtracting the bias score from the interaction score for the interactive target. These three interaction scores are then concatenated and input into the interactive fusion network (DNN) for fusion, yielding the output fused score.
[0157] In the above embodiments, multi-objective prediction is performed using a first multi-objective prediction network to obtain the output fusion recommendation level. Since the first multi-objective prediction network performs multi-objective prediction, it obtains the target interaction probability corresponding to each business objective. Then, the target interaction probabilities are fused through an interaction fusion network, thereby improving the accuracy of the obtained fusion recommendation level.
[0158] In one embodiment, the interaction probabilities of each target are corrected based on the degree of bias of each business objective, thereby obtaining the interaction probabilities of each target, including the following steps:
[0159] The difference between the interaction probability and the corresponding bias degree of each business objective is calculated to obtain the target interaction probability of each business objective.
[0160] Specifically, the server subtracts the bias degree corresponding to the business objective from the interaction probability of the business objective to obtain the target interaction probability of the business objective. By traversing each business objective, the target interaction probability of each business objective is obtained, which improves the accuracy of the obtained target interaction probability. In a specific embodiment, the target interaction probability can be calculated using the following formula (1).
[0161]
[0162] in, It refers to the likelihood of target interactions for business objectives. This refers to the possibility of interaction with business objectives. This refers to the degree of bias towards business objectives. "Task" refers to the business objective itself.
[0163] In one embodiment, the information recommendation method further includes the step of:
[0164] The biases of each business objective are merged to obtain the biased recommendation level corresponding to the information to be recommended; the sum of the biased recommendation level and the merged recommendation level is calculated to obtain the target recommendation level.
[0165] Among them, the biased recommendation degree is the probability of recommending the desired information using the biased recommendation degree. The target recommendation degree is used to characterize the probability of recommending the desired information to the desired object.
[0166] Specifically, the server can further fuse the biases of various business objectives. This fusion can be achieved through deep neural networks or by using pre-set weights of the business objectives to obtain the biased recommendation level corresponding to the information to be recommended.
[0167] In a specific embodiment, the target recommendation level can be calculated using the following formula (2).
[0168] log it = log. bias +log it debias Formula (2)
[0169] Here, log it refers to the degree of recommendation of the target. bias This refers to the degree of bias in recommending. debias This refers to the degree of integration and recommendation.
[0170] Step 210: When the degree of fusion recommendation meets the preset recommendation conditions, the information to be recommended is recommended to the terminal corresponding to the object to be recommended, including the following steps:
[0171] When the target recommendation level meets the preset target recommendation conditions, the information to be recommended will be recommended to the terminal corresponding to the target object.
[0172] Among them, the preset target recommendation conditions refer to the pre-set conditions for recommending the information to be recommended to the terminal corresponding to the object to be recommended, including but not limited to reaching the preset recommendation threshold or reaching the preset sorting position.
[0173] Specifically, the server determines whether the target recommendation level meets the preset target recommendation conditions. When it meets the preset target recommendation conditions, such as when the target recommendation level reaches the preset recommendation threshold, the server recommends the information to be recommended to the terminal corresponding to the target object.
[0174] In one embodiment, the bias of each business objective is fused to obtain the biased recommendation degree corresponding to the information to be recommended, and the method further includes the following steps:
[0175] The biases of each business objective are concatenated to obtain a bias concatenation vector. This bias concatenation vector is then input into a bias fusion network for fusion to obtain the bias recommendation level corresponding to the information to be recommended.
[0176] Among them, biased fusion network refers to a deep neural network that fuses the biases of various business objectives.
[0177] Specifically, the server uses the bias of each business objective as an element in a vector, concatenates the first and last elements to obtain a bias concatenated vector, and then inputs the bias concatenated vector into the bias fusion network to fuse the data using network parameters, thereby obtaining the bias recommendation degree corresponding to the output recommendation information.
[0178] In a specific embodiment, such as Figure 6The diagram illustrates the network structure for obtaining the degree of biased recommendation. The server concatenates the bias scores for click-based business objectives, duration-based business objectives, and interaction-based business objectives. This concatenated bias score is then input into the biased fusion network to obtain the output biased fusion score. The sum of the biased fusion score and the interaction fusion score is then calculated to obtain the target score. Based on the target score, it can be determined whether to recommend the information to the corresponding terminal. For example, if the target score exceeds a threshold of 95 points, the information is recommended to the corresponding terminal. Similarly, if the target score of the information ranks among the top three in the sequence of target scores for all information to be recommended, it is recommended to the corresponding terminal.
[0179] In one embodiment, the server can also directly compare the interaction fusion score with the preset target recommendation conditions. When the interaction fusion score meets the preset target recommendation conditions, the information to be recommended will be recommended to the terminal corresponding to the object to be recommended.
[0180] In the above embodiments, by separating the recommendation object attributes and the information to be recommended during multi-target fusion, and performing multi-target recommendation fusion on the recommendation object attributes and the information to be recommended separately, the biased recommendation degree and the fused recommendation degree are finally obtained. This enhances the representation ability of the information to be recommended and improves the accuracy of the obtained biased recommendation degree and fused recommendation degree. Then, the sum of the biased recommendation degree and the fused recommendation degree is calculated to obtain the target recommendation degree, which further improves the accuracy of the obtained target recommendation degree.
[0181] In one embodiment, such as Figure 7 As shown, a method for training a multi-objective recommendation model is provided, which is then applied to... Figure 1 Taking a server as an example, it can be understood that this method can also be applied to servers, and also to systems including terminals and servers, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0182] Step 702: Obtain the positive and negative samples corresponding to the training recommendation objects. The positive samples include the training recommendation object features and positive recommendation information features, and the negative samples include the training recommendation object features and negative recommendation information features.
[0183] Here, "training recommendation object" refers to the object to which information will be recommended during training. "Training recommendation object features" refers to the features corresponding to the attribute information of the training recommendation object. "Positive recommendation information features" refers to the features of information that has already been recommended to the training recommendation object. "Negative recommendation information features" refers to the features of information that has not been recommended to the training recommendation object; the recommendation level of this unrecommended information does not meet the preset recommendation criteria.
[0184] Specifically, the server can obtain the positive and negative samples corresponding to the training recommendation objects from a database, or from a data service provider. Alternatively, the server can obtain the positive and negative samples from the terminal.
[0185] Step 704: Input the training recommendation object features and positive recommendation information features into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, obtain the training bias degree and positive interaction probability of the training recommendation object for each business objective, and correct the corresponding positive interaction probability by the training bias degree of each business objective to obtain the positive interaction probability of each objective, and fuse the positive interaction probabilities of each objective to obtain the positive fusion recommendation degree corresponding to the positive recommendation information.
[0186] Here, the first initial multi-objective recommendation model refers to the first multi-objective recommendation model with initialized model parameters. Training bias refers to the degree of bias predicted during training; positive interaction probability refers to the interaction probability predicted using positive samples during training; target positive interaction probability refers to the target interaction probability predicted using positive samples during training; and positive fusion recommendation degree refers to the fusion recommendation degree predicted using positive samples during training.
[0187] Specifically, the training recommendation object features and positive recommendation information features are input into the first initial multi-objective recommendation model. The first initial multi-objective recommendation model uses the training recommendation object features to extract underlying features, obtaining training object extracted features. Based on these training object extracted features, bias prediction is performed on each business objective, obtaining the training bias degree of the training recommendation object towards each business objective. Then, the training recommendation object features and positive recommendation information features are combined to obtain positive combined features. Next, the positive combined features, training object extracted features, and positive recommendation information features are used to predict interactions between each business objective, obtaining the positive interaction probability of the training recommendation object towards each business objective. Then, the training bias degree of each business objective is used to correct its corresponding positive interaction probability, obtaining the positive interaction probability of each objective. Finally, the positive interaction probabilities of each objective are fused to obtain the positive fused recommendation degree corresponding to the positive recommendation information.
[0188] Step 706: Input the training recommendation object features and negative recommendation information features into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, obtain the training bias degree and negative interaction probability of the training recommendation object for each business objective, and correct the corresponding negative interaction probability by the training bias degree of each business objective to obtain the negative interaction probability of each objective, and fuse the negative interaction probabilities of each objective to obtain the negative fused recommendation degree corresponding to the negative recommendation information.
[0189] Here, negative interaction probability refers to the interaction probability predicted using negative samples during training. Target negative interaction probability refers to the target interaction probability predicted using negative samples during training. Negative fusion recommendation degree refers to the fusion recommendation degree predicted using negative samples during training.
[0190] Specifically, the training recommendation object features and negative recommendation information features are input into the first initial multi-objective recommendation model. The first initial multi-objective recommendation model uses the training recommendation object features to extract underlying features, obtaining training object extracted features. Based on these training object extracted features, bias prediction is performed on each business objective, obtaining the training bias degree of the training recommendation object towards each business objective. Then, the training recommendation object features and negative recommendation information features are combined to obtain negative combined features. Next, the negative combined features, training object extracted features, and negative recommendation information features are used to predict interactions between each business objective, obtaining the negative interaction probability of the training recommendation object towards each business objective. Then, the training bias degree of each business objective is used to correct its corresponding negative interaction probability, obtaining the negative interaction probability of each objective. Finally, the negative interaction probabilities of each objective are fused to obtain the negative fused recommendation degree corresponding to the negative recommendation information.
[0191] Step 708: Calculate the model loss based on the positive fusion recommendation degree and the negative fusion recommendation degree to obtain the model loss information.
[0192] Among them, the model loss information is used to characterize the error between the positive fusion recommendation degree and the negative fusion recommendation degree.
[0193] Specifically, the server uses a loss function to calculate the error between the positive fusion recommendation level and the negative fusion recommendation level, thus obtaining model loss information.
[0194] Step 710: Based on the model loss information, the first initial multi-objective recommendation model is updated in reverse to obtain the first updated multi-objective recommendation model. The first updated multi-objective recommendation model is used as the first initial multi-objective recommendation model, and the steps of obtaining the positive and negative samples corresponding to the training recommendation objects are returned to be executed until the first training completion condition is met, and the first multi-objective recommendation model is obtained.
[0195] The first training completion condition refers to the conditions under which the first multi-objective recommendation model is trained, including but not limited to reaching the maximum iteration limit, the model loss reaching a preset threshold, or the model parameters no longer changing. The first multi-objective recommendation model is the trained multi-objective recommendation model used to recommend information to the target audience.
[0196] Specifically, when the server receives the model loss information, it determines whether the first training completion condition has been met. If the training completion condition has not been met, the gradient descent algorithm is used to update the initial model parameters in reverse. That is, the model parameters in the first initial multi-objective recommendation model are updated using the model loss information to obtain the updated first initial multi-objective recommendation model, i.e., the first updated multi-objective recommendation model. Then, the first updated multi-objective recommendation model is used as the first initial multi-objective recommendation model, and the steps of obtaining the positive and negative samples corresponding to the training recommendation objects are returned for execution until the first training completion condition is met. At this point, the first initial multi-objective recommendation model that meets the first training completion condition is used as the first multi-objective recommendation model.
[0197] The aforementioned multi-objective recommendation model training method improves the accuracy of obtaining the positive fusion recommendation degree by performing multi-objective prediction on positive samples and then performing multi-objective prediction on negative samples to obtain the negative fusion recommendation degree. The positive and negative fusion recommendation degrees are then used to calculate the model loss, improving the accuracy of the model loss information. This model loss information is then used to train a first initial multi-objective recommendation model. When the first training completion condition is met, a first multi-objective recommendation model is obtained, thereby improving the accuracy of the first multi-objective recommendation model and ultimately enhancing the accuracy of information recommendation.
[0198] In one embodiment, the first initial multi-objective recommendation model includes a first trained multi-objective prediction network and a first trained interactive fusion network;
[0199] Step 704: Input the training recommendation object features and positive recommendation information features into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, obtain the training bias degree and positive interaction probability of the training recommendation object for each business objective, and correct the corresponding positive interaction probability by using the training bias degree of each business objective to obtain the positive interaction probability of each objective. Then, fuse the positive interaction probabilities of each objective to obtain the positive fused recommendation degree corresponding to the positive recommendation information, including the following steps:
[0200] The features of the training recommendation objects and the features of positive recommendation information are input into the first training multi-objective prediction network to perform bias prediction and interaction prediction for each business objective, thereby obtaining the training bias degree and positive interaction probability of the training recommendation objects towards each business objective. The training bias degree of each business objective is used to correct the corresponding positive interaction probability, thereby obtaining the positive interaction probability of each objective. The positive interaction probabilities of each objective are then input into the first training interaction fusion network for fusion, thereby obtaining the positive fused recommendation degree corresponding to the positive recommendation information.
[0201] Here, the first training multi-objective prediction network refers to the first multi-objective prediction network that needs to be trained. The first training interactive fusion network refers to the interactive fusion network that needs to be trained.
[0202] Specifically, the server inputs the training recommendation object features and positive recommendation information features into the first training multi-objective prediction network to perform bias prediction and interaction prediction for each business objective, thereby obtaining the training bias degree and positive interaction probability of the training recommendation object for each business objective. Then, the positive interaction probability of each objective is fused in the first training interaction fusion network to obtain the positive fusion recommendation degree corresponding to the positive recommendation information.
[0203] In one specific embodiment, the network structure of the first trained multi-objective prediction network can be as follows: Figure 4 As shown. The network structure of the first training interactive fusion network can be as follows: Figure 5 As shown.
[0204] In one embodiment, such as Figure 8 As shown, in step 708, the model loss is calculated based on the positive and negative fusion recommendation levels to obtain model loss information, including:
[0205] Step 802: Determine the positive sequence position corresponding to the positive sample based on the positive fusion recommendation degree, and calculate the cumulative gain of loss based on the positive sequence position to obtain the positive ranking quality corresponding to the positive sample.
[0206] Here, "positive sequence position" refers to the position of the positive fusion recommendation degree corresponding to a positive sample within the fusion recommendation degree sequence. This fusion recommendation degree sequence is obtained by ranking the fusion recommendation degrees of a batch of training samples. This batch of training samples includes individual training samples, which can be positive or negative. Positive ranking quality is used to characterize the accuracy of ranking positive samples. Ranking quality can be calculated using ranking evaluation metrics, such as NDCG (Normalized Discounted Cumulative Gain).
[0207] Specifically, the server uses the positive fusion recommendation score to obtain the positive sequence position corresponding to the positive sample. The model can be trained in batches, with each batch using multiple training samples simultaneously. Each training sample generates a corresponding fusion recommendation score, which is then sorted (from largest to smallest) to obtain a fusion recommendation score sequence. The position of the positive fusion recommendation score corresponding to the positive sample within this sequence is then obtained, yielding the positive sequence position. Finally, the positive sequence position is used to calculate the cumulative gain with depreciation, thus obtaining the positive ranking quality corresponding to the positive sample.
[0208] Step 804: Determine the negative sequence position corresponding to the negative sample based on the negative fusion recommendation degree, and calculate the cumulative gain of loss based on the negative sequence position to obtain the negative ranking quality corresponding to the negative sample.
[0209] Here, negative sequence position refers to the position of the negative fusion recommendation degree corresponding to the negative sample in the fusion recommendation degree sequence. Negative ranking quality is used to characterize the accuracy of negative sample ranking.
[0210] Specifically, the server obtains the negative sequence position corresponding to the negative sample based on the negative fusion recommendation degree. This can be achieved by determining the position of the negative fusion recommendation degree from the fusion recommendation degree sequence. Then, the negative sequence position is used to calculate the cumulative gain of the loss, thus obtaining the negative ranking quality corresponding to the negative sample.
[0211] Step 806: Calculate the difference between positive and negative sorting quality to obtain relative quality.
[0212] Among them, relative quality is used to characterize the difference in ranking accuracy between positive and negative samples; the smaller the difference, the higher the evaluation.
[0213] Specifically, the server calculates the difference between the quality of positive sorting and the quality of negative sorting to obtain the relative quality.
[0214] Step 808: Calculate the sample pair loss based on the positive fusion recommendation degree and the negative fusion recommendation degree to obtain the sample pair loss information, and use relative quality to weight the sample pair loss information to obtain the model loss information.
[0215] Among them, the sample pair loss information is used to characterize the error between the positive fusion recommendation degree and the negative fusion recommendation degree.
[0216] Specifically, the server can use a loss function to calculate the loss between the positive and negative fusion recommendation levels, obtaining sample pair loss information. The loss function can be either a logarithmic loss function or a cross-entropy loss function. Then, the sample pair loss information is weighted using relative quality to obtain the model loss information.
[0217] In one specific embodiment, the relative quality can be calculated using the formula (3) shown below, and the sample pair loss information can be calculated using the formula (4) shown below.
[0218]
[0219] Among them, w ndcg It refers to relative quality, rank positive It refers to the position in the positive sequence. This refers to the quality of the positive sorting, rank. negative This refers to the position of the negative sequence. This refers to negative sorting quality.
[0220] loss pair = -logsigmoid(log it) positive -logit negative )Formula (4)
[0221] Where, loss pair This refers to the loss information of the sample pair, log it positive This refers to the degree of positive fusion recommendation, logit negative This refers to the degree of negative fusion recommendation.
[0222] In the above embodiments, by calculating the relative quality and sample pair loss, and then using the relative quality to weight the sample pair loss, the model loss information is obtained, thereby improving the accuracy of the obtained model loss information.
[0223] In one embodiment, step 808 involves weighting the sample loss information using relative quality to obtain model loss information, including:
[0224] Obtain the sample types corresponding to positive and negative samples, and obtain the sample pair weights based on the sample types; use the sample pair weights and relative quality to weight the sample pair loss information to obtain the model loss information.
[0225] Among them, sample type refers to the type to which a sample pair belongs, and the highest priority business objective corresponding to the positive sample and the negative sample is the same.
[0226] Specifically, the server obtains the highest-priority business objective corresponding to the positive and negative samples, and uses this business objective as the sample type for the positive and negative samples, respectively. Higher priority indicates a more important business objective. Priorities can be set according to requirements. For example, interaction objectives can be set to have the highest priority, followed by duration objectives and then click objectives. If the highest-priority business objective for a sample pair is interaction, then the sample type for that sample pair is the interaction type. Then, the server obtains the pre-set weights corresponding to the highest-priority business objective, resulting in sample pair weights. Finally, the server calculates the product of the sample pair weights, relative quality, and sample pair loss information to obtain the model loss information.
[0227] In a specific embodiment, the sample pair weights are obtained using the following formula (5), and the model loss information is calculated using the following formula (6).
[0228]
[0229] Among them, w pair This represents the weight of the sample pair. When the sample pair is an interactive sample type, the weight is 8. When the sample pair is a duration sample type, the weight is 2. When the sample pair is a click sample type, the weight is 1.
[0230] loss1 = w ndcg w pair loss pair Formula (6)
[0231] Here, loss1 represents the model loss information. The model loss information is obtained by calculating the product of relative quality, sample pair weights, and sample pair loss, which improves the accuracy of the obtained model loss information.
[0232] In one embodiment, such as Figure 9 As shown, a method for training a multi-objective recommendation model is provided, which is then applied to... Figure 1 Taking a server as an example, it can be understood that this method can also be applied to servers, and also to systems including terminals and servers, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0233] Step 902: Obtain training samples, which include training recommendation object information, training recommendation information, and various interaction tags.
[0234] Training samples are the samples used during training, and there can be multiple training samples. Training recommendation information refers to information that has been recommended to the training recommendation objects in the past, including but not limited to videos, images, text, and audio. Training recommendation object information refers to the information used during training that can characterize the training recommendation objects, including but not limited to the basic attribute information, social relationship information, spending power information, and behavioral information of the recommendation objects. The basic attribute information is used to characterize the basic attributes of the recommendation objects, the social relationship information is used to characterize the social relationships of the recommendation objects, the spending power information is used to characterize the spending power of the recommendation objects, and the behavioral information is used to characterize the behavior of the recommendation objects. Interaction refers to the interaction of the recommendation objects with the recommendation information, including but not limited to clicking, reading, and interacting. Training recommendation objects refer to the recommendation objects used during training, which can include real objects or virtual objects. Real objects include anything from people to animals. Virtual objects can refer to objects using virtual avatars for online live streaming.
[0235] Interaction labels are tags used during training to characterize the interaction results of the training recommendation object with the training recommendation information. The training recommendation object will have different interactions with the training recommendation information, and different interactions will produce different interaction results, i.e., different interaction labels. The interaction can be a business objective; for example, if the interaction label is a click label, then the business objective is for the training recommendation object to click on the training recommendation information.
[0236] Specifically, the server can obtain training samples from a database. These training samples include training recommendation object information, corresponding training recommendation information, and various interaction tags corresponding to the training recommendation object information. These interaction tags are obtained after acquiring the interaction results of the training recommendation object with the training recommendation information. The server can also obtain training samples from a training sample set. The server can also obtain training samples from data service providers, business users, or from users' data uploads.
[0237] Step 904: Input the training recommendation object features and training recommendation information features into the second initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, obtain the training bias degree and training interaction probability of the training recommendation object for each business objective, and correct the corresponding training interaction probability through the training bias degree of each business objective to obtain the training interaction probability of each objective. Fuse the training interaction probabilities of each objective to obtain the training fusion recommendation degree corresponding to the training recommendation information. Fuse the training bias degrees of each business objective to obtain the training bias recommendation degree corresponding to the training recommendation information. Calculate the sum of the training bias recommendation degree and the training fusion recommendation degree to obtain the training objective recommendation degree.
[0238] The second initial information recommendation model refers to the information recommendation model with initialized model parameters. This model is used for information recommendation prediction and is an artificial intelligence model built using a neural network. Model parameter initialization can be random initialization, zero initialization, or Gaussian distribution initialization, etc. The network structure of the second initial information recommendation model differs from that of the first initial information recommendation model. Training bias refers to the degree of bias obtained during multi-business objective prediction; training interaction probability refers to the interaction probability obtained during multi-business objective prediction; training fusion recommendation degree refers to the fusion recommendation degree obtained when training the second initial information recommendation model; training biased recommendation degree refers to the biased recommendation degree obtained during training; and training target recommendation degree refers to the target recommendation degree obtained during training.
[0239] Specifically, the server inputs the training recommendation object features and training recommendation information features into the second initial information recommendation model. The second initial information recommendation model performs multi-task learning, that is, it uses the training recommendation object features to predict interaction bias, obtaining the training bias degree of the training recommendation object for each business goal, and then uses the training recommendation object features and training recommendation information features to predict each interaction, obtaining the training interaction probability of the training recommendation object for each business goal. Finally, it corrects the corresponding training interaction probability by the training bias degree of each business goal, obtaining the training interaction probability of each goal, and fuses the training interaction probabilities of each goal to obtain the training fusion recommendation degree corresponding to the training recommendation information. It also fuses the training bias degrees of each business goal to obtain the training bias recommendation degree corresponding to the training recommendation information. Finally, it calculates the sum of the training bias recommendation degree and the training fusion recommendation degree to obtain the training target recommendation degree.
[0240] Step 906: Calculate the fusion loss based on each interactive label and the recommendation degree of the training target to obtain the fusion loss information, and adjust the fusion loss information with the bias of each training bias to obtain the target loss information.
[0241] The fusion loss information is used to characterize the error between the target recommendation level of the training output and the actual recommendation result. The target loss information is obtained by weighting the fusion loss information using the training bias.
[0242] Specifically, the server can use a loss function to calculate the error between each interaction label and the degree of fusion recommendation, obtaining the loss for each learning task during multi-task learning. Then, it performs fusion loss calculation to obtain fusion loss information. The server then uses the various training biases to obtain the weights of the training samples, and uses these weights to weight the fusion loss information to obtain the target loss information.
[0243] Step 908: Based on the target loss information, the second initial multi-objective recommendation model is updated in reverse to obtain the second updated multi-objective recommendation model. The second updated multi-objective recommendation model is used as the second initial multi-objective recommendation model, and the step of obtaining training samples is returned to execute until the second training completion condition is met, thus obtaining the second multi-objective recommendation model.
[0244] The second training completion condition refers to the conditions under which the second initial information recommendation model is trained, which may include reaching the maximum number of iterations, the maximum threshold of training loss information, and the model parameters no longer changing, etc. The second updated information recommendation model refers to the information recommendation model after the model parameters are updated. The second multi-objective recommendation model refers to the trained second initial information recommendation model, used for information recommendation.
[0245] Specifically, the server can first determine whether the second training completion condition has been met. If not, it uses gradient descent based on the target loss information to back-update the model parameters in the second initial information recommendation model, obtaining the second updated information recommendation model. Then, the second updated information recommendation model is used as the initial information recommendation model, and the step of obtaining training samples is iteratively executed until the second training completion condition is met. At this point, the second initial information recommendation model that meets the second training completion condition is used as the second multi-objective recommendation model. The trained second multi-objective recommendation model can then be used for information recommendation; that is, the second multi-objective recommendation model is deployed, and when needed, it is directly called.
[0246] The aforementioned multi-objective recommendation model training method inputs the training recommendation object features and training recommendation information features into a second initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective. This yields the training bias degree and training interaction probability of the training recommendation object towards each business objective. The training bias degree of each business objective is used to correct the corresponding training interaction probability, resulting in the training interaction probability of each objective. These training interaction probabilities are then fused to obtain the training fusion recommendation degree corresponding to the training recommendation information. The training bias degrees of each business objective are also fused to obtain the training bias recommendation degree corresponding to the training recommendation information. The sum of the training bias recommendation degree and the training fusion recommendation degree is calculated to obtain the training target recommendation degree, thus improving the accuracy of the training target recommendation degree. This training target recommendation degree is then used to calculate the fusion loss information, and the training bias degree is used to adjust the fusion loss information to obtain the target loss information, thereby improving the accuracy of the obtained target loss information. Finally, the target loss information is used to train the second initial multi-objective recommendation model. When the second training completion condition is met, the second multi-objective recommendation model is obtained, improving the accuracy of the obtained second multi-objective recommendation model.
[0247] In one embodiment, the second initial multi-objective recommendation model includes a second trained multi-objective prediction network, a second trained interactive fusion network, and a trained biased fusion network;
[0248] like Figure 10 As shown, step 904 involves inputting the training recommendation object features and training recommendation information features into the second initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, obtaining the training bias degree and training interaction probability of the training recommendation object for each business objective, and correcting the corresponding training interaction probability through the training bias degree of each business objective to obtain the training interaction probability of each objective. The training interaction probabilities of each objective are then fused to obtain the training fusion recommendation degree corresponding to the training recommendation information. The training bias degrees of each business objective are then fused to obtain the training bias recommendation degree corresponding to the training recommendation information. The sum of the training bias recommendation degree and the training fusion recommendation degree is calculated to obtain the training objective recommendation degree, including:
[0249] Step 1002: Input the training recommendation object features and training recommendation information features into the second training multi-objective prediction network to perform bias prediction and interaction prediction for each business objective, and obtain the training bias degree and training interaction probability of the training recommendation object for each business objective.
[0250] The second training multi-object prediction network refers to the second multi-object prediction network that needs to be trained. The network structure of the second multi-object prediction network is the same as that of the first multi-object prediction network, but the network parameters are different.
[0251] Specifically, the server can input the training recommendation object features and training recommendation information features as the overall features of the training samples into the second training multi-objective prediction network to perform bias prediction and interaction prediction for each business objective, and obtain the training bias degree and training interaction probability of the training recommendation object for each business objective.
[0252] Step 1004: Correct the training interaction probability corresponding to each business objective by the training bias degree of each objective, obtain the training interaction probability of each objective, and input the training interaction probability of each objective into the second training interaction fusion network for fusion to obtain the training fusion recommendation degree corresponding to the training recommendation information.
[0253] The second training interaction fusion network refers to the second interaction fusion network that needs to be trained, and its network structure differs from that of the first training interaction fusion network. The target training interaction probability refers to the target interaction probability obtained during training.
[0254] Specifically, the server first performs input correction before fusion, that is, it corrects the corresponding training interaction probability by the training bias degree of each business objective to obtain the training interaction probability of each objective. Then, the training interaction probability of each objective is input into the second training interaction fusion network for fusion to obtain the training fusion recommendation degree corresponding to the training recommendation information.
[0255] Step 1006: Input the training bias degree of each business objective into the training bias fusion network for fusion to obtain the training bias recommendation degree corresponding to the training recommendation information, and calculate the sum of the training bias recommendation degree and the training fusion recommendation degree to obtain the training objective recommendation degree.
[0256] Among them, the biased fusion network refers to the biased fusion network used when training the second multi-objective prediction network.
[0257] Specifically, the server directly inputs the training bias of each business objective into the training bias fusion network for fusion to obtain the training bias recommendation degree corresponding to the training recommendation information. Then, it performs output correction after fusion, that is, calculates the sum of the training bias recommendation degree and the training fusion recommendation degree to obtain the training objective recommendation degree. In a specific embodiment, the network structure of the second training multi-objective prediction network can be as follows: Figure 4 As shown, the network structure of the second training interactive fusion network can be as follows: Figure 6 As shown.
[0258] In the above embodiments, multi-object prediction is performed by using a second trained multi-object prediction network, then interactive fusion is performed by using a second trained interactive fusion network, and biased fusion is performed by using a trained biased fusion network, finally obtaining the training target recommendation degree, which improves the accuracy of the obtained training target recommendation degree.
[0259] In one embodiment, such as Figure 11 As shown, in step 906, the fusion loss is calculated based on the recommendation level of each interaction label and the training target to obtain the fusion loss information. Then, the fusion loss information is adjusted using the various training biases to obtain the target loss information, including:
[0260] Step 1102: Calculate the interaction loss for each business objective based on each interaction tag and the degree of fusion recommendation, and obtain the interaction loss information for each interaction.
[0261] Among them, the interaction loss information is used to characterize the error between the degree of fusion recommendation and the interaction label. The smaller the interaction loss information, the closer the degree of fusion recommendation is to the actual interaction result corresponding to the interaction label.
[0262] Specifically, the server can use the cross-entropy loss function to calculate the loss between each interaction tag and the fusion recommendation level, obtaining the interaction loss information corresponding to each interaction tag. Each business objective has a corresponding interaction tag, and the interaction loss information corresponding to each business objective is calculated separately. Different cross-entropy loss functions can be used for different tasks. When the interaction tag is a tag for a binary classification task, the binary classification cross-entropy loss function can be used to calculate the loss information for that task. For example, if the interaction tag is a tag for a click prediction task, which is a binary classification task, the binary classification cross-entropy loss function can be used to calculate the loss information. When the interaction tag is a tag for a multi-class classification task, the multi-class cross-entropy loss function can be used to calculate the loss information for that task. For example, if the interaction tag is a tag for a duration prediction task, which is a multi-class classification task, the multi-class cross-entropy loss function can be used to calculate the loss information. In one embodiment, the server can also use a linear regression loss function, a minimum loss function, etc., to calculate the loss information between the interaction tag and the fusion recommendation level. The server can calculate the cross-entropy loss between the interaction tag and the fusion recommendation level corresponding to each business objective, obtaining the loss information corresponding to each business objective.
[0263] Step 1104: Obtain the interaction weights corresponding to each interaction label, and calculate the fusion loss based on each interaction weight and each interaction loss information to obtain the fusion loss information.
[0264] The interaction weight is used to characterize the importance of the business objective corresponding to the interaction tag. Different business objectives have different interaction weights, which are pre-set and can be adjusted as needed.
[0265] Specifically, the server can pre-store interaction weights in the database. When needed, the server directly retrieves the interaction weights corresponding to each interaction tag from the database. The server can also obtain the interaction weights corresponding to each interaction tag uploaded by the terminal in real time. Furthermore, the server can obtain the interaction weights corresponding to each interaction tag from the business side. Then, the interaction weights are used to weight the corresponding interaction loss information, and finally, a weighted sum is calculated to obtain the fusion loss information.
[0266] In one specific embodiment, interaction tags include, but are not limited to, click tags, reading duration tags, interaction tags, and exposure tags. The server uses a binary cross-entropy loss function to calculate the cross-entropy loss between the click tag and the fusion recommendation level, obtaining the click prediction loss information corresponding to the click tag. Simultaneously, the server uses the same binary cross-entropy loss function to calculate the cross-entropy loss between the reading duration tag and the fusion recommendation level, obtaining the reading duration prediction loss information corresponding to the reading duration tag. Similarly, the server uses the same binary cross-entropy loss function to calculate the cross-entropy loss between the interaction tag and the fusion recommendation level, obtaining the interaction prediction loss information corresponding to the interaction tag; and the server uses the same binary cross-entropy loss function to calculate the cross-entropy loss between the exposure tag and the fusion recommendation level, obtaining the exposure prediction loss information corresponding to the exposure tag. Then, the loss information corresponding to the click tag, reading duration tag, interaction tag, and exposure tag is weighted using the corresponding interaction weights, and the weighted sum is calculated to obtain the fusion loss information.
[0267] Step 1106: Determine the sample weights corresponding to the training samples based on each interaction label and the degree of each interaction bias, and adjust the fusion loss information based on the sample weights to obtain the target loss information.
[0268] Among them, the sample weight is used to characterize the importance of the training sample relative to the training recommendation object. The higher the sample weight, the more representative the interaction tendency of the training sample is of the interaction tendency of the training recommendation object.
[0269] Specifically, the server uses each interaction label and the degree of bias of each interaction to determine the sample weights corresponding to the training samples. Then, the fusion loss information is weighted and calculated using the sample weights to obtain the target loss information corresponding to the training sample.
[0270] In the above embodiment, the fusion loss information is obtained by calculating the fusion loss through each interaction weight and each interaction loss information. Then, the sample weights corresponding to the training samples are determined by each interaction label and each interaction bias degree. The fusion loss information is biased and adjusted based on the sample weights to obtain the target loss information, thereby improving the accuracy of the target loss information.
[0271] In one embodiment, such as Figure 12 As shown, in step 1104, the interaction weights corresponding to each interaction label are obtained, and the fusion loss is calculated based on each interaction weight and each interaction loss information to obtain the fusion loss information, including:
[0272] Step 1202: Obtain the interaction priority corresponding to each interaction label, and determine the sample type corresponding to the training sample based on the interaction priority.
[0273] Interaction priority is used to characterize the importance of each business objective relative to the information recommendation business. A higher interaction priority indicates greater importance of that business objective. This interaction priority can be set according to requirements. For example, interactive tasks can be set to have the highest priority, followed by time-based tasks and then click-based tasks.
[0274] Specifically, the server can retrieve the interaction priorities corresponding to each interaction tag from the database, or it can retrieve the interaction priorities corresponding to each interaction tag uploaded by the terminal. Then, it uses the interaction task corresponding to the interaction tag with the highest interaction priority as the sample type corresponding to the training sample. For example, if the interaction task has the highest priority, then the sample type corresponding to that training sample is an interaction sample.
[0275] Step 1204: Find the corresponding interaction weights for each sample type from the preset interaction matrix based on the sample type.
[0276] The preset interaction matrix refers to a pre-set matrix that stores interaction weights. The rows of this matrix represent sample types, and the columns represent various business objectives. For example, if there are 4 sample types and 3 business objectives, then the preset interaction matrix is a 4x3 matrix.
[0277] Specifically, the server searches for the interaction weights corresponding to the same sample type from the prediction interaction matrix based on the sample type.
[0278] In one specific embodiment, the preset interaction matrix can be as shown in Table 1 below.
[0279] Table 1 Preset Interaction Matrix
[0280] W Click W duration W Interactive interactive 0 0 1.0 Duration 0 0.7 0.6 Click 0.1 0.9 0 exposure 1.0 0 0
[0281] The interaction weights for the interactive samples include a click weight of 0, a duration weight of 1, and an interaction weight of 1. The interaction weights for the duration samples include a click weight of 0, a duration weight of 0.7, and an interaction weight of 0.6. The interaction weights for the click samples include a click weight of 0.1, a duration weight of 0.9, and an interaction weight of 0. The interaction weights for the exposure samples include a click weight of 1, a duration weight of 0, and an interaction weight of 0.
[0282] Step 1206: Use each interaction weight to weight the corresponding interaction loss information to obtain each weighted loss information, and calculate the sum of information of each weighted loss information to obtain the fusion loss information.
[0283] Among them, weighted loss information refers to the loss information obtained by weighting the interaction loss information using interaction weights.
[0284] Specifically, the server performs weighted calculations on each interaction loss information using the corresponding interaction weight. For example, for a click sample, the interaction loss information for the click business objective is calculated by multiplying it by the click weight of 0.1 to obtain the weighted loss information for the click business objective. Then, the interaction loss information for the duration business objective is calculated by multiplying it by the duration weight of 0.9 to obtain the weighted loss information for the duration business objective. Finally, the interaction loss information for the interaction business objective is calculated by multiplying it by the interaction weight of 0 to obtain the weighted loss information for the interaction business objective. Finally, the sum of all weighted loss information is calculated to obtain the fused loss information.
[0285] In a specific embodiment, the fusion loss information can be calculated using the formula (7) shown below.
[0286]
[0287] Where, loss matrix This refers to fusing loss information. task This refers to the interaction loss information corresponding to the business objectives. This refers to the interaction weight corresponding to the business objective.
[0288] In the above embodiment, the sample type corresponding to the training sample is determined by using the interaction priority, then the interaction weights corresponding to the sample type are found from the preset interaction matrix, and finally the interaction weights are used to weight the interaction loss information to obtain the weighted loss information. The sum of the information of the weighted loss information is calculated to obtain the fusion loss information, which improves the accuracy of the obtained fusion loss information.
[0289] In one embodiment, such as Figure 13 As shown, in step 1106, the sample weights corresponding to the training samples are determined based on each interaction label and the degree of bias of each interaction. The fusion loss information is then biased and adjusted based on these sample weights to obtain the target loss information, including:
[0290] Step 1302: Obtain the interaction priority corresponding to each interaction label, and determine the sample type corresponding to the training sample based on the interaction priority.
[0291] Step 1304: Obtain the corresponding training sample sequence based on each interactive label, and determine the target sample sequence from each training sample sequence based on the sample type.
[0292] The training sample sequence refers to the sequence obtained by sorting the training samples according to the degree of interaction bias corresponding to each training sample. The training samples are sorted according to different business objectives to obtain the training sample sequence corresponding to each business objective. Each business objective also corresponds to an interaction label. The target sample sequence refers to the training sample sequence determined from the training sample sequences according to the sample type. That is, the target sample sequence is obtained by sorting according to the degree of interaction bias of the business objective corresponding to the sample type.
[0293] Specifically, the server can retrieve the interaction priority corresponding to each interaction tag from the database, and then determine the sample type corresponding to the training sample based on the interaction priority. Then, it obtains the corresponding training sample sequence based on each interaction tag, and determines the target sample sequence from each training sample sequence based on the sample type.
[0294] Step 1306: Determine the sequence position corresponding to the training sample from the target sample sequence. When the sequence position exceeds the preset position threshold, the sample weight corresponding to the training sample is obtained as the first target sample weight.
[0295] Step 1308: When the sequence position does not exceed the preset position threshold, the sample weight corresponding to the training sample is obtained as the second target sample weight.
[0296] Here, the weights of the first and second target samples are pre-set sample weights. The weight of the first target sample is greater than the weight of the second target sample. The preset position threshold refers to a pre-set position threshold for determining the sample weights. For example, the preset position threshold can be 20% of the sequence position sorting.
[0297] Specifically, the server determines the sequence position of the training sample from the target sample sequence, i.e., determines the ranking position of the training sample in the target sample sequence. Then, it compares the sequence position of the training sample with a preset position threshold. When the sequence position exceeds the preset position threshold, the training sample is ranked higher, and the weight of the first target sample is used as the weight of the corresponding training sample. When the sequence position exceeds the preset position threshold, it indicates that the training sample has hit the bias of the recommended object; therefore, the training sample better represents the bias of the recommended object and is assigned a higher weight. When the sequence position does not exceed the preset position threshold, the training sample is ranked lower, indicating that the training sample is a normal sample, and is assigned a lower weight.
[0298] In a specific embodiment, the fusion loss information can be calculated using the formula (8) shown below.
[0299] loss2 = w bias loss matrixFormula (8)
[0300] Where loss2 refers to the target loss information, loss matrix This refers to fusing loss information, w bias This refers to the sample weights corresponding to the training samples. The values of these sample weights can be represented by the following formula (9):
[0301]
[0302] in, The sample weight is 4 when the sequence position of the training sample is within the first 20% of the training sample sequence. The sample weight is 1 when the sequence position of the training sample is not within the first 20% of the training sample sequence.
[0303] In the above embodiments, the target sample sequence is determined from each training sample sequence by sample type, then the sequence position corresponding to the training sample is determined from the target sample sequence, and finally the sample weight corresponding to the training sample is determined according to the sequence position. In this way, the obtained sample weight can enhance the tendency of the training recommendation object itself, thereby improving the model training effect and improving the accuracy of information recommendation of the trained model.
[0304] In one embodiment, step 1304, which involves obtaining the corresponding training sample sequence based on each interaction label, includes the following steps:
[0305] Obtain each training sample and the degree of each interaction bias corresponding to each training sample; sort each training sample according to the degree of each interaction bias corresponding to each training sample based on each interaction label to obtain the training sample sequence corresponding to each interaction label.
[0306] Specifically, the server can obtain each training sample and input them sequentially into the second initial information recommendation model to obtain the interaction bias degree corresponding to each training sample. Alternatively, the server can directly obtain each training sample and its corresponding interaction bias degree from the database. Then, it sorts all the training samples according to all interaction bias degrees corresponding to the business objective, either from largest to smallest, finally obtaining the training sample sequence for each interaction label's business objective. In other words, by using interaction bias degrees to sort the training samples, the accuracy of the obtained training sample sequence is improved.
[0307] In a specific embodiment, this information recommendation method is applied to a news recommendation scenario. Specifically, when a news platform makes news recommendations, it obtains information about the news to be recommended and the object to be recommended, and inputs this information recommendation model into a multi-objective recommendation model, such as... Figure 14 The diagram illustrates the bias correction process of the second multi-objective recommendation model. By inputting the information of the news to be recommended and the object to be recommended into the multi-objective prediction network of the second multi-objective recommendation model, the model outputs click bias scores and click recommendation scores for the click business objective, duration bias scores and duration recommendation scores for the duration business objective, and interaction bias scores and interaction recommendation scores for the interaction business objective. Then, the click bias score is subtracted from the click recommendation score to obtain the click target recommendation score; the duration bias score is subtracted from the duration recommendation score to obtain the duration target recommendation score; and the interaction bias score is subtracted from the interaction recommendation score to obtain the interaction target recommendation score. This is the bias correction before fusion. The click target recommendation score, duration target recommendation score, and interaction target recommendation score are then fused to obtain the interaction fusion score. Simultaneously, the click bias score, duration bias score, and interaction bias score are fused to obtain the bias fusion score. Finally, the fused output is corrected by calculating the sum of the interaction fusion score and the bias fusion score to obtain the final fused recommendation score. When the fusion recommendation score exceeds a pre-set recommendation threshold, the news platform recommends the news to the corresponding recommendation target terminal. Alternatively, a second multi-objective recommendation model can be used to obtain the fusion recommendation score for each candidate news item, then sort the candidate news items from highest to lowest fusion recommendation score to obtain a candidate news sequence. The top-ranked candidate news items are then selected as the news to be recommended, for example, the top 10. The selected news is then sent to the recommendation target terminal. When the recommendation target terminal receives the recommended news, it displays it on the news page, such as... Figure 15 The image shown is a schematic diagram of a news recommendation page, where recommended news is displayed in the form of an image and text information stream.
[0308] In one specific embodiment, the information recommendation method is applied to the news function of an instant messaging application, specifically to the recommendation of news highlights. The first multi-objective recommendation model is tested using A / B testing, and the test results include, for example: Figure 16-A The diagram showing the comparison of click-through rates for images and text is as follows. Figure 16-B The diagram shows a comparison of the test result indicators. Among them, Figure 16-A The line graph in the image represents a comparison between the experimental group and the control group in terms of image clicks. Figure 16-AThe bar chart in the figure represents the relative difference, which is the ratio of the experimental group's value minus the control group's value to the control group's value. It is clear that after using the first multi-objective recommendation model for news recommendation, the click-through rate and various indicators of the news articles in the "Watchpoints" section significantly improved. In other words, the first multi-objective recommendation model in this application significantly improves the accuracy of news recommendation.
[0309] In one specific embodiment, this information recommendation method is applied to the news function of an instant messaging application, specifically to the recommendation of news highlights. The second multi-objective recommendation model is tested using A / B testing, and the test results include, for example: Figure 17-A The diagram showing the comparison of click-through rates for images and text is as follows. Figure 17-B The diagram shows a comparison of the test result indicators. Among them, Figure 17-A The line graph in the image represents a comparison between the experimental group and the control group in terms of image clicks. Figure 17-A The bar chart in the figure represents the relative difference, which is the ratio of the experimental group's value minus the control group's value to the control group's value. It is clear that after using the second multi-objective recommendation model for news recommendation, the click-through rate and various indicators of the news articles in the "Watchpoints" section significantly improved. In other words, the second multi-objective recommendation model in this application significantly improves the accuracy of news recommendation.
[0310] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0311] Based on the same inventive concept, this application also provides an information recommendation device and a multi-objective recommendation model training device for implementing the information recommendation method described above. The solution provided by this device is similar to the implementation scheme described in the above method. Therefore, the specific limitations of one or more information recommendation device embodiments or multi-objective recommendation model training device embodiments provided below can be found in the limitations of the information recommendation method or multi-objective recommendation model training method described above, and will not be repeated here.
[0312] In one embodiment, such as Figure 18As shown, an information recommendation device 1800 is provided, including: a feature acquisition module 1802, a bias prediction module 1804, an interaction prediction module 1806, a fusion module 1808, and a recommendation module 1810, wherein:
[0313] The feature acquisition module 1802 is used to acquire the object attribute features and the recommendation information features corresponding to the object to be recommended.
[0314] The bias prediction module 1804 is used to extract low-level features based on object attribute features to obtain object extracted features, and to perform bias prediction on each business objective based on the object extracted features to obtain the degree of bias of the object to be recommended to each business objective.
[0315] The interaction prediction module 1806 is used to combine the object attribute features and the information features to be recommended to obtain combined features. Based on the combined features, object extraction features and information features to be recommended, interaction prediction is performed on each business objective to obtain the interaction probability of the object to be recommended with each business objective.
[0316] The fusion module 1808 is used to correct the corresponding interaction possibilities of each business objective by the degree of bias of each objective, obtain the interaction possibilities of each objective, and fuse the interaction possibilities of each objective to obtain the fusion recommendation degree of the information to be recommended.
[0317] The recommendation module 1810 is used to recommend the information to be recommended to the terminal corresponding to the object to be recommended when the degree of fusion recommendation meets the preset recommendation conditions.
[0318] In one embodiment, the information recommendation device 1800 further includes:
[0319] The model recommendation module is used to input object attribute features and features of information to be recommended into the first multi-objective recommendation model. The first multi-objective recommendation model extracts low-level features based on object attribute features to obtain object-extracted features. Based on these object-extracted features, it predicts the bias of each business objective to obtain the degree of bias of the object to be recommended towards each business objective. The first multi-objective recommendation model combines the object attribute features and features of information to be recommended to obtain combined features. Based on these combined features, object-extracted features, and features of information to be recommended, it predicts the interaction probability of each business objective to obtain the interaction probability of the object to be recommended towards each business objective. The first multi-objective recommendation model uses the degree of bias of each business objective to correct the corresponding interaction probability to obtain the interaction probability of each objective. Finally, it fuses the interaction probabilities of each objective to obtain the fused recommendation degree corresponding to the information to be recommended.
[0320] In one embodiment, the first multi-objective recommendation model includes a first multi-objective prediction network and an interactive fusion network, wherein the first multi-objective prediction network includes a first biased prediction sub-network and a first interactive prediction sub-network.
[0321] The model recommendation module is also used to input object attribute features into the first bias prediction subnetwork in the first multi-objective prediction network; the first bias prediction subnetwork extracts low-level features from the object attribute features to obtain object extracted features, and performs bias prediction on each business objective based on the object extracted features to obtain the degree of bias of the object to be recommended to each business objective.
[0322] The model recommendation module is also used to input the combined features, object extraction features, and features of the information to be recommended into the first interactive prediction subnetwork in the first multi-objective prediction network; through the first interactive prediction subnetwork, the combined features, object extraction features, and features of the information to be recommended are used to perform interactive predictions for each business objective, so as to obtain the interaction probability of the object to be recommended with each business objective;
[0323] The model recommendation module is also used to concatenate the interaction possibilities of each target to obtain a concatenated vector, and input the concatenated vector into the interaction fusion network for fusion to obtain the fusion recommendation degree corresponding to the information to be recommended.
[0324] In one embodiment, the fusion module 1808 is further used to calculate the difference between the interaction probability of each business objective and the corresponding degree of bias, so as to obtain the target interaction probability of each business objective.
[0325] In one embodiment, the information recommendation device 1800 further includes:
[0326] The bias fusion module is used to fuse the biases of various business objectives to obtain the bias recommendation degree corresponding to the information to be recommended; the sum of the bias recommendation degree and the fused recommendation degree is calculated to obtain the target recommendation degree.
[0327] The recommendation module 1810 is also used to recommend the information to be recommended to the terminal corresponding to the object to be recommended when the target recommendation level meets the preset target recommendation conditions.
[0328] In one embodiment, the bias fusion module is further used to concatenate the bias of each business objective to obtain a bias concatenation vector, and input the bias concatenation vector into the bias fusion network for fusion to obtain the bias recommendation degree corresponding to the information to be recommended.
[0329] In one embodiment, such as Figure 19As shown, a multi-objective recommendation model training device 1900 is provided, including: a sample pair acquisition module 1902, a positive prediction module 1904, a negative prediction module 1906, a loss calculation module 1908, and a first iteration module 1910, wherein:
[0330] The sample pair acquisition module 1902 is used to acquire positive and negative samples corresponding to the training recommendation object. The positive samples include the training recommendation object features and positive recommendation information features, and the negative samples include the training recommendation object features and negative recommendation information features.
[0331] The positive prediction module 1904 is used to input the training recommendation object features and positive recommendation information features into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, obtain the training bias degree and positive interaction probability of the training recommendation object for each business objective, and correct the corresponding positive interaction probability by the training bias degree of each business objective to obtain the positive interaction probability of each objective, and fuse the positive interaction probabilities of each objective to obtain the positive fusion recommendation degree corresponding to the positive recommendation information;
[0332] The negative prediction module 1906 is used to input the training recommendation object features and negative recommendation information features into the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, obtain the training bias degree and negative interaction probability of the training recommendation object for each business objective, and correct the corresponding negative interaction probability by the training bias degree of each business objective to obtain the negative interaction probability of each objective, and fuse the negative interaction probabilities of each objective to obtain the negative fusion recommendation degree corresponding to the negative recommendation information;
[0333] The loss calculation module 1908 is used to calculate the model loss based on the positive fusion recommendation degree and the negative fusion recommendation degree, and obtain the model loss information.
[0334] The first iteration module 1910 is used to back-update the first initial multi-objective recommendation model based on the model loss information to obtain the first updated multi-objective recommendation model. The first updated multi-objective recommendation model is used as the first initial multi-objective recommendation model, and the step of obtaining the positive and negative samples corresponding to the training recommendation objects is returned to be executed until the first training completion condition is met, and the first multi-objective recommendation model is obtained.
[0335] In one embodiment, the first initial multi-objective recommendation model includes a first trained multi-objective prediction network and a first trained interactive fusion network;
[0336] The positive prediction module 1904 is also used to input the training recommendation object features and positive recommendation information features into the first training multi-objective prediction network to perform bias prediction and interaction prediction for each business objective, so as to obtain the training bias degree and positive interaction probability of the training recommendation object for each business objective; to correct the corresponding positive interaction probability of each business objective by adjusting the training bias degree of each business objective, so as to obtain the positive interaction probability of each objective; and to input the positive interaction probability of each objective into the first training interaction fusion network for fusion, so as to obtain the positive fusion recommendation degree corresponding to the positive recommendation information.
[0337] In one embodiment, the loss calculation module 1908 is further configured to determine the positive sequence position corresponding to a positive sample based on the positive fusion recommendation degree, calculate the cumulative gain with loss based on the positive sequence position to obtain the positive ranking quality corresponding to the positive sample; determine the negative sequence position corresponding to a negative sample based on the negative fusion recommendation degree, calculate the cumulative gain with loss based on the negative sequence position to obtain the negative ranking quality corresponding to the negative sample; calculate the difference between the positive ranking quality and the negative ranking quality to obtain the relative quality; calculate the sample pair loss based on the positive and negative fusion recommendation degrees to obtain the sample pair loss information, and weight the sample pair loss information using the relative quality to obtain the model loss information.
[0338] In one embodiment, the loss calculation module 1908 is further configured to obtain the sample types corresponding to positive and negative samples, obtain sample pair weights based on the sample types, and use the sample pair weights and relative quality to weight the sample pair loss information to obtain the model loss information.
[0339] In one embodiment, such as Figure 20 As shown, a multi-objective recommendation model training device 2000 is provided, including: a sample acquisition module 2002, a training module 2004, an objective loss calculation module 2006, and a second iteration module 2008, wherein:
[0340] The sample acquisition module 2002 is used to acquire training samples, which include training recommendation object information, training recommendation information and various interaction labels.
[0341] The training module 2004 is used to input the features of the training recommendation object and the features of the training recommendation information into the second initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, obtain the training bias degree and training interaction probability of the training recommendation object for each business objective, and correct the corresponding training interaction probability by the training bias degree of each business objective to obtain the training interaction probability of each objective. The training interaction probabilities of each objective are fused to obtain the training fusion recommendation degree corresponding to the training recommendation information. The training bias degree of each business objective is fused to obtain the training bias recommendation degree corresponding to the training recommendation information. The sum of the training bias recommendation degree and the training fusion recommendation degree is calculated to obtain the training objective recommendation degree.
[0342] The target loss calculation module 2006 is used to calculate the fusion loss based on each interactive label and the degree of recommendation of the training target, obtain the fusion loss information, and adjust the fusion loss information with each training bias degree to obtain the target loss information;
[0343] The second iteration module 2008 is used to update the second initial multi-objective recommendation model in reverse based on the target loss information to obtain the second updated multi-objective recommendation model. The second updated multi-objective recommendation model is used as the second initial multi-objective recommendation model, and the step of obtaining training samples is returned to execute until the second training completion condition is met, and the second multi-objective recommendation model is obtained.
[0344] In one embodiment, the second initial multi-objective recommendation model includes a second trained multi-objective prediction network, a second trained interactive fusion network, and a trained biased fusion network;
[0345] The training module 2004 is also used to input the features of the training recommendation object and the features of the training recommendation information into the second training multi-objective prediction network to perform bias prediction and interaction prediction for each business objective, thereby obtaining the training bias degree and training interaction probability of the training recommendation object for each business objective; correcting the corresponding training interaction probability by the training bias degree of each business objective, thereby obtaining the training interaction probability of each objective; inputting the training interaction probability of each objective into the second training interaction fusion network for fusion, thereby obtaining the training fusion recommendation degree corresponding to the training recommendation information; inputting the training bias degree of each business objective into the training bias fusion network for fusion, thereby obtaining the training bias recommendation degree corresponding to the training recommendation information; and calculating the sum of the training bias recommendation degree and the training fusion recommendation degree to obtain the training objective recommendation degree.
[0346] In one embodiment, the target loss calculation module 2006 is further configured to calculate the interaction loss of each business objective based on each interaction label and the degree of fusion recommendation, and obtain each interaction loss information; obtain the interaction weights corresponding to each interaction label, calculate the fusion loss based on each interaction weight and each interaction loss information, and obtain the fusion loss information; determine the sample weights corresponding to the training samples based on each interaction label and each interaction bias, and adjust the fusion loss information based on the sample weights to obtain the target loss information.
[0347] In one embodiment, the target loss calculation module 2006 is further configured to obtain the interaction priority corresponding to each interaction label, determine the sample type corresponding to the training sample based on the interaction priority, find the interaction weight corresponding to the sample type from the preset interaction matrix based on the sample type, use each interaction weight to weight the corresponding interaction loss information to obtain each weighted loss information, and calculate the sum of information of each weighted loss information to obtain the fusion loss information.
[0348] In one embodiment, the target loss calculation module 2006 is further configured to obtain the interaction priority corresponding to each interaction label, determine the sample type corresponding to the training sample based on the interaction priority; obtain the corresponding training sample sequence based on each interaction label, determine the target sample sequence from each training sample sequence based on the sample type; determine the sequence position corresponding to the training sample from the target sample sequence, and when the sequence position exceeds a preset position threshold, obtain the sample weight corresponding to the training sample as the first target sample weight; when the sequence position does not exceed the preset position threshold, obtain the sample weight corresponding to the training sample as the second target sample weight.
[0349] In one embodiment, the target loss calculation module 2006 is further configured to obtain each training sample and the degree of each interaction bias corresponding to each training sample; and to sort each training sample according to the degree of each interaction bias corresponding to each training sample based on each interaction label to obtain the training sample sequence corresponding to each interaction label.
[0350] Each module in the aforementioned information recommendation device and multi-objective recommendation model training device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0351] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 21As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores training sample data, attribute information of recommended objects, and information to be recommended. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When the computer program is executed by the processor, it implements an information recommendation method or a multi-objective recommendation model training method.
[0352] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 22 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an information recommendation method or a multi-objective recommendation model training method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0353] Those skilled in the art will understand that Figure 21 or Figure 22The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0354] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0355] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0356] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0357] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions. Users can refuse push notifications or conveniently refuse push notifications.
[0358] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0359] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0360] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An information recommendation method, characterized in that, The method includes: Obtain the object attribute features corresponding to the object to be recommended and the features of the news to be recommended; The first bias prediction subnetwork in the first multi-objective recommendation model extracts low-level features of the object attribute features to obtain object extraction features, and performs bias prediction on each business objective based on the object extraction features to obtain the degree of bias of the object to be recommended to each business objective. The object attribute features and the news features to be recommended are combined to obtain combined features; Through the first interaction prediction sub-network in the first multi-objective recommendation model, the combined features, the object extraction features and the news to be recommended are used to perform interaction prediction on each business objective, so as to obtain the interaction probability of the object to be recommended with each business objective. Using the first multi-objective recommendation model, the bias of each business objective is used to correct the corresponding interaction probability, thereby obtaining the interaction probability of each objective. The interaction probabilities of each objective are then concatenated to obtain a concatenated vector. The concatenated vector is input into the interactive fusion network of the first multi-objective recommendation model for fusion to obtain the fusion recommendation degree corresponding to the news to be recommended. When the degree of fusion recommendation meets the preset recommendation conditions, the news to be recommended is recommended to the terminal corresponding to the news to be recommended.
2. The method according to claim 1, characterized in that, The step of correcting the corresponding interaction probabilities using the bias degree of each business objective to obtain the interaction probabilities of each objective includes: The difference between the interaction probability and the corresponding bias degree of each business objective is calculated to obtain the target interaction probability of each business objective.
3. The method according to claim 1, characterized in that, The method further includes: The degree of bias of each of the business objectives is combined to obtain the degree of bias recommendation corresponding to the news to be recommended; The target recommendation level is obtained by summing the biased recommendation level and the fused recommendation level. When the degree of fusion recommendation meets the preset recommendation conditions, the news to be recommended is recommended to the terminal corresponding to the news to be recommended, including: When the target recommendation level meets the preset target recommendation conditions, the news to be recommended is recommended to the terminal corresponding to the target object.
4. The method according to claim 3, characterized in that, The step of fusing the biases of the various business objectives to obtain the biased recommendation level corresponding to the news to be recommended also includes: The biases of each business objective are concatenated to obtain a bias concatenation vector. This bias concatenation vector is then input into a bias fusion network for fusion to obtain the bias recommendation level corresponding to the news to be recommended.
5. A method for training a multi-objective recommendation model, characterized in that, The method includes: Obtain positive and negative samples corresponding to the training recommendation object. The positive samples include the training recommendation object features and positive recommendation news features, and the negative samples include the training recommendation object features and negative recommendation news features. The training recommendation object features and positive recommendation news features are input into the first training multi-objective prediction network of the first initial multi-objective recommendation model. Bias prediction and interaction prediction are performed on each business objective to obtain the training bias degree and positive interaction probability of the training recommendation object to each business objective. The positive interaction probability corresponding to each business objective is corrected by the training bias degree of each business objective to obtain the positive interaction probability of each objective. The positive interaction probability of each objective is input into the first training interaction fusion network of the first initial multi-objective recommendation model for fusion to obtain the positive fusion recommendation degree corresponding to the positive recommendation news. The training recommendation object features and negative recommendation news features are input into the first training multi-objective prediction network of the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, thereby obtaining the training bias degree and negative interaction probability of the training recommendation object for each business objective. The negative interaction probability corresponding to each business objective is corrected by the training bias degree of each business objective to obtain the negative interaction probability of each objective. The negative interaction probability of each objective is input into the first training interaction fusion network of the first initial multi-objective recommendation model for fusion to obtain the negative fusion recommendation degree corresponding to the negative recommendation news. The model loss is calculated based on the positive fusion recommendation degree and the negative fusion recommendation degree to obtain model loss information; Based on the model loss information, the first initial multi-objective recommendation model is updated in reverse to obtain the first updated multi-objective recommendation model. The first updated multi-objective recommendation model is used as the first initial multi-objective recommendation model, and the step of obtaining the positive and negative samples corresponding to the training recommendation objects is returned to be executed until the first training completion condition is met, and the first multi-objective recommendation model is obtained.
6. The method according to claim 5, characterized in that, The model loss is calculated based on the positive fusion recommendation degree and the negative fusion recommendation degree to obtain model loss information, including: Based on the positive fusion recommendation degree, the positive sequence position corresponding to the positive sample is determined, and the cumulative gain of loss is calculated based on the positive sequence position to obtain the positive ranking quality corresponding to the positive sample; Based on the negative fusion recommendation degree, the negative sequence position corresponding to the negative sample is determined, and the cumulative gain of loss is calculated based on the negative sequence position to obtain the negative ranking quality corresponding to the negative sample. The difference between the positive and negative sorting qualities is calculated to obtain the relative quality. The sample pair loss is calculated based on the positive fusion recommendation degree and the negative fusion recommendation degree to obtain the sample pair loss information. The sample pair loss information is then weighted using the relative quality to obtain the model loss information.
7. The method according to claim 6, characterized in that, The step of weighting the sample pairs' loss information using the relative quality to obtain the model loss information includes: Obtain the sample types corresponding to the positive and negative samples, and obtain the sample pair weights based on the sample types; The loss information of the sample pair is obtained by weighting the sample pair weights and the relative quality.
8. A method for training a multi-objective recommendation model, characterized in that, The method includes: Obtain training samples, which include training recommendation object information, training recommendation news, and various interaction tags; The training recommendation object features extracted from the training recommendation object information and the training recommendation news features extracted from the training recommendation news are input into the second training multi-objective prediction network of the second initial multi-objective recommendation model. Bias prediction and interaction prediction are performed on each business objective to obtain the training bias degree and training interaction probability of the training recommendation object to each business objective. The training bias of each business objective is used to correct the corresponding training interaction probability, thereby obtaining the training interaction probability of each objective. The training interaction probability of each objective is then input into the second training interaction fusion network of the second initial multi-objective recommendation model for fusion, thereby obtaining the training fusion recommendation degree corresponding to the training recommended news. The training bias of each business objective is input into the training bias fusion network for fusion to obtain the training bias recommendation degree corresponding to the training recommended news. The sum of the training bias recommendation degree and the training fusion recommendation degree is calculated to obtain the training objective recommendation degree. The fusion loss is calculated based on the various interactive labels and the recommendation degree of the training target to obtain fusion loss information. The fusion loss information is then adjusted using the various training bias degrees to obtain target loss information. The second initial multi-objective recommendation model is updated in reverse based on the target loss information to obtain the second updated multi-objective recommendation model. The second updated multi-objective recommendation model is used as the second initial multi-objective recommendation model, and the step of obtaining training samples is returned to execute until the second training completion condition is met, thus obtaining the second multi-objective recommendation model.
9. The method according to claim 8, characterized in that, The step of calculating the fusion loss based on the various interactive tags and the recommendation level of the training target to obtain fusion loss information, and then adjusting the fusion loss information using various training biases to obtain target loss information, includes: Based on the interaction tags and the degree of fusion recommendation, the interaction loss of each business objective is calculated to obtain the interaction loss information. Obtain the interaction weights corresponding to each interaction label, and calculate the fusion loss based on each interaction weight and each interaction loss information to obtain the fusion loss information; Based on the interaction labels and the degree of bias of each interaction, the sample weights corresponding to the training samples are determined, and the fusion loss information is biased and adjusted based on the sample weights to obtain the target loss information.
10. The method according to claim 9, characterized in that, The step of obtaining the interaction weights corresponding to each interaction tag, and calculating the fusion loss based on each interaction weight and each interaction loss information to obtain the fusion loss information includes: Obtain the interaction priority corresponding to each interaction label, and determine the sample type corresponding to the training sample based on the interaction priority; Based on the sample type, find the corresponding interaction weights for each sample type from the preset interaction matrix; The interaction loss information is weighted according to the interaction weights to obtain the weighted loss information, and the sum of the information of the weighted loss information is calculated to obtain the fusion loss information.
11. The method according to claim 9, characterized in that, The process of determining the sample weights corresponding to the training samples based on the various interaction labels and the degree of bias of each interaction, and adjusting the fusion loss information based on the sample weights to obtain the target loss information includes: Obtain the interaction priority corresponding to each interaction label, and determine the sample type corresponding to the training sample based on the interaction priority; Based on each interactive label, obtain the corresponding training sample sequence, and based on the sample type, determine the target sample sequence from each training sample sequence; The sequence position corresponding to the training sample is determined from the target sample sequence. When the sequence position exceeds a preset position threshold, the sample weight corresponding to the training sample is obtained as the first target sample weight. When the sequence position does not exceed the preset position threshold, the sample weight corresponding to the training sample is obtained as the second target sample weight.
12. The method according to claim 11, characterized in that, The step of obtaining the corresponding training sample sequence based on each interaction label includes: Obtain each training sample and the degree of each interaction bias corresponding to each training sample; Based on the interaction labels, the training samples are sorted according to the degree of interaction bias corresponding to each training sample to obtain the training sample sequence corresponding to each interaction label.
13. An information recommendation device, characterized in that, The device includes: The feature acquisition module is used to acquire the object attribute features corresponding to the object to be recommended and the features of the news to be recommended. The bias prediction module is used to extract low-level features from the object attribute features through the first bias prediction sub-network in the first multi-objective recommendation model to obtain object extraction features, and to perform bias prediction on each business objective based on the object extraction features to obtain the degree of bias of the object to be recommended to each business objective. The interaction prediction module is used to combine the object attribute features and the news features to be recommended to obtain combined features; through the first interaction prediction sub-network in the first multi-objective recommendation model, the combined features, the object extraction features and the news features to be recommended are used to perform interaction prediction on each business objective to obtain the interaction probability of the object to be recommended with each business objective. The fusion module is used to correct the corresponding interaction probabilities of each business objective using the bias degree of each business objective through the first multi-objective recommendation model, obtain the interaction probabilities of each objective, and concatenate the interaction probabilities of each objective to obtain a concatenated vector; the concatenated vector is input into the interaction fusion network in the first multi-objective recommendation model for fusion to obtain the fusion recommendation degree corresponding to the news to be recommended; The recommendation module is used to recommend the news to be recommended to the terminal corresponding to the news to be recommended when the degree of fusion recommendation meets the preset recommendation conditions.
14. The information recommendation device according to claim 13, characterized in that, The fusion module is also used to calculate the difference between the interaction probability and the corresponding bias degree of each business objective, so as to obtain the target interaction probability of each business objective.
15. The information recommendation device according to claim 13, characterized in that, The device further includes a bias fusion module, which is used to fuse the bias of each business objective to obtain the bias recommendation degree corresponding to the news to be recommended. The target recommendation level is obtained by summing the biased recommendation level and the fused recommendation level. The recommendation module is also used to recommend the news to be recommended to the terminal corresponding to the news to be recommended when the target recommendation level meets the preset target recommendation conditions.
16. The information recommendation device according to claim 15, characterized in that, The biased fusion module is further used to concatenate the biases of each business objective to obtain a biased concatenation vector, and input the biased concatenation vector into the biased fusion network for fusion to obtain the biased recommendation degree corresponding to the news to be recommended.
17. A multi-objective recommendation model training device, characterized in that, The device includes: The sample pair acquisition module is used to acquire positive and negative samples corresponding to the training recommendation object. The positive samples include the training recommendation object features and positive recommendation news features, and the negative samples include the training recommendation object features and negative recommendation news features. The positive prediction module is used to input the features of the training recommendation object and the features of the positive recommendation news into the first training multi-objective prediction network of the first initial multi-objective recommendation model, perform bias prediction and interaction prediction for each business objective, obtain the training bias degree and positive interaction probability of the training recommendation object to each business objective, and correct the corresponding positive interaction probability through the training bias degree of each business objective to obtain the positive interaction probability of each objective. The positive interaction probability of each objective is then input into the first training interaction fusion network of the first initial multi-objective recommendation model for fusion to obtain the positive fusion recommendation degree corresponding to the positive recommendation news. The negative prediction module is used to input the features of the training recommendation object and the features of the negative recommendation news into the first training multi-objective prediction network of the first initial multi-objective recommendation model to perform bias prediction and interaction prediction for each business objective, so as to obtain the training bias degree and negative interaction probability of the training recommendation object to each business objective, and correct the corresponding negative interaction probability by the training bias degree of each business objective to obtain the negative interaction probability of each objective, and input the negative interaction probability of each objective into the first training interaction fusion network of the first initial multi-objective recommendation model for fusion to obtain the negative fusion recommendation degree corresponding to the negative recommendation news; The loss calculation module is used to calculate the model loss based on the positive fusion recommendation degree and the negative fusion recommendation degree to obtain model loss information; The first iteration module is used to back-update the first initial multi-objective recommendation model based on the model loss information to obtain a first updated multi-objective recommendation model. The first updated multi-objective recommendation model is used as the first initial multi-objective recommendation model, and the step of obtaining the positive and negative samples corresponding to the training recommendation objects is returned to be executed until the first training completion condition is met, and the first multi-objective recommendation model is obtained.
18. The multi-objective recommendation model training device according to claim 17, characterized in that, The loss calculation module is also used to determine the positive sequence position corresponding to the positive sample based on the positive fusion recommendation degree, and to perform loss cumulative gain calculation based on the positive sequence position to obtain the positive ranking quality corresponding to the positive sample. Based on the negative fusion recommendation degree, the negative sequence position corresponding to the negative sample is determined, and the cumulative gain of loss is calculated based on the negative sequence position to obtain the negative ranking quality corresponding to the negative sample. The difference between the positive ranking quality and the negative ranking quality is calculated to obtain the relative quality; the sample pair loss is calculated based on the positive fusion recommendation degree and the negative fusion recommendation degree to obtain the sample pair loss information, and the sample pair loss information is weighted using the relative quality to obtain the model loss information.
19. The multi-objective recommendation model training device according to claim 18, characterized in that, The loss calculation module is also used to obtain the sample types corresponding to the positive samples and the negative samples, obtain the sample pair weights based on the sample types, and use the sample pair weights and the relative quality to weight the sample pair loss information to obtain the model loss information.
20. A multi-objective recommendation model training device, characterized in that, The device includes: The sample acquisition module is used to acquire training samples, which include training recommendation object information, training recommendation news, and various interaction tags; The training module is used to input the features of the training recommendation objects extracted from the information of the training recommendation objects and the features of the training recommendation news extracted from the training recommendation news into the second training multi-objective prediction network of the second initial multi-objective recommendation model. It performs bias prediction and interaction prediction for each business objective to obtain the training bias degree and training interaction probability of the training recommendation object towards each business objective. It then corrects the corresponding training interaction probability based on the training bias degree of each business objective to obtain the training interaction probability of each objective. The training interaction probability of each objective is then input into the second training interaction fusion network of the second initial multi-objective recommendation model for fusion to obtain the training fusion recommendation degree corresponding to the training recommendation news. Finally, it inputs the training bias degree of each business objective into the training bias fusion network for fusion to obtain the training bias recommendation degree corresponding to the training recommendation news. The sum of the training bias recommendation degree and the training fusion recommendation degree is calculated to obtain the training objective recommendation degree. The target loss calculation module is used to calculate the fusion loss based on the various interactive labels and the recommendation degree of the training target, obtain the fusion loss information, and adjust the fusion loss information with the various training bias degrees to obtain the target loss information; The second iteration module is used to update the second initial multi-objective recommendation model in reverse based on the target loss information to obtain the second updated multi-objective recommendation model. The second updated multi-objective recommendation model is used as the second initial multi-objective recommendation model, and the step of obtaining training samples is returned to be executed until the second training completion condition is met, and the second multi-objective recommendation model is obtained.
21. The multi-objective recommendation model training device according to claim 20, characterized in that, The target loss calculation module is also used to calculate the interaction loss of each business objective based on each interaction tag and the degree of fusion recommendation, and obtain each interaction loss information; obtain the interaction weight corresponding to each interaction tag, and calculate the fusion loss based on each interaction weight and each interaction loss information, and obtain the fusion loss information. Based on the interaction labels and the degree of bias of each interaction, the sample weights corresponding to the training samples are determined, and the fusion loss information is biased and adjusted based on the sample weights to obtain the target loss information.
22. The multi-objective recommendation model training device according to claim 21, characterized in that, The target loss calculation module is also used to obtain the interaction priority corresponding to each interaction label, and determine the sample type corresponding to the training sample based on the interaction priority. Based on the sample type, find the corresponding interaction weights from the preset interaction matrix; use each interaction weight to weight the corresponding interaction loss information to obtain each weighted loss information, and calculate the sum of the information of each weighted loss information to obtain the fusion loss information.
23. The multi-objective recommendation model training device according to claim 21, characterized in that, The target loss calculation module is further configured to obtain the interaction priority corresponding to each interaction label, determine the sample type corresponding to the training sample based on the interaction priority; obtain the corresponding training sample sequence based on each interaction label, determine the target sample sequence from each training sample sequence based on the sample type; determine the sequence position corresponding to the training sample from the target sample sequence, and when the sequence position exceeds a preset position threshold, obtain the sample weight corresponding to the training sample as the first target sample weight. When the sequence position does not exceed the preset position threshold, the sample weight corresponding to the training sample is obtained as the second target sample weight.
24. The multi-objective recommendation model training device according to claim 23, characterized in that, The target loss calculation module is also used to obtain each training sample and the degree of each interaction bias corresponding to each training sample; based on each interaction label, the training samples are sorted according to the degree of each interaction bias corresponding to each training sample to obtain the training sample sequence corresponding to each interaction label.
25. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.
26. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.
27. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.