Object recommendation model training method and device, electronic equipment and storage medium
By combining predicted satisfactory information and real satisfactory information in the object recommendation model for training, the problem of inaccurate user object recommendation is solved, and more accurate user interest modeling and recommendation effects are achieved.
Patent Information
- Application Number
- CN202411959389.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is not accurate enough when recommending user objects, mainly due to the noise in the historical data of user interaction and the sparse negative feedback data, which makes it impossible to accurately reflect the user's true preferences.
By entering the account attribute information, object attribute information and interaction history of the sample account into the object recommendation model to be trained, the satisfactory information and click information are predicted. At the same time, the pre-trained satisfaction model is used to obtain real satisfaction information, and the object recommendation model is trained based on real click information.
It improves the accuracy of the object recommendation model, can better model account interests, accurately predict user clicks and satisfaction, thereby improving the accuracy of recommendations.
Smart Images

Figure CN120045775A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, electronic device, and storage medium for training an object recommendation model. Background Art
[0002] With the development of Internet technology, various application programs (abbreviated as APPs) have emerged in people's lives. Through various APPs, people can complete activities such as shopping, learning, and entertainment.
[0003] Currently, when a user uses an APP, the APP often recommends objects to the user. For example, when a user uses a shopping APP, the shopping APP recommends products that the user may like based on the user's click history. When a user uses a video APP, the video APP recommends videos that the user may want to watch based on the user's click history. However, the user may have interaction behaviors for many reasons, resulting in a lot of noise in the user's interaction history data, and the negative feedback data in the interaction history data is also relatively sparse, making the user's click history data unable to reflect the user's true preferences, resulting in an inability to accurately predict the user's satisfaction with the recommended objects, and thus the APP cannot accurately recommend objects to the user.
[0004] Therefore, there is a problem in the traditional technology that the object recommendation for users is not accurate enough. Summary of the Invention
[0005] The present disclosure provides a method, apparatus, electronic device, storage medium, and computer program product for training an object recommendation model to at least solve the problem that the object recommendation for users in the related technology is not accurate enough. The technical solutions of the present disclosure are as follows:
[0006] According to a first aspect of an embodiment of the present disclosure, there is provided a method for training an object recommendation model, including:
[0007] Inputting the account attribute information of a sample account, the object attribute information of each sample object, and the interaction history record between the sample account and each sample object into an object recommendation model to be trained, and obtaining predicted satisfaction information and predicted click information; the predicted satisfaction information is used to represent the satisfaction situation of the sample account with respect to each sample object; the predicted click information is used to represent the click situation of the sample account with respect to each sample object;
[0008] Inputting the account attribute information of the sample account and the object attribute information of each sample object into a pre-trained satisfaction model to obtain the satisfaction information corresponding to each sample object for the sample account; the pre-trained satisfaction model is trained using account feedback data; the account feedback data includes the satisfaction feedback data of the sample account with respect to each sample object;
[0009] Determine the true click information of the sample accounts for each sample object. Train the object recommendation model to be trained based on the true click information, predicted click information, satisfaction information, and predicted satisfaction information of the sample accounts for each sample object, and obtain the trained object recommendation model.
[0010] In a possible implementation, training the object recommendation model to be trained based on the true click information, predicted click information, satisfaction information, and predicted satisfaction information of the sample accounts for each sample object to obtain the trained object recommendation model includes:
[0011] Perform weighted mixing on the predicted satisfaction information and the satisfaction information to obtain the weighted satisfaction information;
[0012] Determine the satisfaction loss information based on the satisfaction information and the weighted satisfaction information, and determine the click loss information based on the true click information and the predicted click information;
[0013] Train the object recommendation model to be trained based on the satisfaction loss information and the click loss information to obtain the trained object recommendation model.
[0014] In a possible implementation, the training process of the pre-trained satisfaction model is as follows:
[0015] Input the account attribute information of the sample account, the object attribute information of the sample object, and the satisfaction feedback data of the sample account for the sample object into the propensity score model to obtain the propensity score information, and input the account attribute information of the sample account and the object attribute information of the sample object into the satisfaction model to be trained to obtain the satisfaction score information; the propensity score information represents the probability that the sample account submits satisfaction feedback data for the sample object; the satisfaction score information represents the satisfaction degree of the sample account for the sample object;
[0016] Determine the true satisfaction information of the sample account for the sample object, and determine the satisfaction score loss information based on the true satisfaction information and the satisfaction score information;
[0017] Weight the satisfaction score loss information using the propensity score information to obtain the weighted satisfaction score loss information;
[0018] Train the satisfaction model to be trained based on the weighted satisfaction score loss information.
[0019] In a possible implementation, input the account features of the sample account, the object features of each sample object, and the interaction history records between the sample account and each sample object into the object recommendation model to be trained to obtain the predicted satisfaction information and the predicted click information, including:
[0020] Input the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history record between the sample account and each sample object into the encoding layer of the object recommendation model to obtain the account feature representation corresponding to the sample account, the object feature representation corresponding to each sample object, and the interaction history feature representation corresponding to the sample account;
[0021] Input the interaction history feature representation corresponding to the sample account into the satisfaction feature extraction network of the object recommendation model to obtain the satisfaction feature representation corresponding to the sample account;
[0022] Input the satisfaction feature representation corresponding to the sample account, the account feature representation corresponding to the sample account, and the object feature representation corresponding to each sample object into the attention network of the object recommendation model to obtain the predicted satisfaction information and the predicted click information.
[0023] In a possible implementation, inputting the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history record between the sample account and each sample object into the encoding layer of the object recommendation model to obtain the account feature representation corresponding to the sample account, the object feature representation corresponding to each sample object, and the interaction history feature representation corresponding to the sample account includes:
[0024] According to the satisfaction information corresponding to each sample object of the sample account, separate the satisfied interaction history record and the dissatisfied interaction history record corresponding to the sample account from the interaction history record;
[0025] Input the interaction history record, the satisfied interaction history record, and the dissatisfied interaction history record into the encoding layer of the object recommendation model to obtain the interaction feature representation, the satisfied interaction feature representation, and the dissatisfied interaction feature representation corresponding to the sample account;
[0026] Use the interaction feature representation, the satisfied interaction feature representation, and the dissatisfied interaction feature representation as the interaction history feature representation.
[0027] In a possible implementation, the satisfaction feature extraction network includes a satisfaction information disentanglement network and a satisfaction information enhancement network. Inputting the interaction history feature representation corresponding to the sample account into the satisfaction feature extraction network of the object recommendation model to obtain the satisfaction feature representation corresponding to the sample account includes:
[0028] Input the interaction feature representation, the satisfied interaction feature representation, and the dissatisfied interaction feature representation into the sequence feature extraction network of the object recommendation model to obtain the sequence feature representations corresponding to the interaction history record, the satisfied interaction history record, and the dissatisfied interaction history record respectively;
[0029] Input the sequence feature representation corresponding to the interaction history into the satisfaction information enhancement network of the object recommendation model to obtain the satisfaction information enhanced representation corresponding to the sample account. And input the sequence feature representations corresponding to the interaction history, the satisfied interaction history, and the dissatisfied interaction history into the satisfaction information disentanglement network in the object recommendation model. Through the satisfaction information disentanglement network, perform satisfaction information disentanglement on the interaction history to obtain the satisfaction information representation and the dissatisfied information representation corresponding to the sample account;
[0030] Generate the satisfaction feature representation corresponding to the sample account from the satisfaction information enhanced representation, the satisfaction information representation, the dissatisfied information representation corresponding to the sample account, and the sequence feature representation corresponding to the interaction history.
[0031] In a possible implementation manner, inputting the sequence feature representation corresponding to the interaction history into the satisfaction information enhancement network of the object recommendation model to obtain the satisfaction information enhanced representation corresponding to the sample account includes:
[0032] Group the sample objects in the interaction history according to the satisfaction information corresponding to each sample object of the sample account to obtain multiple sample object groups; the satisfaction information between the sample objects belonging to the same sample object group satisfies a preset similarity condition;
[0033] For any sample object group, perform weighted aggregation on the satisfaction information of each sample object in the sample object group to obtain the satisfaction information representation corresponding to the sample object group;
[0034] Perform weighted aggregation on the satisfaction information representations corresponding to each sample object group to obtain the satisfaction information enhanced representation corresponding to the sample account.
[0035] In a possible implementation manner, inputting the sequence feature representations corresponding to the interaction history, the satisfied interaction history, and the dissatisfied interaction history into the satisfaction information disentanglement network in the object recommendation model, and performing satisfaction information disentanglement on the interaction history through the satisfaction information disentanglement network to obtain the satisfaction information representation and the dissatisfied information representation corresponding to the sample account includes:
[0036] Through the attention mechanism in the satisfaction information disentanglement network, based on the sequence feature representations corresponding to the interaction history and the satisfied interaction history, separate the satisfaction information corresponding to the sample account from the interaction history. And through the attention mechanism in the satisfaction information disentanglement network, based on the sequence feature representations corresponding to the interaction history and the dissatisfied interaction history, separate the dissatisfied information of the sample account from the interaction history;
[0037] Perform weighted aggregation on the sequence feature representation corresponding to the satisfactory interaction history and the satisfactory information corresponding to the sample account to obtain the satisfactory information representation corresponding to the sample account, and perform weighted aggregation on the sequence feature representation corresponding to the unsatisfactory interaction history and the unsatisfactory information corresponding to the sample account to obtain the unsatisfactory information representation corresponding to the sample account.
[0038] According to a second aspect of the embodiments of the present disclosure, there is provided a training device for an object recommendation model, including:
[0039] A prediction module, configured to input the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history between the sample account and each sample object into the object recommendation model to be trained, and obtain predicted satisfactory information and predicted click information; the predicted satisfactory information is used to characterize the satisfaction of the sample account with respect to each sample object; the predicted click information is used to characterize the click situation of the sample account with respect to each sample object;
[0040] A determination module, configured to input the account attribute information of the sample account and the object attribute information of each sample object into a pre-trained satisfaction model to obtain the satisfactory information corresponding to each sample object for the sample account; the pre-trained satisfaction model is trained using account feedback data; the account feedback data includes satisfactory feedback data of the sample account with respect to each sample object;
[0041] A training module, configured to determine the true click information of the sample account with respect to each sample object, and train the object recommendation model to be trained according to the true click information, predicted click information, satisfactory information, and predicted satisfactory information of the sample account with respect to each sample object, to obtain a trained object recommendation model.
[0042] In a possible implementation manner, the training module is specifically configured to: perform weighted mixing on the predicted satisfactory information and the satisfactory information to obtain weighted satisfactory information; determine satisfactory loss information according to the satisfactory information and the weighted satisfactory information, and determine click loss information according to the true click information and the predicted click information; train the object recommendation model to be trained according to the satisfactory loss information and the click loss information to obtain a trained object recommendation model.
[0043] In a possible implementation, the device further includes: a pre-training module, configured to input the account attribute information of a sample account, the object attribute information of a sample object, and the satisfaction feedback data of the sample account for the sample object into a propensity score model to obtain propensity score information, and input the account attribute information of the sample account and the object attribute information of the sample object into a satisfaction model to be trained to obtain satisfaction score information; the propensity score information represents the probability that the sample account submits satisfaction feedback data for the sample object; the satisfaction score information represents the satisfaction degree of the sample account for the sample object; determine the true satisfaction information of the sample account for the sample object, and determine satisfaction score loss information according to the true satisfaction information and the satisfaction score information; weight the satisfaction score loss information by using the propensity score information to obtain weighted satisfaction score loss information; and train the satisfaction model to be trained according to the weighted satisfaction score loss information.
[0044] In a possible implementation, the prediction module is specifically configured to: input the account attribute information of a sample account, the object attribute information of each sample object, and the interaction history record between the sample account and each sample object into the encoding layer of an object recommendation model to obtain an account feature representation corresponding to the sample account, an object feature representation corresponding to each sample object, and an interaction history feature representation corresponding to the sample account; input the interaction history feature representation corresponding to the sample account into the satisfaction feature extraction network of the object recommendation model to obtain a satisfaction feature representation corresponding to the sample account; and input the satisfaction feature representation corresponding to the sample account, the account feature representation corresponding to the sample account, and the object feature representations corresponding to each sample object into the attention network of the object recommendation model to obtain predicted satisfaction information and predicted click information.
[0045] In a possible implementation, the prediction module is specifically configured to: separate the satisfaction interaction history record and the dissatisfaction interaction history record corresponding to the sample account from the interaction history record according to the satisfaction information of the sample account for each sample object corresponding sample object; input the interaction history record, the satisfaction interaction history record, and the dissatisfaction interaction history record into the encoding layer of the object recommendation model to obtain an interaction feature representation corresponding to the sample account, a satisfaction interaction feature representation, and a dissatisfaction interaction feature representation; and use the interaction feature representation, the satisfaction interaction feature representation, and the dissatisfaction interaction feature representation as the interaction history feature representation.
[0046] In a possible implementation, the satisfaction feature extraction network includes a satisfaction information disentanglement network, a satisfaction information enhancement network, and a prediction module, which is specifically configured to: input the interaction feature representation, the satisfied interaction feature representation, and the dissatisfied interaction feature representation into the sequence feature extraction network of the object recommendation model to obtain the sequence feature representations corresponding to the interaction history, the satisfied interaction history, and the dissatisfied interaction history respectively; input the sequence feature representation corresponding to the interaction history into the satisfaction information enhancement network of the object recommendation model to obtain the satisfaction information enhancement representation corresponding to the sample account, and input the sequence feature representations corresponding to the interaction history, the satisfied interaction history, and the dissatisfied interaction history into the satisfaction information disentanglement network in the object recommendation model, and perform satisfaction information disentanglement on the interaction history through the satisfaction information disentanglement network to obtain the satisfaction information representation and the dissatisfied information representation corresponding to the sample account; generate the satisfaction feature representation corresponding to the sample account from the satisfaction information enhancement representation, the satisfaction information representation, the dissatisfied information representation corresponding to the sample account, and the sequence feature representation corresponding to the interaction history.
[0047] In a possible implementation, the prediction module is specifically configured to: group the sample objects in the interaction history according to the satisfaction information corresponding to each sample object for the sample account to obtain a plurality of sample object groups; the satisfaction information between the sample objects belonging to the same sample object group satisfies a preset similarity condition; for any sample object group, perform weighted aggregation on the satisfaction information of each sample object in the sample object group to obtain the satisfaction information representation corresponding to the sample object group; perform weighted aggregation on the satisfaction information representations corresponding to each sample object group to obtain the satisfaction information enhancement representation corresponding to the sample account.
[0048] In a possible implementation, the prediction module is specifically configured to: separate the satisfaction information corresponding to the sample account from the interaction history through the attention mechanism in the satisfaction information disentanglement network based on the sequence feature representations corresponding to the interaction history and the satisfied interaction history respectively, and separate the dissatisfied information corresponding to the sample account from the interaction history through the attention mechanism in the satisfaction information disentanglement network based on the sequence feature representations corresponding to the interaction history and the dissatisfied interaction history respectively; perform weighted aggregation on the sequence feature representation corresponding to the satisfied interaction history and the satisfaction information corresponding to the sample account to obtain the satisfaction information representation corresponding to the satisfied interaction history, and perform weighted aggregation on the sequence feature representation corresponding to the dissatisfied interaction history and the dissatisfied information corresponding to the sample account to obtain the dissatisfied information representation corresponding to the dissatisfied interaction history.
[0049] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the training method of the object recommendation model according to the first aspect or any possible implementation manner of the first aspect is implemented.
[0050] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the training method of the object recommendation model according to the first aspect or any possible implementation manner of the first aspect is implemented.
[0051] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including instructions. When the instructions are executed by a processor of an electronic device, the electronic device can execute the training method of the object recommendation model according to the first aspect or any possible implementation manner of the first aspect.
[0052] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: By inputting the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history records between the sample account and each sample object into the object recommendation model to be trained, predicted satisfaction information and predicted click information are obtained; the predicted satisfaction information is used to characterize the satisfaction of the sample account with respect to each sample object; the predicted click information is used to characterize the click situation of the sample account with respect to each sample object; inputting the account attribute information of the sample account and the object attribute information of each sample object into the pre-trained satisfaction model to obtain the satisfaction information of the sample account corresponding to each sample object; the pre-trained satisfaction model is trained using account feedback data; the account feedback data includes the satisfaction feedback data of the sample account with respect to each sample object; determining the true click information of the sample account with respect to each sample object, and training the object recommendation model to be trained according to the true click information, predicted click information, satisfaction information, and predicted satisfaction information of the sample account with respect to each sample object to obtain the trained object recommendation model; in this way, an unbiased satisfaction model can be pre-trained in advance using account feedback data that captures account satisfaction information, and the satisfaction information of the sample account with respect to the sample object output by the unbiased satisfaction model and the true click information of the sample account with respect to the sample object are used together as the true information in the training stage of the object recommendation model, that is, as a supervision signal, and compared with the prediction information in the training stage of the object recommendation model (that is, the predicted satisfaction information and predicted click information of the sample account with respect to the sample object output by the object recommendation model in the training stage), and then the object recommendation model is trained to obtain an object recommendation model that can better model the account interest, which can accurately predict the click situation and satisfaction situation of the account, so as to improve the accuracy of object recommendation for the account.
[0053] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation to the present disclosure.
[0055] Figure 1 is an application environment diagram of a method for training an object recommendation model shown according to an exemplary embodiment.
[0056] Figure 2 is a flowchart of a method for training an object recommendation model shown according to an exemplary embodiment.
[0057] Figure 3 is a schematic diagram of a method for training an object recommendation model shown according to an exemplary embodiment.
[0058] Figure 4 is a flowchart of training steps of an object recommendation model shown according to an exemplary embodiment.
[0059] Figure 5 is a schematic diagram of a model framework of an object recommendation model shown according to an exemplary embodiment.
[0060] Figure 6 is a flowchart of a method for training an object recommendation model shown according to another exemplary embodiment.
[0061] Figure 7 is a block diagram of an apparatus for training an object recommendation model shown according to an exemplary embodiment.
[0062] Figure 8 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings.
[0064] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0065] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.
[0066] The training method for the object recommendation model provided by the present disclosure can be applied to an application environment as Figure 1 shown. Among them, the terminal 110 interacts with the server 120 through the network. The server 120 first inputs the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history record between the sample account and each sample object into the object recommendation model to be trained, and obtains the predicted satisfaction information and the predicted click information; the predicted satisfaction information is used to characterize the satisfaction of the sample account with each sample object; the predicted click information is used to characterize the click situation of the sample account with each sample object; then, the server 120 inputs the account attribute information of the sample account and the object attribute information of each sample object into the pre-trained satisfaction model to obtain the satisfaction information corresponding to each sample object by the sample account; the pre-trained satisfaction model is trained using account feedback data; the account feedback data includes the satisfaction feedback data of the sample account with each sample object; finally, the server 120 determines the true click information of the sample account with respect to each sample object, and trains the object recommendation model to be trained according to the true click information, predicted click information, satisfaction information, and predicted satisfaction information of the sample account with respect to each sample object, to obtain the trained object recommendation model. The trained object recommendation model can be used to recommend objects of interest to the users of the terminal 120. In practical applications, the terminal 110 can be a computer, a mobile phone, a tablet computer or other terminals, and the server 120 can be implemented by an independent server or a server cluster composed of multiple servers.
[0067] Figure 2 is a flowchart of a training method for an object recommendation model shown according to an exemplary embodiment. As Figure 2 shown, the training method for the object recommendation model is used in the server 120 and includes the following steps:
[0068] In step S210, the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history records between the sample account and each sample object are input into the object recommendation model to be trained, and predicted satisfaction information and predicted click information are obtained.
[0069] Among them, the sample account refers to the account in any training sample in the training sample set. The training sample set includes multiple training samples, and each training sample includes the account attribute information of a sample account, the interaction history record of a sample account, and the object attribute information of each sample object interacted with by a sample account.
[0070] Among them, the account attribute information of the sample account refers to various account features of the sample account, which can be represented as u. For example, account id, account level.
[0071] Among them, the sample object can be an object interacted with by the sample account within a preset time period. For example, the sample object can be an item, video, tweet, etc. clicked by the sample account in the historical time period.
[0072] Among them, the object attribute information of the sample object refers to various object features of the sample object, which can be represented as i. For example, when the sample object is an item, the object features can refer to item id, item category, and the item category can be sports items, game items; when the sample object is a video, the object features can refer to video id, video category, and the video category can be popular science videos, competition videos, cute pet videos.
[0073] Among them, the interaction history record between the sample account and each sample object refers to the click sequence formed by the sample account clicking on the sample object within a preset time period, which is composed of the object identifiers of the sample objects clicked by the sample account within the preset time period. The interaction history record of the sample account can be represented as r.
[0074] Among them, the object recommendation model refers to a recommendation model used to recommend objects to a user account. For example, the object recommendation model can be a recommendation model that recommends videos, products, and graphics to a user account.
[0075] Among them, the predicted satisfaction information refers to the satisfaction of the sample account predicted by the object recommendation model for each sample object.
[0076] Among them, the predicted click information refers to the click-through rate of the sample account predicted by the object recommendation model for each sample object.
[0077] Specifically, the server inputs the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history records between the sample account and each sample object into the object recommendation model to be trained, and predicts the satisfaction degree of the sample account with respect to each sample object through the object recommendation model to be trained. and click-through rate
[0078] In step S220, the account attribute information of the sample account and the object attribute information of each sample object are input into the pre-trained satisfaction model to obtain the satisfaction information corresponding to each sample object for the sample account.
[0079] Among them, the pre-trained satisfaction model is pre-trained in advance using account feedback data.
[0080] Among them, the account feedback data includes the satisfaction feedback data of the sample account for each sample object.
[0081] In practical applications, the account feedback data can be questionnaire feedback data collected after distributing questionnaires of each sample object to the sample account. The questionnaire feedback data can include both the questionnaire data submitted when the sample account is satisfied with the sample object and the questionnaire data submitted when the sample account is not satisfied with the sample object. The questionnaire data submitted when the sample account is not satisfied with the sample object may be empty, which indicates that the sample account does not feedback the specific reason for dissatisfaction. The questionnaire data submitted when the sample account is not satisfied with the sample object can be contentful, which indicates that the sample account feedbacks the specific reason for dissatisfaction. Whether the questionnaire data submitted by the sample account is empty or not can represent the behavioral characteristics of the sample account's feedback on the questionnaire. Therefore, the account feedback data clearly captures the information of the sample account's satisfaction and dissatisfaction, and the sample account filling out the questionnaire is not affected by other accounts, which can better express the true preferences of the sample account. Therefore, using the account feedback data to pre-train the satisfaction model can obtain an unbiased satisfaction model.
[0082] Specifically, the server inputs the account attribute information u of the sample account and the object attribute information i of each sample object into the pre-trained satisfaction model. Since the satisfaction model is an unbiased satisfaction model pre-trained by account feedback data, the satisfaction information of the sample account for each sample object output by the pre-trained satisfaction model can be used as the true satisfaction degree of the object recommendation model in the training stage.
[0083] In step S230, the true click information of the sample account for each sample object is determined, and the object recommendation model to be trained is trained according to the true click information, predicted click information, satisfaction information, and predicted satisfaction information of the sample account for each sample object to obtain the trained object recommendation model.
[0084] Specifically, referring to Figure 3 , the server determines the true click-through rate y of the sample account for each sample object. Then, based on the true click-through rate y of the sample account for each sample object, the predicted click-through rate true satisfaction and the predicted satisfaction train the object recommendation model to be trained to obtain a trained object recommendation model.
[0085] In the above training method of the object recommendation model, by inputting the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history record between the sample account and each sample object into the object recommendation model to be trained, predicted satisfaction information and predicted click information are obtained; the predicted satisfaction information is used to characterize the satisfaction of the sample account with each sample object; the predicted click information is used to characterize the click situation of the sample account with each sample object; input the account attribute information of the sample account and the object attribute information of each sample object into the pre-trained satisfaction model to obtain the satisfaction information corresponding to each sample object by the sample account; the pre-trained satisfaction model is trained using account feedback data; the account feedback data includes the satisfaction feedback data of the sample account for each sample object; determine the true click information of the sample account for each sample object, and based on the true click information, predicted click information, satisfaction information, and predicted satisfaction information of the sample account for each sample object, train the object recommendation model to be trained to obtain a trained object recommendation model; in this way, it is possible to pre-train an unbiased satisfaction model using account feedback data that captures account satisfaction information, and use the satisfaction information of the sample account for the sample object output by this unbiased satisfaction model and the true click information of the sample account for the sample object together as the true information in the training stage of the object recommendation model, that is, as a supervision signal, and compare it with the prediction information in the training stage of the object recommendation model (that is, the predicted satisfaction information and predicted click information output by the object recommendation model in the training stage for the sample account with respect to the sample object), and then train the object recommendation model to obtain an object recommendation model that can better model account interests, which can accurately predict the click situation and satisfaction situation of the account, so as to improve the accuracy of object recommendation for the account.
[0086] In an exemplary embodiment, as Figure 4 shown, step S230 can be specifically implemented through the following steps:
[0087] In step S410, the predicted satisfaction information and the satisfaction information are weighted and mixed to obtain weighted satisfaction information.
[0088] Among them, the weighted satisfaction information refers to, in the training stage of the object recommendation model, the satisfaction information output by the pre-trained satisfaction model and the predicted satisfaction information output by the object recommendation model After weighted mixing, a new satisfaction information is obtained.
[0089] Specifically, the server weights and mixes the predicted satisfaction degree output by the object recommendation model during the training phase with the true satisfaction degree output by the pre-trained satisfaction model to obtain the weighted satisfaction degree.
[0090] In this application, the satisfaction information output by the pre-trained satisfaction model is used as a supervision signal for the object recommendation model to supervise the training of the object recommendation model. During the training process of the object recommendation model, the predicted satisfaction information output by the object recommendation model may gradually be more accurate than the satisfaction information output by the pre-trained satisfaction model Therefore, an adaptive label correction technology can be adopted to weight and mix the predicted satisfaction information output by the object recommendation model during the training phase with the satisfaction information output by the pre-trained satisfaction model to form a new satisfaction information label as a supervision signal to more accurately guide the training of the object recommendation model.
[0091] In step S420, according to the satisfaction information and the weighted satisfaction information, the satisfaction loss information is determined, and according to the true click information and the predicted click information, the click loss information is determined.
[0092] Among them, the satisfaction loss information refers to the loss between the weighted satisfaction information and the satisfaction information, and the click loss information refers to the loss between the predicted click information and the true click information.
[0093] Specifically, the server determines the satisfaction loss according to the true satisfaction degree output by the pre-trained satisfaction model and the weighted satisfaction degree (obtained by weighted mixing of and ), and determines the click loss according to the true click-through rate y and the predicted click-through rate
[0094] In step S430, according to the satisfaction loss information and the click loss information, the object recommendation model to be trained is trained to obtain the trained object recommendation model.
[0095] Specifically, the server trains the object recommendation model to be trained according to the satisfaction loss and the click loss to obtain the trained object recommendation model.
[0096] In step S430, the object recommendation model is trained according to the two-part loss information. It can be seen that the object recommendation model can not only predict the click situation of the account for the object, but also predict the satisfaction situation of the account for the object. For the prediction of the click situation, binary cross-entropy loss is used as the optimization objective function. For predicting satisfaction, due to the lack of feedback in a large number of user interactions, the satisfaction information provided by the pre-trained satisfaction model is used as pseudo-labels to supervise the training process. In fact, in this application, the satisfaction information provided by the pre-trained satisfaction model is used as a pseudo-label for multi-task learning. These pseudo-labels help the recommendation model better understand and capture the satisfaction of the account with the object, thereby improving the ability of the object recommendation model to be consistent with the true preferences of the user.
[0097] In the technical solution of this embodiment, the predicted satisfaction information and the satisfaction information are weighted and mixed to obtain the weighted satisfaction information; according to the satisfaction information and the weighted satisfaction information, the satisfaction loss information is determined, and according to the true click information and the predicted click information, the click loss information is determined; according to the satisfaction loss information and the click loss information, the object recommendation model to be trained is trained to obtain the trained object recommendation model; in this way, the predicted satisfaction information output by the object recommendation model in the training stage can be weighted and mixed with the satisfaction information output by the pre-trained satisfaction model to form a new satisfaction information label, which is used as the supervision signal of the object recommendation model in the training stage, and can more accurately guide the training of the object recommendation model, thereby facilitating the training of a more accurate object recommendation model.
[0098] For the convenience of those skilled in the art to understand, Figure 5 exemplarily provides a model framework of an object recommendation model based on account feedback data. By adopting the model framework of the object recommendation model, the true preferences of users can be effectively aligned. This model framework solves the problems of data sparsity and selective bias existing in account feedback data by pre-training an unbiased satisfaction model, decouples the satisfaction information and dissatisfaction information of the account from the interaction history by using account feedback data, performs multi-level satisfaction division on the interaction history by using the imputed satisfaction provided by the unbiased satisfaction model, and uses the imputed satisfaction as a supervision signal to better model user interests. Adopting the Figure 5 model framework as shown can improve the sorting accuracy of the prediction results of the object recommendation model, improve the satisfaction of users while enhancing the recommendation performance, and improve the user experience. Figure 5 (a) is the training process of the satisfaction model, Figure 5 (b) is the simplified framework of the object recommendation model, Figure 5 (c) is the detailed framework of the object recommendation model, Figure 5 (d) is the framework of the MLSE satisfaction information enhancement network, Figure 5(e) is the framework of the SID satisfaction information disentanglement network, and the following embodiments are detailed descriptions for Figure 5 .
[0099] In an exemplary embodiment, the training process of the pre-trained satisfaction model is as follows: input the account attribute information of the sample account, the object attribute information of the sample object, and the satisfaction feedback data of the sample account for the sample object into the propensity score model to obtain propensity score information, and input the account attribute information of the sample account and the object attribute information of the sample object into the satisfaction model to be trained to obtain satisfaction score information; the propensity score information represents the probability that the sample account submits satisfaction feedback data for the sample object; the satisfaction score information represents the satisfaction degree of the sample account for the sample object; determine the true satisfaction information of the sample account for the sample object, and determine the satisfaction score loss information according to the true satisfaction information and the satisfaction score information; use the propensity score information to weight the satisfaction score loss information to obtain the weighted satisfaction score loss information; train the satisfaction model to be trained according to the weighted satisfaction score loss information.
[0100] Among them, the propensity score model refers to the Inverse Propensity Score (IPS) model, and the role of the propensity score model is: by setting a loss function to improve the accuracy of the satisfaction model, correct the possible selection bias in the account feedback data, so as to train an unbiased satisfaction model.
[0101] Among them, the propensity score refers to the probability that the sample account submits satisfaction feedback data for the sample object predicted by the propensity score model. In practical applications, the propensity score refers to the probability value that the sample account gives feedback on the questionnaire of the sample object output by the propensity score model.
[0102] Among them, the satisfaction model to be trained refers to the satisfaction model that has not been trained without bias.
[0103] Among them, the satisfaction score information refers to the satisfaction degree of the sample account for the sample object predicted by the satisfaction model during the training stage of the satisfaction model.
[0104] It should be noted that the satisfaction score information in this embodiment is different from the "satisfaction information output by the pre-trained satisfaction model" in the above embodiment. The "satisfaction information output by the pre-trained satisfaction model" in the above embodiment is the satisfaction information output by the pre-trained satisfaction model during the training stage of the object recommendation model, and it is used as a supervision signal during the training stage of the object recommendation model. The satisfaction score information in this embodiment is the satisfaction information output by the satisfaction model to be trained during the training stage of the satisfaction model, and it is used as prediction information during the training stage of the satisfaction model.
[0105] Among them, the true satisfaction information refers to the satisfaction of the sample account with respect to the sample object determined based on the actually collected data.
[0106] It should be noted that the true satisfaction information in this embodiment is different from the "satisfaction information output by the pre-trained satisfaction model" in the above embodiment. The "satisfaction information output by the pre-trained satisfaction model" in the above embodiment is the satisfaction information output by the pre-trained satisfaction model during the training stage of the object recommendation model. It is used as a supervision signal during the training stage of the object recommendation model. In fact, it is used as the true satisfaction during the training stage of the object recommendation model, while the true satisfaction information in this embodiment is the true satisfaction determined based on the actually collected data.
[0107] Among them, the satisfaction score loss information refers to the loss between the satisfaction score information predicted by the satisfaction model and the true satisfaction information during the training stage of the satisfaction model.
[0108] It should be noted that the satisfaction score loss information in this embodiment is different from the satisfaction loss information in the above embodiment. The satisfaction score loss information in this embodiment is the loss information of the satisfaction model during the training stage, while the satisfaction loss information in the above embodiment is the loss information of the object recommendation model during the training stage.
[0109] Among them, the weighted satisfaction score loss information is obtained by weighting the loss information of the satisfaction model during the training stage using the propensity score output by the propensity score model.
[0110] In this application, there is a problem of selection bias in the account feedback data, which may be caused by the platform questionnaire distribution mechanism of the interaction platform or by the users. Specifically, when the platform questionnaire distribution mechanism actually distributes questionnaires, it often targets specific objects to distribute questionnaires. For example, objects that may have a certain impact after a large number of pushes or objects that the platform plans to push a large number of times but is not sure whether there will be traffic; when users give questionnaire feedback, they can choose to be satisfied or not satisfied, but if users do not want to give questionnaire feedback, they can also directly swipe away the questionnaire. Therefore, the selection of the platform questionnaire distribution mechanism and users leads to selection bias in the account feedback data.
[0111] This embodiment can achieve unbiased learning based on the account feedback data, solve the selection bias existing in the account feedback data, see Figure 5 (a), by training a propensity score model to predict the probability that a sample account will give feedback on a questionnaire for a sample object Using the trained propensity score model to train the satisfaction model, using the propensity score output by the propensity score model Weight the loss function of the satisfaction model during the training phase (the loss between the satisfaction score output by the satisfaction model during the training phase and the true satisfaction), and increase the weight for account-object pairs with a small probability of questionnaire feedback for the account. (The feature pair composed of a sample account and a sample object) can achieve unbiased learning on the observed account feedback data and solve the selection bias existing in the account feedback data.
[0112] Specifically, the server inputs the account attribute information u of the sample account, the object attribute information i of each sample object, and the satisfaction feedback data of the sample account for the sample object into the propensity score model to obtain the propensity score. Moreover, the server inputs the account attribute information u of the sample account and the object attribute information i of each sample object into the satisfaction model to be trained to obtain the satisfaction score The server determines the true satisfaction of the sample account for the sample object. The server determines the satisfaction score loss based on the true satisfaction and the satisfaction score predicted by the satisfaction model during the training phase The server weights the satisfaction score loss using the propensity score to obtain the weighted satisfaction score loss. The server trains the satisfaction model to be trained based on the weighted satisfaction score loss.
[0113] During the training process of the satisfaction model, use the propensity gradient clipping technique to alleviate the problem of high variance, and use the object recommendation model to initialize the encoding layer of the satisfaction model, which helps the model generalize to accounts and objects with satisfactory questionnaire feedback. After training the satisfaction model, use the pre-trained satisfaction model to guide the feature representation and training process of the object recommendation model.
[0114] The technical solution of this embodiment trains a propensity score model to predict the probability score of a sample account giving a satisfactory feedback to a sample object when pre-training the satisfaction model, and uses this probability score to weight the loss function of the satisfaction model training. Increasing the weight for sample accounts and sample objects with a small observed satisfactory feedback probability can enable the satisfaction model to be trained to achieve unbiased learning on the basis of sparse account feedback data. By using the propensity score model, the selection bias problem existing in the account feedback data is solved, so that an unbiased satisfaction model can be trained, and this unbiased satisfaction model can be used to guide the object recommendation model for more accurate feature representation and model training.
[0115] In an exemplary embodiment, the account features of the sample account, the object features of each sample object, and the interaction history between the sample account and each sample object are input into the object recommendation model to be trained, and prediction satisfaction information and prediction click information are obtained, including: inputting the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history between the sample account and each sample object into the encoding layer of the object recommendation model to obtain the account feature representation corresponding to the sample account, the object feature representation corresponding to each sample object, and the interaction history feature representation corresponding to the sample account; inputting the interaction history feature representation corresponding to the sample account into the satisfaction feature extraction network of the object recommendation model to obtain the satisfaction feature representation corresponding to the sample account; inputting the satisfaction feature representation corresponding to the sample account, the account feature representation corresponding to the sample account, and the object feature representation corresponding to each sample object into the attention network of the object recommendation model to obtain the prediction satisfaction information and the prediction click information.
[0116] Among them, the encoding layer refers to the network structure in the object recommendation model used to encode the input information of the object recommendation model, which can be expressed as Embedding Layer, and the embedding technology is used for encoding in the encoding layer.
[0117] For example, various account features of the sample account (such as account id, account level, etc.) are mapped from the discrete space to vectors in the dense space, so as to obtain the account feature representation of the sample account. The sample account can be represented as u, and the account feature representation can be represented as V u 。
[0118] For example, various object features of the sample object (such as object id, object type, etc.) are mapped from the discrete space to vectors in the dense space, so as to obtain the object feature representation of the sample object. The sample object can be represented as i, and the object feature representation can be represented as V i 。
[0119] For example, each element in the interaction history record is converted into a vector, so as to obtain the vector corresponding to each element in the interaction history record. In fact, an element in the interaction history record of the sample account represents a sample object clicked by the sample account. By converting each element in the interaction history record into a vector, the interaction history feature representation corresponding to the interaction history record can be obtained.
[0120] Among them, the satisfaction feature extraction network refers to the network structure Satisfaction Feature Alignment used to extract the satisfaction information of the sample account from the interaction history record of the sample account.
[0121] Among them, the satisfaction feature representation refers to the satisfaction feature vector corresponding to the sample account extracted by the satisfaction feature extraction network.
[0122] Among them, the attention network refers to a network structure Dot Product that adopts the click attention mechanism. In the personalized recommendation scenario, the click attention mechanism can be used in the feature matching process of the account and the object. By calculating the similarity between the account interest vector and the object feature vector, the degree of interest of the account in the object can be determined.
[0123] Specifically, referring to Figure 5 (b) and (c), the server inputs the account attribute information u of the sample account, the object attribute information i of each sample object, and the interaction history record r between the sample account and each sample object into the encoding layer Embedding Layer of the object recommendation model. Through the encoding layer Embedding Layer, the account feature representation V corresponding to the sample account is encoded using the embedding technology u , the object feature representation V corresponding to each sample object i and the interaction history feature representation corresponding to the sample account. Then, the server inputs the interaction history feature representation corresponding to the sample account into the satisfaction feature extraction network of the object recommendation model, extracts the satisfaction feature representation corresponding to the sample account, and inputs the satisfaction feature representation corresponding to the sample account, the account feature representation V u corresponding to the sample account and the object feature representation V i corresponding to each sample object into the click attention network Dot Product of the object recommendation model to obtain the predicted satisfaction degree and the predicted click-through rate
[0124] The technical solution of this embodiment can first map the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history record between the sample account and each sample object into vectors in the same dense space, and then further perform feature extraction, so as to obtain more accurate predicted satisfaction degree and predicted click-through rate.
[0125] In an exemplary embodiment, the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history record between the sample account and each sample object are input into the encoding layer of the object recommendation model to obtain the account feature representation corresponding to the sample account, the object feature representation corresponding to each sample object, and the interaction history feature representation corresponding to the sample account, including: separating the satisfactory interaction history record and the unsatisfactory interaction history record corresponding to the sample account from the interaction history record according to the satisfaction information corresponding to each sample object by the sample account; inputting the interaction history record, the satisfactory interaction history record, and the unsatisfactory interaction history record into the encoding layer of the object recommendation model to obtain the interaction feature representation, the satisfactory interaction feature representation, and the unsatisfactory interaction feature representation corresponding to the sample account; and using the interaction feature representation, the satisfactory interaction feature representation, and the unsatisfactory interaction feature representation as the interaction history feature representation.
[0126] Among them, the satisfactory interaction history record and the unsatisfactory interaction history record are obtained by dividing the interaction history record of the sample account based on the satisfaction of the sample account with respect to each sample object. The satisfactory interaction history record is a click sequence formed by the object identifiers of each sample object that the sample account clicks and is satisfied with within a preset time period, and the unsatisfactory interaction history record is a click sequence formed by the object identifiers of each sample object that the sample account clicks and is dissatisfied with within a preset time period. In this application, there are three account behavior sequences: the interaction history record, the satisfactory interaction history record, and the unsatisfactory interaction history record.
[0127] In practical applications, the interaction history record of the sample account can be represented as r, the satisfactory interaction history record of the sample account can be represented as s+, and the unsatisfactory interaction history record of the sample account can be represented as s - 。r, s + 、s - All are time series.
[0128] Among them, the interaction feature representation refers to the encoded feature vector obtained by encoding the interaction history record r of the sample account, and can be represented as V r 。
[0129] Among them, the satisfactory interaction feature representation refers to the encoded feature vector obtained by encoding the satisfactory interaction history record s + of the sample account, and can be represented as
[0130] Among them, the unsatisfactory interaction feature representation refers to the encoded feature vector obtained by encoding the unsatisfactory interaction history record s - of the sample account, and can be represented as
[0131] Among them, the interaction history feature representation of the sample account can be represented as
[0132] Specifically, refer to Figure 5 (c), the server separates the satisfaction interaction history record s of the sample account from the interaction history record r of the sample account corresponding to each sample object according to the satisfaction information output by the pre-trained satisfaction model and the dissatisfaction interaction history record s + from the interaction history record r of the sample account. Then, the server inputs the interaction history record r, the satisfaction interaction history record s - and the dissatisfaction interaction history record s + into the encoding layer Embedding Layer of the object recommendation model, and uses the Embedding technology to encode to obtain the interaction feature representation V corresponding to the sample account - , the satisfaction interaction feature representation r and the dissatisfaction interaction feature representation . The server then uses the interaction feature representation V , the satisfaction interaction feature representation r and the dissatisfaction interaction feature representation as the interaction history feature representation.
[0133] In the technical solution of this embodiment, the satisfaction information of the sample account output by the pre-trained satisfaction model can be used to accurately separate the satisfaction interaction history record and the dissatisfaction interaction history record from the interaction history record, and then the interaction history record, the satisfaction interaction history record and the dissatisfaction interaction history record are mapped into vector representations in a unified dense space through the encoding layer, and the interaction history record, the satisfaction interaction history record and the dissatisfaction interaction history record are mapped into vector representations in a unified dense space as the overall interaction history feature representation, which is beneficial to accurately extracting the satisfaction information and dissatisfaction information of the sample account from the interaction history record of the sample account in the subsequent process.
[0134] In an exemplary embodiment, the satisfaction feature extraction network includes a satisfaction information disentanglement network and a satisfaction information enhancement network. Inputting the interaction history feature representation corresponding to the sample account into the satisfaction feature extraction network of the object recommendation model to obtain the satisfaction feature representation corresponding to the sample account, including: inputting the interaction feature representation, satisfaction interaction feature representation, and dissatisfaction interaction feature representation into the sequence feature extraction network of the object recommendation model to obtain the sequence feature representations corresponding to the interaction history record, satisfaction interaction history record, and dissatisfaction interaction history record respectively; inputting the sequence feature representation corresponding to the interaction history record into the satisfaction information enhancement network of the object recommendation model to obtain the satisfaction information enhanced representation corresponding to the sample account, and inputting the sequence feature representations corresponding to the interaction history record, satisfaction interaction history record, and dissatisfaction interaction history record into the satisfaction information disentanglement network in the object recommendation model, and performing satisfaction information disentanglement on the interaction history record through the satisfaction information disentanglement network to obtain the satisfaction information representation and dissatisfaction information representation corresponding to the sample account; generating the satisfaction feature representation corresponding to the sample account from the satisfaction information enhanced representation, satisfaction information representation, dissatisfaction information representation corresponding to the sample account, and the sequence feature representation corresponding to the interaction history record.
[0135] Among them, the sequence feature extraction network refers to the Transformer network. The Transformer can be used to model the behavior sequence of the interaction history record of the sample account in the short term, so that the hidden vector at each moment contains the information of the entire behavior sequence. For example, using the Transformer network to perform context information encoding on the interaction history record to capture the connection between each sample object and other sample objects.
[0136] Among them, the sequence feature representation corresponding to the interaction history record refers to the vector obtained by performing context information encoding on the interaction feature representation V r corresponding to the interaction history record r using the Transformer network, which can be expressed as H r .
[0137] Among them, the sequence feature representation corresponding to the satisfaction interaction history record refers to the vector obtained by performing context information encoding on the satisfaction interaction feature representation + corresponding to the satisfaction interaction history record s using the Transformer network, which can be expressed as
[0138] Among them, the sequence feature representation corresponding to the dissatisfaction interaction history record refers to the vector obtained by performing context information encoding on the dissatisfaction interaction feature representation - corresponding to the dissatisfaction interaction history record s using the Transformer network, which can be expressed as
[0139] Due to the interaction history r and the satisfied interaction history s + and the dissatisfied interaction history s - Both are sequences formed by arranging in the order of clicks. For the sample objects later in the sequence, the sample objects earlier in the sequence are often the objects that the sample account is more interested in. Therefore, by performing context information encoding processing through Transformer, the implicit behavioral features in the three sequences of the interaction history, satisfied interaction history, and dissatisfied interaction history of the sample account can be further extracted.
[0140] Among them, the satisfied feature extraction network includes the satisfied information disentanglement network SID and the satisfied information enhancement network MLSE.
[0141] Among them, the satisfied information disentanglement network SID uses the attention mechanism to disentangle the satisfied information and dissatisfied information of the sample account from the interaction history, as the satisfied information representation and the dissatisfied information representation
[0142] Among them, the satisfied information enhancement network MLSE uses a pre-trained satisfaction model to extract representations of multiple satisfaction levels from the interaction history of the sample account, as the satisfied information enhancement representation u MLSE .
[0143] Among them, the satisfied feature representation characterizes the satisfied information of the sample account as a whole, and the satisfied feature representation is composed of the satisfied information enhancement representation u corresponding to the interaction history MLSE , the satisfied information representation the dissatisfied information representation and the sequence feature representation H corresponding to the interaction history r These four parts of information are jointly represented.
[0144] Specifically, referring to Figure 5 (c), the server will use the interaction feature representation V corresponding to the interaction history r , the satisfied interaction feature representation corresponding to the satisfied interaction history and the dissatisfied interaction feature representation corresponding to the dissatisfied interaction history Input to the sequence feature extraction network of the object recommendation model to extract the sequence features respectively possessed by the interaction history, satisfied interaction history, and dissatisfied interaction history, that is, to obtain the sequence feature representations H corresponding to the interaction history, satisfied interaction history, and dissatisfied interaction history respectively r , and The server will then use the sequence feature representation H corresponding to the extracted interaction history rInput into the satisfaction information enhancement network MLSE of the object recommendation model to obtain the satisfaction information enhancement representation u corresponding to the interaction history record MLSE , and the sequence features corresponding to the interaction history records, satisfactory interaction history records, and unsatisfactory interaction history records are represented by H r , and Input into the satisfaction information disentanglement network SID in the object recommendation model, and use the satisfaction information disentanglement network to disentangle the interaction history r to obtain the satisfaction information representation and dissatisfied information The server then enhances the satisfaction information corresponding to the interaction history record to represent u MLsE , Satisfactory information expression Dissatisfied information expression The sequence feature corresponding to the interaction history is represented by H r Generate satisfactory feature representations corresponding to sample accounts.
[0145] In practical applications, the satisfaction information corresponding to the interaction history record is enhanced to represent u MLsE , Satisfactory information expression Dissatisfied information expression The sequence feature corresponding to the interaction history is represented by H r It includes the interaction information, satisfaction information, dissatisfaction information, multi-level satisfaction information and the attribute information of the sample account itself. The weight of each vector is obtained through a multi-layer perceptron (MLP), and then weighted aggregation is performed to obtain the final sample account representation. The object recommendation model is used to predict the click rate and satisfaction of the sample account with the sample object.
[0146] In the technical solution of this embodiment, the two structures of the satisfactory information enhancement network and the satisfactory information disentanglement network can be used to achieve the enhancement of satisfactory information while achieving satisfactory information disentanglement, and the satisfactory features in the object recommendation model can be accurately aligned.
[0147] In an exemplary embodiment, a sequence feature representation corresponding to an interaction history record is input into a satisfaction information enhancement network of an object recommendation model to obtain an enhanced representation of the satisfaction information corresponding to the interaction history record, including: grouping each sample object in the interaction history record according to the satisfaction information corresponding to each sample object in the sample account to obtain a plurality of sample object groups; the satisfaction information between sample objects belonging to the same sample object group satisfies a preset similarity condition; for any sample object group, weighted aggregation is performed on the satisfaction information of each sample object in the sample object group to obtain a satisfaction information representation corresponding to the sample object group; weighted aggregation is performed on the satisfaction information representation corresponding to each sample object group to obtain an enhanced representation of the satisfaction information corresponding to the sample account.
[0148] Among them, the sample object group refers to a set of object identifiers composed of multiple sample objects. In practical applications, multiple satisfaction levels can be preset in advance, and different satisfaction levels correspond to different satisfaction value ranges. For example, the first-level satisfaction means that the satisfaction value is in the range of [91, 100], the second-level satisfaction means that the satisfaction value is in the range of [81 - 90], and the third-level satisfaction means that the satisfaction value is in the range of [71, 80].
[0149] Among them, the satisfaction information between the sample objects in the same sample object group satisfying the preset similarity condition means that the satisfaction values of the sample account for any two sample objects in the same sample object group belong to the numerical range corresponding to the same level of satisfaction.
[0150] Among them, the enhanced representation of the satisfaction information corresponding to the interaction history record can be expressed as u MLSE , u MLSE is a multi-level satisfaction representation.
[0151] Since the user's interests will naturally be reflected in multiple aspects, the satisfaction of the sample account with each sample object in the interaction history record will also be different. For example, taking a short video platform as an example, users may frequently browse certain categories of videos, which indicates that users have stronger interests and higher satisfaction with these categories; however, it is also possible that due to misleading titles, thumbnails, inconvenient positions, or other reasons, the videos clicked by users are not what they really are interested in, resulting in users being dissatisfied. For other videos, users may only browse occasionally, which is an intermediate zone between user satisfaction and dissatisfaction, which means that the sample account has different degrees of satisfaction with the sample objects in the interaction history record. It can be known that exploring the different levels of satisfaction of users with the interaction history record can help better simulate the real preferences of users. Therefore, in this embodiment, a pre-trained satisfaction model is used to assign satisfaction scores to the sample objects in the interaction history record, so as to model the multi-level satisfaction of the sample account. In practical applications, by inputting the sample account and its interaction history record into the satisfaction model, the satisfaction score of each sample object in the interaction history record is obtained, and the sample objects in the interaction history record of the sample account are divided into multiple groups according to the satisfaction scores. The sample objects with similar satisfaction become a group, and for each group, a weighted aggregation is performed based on their satisfaction scores to obtain the representation of the corresponding satisfaction level of the group.
[0152] Specifically, referring to Figure 5 (d), the server inputs the account attribute information corresponding to the sample account and the interaction history record of the sample account into the pre-trained satisfaction model to obtain the satisfaction score of each sample object in the interaction history record. The server groups the satisfaction scores of each sample object in the interaction history record, and divides those with similar satisfaction scores into a group, so as to obtain N IThe number of satisfaction levels is N I satisfaction groupings, and the grouping results of satisfaction are expressed as Then, based on the grouping results of satisfaction, the server groups each sample object in the interaction history record r, and the sample object grouping results are Then, the server performs weighted aggregation on the satisfaction scores of each sample object belonging to the same sample object group to obtain the satisfaction information representation corresponding to the satisfaction level of this sample object group. The satisfaction level representations of each sample object group can be expressed as Finally, the server, through the satisfaction information representation corresponding to the satisfaction level of each sample object group continues to perform weighted aggregation to obtain the enhanced satisfaction information representation u of this sample account MLsE .
[0153] In the technical solution of this embodiment, it is possible to extract the satisfaction information representations of multiple satisfaction levels from the interaction history record of the sample account based on the pre-trained satisfaction model, realizing the modeling of the multi-level satisfaction of the sample account and being able to better simulate the true preferences of the sample account.
[0154] In an exemplary embodiment, the sequence feature representations corresponding to the interaction history record, the satisfied interaction history record, and the dissatisfied interaction history record are respectively input into the satisfaction information disentangling network in the object recommendation model. Through the satisfaction information disentangling network, the satisfaction information disentangling of the interaction history record is performed to obtain the satisfaction information representation and the dissatisfied information representation corresponding to the sample account, including: separating the satisfaction information corresponding to the sample account from the interaction history record through the attention mechanism in the satisfaction information disentangling network based on the sequence feature representations corresponding to the interaction history record and the satisfied interaction history record, and separating the dissatisfied information of the sample account from the interaction history record through the attention mechanism in the satisfaction information disentangling network based on the sequence feature representations corresponding to the interaction history record and the dissatisfied interaction history record; performing weighted aggregation on the sequence feature representation corresponding to the satisfied interaction history record and the satisfaction information corresponding to the sample account to obtain the satisfaction information representation corresponding to the sample account, and performing weighted aggregation on the sequence feature representation corresponding to the dissatisfied interaction history record and the dissatisfied information corresponding to the sample account to obtain the dissatisfied information representation corresponding to the sample account.
[0155] Among them, the satisfaction information corresponding to the sample account refers to the satisfaction information of the sample account separated from the interaction history record by the attention mechanism in the satisfaction information disentangling network SID based on the satisfied interaction history record and the interaction history record.
[0156] Among them, the dissatisfied information of the sample account refers to the dissatisfied information of the sample account separated from the interaction history record by the attention mechanism in the satisfaction information disentangling network SID based on the dissatisfied interaction history record and the interaction history record.
[0157] Among them, the satisfaction information representation is a vector that integrates the satisfaction information of the sample account and the sequence feature representation of the satisfaction interaction history record of the sample account The satisfaction information representation can be expressed as
[0158] Among them, the dissatisfaction information representation is a vector that integrates the dissatisfaction information of the sample account and the sequence feature representation of the dissatisfaction interaction history record of the sample account The dissatisfaction information representation can be expressed as
[0159] Since the sample account may not always be satisfied with the sample object it interacts with, because the interaction may be caused by accidental touch or other reasons, and the interaction history record of the sample account includes satisfaction information and dissatisfaction information, the goal of this embodiment is to use the account feedback data to separate the satisfaction information and dissatisfaction information of the sample account from the interaction history record.
[0160] Specifically, referring to Figure 5 (e), the server separates the satisfaction information corresponding to the sample account from the interaction history record r through the attention mechanism in the satisfaction information disentanglement network SID based on the sequence feature representation H r corresponding to the satisfaction interaction history record and And, the server separates the dissatisfaction information corresponding to the sample account from the interaction history record r through the attention mechanism in the satisfaction information disentanglement network SID based on the sequence feature representation H corresponding to the dissatisfaction interaction history record and r The server performs weighted aggregation on the sequence feature representation corresponding to the satisfaction interaction history record and the satisfaction information corresponding to the sample account to obtain the satisfaction information representation corresponding to the sample account And, the server performs weighted aggregation on the sequence feature representation corresponding to the dissatisfaction interaction history record and the dissatisfaction information corresponding to the sample account to obtain the dissatisfaction information representation corresponding to the sample account
[0161] In practical applications, the last positions represented by the sequence features of the interaction history, the satisfied interaction history, and the dissatisfied interaction history can be used as the overall representation of the sample account interaction, the overall satisfied representation, and the overall dissatisfied representation, respectively. Then, using the attention mechanism, the overall representations of the satisfied overall representation and the dissatisfied overall representation are used as queries to separate the satisfied information and the dissatisfied information of the sample account from the interaction history. The attention mechanism calculates the similarity between each element in the interaction history and the satisfied (or dissatisfied) overall representation, enabling the identification of elements for which the sample account is more likely to be satisfied (or dissatisfied). In this way, the similarity can be used as a weight to separate the satisfied (or dissatisfied) information of the sample account from the interaction history, rather than treating all elements equally. After separating the satisfied (dissatisfied) information from the interaction history, a multi-layer perceptron (MLP) is used to fuse the satisfied (dissatisfied) information with the feature information expressed in the satisfied (dissatisfied) interaction history of the sample account to obtain the final satisfied (dissatisfied) information of the sample account.
[0162] In the technical solution of this embodiment, the attention mechanism in the satisfied information disentanglement network can accurately disentangle the satisfied information and the dissatisfied information of the sample account from the interaction history of the sample account based on the interaction history, the satisfied interaction history, and the dissatisfied interaction history of the sample account, and fuse the satisfied information with the sequence feature representation of the satisfied interaction history of the sample account and fuse the dissatisfied information with the sequence feature representation of the dissatisfied interaction history of the sample account to obtain the final satisfied information representation and dissatisfied information representation, and can extract more accurate satisfied information and dissatisfied information of the sample account.
[0163] In another embodiment, a flowchart of a training method for an object recommendation model is provided, as Figure 6 shown. This method is applied to Figure 1 server 120 and includes the following steps.
[0164] Step S610: Input the account attribute information of the sample account, the object attribute information of the sample object, and the satisfaction feedback data of the sample account for the sample object into the propensity score model to obtain propensity score information. Also, input the account attribute information of the sample account and the object attribute information of the sample object into the satisfaction model to be trained to obtain satisfaction score information. The propensity score information represents the probability that the sample account submits satisfaction feedback data for the sample object. The satisfaction score information represents the degree of satisfaction of the sample account with the sample object. Step S620: Determine the true satisfaction information of the sample account for the sample object, and based on the true satisfaction information and the satisfaction score information, determine the satisfaction score loss information. Step S630: Use the propensity score information to weight the satisfaction score loss information to obtain the weighted satisfaction score loss information. Step S640: Train the satisfaction model to be trained according to the weighted satisfaction score loss information. Step S650: Input the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history records between the sample account and each sample object into the object recommendation model to be trained to obtain predicted satisfaction information and predicted click information. The predicted satisfaction information is used to represent the satisfaction of the sample account with each sample object. The predicted click information is used to represent the click situation of the sample account with each sample object. Step S660: Input the account attribute information of the sample account and the object attribute information of each sample object into the pre-trained satisfaction model to obtain the satisfaction information corresponding to each sample object for the sample account. Step S670: Determine the true click information of the sample account for each sample object, and based on the true click information, predicted click information, satisfaction information, and predicted satisfaction information of the sample account for each sample object, train the object recommendation model to be trained to obtain the trained object recommendation model. The specific limitations of the above steps can refer to the specific limitations of a training method for an object recommendation model described above, and will not be elaborated here.
[0165] It should be understood that although Figures 2 - 6 the steps in the flowchart of Figures 2 - 6 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,
[0166] It should be understood that for the same / similar parts among the various embodiments of the above methods in this specification, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. For related parts, reference can be made to the descriptions of other method embodiments.
[0167] Figure 7 It is a block diagram of a training device for an object recommendation model shown according to an exemplary embodiment. Refer to Figure 7 This device includes a prediction module 702, a determination module 704, and a training module 706.
[0168] The prediction module 702 is configured to input the account attribute information of the sample account, the object attribute information of each sample object, and the interaction history records between the sample account and each sample object into the object recommendation model to be trained, and obtain predicted satisfaction information and predicted click information. The predicted satisfaction information is used to characterize the satisfaction of the sample account with respect to each sample object. The predicted click information is used to characterize the click situation of the sample account with respect to each sample object.
[0169] The determination module 704 is configured to input the account attribute information of the sample account and the object attribute information of each sample object into the pre-trained satisfaction model to obtain the satisfaction information corresponding to each sample object for the sample account. The pre-trained satisfaction model is trained using account feedback data. The account feedback data includes the satisfaction feedback data of the sample account with respect to each sample object.
[0170] The training module 706 is configured to determine the true click information of the sample account with respect to each sample object, and train the object recommendation model to be trained based on the true click information, predicted click information, satisfaction information, and predicted satisfaction information of the sample account with respect to each sample object, to obtain the trained object recommendation model.
[0171] In an exemplary embodiment, the training module 706 is specifically configured to perform weighted mixing of the predicted satisfaction information and the satisfaction information to obtain weighted satisfaction information; determine satisfaction loss information based on the satisfaction information and the weighted satisfaction information, and determine click loss information based on the true click information and the predicted click information; and train the object recommendation model to be trained based on the satisfaction loss information and the click loss information to obtain the trained object recommendation model.
[0172] In an exemplary embodiment, the device further includes: a pre-training module, specifically configured to execute inputting the account attribute information of a sample account, the object attribute information of a sample object, and the satisfaction feedback data of the sample account for the sample object into a propensity score model to obtain propensity score information, and inputting the account attribute information of the sample account and the object attribute information of the sample object into a satisfaction model to be trained to obtain satisfaction score information; the propensity score information represents the probability that the sample account submits satisfaction feedback data for the sample object; the satisfaction score information represents the satisfaction degree of the sample account for the sample object; determining the true satisfaction information of the sample account for the sample object, and determining satisfaction score loss information according to the true satisfaction information and the satisfaction score information; weighting the satisfaction score loss information by using the propensity score information to obtain weighted satisfaction score loss information; and training the satisfaction model to be trained according to the weighted satisfaction score loss information.
[0173] In an exemplary embodiment, the prediction module 706 is specifically configured to execute inputting the account attribute information of a sample account, the object attribute information of each sample object, and the interaction history record between the sample account and each sample object into the encoding layer of an object recommendation model to obtain an account feature representation corresponding to the sample account, an object feature representation corresponding to each sample object, and an interaction history feature representation corresponding to the sample account; inputting the interaction history feature representation corresponding to the sample account into the satisfaction feature extraction network of the object recommendation model to obtain a satisfaction feature representation corresponding to the sample account; and inputting the satisfaction feature representation corresponding to the sample account, the account feature representation corresponding to the sample account, and the object feature representations corresponding to each sample object into the attention network of the object recommendation model to obtain predicted satisfaction information and predicted click information.
[0174] In an exemplary embodiment, the prediction module 706 is specifically configured to execute separating the satisfaction interaction history record and the dissatisfaction interaction history record corresponding to the sample account from the interaction history record according to the satisfaction information of the sample account for each sample object corresponding sample object; inputting the interaction history record, the satisfaction interaction history record, and the dissatisfaction interaction history record into the encoding layer of the object recommendation model to obtain an interaction feature representation corresponding to the sample account, a satisfaction interaction feature representation, and a dissatisfaction interaction feature representation; and using the interaction feature representation, the satisfaction interaction feature representation, and the dissatisfaction interaction feature representation as the interaction history feature representation.
[0175] In an exemplary embodiment, the satisfaction feature extraction network includes a satisfaction information disentanglement network and a satisfaction information enhancement network, and a prediction module 706, which is specifically configured to perform operations including inputting the interaction feature representation, the satisfied interaction feature representation, and the dissatisfied interaction feature representation into the sequence feature extraction network of the object recommendation model to obtain the sequence feature representations corresponding to the interaction history, the satisfied interaction history, and the dissatisfied interaction history respectively; inputting the sequence feature representation corresponding to the interaction history into the satisfaction information enhancement network of the object recommendation model to obtain the satisfaction information enhanced representation corresponding to the sample account, and inputting the sequence feature representations corresponding to the interaction history, the satisfied interaction history, and the dissatisfied interaction history into the satisfaction information disentanglement network in the object recommendation model, and performing satisfaction information disentanglement on the interaction history through the satisfaction information disentanglement network to obtain the satisfaction information representation and the dissatisfied information representation corresponding to the sample account; generating the satisfaction feature representation corresponding to the sample account from the satisfaction information enhanced representation, the satisfaction information representation, the dissatisfied information representation corresponding to the sample account, and the sequence feature representation corresponding to the interaction history.
[0176] In an exemplary embodiment, the prediction module 706 is specifically configured to perform operations including grouping the sample objects in the interaction history according to the satisfaction information corresponding to each sample object for the sample account to obtain a plurality of sample object groups; the satisfaction information between the sample objects belonging to the same sample object group satisfies a preset similarity condition; for any sample object group, performing weighted aggregation on the satisfaction information of each sample object in the sample object group to obtain the satisfaction information representation corresponding to the sample object group; and performing weighted aggregation on the satisfaction information representations corresponding to each sample object group to obtain the satisfaction information enhanced representation corresponding to the sample account.
[0177] In an exemplary embodiment, the prediction module 706 is specifically configured to perform operations including separating the satisfaction information corresponding to the sample account from the interaction history through the attention mechanism in the satisfaction information disentanglement network based on the sequence feature representations corresponding to the interaction history and the satisfied interaction history respectively, and separating the dissatisfied information corresponding to the sample account from the interaction history through the attention mechanism in the satisfaction information disentanglement network based on the sequence feature representations corresponding to the interaction history and the dissatisfied interaction history respectively; performing weighted aggregation on the sequence feature representation corresponding to the satisfied interaction history and the satisfaction information corresponding to the sample account to obtain the satisfaction information representation corresponding to the satisfied interaction history, and performing weighted aggregation on the sequence feature representation corresponding to the dissatisfied interaction history and the dissatisfied information corresponding to the sample account to obtain the dissatisfied information representation corresponding to the dissatisfied interaction history.
[0178] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0179] Figure 8 FIG. 800 is a block diagram of an electronic device for training an object recommendation model according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0180] Referring to Figure 8 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0181] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-described method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0182] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, an optical disk, or a graphene memory.
[0183] The power component 806 provides power to the various components of the electronic device 800. The power component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0184] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0185] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0186] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0187] The sensor component 814 includes one or more sensors for providing status assessments of various aspects of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor component 814 can also detect a change in the position of the electronic device 800 or an electronic device 800 component, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the device 800, and a change in the temperature of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0188] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0189] In an exemplary embodiment, the electronic device 800 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0190] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as the memory 804 including instructions, and the above instructions can be executed by the processor 820 of the electronic device 800 to complete the above method. For example, the computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0191] In an exemplary embodiment, a computer program product is also provided, and the computer program product includes instructions, and the above instructions can be executed by the processor 820 of the electronic device 800 to complete the above method.
[0192] It should be noted that the above-mentioned devices, electronic devices, computer-readable storage media, computer program products, etc. may also include other implementation manners according to the description of the method embodiments. The specific implementation manners can refer to the description of the relevant method embodiments and will not be elaborated here one by one.
[0193] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0194] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A training method for an object recommendation model, characterized in that: include: Input the account attribute information of the sample account and the object attribute information of each sample object, as well as the interaction history records between the sample account and each sample object, into the object recommendation model to be trained to obtain predicted satisfaction information and predicted click information; the predicted satisfaction information is used to characterize the satisfaction of the sample account with respect to each sample object; the predicted click information is used to characterize the click status of the sample account with respect to each sample object; Inputting the account attribute information of the sample account and the object attribute information of each of the sample objects into a pre-trained satisfaction model to obtain the satisfaction information corresponding to each of the sample objects of the sample account; the pre-trained satisfaction model is trained using account feedback data; the account feedback data includes the satisfaction feedback data of the sample account for each of the sample objects; Determine the actual click information of the sample account for each of the sample objects, and train the object recommendation model to be trained according to the actual click information, the predicted click information, the satisfaction information and the predicted satisfaction information of the sample account for each of the sample objects to obtain a trained object recommendation model.
2. The object recommendation model training method according to claim 1, characterized in that: The step of training the object recommendation model to be trained according to the real click information, the predicted click information, the satisfaction information and the predicted satisfaction information of each sample object by the sample account to obtain the trained object recommendation model comprises: The predicted satisfaction information and the satisfaction information are weighted and mixed to obtain weighted satisfaction information; Determine satisfaction loss information according to the satisfaction information and the weighted satisfaction information, and determine click loss information according to the real click information and the predicted click information; The object recommendation model to be trained is trained according to the satisfaction loss information and the click loss information to obtain the trained object recommendation model.
3. The object recommendation model training method according to claim 1, characterized in that: The training process of the pre-trained satisfaction model is: Inputting the account attribute information of the sample account and the object attribute information of the sample object, as well as the satisfactory feedback data of the sample account for the sample object into a propensity score model to obtain propensity score information, and inputting the account attribute information of the sample account and the object attribute information of the sample object into a satisfaction model to be trained to obtain satisfaction score information; the propensity score information represents the probability of the sample account submitting the satisfactory feedback data for the sample object; the satisfaction score information represents the satisfaction degree of the sample account for the sample object; Determining the real satisfaction information of the sample account for the sample object, and determining satisfaction score loss information according to the real satisfaction information and the satisfaction score information; weighting the satisfaction score loss information by using the propensity score information to obtain weighted satisfaction score loss information; The satisfaction model to be trained is trained according to the weighted satisfaction score loss information.
4. The object recommendation model training method according to claim 1, characterized in that: The step of inputting the account features of the sample accounts and the object features of each sample object, as well as the interaction history records between the sample accounts and each sample object into the object recommendation model to be trained to obtain predicted satisfaction information and predicted click information includes: Input the account attribute information of the sample account and the object attribute information of each sample object, as well as the interaction history records between the sample account and each sample object into the encoding layer of the object recommendation model, and obtain the account feature representation corresponding to the sample account, the object feature representation corresponding to each sample object, and the interaction history feature representation corresponding to the sample account; Inputting the interaction history feature representation corresponding to the sample account into the satisfaction feature extraction network of the object recommendation model to obtain the satisfaction feature representation corresponding to the sample account; The satisfaction feature representation corresponding to the sample account, the account feature representation corresponding to the sample account, and the object feature representation corresponding to each of the sample objects are input into the attention network of the object recommendation model to obtain the predicted satisfaction information and the predicted click information.
5. The object recommendation model training method according to claim 4, characterized in that: Inputting the account attribute information of the sample account and the object attribute information of each sample object, as well as the interaction history records between the sample account and each sample object into the encoding layer of the object recommendation model, obtaining the account feature representation corresponding to the sample account, the object feature representation corresponding to each sample object, and the interaction history feature representation corresponding to the sample account, including: According to the satisfaction information corresponding to each of the sample objects by the sample account, separating the satisfactory interaction history records and the unsatisfactory interaction history records corresponding to the sample account from the interaction history records; Inputting the interaction history records, the satisfactory interaction history records and the unsatisfactory interaction history records into the encoding layer of the object recommendation model to obtain the interaction feature representation, the satisfactory interaction feature representation and the unsatisfactory interaction feature representation corresponding to the sample account; The interaction feature representation, the satisfactory interaction feature representation and the unsatisfactory interaction feature representation are used as the interaction history feature representation.
6. The object recommendation model training method according to claim 5, characterized in that: The satisfactory feature extraction network includes a satisfactory information disentanglement network and a satisfactory information enhancement network. The step of inputting the interactive history feature representation corresponding to the sample account into the satisfactory feature extraction network of the object recommendation model to obtain the satisfactory feature representation corresponding to the sample account includes: Inputting the interaction feature representation, the satisfactory interaction feature representation and the unsatisfactory interaction feature representation into the sequence feature extraction network of the object recommendation model to obtain the sequence feature representations corresponding to the interaction history record, the satisfactory interaction history record and the unsatisfactory interaction history record respectively; Inputting the sequence feature representation corresponding to the interaction history record into the satisfaction information enhancement network of the object recommendation model to obtain the satisfaction information enhancement representation corresponding to the sample account, and inputting the sequence feature representation corresponding to the interaction history record, the satisfactory interaction history record and the unsatisfactory interaction history record into the satisfaction information disentanglement network in the object recommendation model, performing satisfaction information disentanglement on the interaction history record through the satisfaction information disentanglement network to obtain the satisfaction information representation and the unsatisfactory information representation corresponding to the sample account; The satisfaction information enhanced representation, the satisfaction information representation, the dissatisfaction information representation and the sequence feature representation corresponding to the interaction history records corresponding to the sample account are generated into a satisfaction feature representation corresponding to the sample account.
7. The object recommendation model training method according to claim 6, characterized in that: Inputting the sequence feature representation corresponding to the interaction history record into the satisfaction information enhancement network of the object recommendation model to obtain the satisfaction information enhancement representation corresponding to the sample account, including: According to the satisfaction information corresponding to each sample object by the sample account, each sample object in the interaction history record is grouped to obtain a plurality of sample object groups; the satisfaction information between the sample objects belonging to the same sample object group meets a preset similarity condition; For any of the sample object groups, weighted aggregation is performed on the satisfaction information of each of the sample objects in the sample object group to obtain a satisfaction information representation corresponding to the sample object group; The satisfaction information representations corresponding to each of the sample object groups are weightedly aggregated to obtain enhanced satisfaction information representations corresponding to the sample accounts.
8. The object recommendation model training method according to claim 6, characterized in that: The step of inputting the sequence feature representations corresponding to the interaction history records, the satisfactory interaction history records, and the unsatisfactory interaction history records into the satisfactory information disentanglement network in the object recommendation model, performing satisfactory information disentanglement on the interaction history records through the satisfactory information disentanglement network, and obtaining satisfactory information representations and unsatisfactory information representations corresponding to the sample account, comprises: Separating the satisfactory information corresponding to the sample account from the interaction history records based on the sequence feature representations corresponding to the interaction history records and the satisfactory interaction history records respectively through the attention mechanism in the satisfactory information disentanglement network, and separating the dissatisfied information of the sample account from the interaction history records based on the sequence feature representations corresponding to the interaction history records and the dissatisfied interaction history records respectively through the attention mechanism in the satisfactory information disentanglement network; The sequence feature representation corresponding to the satisfactory interaction history record and the satisfactory information corresponding to the sample account are weightedly aggregated to obtain the satisfactory information representation corresponding to the satisfactory interaction history record, and the sequence feature representation corresponding to the unsatisfactory interaction history record and the unsatisfactory information corresponding to the sample account are weightedly aggregated to obtain the unsatisfactory information representation corresponding to the unsatisfactory interaction history record.
9. A training device for an object recommendation model, characterized in that: include: A prediction module, used to input the account attribute information of the sample account and the object attribute information of each sample object, as well as the interaction history records between the sample account and each sample object, into the object recommendation model to be trained, to obtain predicted satisfaction information and predicted click information; the predicted satisfaction information is used to characterize the satisfaction of the sample account with respect to each sample object; the predicted click information is used to characterize the click status of the sample account with respect to each sample object; A determination module, configured to input the account attribute information of the sample account and the object attribute information of each of the sample objects into a pre-trained satisfaction model, and obtain satisfaction information corresponding to each of the sample objects of the sample account; the pre-trained satisfaction model is obtained by training using account feedback data; the account feedback data includes satisfaction feedback data of the sample account for each of the sample objects; A training module is used to determine the actual click information of the sample account for each of the sample objects, and train the object recommendation model to be trained according to the actual click information, the predicted click information, the satisfaction information and the predicted satisfaction information of the sample account for each of the sample objects to obtain a trained object recommendation model.
10. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the object recommendation model training method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the training method for the object recommendation model as described in any one of claims 1 to 8.
12. A computer program product, comprising instructions, characterized in that: When the instruction is executed by a processor of an electronic device, the electronic device is enabled to execute the training method of the object recommendation model as described in any one of claims 1 to 8.