Model training method, target object selection method, device and electronic equipment
By focusing on key user features with high coverage among new users in the recommender system and increasing their weight in the model, the problem of insufficient accuracy of recommender systems in complex scenarios in existing technologies is solved, and more accurate target object recommendations are achieved.
Patent Information
- Application Number
- CN202210665186.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2042-06-13
AI Technical Summary
In existing recommendation systems, target object selection methods based on user click-through rates cannot accurately provide recommendation services to users in complex scenarios, thus affecting user experience.
By acquiring training data including user features, target object features, and scene features, the model to be trained is trained. The focus is on key user features with high coverage on new users, and their weight in the model is increased to improve the accuracy of the parameter prediction model.
This improves the accuracy of the parameter prediction model, enabling more precise recommendations of target objects to users and enhancing the user experience.
Smart Images

Figure CN114996578B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computers, and particularly relates to a model training method, a target object selection method, a device, electronic equipment and a storage medium. BACKGROUND
[0002] With the rapid development of the Internet and the advent of the big data era, people are surrounded by a large amount of information. In order to accurately push information to each user, a recommendation system has become a research hotspot. However, the recommendation accuracy of the related target object selection method still needs to be improved. SUMMARY
[0003] In view of the above problems, the present application provides a model training method, a target object selection method, a device, electronic equipment and a storage medium to improve the above problems.
[0004] In a first aspect, an embodiment of the present application provides a model training method, which comprises: obtaining training data, the training data comprising first training data and second training data, the first training data comprising user features of a user, target object features corresponding to the user features, and scene features corresponding to the target object features, the second training data comprising key user features of the user, the user features comprising the key user features; training a to-be-trained model based on the training data until a training end condition is met, to obtain a parameter prediction model.
[0005] In a second aspect, an embodiment of the present application provides a target object selection method, which comprises: obtaining target user features and candidate target object features; inputting the target user features and the candidate target object features into a parameter prediction model to obtain a recommendation score corresponding to the candidate target object output by the parameter prediction model, the parameter prediction model being obtained based on any one of the methods of claims 1-5; determining a target object corresponding to the target user based on the recommendation score.
[0006] In a third aspect, an embodiment of the present application provides a model training device, which comprises: a data obtaining unit configured to obtain training data, the training data comprising first training data and second training data, the first training data comprising user features of a user, target object features corresponding to the user features, and scene features corresponding to the target object features, the second training data comprising key user features of the user, the user features comprising the key user features; a training unit configured to train a to-be-trained model based on the training data until a training end condition is met, to obtain a parameter prediction model.
[0007] In a fourth aspect, an embodiment of the present application provides a target object selection device, the device comprising: a feature acquisition unit configured to acquire target user features and candidate target object features; a score determination unit configured to input the target user features and the candidate target object features into a parameter prediction model, and acquire a recommendation score corresponding to the candidate target object output by the parameter prediction model, the parameter prediction model being obtained based on any one of the methods of claims 1-5; and a target object determination unit configured to determine a target object corresponding to the target user based on the recommendation score.
[0008] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising one or more processors and a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to perform the method described above.
[0009] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium storing program code, wherein the program code performs the method described above when running.
[0010] The embodiments of the present application provide a model training method, a target object selection method, a device, an electronic device, and a storage medium. After obtaining training data including first training data and second training data, a to-be-trained model is trained by using the obtained training data to obtain a parameter prediction model, wherein the first training data includes user features of each user, target object features corresponding to the user features, and scene features corresponding to the target object features, and the second training data includes key user features of each user. By using the above method, when the to-be-trained model is trained by using the first training data and the second training data, the key user features with relatively high coverage on new users are focused on, so that the weights of the features with relatively high coverage on new users in the model are larger, thereby the influence degree of the features on the model is strengthened, and the accuracy of the parameters predicted by the parameter prediction model is improved, so that the target objects can be more accurately recommended to the user. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0012] Figure 1 An application scenario schematic diagram of a model training method and a target object selection method according to an embodiment of the present application is shown;
[0013] Figure 2 An application scenario of a model training method and a target object selection method according to an embodiment of the present application is shown in the figure;
[0014] Figure 3 A flowchart of a model training method according to an embodiment of the present application is shown in the figure;
[0015] Figure 4 A flowchart of a model training method according to another embodiment of the present application is shown in the figure;
[0016] Figure 5 A working schematic diagram of a SENet according to another embodiment of the present application is shown in the figure;
[0017] Figure 6 A flowchart of step S220 according to another embodiment of the present application is shown in the figure;
[0018] Figure 7 A flowchart of step S230 according to another embodiment of the present application is shown in the figure;
[0019] Figure 8 A flowchart of step S240 according to another embodiment of the present application is shown in the figure;
[0020] Figure 9 A network structure schematic diagram of a parameter prediction model according to another embodiment of the present application is shown in the figure;
[0021] Figure 10 A flowchart of a target object selection method according to yet another embodiment of the present application is shown in the figure;
[0022] Figure 11 A structural block diagram of a model training apparatus according to an embodiment of the present application is shown in the figure;
[0023] Figure 12 A structural block diagram of a recommendation apparatus according to an embodiment of the present application is shown in the figure;
[0024] Figure 13 A structural block diagram of an electronic device or a server for executing a model training method and a target object selection method according to embodiments of the present application is shown in the figure;
[0025] Figure 14 A storage unit for storing or carrying program codes for implementing a model training method and a target object selection method according to embodiments of the present application is shown in the figure. DETAILED DESCRIPTION
[0026] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work are within the scope of protection of the present application.
[0027] With the advent of the big data era, people are surrounded by a large amount of information, therefore, how to accurately push information to each user is particularly important, so the recommendation system has become a research hotspot, for example, more and more Internet companies begin to introduce the recommendation system, solve the problem of showing the same content to users, so that information pushing can achieve the purpose of thousands of people with different faces.
[0028] However, the inventors found in the research on related target object selection methods that general recommendation behavior is set based on the click rate of users on target objects, for example, based on the Listing-Embedding algorithm, the device analyzes and counts the click behavior of users, and then establishes a related model to provide recommendation services for users based on the modeling. However, the user portrait generated based on a single click rate is one-sided, and in a complex scenario, it cannot accurately provide recommendation services for users, affecting user experience.
[0029] Therefore, the inventors proposed the model training method, target object selection method, device, electronic equipment and storage medium in the present application. After obtaining training data including first training data and second training data, the training data obtained is used to train a to-be-trained model to obtain a parameter prediction model, wherein the first training data includes user features of each user, target object features corresponding to the user features, and scene features corresponding to the target object features, and the second training data includes key user features of each user. Through the above method, when the to-be-trained model is trained by the first training data and the second training data, the key user features with high coverage on new users are focused on, so that the weight of these features with high coverage on new users in the model is larger, thereby strengthening the influence degree of these features on the model, and further improving the accuracy of the parameters predicted by the parameter prediction model, so that the target objects can be more accurately recommended to users.
[0030] In the embodiments of the present application, the model training method and the target object selection method provided can be executed by an electronic device. In this way executed by the electronic device, all steps in the model training method and the target object selection method provided in the embodiments of the present application can be executed by the electronic device. For example, as Figure 1As shown, the processor of the electronic device 100 performs the obtaining training data in the model training method; the training data is used to train the to-be-trained model until the training end condition is met, and a parameter prediction model is obtained. In addition, the processor of the electronic device 100 performs the obtaining target user features and candidate target object features; the target user features and the candidate target object features are input into the parameter prediction model to obtain a recommendation score corresponding to the candidate target object; and the target object corresponding to the target user is determined based on the recommendation score.
[0031] Furthermore, the model training method and the target object selection method provided by the embodiments of the present application can also be executed by a server (cloud). Correspondingly, in this way executed by the server, when the model training method is executed, the server can obtain training data in real time, the training data includes first training data and second training data, the first training data includes user features of a user, target object features corresponding to the user features, and scene features corresponding to the target object features, the second training data includes key user features of the user, the user features include the key user features; the to-be-trained model is trained based on the training data until the training end condition is met, and a parameter prediction model is obtained. In addition, when the target object selection method is executed, the server can obtain target user features and candidate target object features in real time; the target user features and the candidate target object features are input into the parameter prediction model to obtain a recommendation score corresponding to the candidate target object output by the parameter prediction model; and the target object corresponding to the target user is determined based on the recommendation score.
[0032] In addition, the model training method and the target object selection method provided by the embodiments of the present application can also be executed by the electronic device and the server cooperatively.
[0033] For example, as shown in the above examples, the electronic device 100 can execute the target object selection method including the following steps: obtaining target user features and candidate target object features, and then the server 200 executes the following steps: inputting the target user features and the candidate target object features into the parameter prediction model to obtain a recommendation score corresponding to the candidate target object output by the parameter prediction model; and determining the target object corresponding to the target user based on the recommendation score. Figure 2
[0034] It should be noted that in this way executed by the electronic device and the server cooperatively, the steps executed by the electronic device and the server are not limited to the above examples, and in actual applications, the steps executed by the electronic device and the server can be adjusted dynamically according to actual situations.
[0035] It should be noted that the electronic device 100 can be a smart phone as shown in Figure 1 and Figure 2 a car machine device, a wearable device, a tablet computer, a notebook computer, a smart speaker, etc. The server 120 can be a standalone physical server, or a server cluster or a distributed system composed of multiple physical servers.
[0036] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0037] Please refer to Figure 3 The model training method provided by the embodiments of the present application is applied to an electronic device or a server as shown in Figure 1 or Figure 2 The method comprises the following steps:
[0038] Step S110: acquiring training data, wherein the training data comprises first training data and second training data, the first training data comprises user features of a user, target object features corresponding to the user features, and scene features corresponding to the target object features, and the second training data comprises key user features of the user, and the user features comprise the key user features.
[0039] In the embodiments of the present application, the user features of the user are features representing the attributes of the user. Optionally, the user features of the user can comprise a user ID (Identity document, identity number), a gender of the user, an age of the user, a click history of the user, a purchase history of the user, and interests, etc. Among them, the click history of the user can be historical click data of the user on the target object; the purchase history of the user can be historical purchase data of the user on the target object; and the interests can be information of interest to the user, etc. For example, taking a recommendation system for shopping of goods as an example, the user ID can be an account of the user when logging in to the recommendation system; the gender and age of the user can be information filled by the user when registering the user ID; the click history of the user can be the historical click times of the user on a certain product; the purchase history of the user can be historical data of the user on the product; and the interests can be the product types (large category, small category) liked by the user, wherein the product types can be divided into several large categories (such as food, medicine, clothing, etc.), and each large category can be divided into several small categories (such as shirts, skirts, sweaters, etc.).
[0040] The target object feature is a feature representing an attribute of the target object. Optionally, the target object features are different for different categories of target objects, wherein the target object can be understood as a recommended object to be pushed to the user in a specified recommendation scenario. For example, in an information recommendation system, the target object can be various information, and the target object feature can be the source of the information, the category of the information, the time of the information, etc. In a recommendation system for shopping for goods, the target object can be various goods, and the target object feature can be the price of the goods, the goods ID, the name of the goods, the Stock Keeping Unit Identity document (SKU ID), the Standard Product Unit Identity document (SPU ID), the category to which the goods belong, the subcategory to which the goods belong, the Click Through Rate (CTR) of the goods, etc.
[0041] The scenario feature is a feature representing a recommendation scenario in which the target object is located. The scenario feature can include time information and weather information of each recommendation scenario. For example, in a goods recommendation system, the scenario feature can be time information and weather information when the target goods are recommended.
[0042] The key user feature of the user is a feature in the user feature that has a relatively high coverage on new users. For example, the key user feature can include the age of the user, the gender of the user, and some user information provided when the user registers the user ID. The key user feature can also include time information and weather information of each recommendation scenario.
[0043] As one way, the training data can be pre-stored in a specified storage area or a cloud server, and when the training data is needed, the training data can be obtained from the specified storage area or the cloud server. In the training data, the target object features corresponding to the user features can include multiple target objects, that is, one user can correspond to multiple target objects, and each target object corresponds to a target object feature and a scenario feature.
[0044] Step S120: training the to-be-trained model based on the training data until a training end condition is met, to obtain a parameter prediction model.
[0045] In the embodiments of the present application, the training end condition can be that the loss value of a preset loss function meets a preset loss value, or the number of training iterations reaches a preset iteration number, or the network parameters of the model are updated to a preset network parameter, etc., which is not limited here.
[0046] As a manner, after the training data is acquired, a preset loss function is acquired, and then the preset loss function and the training data can be used to train the to-be-trained model. When it is detected that a loss value of the preset loss function meets a preset loss value, or a training iteration number reaches a preset iteration number, or a network parameter of the model is updated to a preset network parameter, it is determined that a training end condition is met, and a parameter prediction model is obtained.
[0047] Optionally, the parameter prediction model is used to predict a recommendation score of the target object.
[0048] The target object selection method provided in the application includes the following steps: acquiring training data including first training data and second training data; and training a to-be-trained model by using the acquired training data to obtain a parameter prediction model, wherein the first training data includes user features of each user, target object features corresponding to the user features, and scene features corresponding to the target object features, and the second training data includes key user features of each user. In the above method, when the to-be-trained model is trained by using the first training data and the second training data, the key user features with high coverage on new users are focused on, so that the weights of the features with high coverage on new users in the model are larger, thereby the influence degree of the features on the model is strengthened, and the accuracy of the parameters predicted by the parameter prediction model is improved, so that the target object can be more accurately recommended to the user.
[0049] Please refer to Figure 4 The model training method provided in the embodiments of the application is applied to an electronic device or a server as shown in Figure 1 or Figure 2 The method includes the following steps.
[0050] Step S210: acquiring training data, wherein the training data includes first training data and second training data, the first training data includes user features of a user, target object features corresponding to the user features, and scene features corresponding to the target object features, and the second training data includes key user features of the user, and the user features include the key user features.
[0051] In the embodiments of the application, the user features of the user in the first training data, the target object features corresponding to the user features, and the scene features corresponding to the target object features form a feature vector, and the key user features in the second training data form a feature vector.
[0052] Step S220: inputting the training data into the feature fusion module to acquire first fusion features output by the feature fusion module.
[0053] In this embodiment, the training data further includes labels corresponding to the target object, and the model to be trained includes a feature fusion module, a click-through rate (CTR) prediction module, and a conversion rate prediction module. The feature fusion module fuses features contained in the training data to obtain a first fused feature; the CTR prediction model predicts the probability that the target object will be clicked; and the conversion rate prediction module predicts the probability of a user's further behavioral conversion after clicking the target object. Optionally, the feature fusion module is connected to both the CTR prediction module and the conversion rate prediction model.
[0054] The tags corresponding to the target object can include click score, conversion score, and recommendation score. Click score represents the probability that the target object is clicked, conversion score represents the probability that the user will take further action after clicking the target object, and recommendation score represents the degree to which the target object is recommended.
[0055] In one approach, the feature fusion module includes a first attention network, a second attention network, a first expert network, a second expert network, a third expert network, and a gating network.
[0056] The first and second attention networks are SENet (Squeeze-and-Excitation Networks). SENet is a network that can explicitly model the dependencies between feature channels. SENet can recalibrate features; feature recalibration refers to automatically learning the importance of each feature channel and then boosting useful features while suppressing features that are less useful for the current task. For example... Figure 5 As shown, the input to SENet is a vector of dimension H*W*C, where H is the height, W is the width, and C is the number of channels. First, SENet performs pooling on the H*W*C vector to obtain a 1*1*C vector. This vector then passes through a fully connected (FC) layer to predict the importance of each channel. The resulting importance values are then applied (enhanced) to the corresponding channels of the original H*W*C vector, assigning different weights to each feature. This enhances useful features and suppresses ineffective features, fully leveraging the potential of each feature.
[0057] The first, second, and third expert networks are network layers that perform different transformations on the input feature vectors. Each expert network can have different effects on different tasks.
[0058] The gating network is used to control the variable of the weight of each expert network. The weight of different expert networks can be different for each task. Therefore, the gating network is used to control the weight of each expert network, and the results of multiple expert networks are combined by the gating network.
[0059] In the embodiment of the present application, the first attention network is connected with the first expert network and the second expert network respectively, the second attention network is connected with the third expert network, and the first expert network, the second expert network and the third expert network are connected with the gating network.
[0060] As shown in Figure 6 The step S220 can specifically include the following steps.
[0061] Step S221: inputting the first training data into the first attention network to obtain first attention features output by the first attention network.
[0062] In the embodiment of the present application, the first training data is input into the first attention network, and different weights are given to the features included in the first training data by the first attention network, so as to obtain features with different weights.
[0063] As a way, when the first training data is input into the first attention network, one user corresponds to a group of first training data. That is, the training data input into the model each time is the user features corresponding to one user, the features of one target object corresponding to one user, and the scene features corresponding to one target object. Similarly, the second training data corresponds to the first training data, and when the first training data and the second training data are input, the features corresponding to the same user are input.
[0064] Step S222: inputting the second training data into the second attention network to obtain second attention features output by the second attention network.
[0065] In the embodiment of the present application, the second training data is some features with relatively high coverage on new users. The second training data is input into the second attention network, and different weights are given to the features included in the second training data by the second attention network, so that the weights of these features in the model are larger, thereby enhancing the influence degree of these features on the model.
[0066] Step S223: inputting the first attention features into the first expert network and the second expert network respectively to obtain first reference attention features output by the first expert network and second reference attention features output by the second expert network respectively.
[0067] In the embodiments of the present application, each expert network has a data area that it is good at, and the expert network is "authoritative" on this set of areas and performs better than other expert networks.
[0068] The different feature areas in the first attention feature can be processed in different ways by the first expert network and the second expert network to obtain respective corresponding reference attention features.
[0069] Step S224: input the second attention feature into the third expert network to obtain a third reference attention feature output by the third expert network.
[0070] In the embodiments of the present application, when training the to-be-trained model, it is considered that the new user has less interaction with the recommendation system, and there is a problem that some features are empty. In addition, the amount of new user data is small. If the new user is not specially processed, the to-be-trained model will be biased towards the old user, and the old user will be learned more fully, and the new user will not be friendly. If the to-be-trained model can serve the new user well, it can not only improve the effect of the to-be-trained model in the short term, but also improve the user experience and increase the user retention in the long term.
[0071] From the perspective of feature missing, some features with high coverage on new users can be input into the third expert network to train the to-be-trained model, and the modeling ability of the to-be-trained model for new users can be strengthened.
[0072] As a way, the second attention feature can be input into the third expert network, and the second attention feature can be processed by the third expert network to obtain a third reference attention feature.
[0073] Step S225: input the first reference attention feature, the second reference attention feature, and the third reference attention feature into the gating network to obtain a first fusion feature output by the gating network.
[0074] In the embodiments of the present application, the weights of the first expert network, the second expert network, and the third expert network can be controlled by the gating network, and then the outputs of the first expert network, the second expert network, and the third expert network are weighted and fused to obtain the first fusion feature.
[0075] Step S230: input the first fusion feature into the click rate prediction module to obtain a click score output by the click rate prediction module, the click score representing a click probability of the user on the target object.
[0076] In one approach, the training data also includes a preset click-through rate vector and scene information corresponding to the target object; the click-through rate prediction module includes multiple expert networks, gating networks, and click-through rate tower networks.
[0077] In this embodiment, the preset click-through rate vector is a pre-set click-through rate vector, and the scene information corresponding to the target object represents the recommendation scene in which the target object is located, such as homepage recommendation, end page recommendation, etc. Different recommendation scenes correspond to different scene information, which are not specifically limited here.
[0078] In this model, multiple expert networks in the click-through rate (CTR) prediction model are connected to the gating network in the feature fusion module, multiple expert networks in the CTR prediction module are connected to the gating network in the CTR prediction module, and the gating network in the CTR prediction module is connected to the click-through rate tower network.
[0079] like Figure 7 As shown, step S230 may specifically include:
[0080] Step S231: Connect the preset click-through rate vector and the first fusion feature to obtain the first connection feature.
[0081] In this embodiment of the application, connecting the preset click-through rate vector and the first fusion feature means adding the preset click-through rate vector and the first fusion feature by channel addition to obtain the first connection feature.
[0082] As one approach, the click-through rate prediction module may also include a connection layer that connects a preset click-through rate vector to a first fused feature.
[0083] Step S232: Input the first connection features into the plurality of expert networks respectively, and obtain the plurality of first reference connection features output by the plurality of expert networks.
[0084] In this embodiment of the application, multiple expert networks are also used to perform different transformation processes on different feature regions in the first connection feature to obtain the first reference connection feature output by each expert network.
[0085] Step S233: Input the scene information and the plurality of first reference connection features into the gating network to obtain the second fusion feature output by the gating network.
[0086] In this embodiment, scene information represents the recommended scene corresponding to the input target object features. To achieve the goal of extracting different information for each recommended scene, the scene information can be used as the input of a gating network, thereby enabling the extraction of different information from multiple expert networks for different scenes, thus achieving the goal of multi-scene modeling.
[0087] As a manner, the gating network is used to control the weights of the plurality of expert networks, and then the outputs of the plurality of expert networks and the scene information are weighted and fused to obtain the second fusion feature.
[0088] Step S234: inputting the second fusion feature into the click rate tower network to obtain a click score output by the click rate tower network.
[0089] In the embodiment of the present application, the click rate tower network is used to predict the probability of user clicking the target object. The second fusion feature is input into the click rate tower network, and the click rate tower network can output a click score corresponding to the target object.
[0090] Step S240: inputting the first fusion feature into the conversion rate prediction module to obtain a conversion score output by the conversion rate prediction module, the conversion score representing a further behavior conversion probability of the user after clicking the target object.
[0091] As a manner, the training data further includes a preset conversion rate vector and scene information corresponding to the target object; the conversion rate prediction module includes a plurality of expert networks, a gating network and a conversion rate tower network.
[0092] Among them, the plurality of expert networks in the click rate prediction model are connected with the gating network in the feature fusion module, the plurality of expert networks in the click rate prediction module are connected with the gating network in the click rate prediction module, and the gating network in the click rate prediction module is connected with the click rate tower network.
[0093] As shown in Figure 8 S240 can specifically include:
[0094] Step S241: connecting the preset conversion rate vector and the first fusion feature to obtain a second connection feature.
[0095] In the embodiment of the present application, connecting the preset conversion rate vector and the first fusion feature means adding the preset conversion rate vector and the first fusion feature in the channel to obtain the second connection feature.
[0096] As a manner, the conversion rate prediction module can further include a connection layer, which is used to connect the preset point conversion rate vector and the first fusion feature.
[0097] Step S242: inputting the second connection feature into the plurality of expert networks respectively to obtain a plurality of second reference connection features output by the plurality of expert networks.
[0098] In the embodiments of the present application, the plurality of expert networks are also used to perform different manner transformation processing on different feature regions in the second connection feature, to obtain the second reference connection feature output by each expert network.
[0099] Step S243: inputting the scene information and the plurality of second reference connection features into the gating network to obtain the third fusion feature output by the gating network.
[0100] In the embodiments of the present application, the scene information represents a recommended scene corresponding to the input target object feature. In order to achieve the purpose of extracting different information for each recommended scene, the scene information can be used as the input of the gating network, so that different information is extracted from the plurality of expert networks for different scenes, thereby achieving the purpose of multi-scene modeling.
[0101] As a manner, the gating network is used to control the weights of the plurality of expert networks, and then the outputs of the plurality of expert networks and the scene information are weighted and fused to obtain the third fusion feature.
[0102] Optionally, the training of the click rate prediction module and the conversion rate prediction module is performed simultaneously, that is, the scene information input into the gating network in the click rate prediction module and the gating network in the conversion rate prediction module is the same, and can be input simultaneously.
[0103] Step S244: inputting the third fusion feature into the conversion rate tower network to obtain the conversion score output by the conversion rate tower network.
[0104] In the embodiments of the present application, the conversion rate tower network is used to perform prediction on the further behavior conversion probability of the user after clicking the target object. The third fusion feature is input into the conversion rate tower network, and the conversion rate tower network can output the conversion score corresponding to the target object.
[0105] Step S250: obtaining a recommendation score based on the click score and the conversion score, the recommendation score representing a recommended degree corresponding to the target object.
[0106] In the embodiments of the present application, the click score and the conversion score corresponding to the target object can be weighted and calculated according to a preset weight value to obtain the recommendation score corresponding to the target object.
[0107] Step S260: training the to-be-trained model based on the recommendation score and the label until a training end condition is met, to obtain the parameter prediction model.
[0108] In the embodiment of the present application, the recommendation scores corresponding to the target objects can be compared with the labels respectively, and the to-be-trained model is trained based on the comparison results until the training end condition is met, so as to obtain the parameter prediction model.
[0109] In the embodiment of the present application, the network structure of the parameter prediction model can be as shown in Figure 9 .
[0110] The model training method provided in the present application inputs the training data into the feature fusion module to obtain first fusion features, then inputs the first fusion features into the click rate prediction module and the conversion rate prediction module to obtain click scores and conversion scores, and then obtains recommendation scores based on the click scores and the conversion scores. Finally, the to-be-trained model is trained based on the recommendation scores and the labels until the training end condition is met, and the parameter prediction model is obtained. Through the above method, when the to-be-trained model is trained by the first training data and the second training data, the key user features with high coverage on new users are focused on, so that the weights of these features with high coverage on new users in the model are larger, thereby strengthening the influence degree of these features on the model, and further improving the accuracy of the parameters predicted by the parameter prediction model, so that the target objects can be more accurately recommended to the users.
[0111] Referring to Figure 10 , the target object selection method provided in the embodiment of the present application is applied to an electronic device or a server as shown in Figure 1 or Figure 2 , and the method comprises the following steps.
[0112] Step S310: Obtain target user features and candidate target object features.
[0113] In the embodiment of the present application, the target user features can refer to the user features corresponding to the target user, and the target user can be a user who logs in to the display interface of the recommendation system. The candidate target object features can be target object features of the target objects to be recommended in the recommendation system, wherein the target objects to be recommended in the recommendation system can refer to all target objects corresponding to the recommendation system, or can refer to the target objects obtained after the recall processing of the recommendation system.
[0114] As a kind of mode, the database of recommendation system can be queried based on the user ID of target user to obtain target user features. After determining target user, the target object features of all target objects in the database of recommendation system can be obtained to obtain the target object features as candidate target object features;The corresponding target object features can be obtained from the database of recommendation system in real time based on the result of recall processing, to obtain the target object features as candidate target object features.
[0115] Step S320: inputting the target user feature and the candidate target object feature into a parameter prediction model, and obtaining a recommendation score corresponding to the candidate target object output by the parameter prediction model.
[0116] In the embodiments of the present application, the parameter prediction model comprises a feature fusion module, a click rate prediction module, and a conversion rate prediction module.
[0117] As a manner, the target user object feature and the candidate target object feature are input into the feature fusion module of the parameter prediction module, to obtain corresponding fusion features; then the fusion features are input into the click rate prediction module and the conversion rate prediction module respectively, to obtain corresponding click scores and conversion scores; and then based on the click scores and the conversion scores, the recommendation score corresponding to the candidate target object is obtained.
[0118] The candidate target object can be multiple, and the target user feature can also be multiple. When the target user feature and the candidate target object feature are input into the parameter prediction model, one target user feature and one candidate target object feature are input into the parameter prediction model as a group of data, to obtain the recommendation score corresponding to each group of data.
[0119] Step S330: determining the target object corresponding to the target user based on the recommendation score.
[0120] As a manner, the recommendation scores corresponding to the candidate target object features can be sorted in descending order, and the candidate target objects corresponding to the first N (N is a positive integer) recommendation scores are taken as the target objects corresponding to the target user.
[0121] Optionally, the value of N can be pre-set based on the capacity of the target objects to be recommended included in the recommendation system. The larger the capacity, the larger the value of N.
[0122] Optionally, after obtaining the target object corresponding to the target user, the target object can be directly displayed to the user through the display interface corresponding to the recommendation system.
[0123] The target object selection method provided by the present application firstly obtains a target user feature and a candidate target object feature, then inputs the target user feature and the candidate target object feature into a parameter prediction model, obtains a recommendation score corresponding to the candidate target object output by the parameter prediction model, and determines the target object corresponding to the target user based on the recommendation score. Through the above method, the target object can be accurately recommended to the user.
[0124] Please refer to Figure 11 The model training device 400 provided by the embodiments of the present application comprises:
[0125] The data acquisition unit 410 is configured to acquire training data, the training data including first training data and second training data, the first training data including user features of a user, target object features corresponding to the user features, and scene features corresponding to the target object features, and the second training data including key user features of the user, the user features including the key user features.
[0126] The training unit 420 is configured to train a to-be-trained model based on the training data until a training end condition is met, to obtain a parameter prediction model.
[0127] As an implementation, the training data further includes labels corresponding to the target objects, and the to-be-trained model includes a feature fusion module, a click rate prediction module, and a conversion rate prediction module. Optionally, the training unit 420 is configured to input the training data into the feature fusion module, to acquire first fusion features output by the feature fusion module; input the first fusion features into the click rate prediction module, to acquire a click score output by the click rate prediction module, the click score representing a click probability of the user on the target object; input the first fusion features into the conversion rate prediction module, to acquire a conversion score output by the conversion rate prediction module, the conversion score representing a further behavior conversion probability of the user after clicking the target object; based on the click score and the conversion score, obtain a recommendation score, the recommendation score representing a recommended degree corresponding to the target object; and based on the recommendation score and the labels, train the to-be-trained model until the training end condition is met, to obtain the parameter prediction model.
[0128] In this implementation, the feature fusion module includes a first attention network, a second attention network, a first expert network, a second expert network, a third expert network, and a gating network. The training unit 420 is further configured to input the first training data into the first attention network, to acquire first attention features output by the first attention network; input the second training data into the second attention network, to acquire second attention features output by the second attention network; input the first attention features into the first expert network and the second expert network respectively, to acquire first reference attention features output by the first expert network and second reference attention features output by the second expert network respectively; input the second attention features into the third expert network, to acquire third reference attention features output by the third expert network; and input the first reference attention features, the second reference attention features, and the third reference attention features into the gating network, to acquire the first fusion features output by the gating network.
[0129] The training data further includes a preset click rate vector and scene information corresponding to the target object; the click rate prediction module includes multiple expert networks, a gating network, and a click rate tower network; and the training unit 420 is further configured to connect the preset click rate vector and the first fusion feature to obtain a first connection feature; input the first connection feature into the multiple expert networks respectively to obtain multiple first reference connection features output by the multiple expert networks; input the scene information and the multiple first reference connection features into the gating network to obtain a second fusion feature output by the gating network; and input the second fusion feature into the click rate tower network to obtain a click score output by the click rate tower network.
[0130] The training data further includes a preset conversion rate vector and scene information corresponding to the target object; the conversion rate prediction module includes multiple expert networks, a gating network, and a conversion rate tower network; and the training unit 420 is further configured to connect the preset conversion rate vector and the first fusion feature to obtain a second connection feature; input the second connection feature into the multiple expert networks respectively to obtain multiple second reference connection features output by the multiple expert networks; input the scene information and the multiple second reference connection features into the gating network to obtain a third fusion feature output by the gating network; and input the third fusion feature into the conversion rate tower network to obtain a conversion score output by the conversion rate tower network.
[0131] Please refer to Figure 12 The embodiment of the application provides a target object selection device 500, the device 500 comprises:
[0132] The feature acquisition unit 510 is configured to acquire target user features and candidate target object features.
[0133] The score determination unit 520 is configured to input the target user features and the candidate target object features into a parameter prediction model to obtain a recommendation score corresponding to the candidate target object output by the parameter prediction model, and the parameter prediction model is obtained based on the method in any one of claims 1-5.
[0134] The target object determination unit 530 is configured to determine a target object corresponding to the target user based on the recommendation score.
[0135] It should be noted that the device embodiments in the application correspond to the foregoing method embodiments, and the specific principles of the device embodiments can be referred to the content in the foregoing method embodiments, which will not be described here.
[0136] The following will be combined Figure 13 An electronic device or a server provided by the application will be described.
[0137] Referring to Figure 13 Based on the above model training method, target object selection method and device, another electronic device or server 800 capable of executing the above model training method and target object selection method is provided. The electronic device or server 800 includes one or more (only one is shown in the figure) processors 802, a memory 804 and a network module 806 coupled with each other. The memory 804 stores programs capable of executing the above embodiments, and the processor 802 can execute the programs stored in the memory 804.
[0138] The processor 802 can include one or more processing cores. The processor 802 connects various parts of the electronic device or server 800 through various interfaces and lines, executes various functions and processes data of the electronic device or server 800 by running or executing instructions, programs, code sets or instruction sets stored in the memory 804, and calling data stored in the memory 804. Alternatively, the processor 802 can be implemented in at least one of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 802 can be integrated with a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes operating systems, user interfaces and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 802, but can be realized by a separate communication chip.
[0139] The memory 804 can include a Random Access Memory (RAM) and can also include a Read-Only Memory (ROM). The memory 804 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 804 can include a program storage area and a data storage area, where the program storage area can store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch control function, a sound playing function, an image playing function, etc.), instructions for implementing each of the methods described below, and the like. The data storage area can also store data (such as a phone book, audio and video data, chat record data) created by the electronic device or the server 800 in use, and the like.
[0140] The network module 806 is configured to receive and send electromagnetic waves, and to convert the electromagnetic waves and electrical signals to each other, so as to communicate with a communication network or other devices, for example, to communicate with an audio playing device. The network module 806 can include various existing circuit elements for performing these functions, for example, an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a Subscriber Identity Module (SIM) card, a memory, and the like. The network module 806 can communicate with various networks such as the Internet, an intranet, a wireless network, or other devices through the wireless network. The wireless network can include a cellular phone network, a wireless local area network or metropolitan area network. For example, the network module 806 can interact with a base station to exchange information.
[0141] Reference is made to Figure 14 which shows a structural block diagram of a computer readable storage medium provided by an embodiment of the present application. The computer readable storage medium 900 stores program codes therein, which can be invoked by a processor to execute the methods described in the above method embodiments.
[0142] The computer readable storage medium 900 can be an electronic storage such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer readable storage medium 900 includes a non-transitory computer readable storage medium. The computer readable storage medium 900 has a storage space for program codes 910 to execute any of the method steps described above. These program codes can be read from or written to one or more computer program products. The program codes 910 can be compressed in an appropriate form, for example.
[0143] The model training method, the target object selection method, the device, the electronic equipment and the storage medium provided in the application train a to-be-trained model through the obtained training data after obtaining training data including first training data and second training data, and obtain a parameter prediction model, wherein the first training data includes user features of each user, target object features corresponding to the user features, and scene features corresponding to the target object features, and the second training data includes key user features of each user. Through the above method, when the to-be-trained model is trained through the first training data and the second training data, the key user features with relatively high coverage on new users are focused on, so that the weights of the features with high coverage on new users in the model are larger, thereby the influence degree of the features on the model is strengthened, and the accuracy of the parameters predicted by the parameter prediction model is improved, so that the target objects can be more accurately recommended to the users.
[0144] The embodiments of the application are described above in combination with the drawings, but the application is not limited to the specific embodiments described above, and the specific embodiments described above are only illustrative but not restrictive. Those skilled in the art can make many forms under the inspiration of the application without departing from the purpose of the application and the scope protected by the claims, and all belong to the protection of the application.
Claims
1. A model training method, characterized in that, The method comprises: obtaining training data, the training data comprising first training data and second training data, the first training data comprising user features of a user, target object features corresponding to the user features, and scene features corresponding to the target object features, the second training data comprising key user features of the user, the user features comprising the key user features; the training data further comprising a label corresponding to the target object; the training data further comprising a preset click rate vector and scene information corresponding to the target object; training a to-be-trained model based on the training data until a training end condition is met, to obtain a parameter prediction model, the to-be-trained model comprising a feature fusion module, a click rate prediction module, and a conversion rate prediction module; the training of the to-be-trained model based on the training data until the training end condition is met, to obtain the parameter prediction model, comprises: inputting the training data into the feature fusion module to obtain first fusion features output by the feature fusion module, the feature fusion module comprising a first attention network, a second attention network, a first expert network, a second expert network, a third expert network, and a gating network; the first expert network, the second expert network, and the third expert network are network layers that perform different transformation processes on input feature vectors; each expert network can have different effects on different tasks; the gating network is used to control the weights of each expert network and to combine the results of multiple expert networks by weighting; inputting the first fusion features into the click rate prediction module to obtain a click score output by the click rate prediction module, the click score representing a click probability of the user on the target object, the click rate prediction module comprising multiple expert networks, a gating network, and a click rate tower network; the inputting of the first fusion features into the click rate prediction module to obtain the click score output by the click rate prediction module comprises: connecting the preset click rate vector and the first fusion features to obtain first connection features; inputting the first connection features into the multiple expert networks respectively to obtain multiple first reference connection features output by the multiple expert networks; inputting the scene information and the multiple first reference connection features into the gating network to obtain second fusion features output by the gating network; inputting the second fusion features into the click rate tower network to obtain the click score output by the click rate tower network; inputting the first fusion features into the conversion rate prediction module to obtain a conversion score output by the conversion rate prediction module, the conversion score representing a further behavior conversion probability of the user after clicking the target object; based on the click score and the conversion score, obtaining a recommendation score representing a recommended degree corresponding to the target object; training the to-be-trained model based on the recommendation score and the label until the training end condition is met, to obtain the parameter prediction model.
2. The method of claim 1, wherein, The method comprises: The first training data is input into the first attention network, and first attention features output by the first attention network are obtained; The second training data is input into the second attention network, and second attention features output by the second attention network are obtained; The first attention features are input into the first expert network and the second expert network respectively, and first reference attention features output by the first expert network and second reference attention features output by the second expert network are obtained respectively; The second attention features are input into the third expert network, and third reference attention features output by the third expert network are obtained; The first reference attention features, the second reference attention features and the third reference attention features are input into the gating network, and first fusion features output by the gating network are obtained.
3. The method of claim 1, wherein, The training data further comprises a preset conversion rate vector and scene information corresponding to the target object; the conversion rate prediction module comprises a plurality of expert networks, a gating network and a conversion rate tower network; The first fusion features are input into the conversion rate prediction module, and a conversion score output by the conversion rate prediction module is obtained, comprising: The preset conversion rate vector and the first fusion features are connected to obtain second connection features; The second connection features are input into the plurality of expert networks respectively, and a plurality of second reference connection features output by the plurality of expert networks are obtained; The scene information and the plurality of second reference connection features are input into the gating network, and third fusion features output by the gating network are obtained; The third fusion features are input into the conversion rate tower network, and a conversion score output by the conversion rate tower network is obtained.
4. A target object selection method characterized by comprising: The method comprises: Obtaining target user features and candidate target object features; Inputting the target user features and the candidate target object features into a parameter prediction model to obtain a recommendation score corresponding to the candidate target object output by the parameter prediction model, the parameter prediction model being obtained based on any one of the methods of claims 1-3; Determining a target object corresponding to the target user based on the recommendation score.
5. A model training apparatus characterized by comprising: The device comprises: A data acquisition unit configured to acquire training data, the training data comprising first training data and second training data, the first training data comprising user features of a user, target object features corresponding to the user features, and scene features corresponding to the target object features, the second training data comprising key user features of the user, the user features comprising the key user features; the training data further comprising labels corresponding to the target object; the training data further comprising a preset click rate vector and scene information corresponding to the target object; The training unit is configured to train a to-be-trained model based on the training data until a training end condition is met, to obtain a parameter prediction model, the to-be-trained model comprising a feature fusion module, a click rate prediction module, and a conversion rate prediction module; the training of the to-be-trained model based on the training data until the training end condition is met, to obtain the parameter prediction model, comprises: inputting the training data into the feature fusion module, to obtain first fused features output by the feature fusion module, the feature fusion module comprising a first attention network, a second attention network, a first expert network, a second expert network, a third expert network, and a gating network; the first expert network, the second expert network, and the third expert network are network layers that perform different transformation processes on input feature vectors; each expert network can have different effects on different tasks; the gating network is configured to control the weights of each expert network, and to combine the results of multiple expert networks by weighting; inputting the first fused features into the click rate prediction module, to obtain a click score output by the click rate prediction module, the click score representing a probability of a click on a target object by the user, the click rate prediction module comprising multiple expert networks, a gating network, and a click rate tower network; the inputting of the first fused features into the click rate prediction module, to obtain the click score output by the click rate prediction module, comprises: connecting the preset click rate vector and the first fused features to obtain first connected features; inputting the first connected features into the multiple expert networks respectively, to obtain multiple first reference connected features output by the multiple expert networks; inputting the scene information and the multiple first reference connected features into the gating network, to obtain second fused features output by the gating network; inputting the second fused features into the click rate tower network, to obtain the click score output by the click rate tower network; inputting the first fused features into the conversion rate prediction module, to obtain a conversion score output by the conversion rate prediction module, the conversion score representing a probability of further behavior conversion by the user after clicking on the target object; based on the click score and the conversion score, obtaining a recommendation score representing a recommended degree corresponding to the target object; based on the recommendation score and the label, training the to-be-trained model until the training end condition is met, to obtain the parameter prediction model.
6. A target object selection apparatus characterized by comprising: The apparatus comprises: a feature acquisition unit configured to acquire target user features and candidate target object features; a score determination unit configured to input the target user features and the candidate target object features into a parameter prediction model, to obtain a recommendation score corresponding to the candidate target object output by the parameter prediction model, the parameter prediction model being obtained based on any one of the methods of claims 1-4; a target object determination unit configured to determine a target object corresponding to the target user based on the recommendation score.
7. An electronic device, comprising: A computer program product comprising a computer readable storage medium having program code stored therein, wherein the program code, when executed by a processor, causes the performance of the method of any one of claims 1-3.
8. A computer-readable storage medium, characterized in that, A computer program product comprising a computer readable storage medium having program code stored therein, wherein the program code, when executed by a processor, causes the performance of the method of any one of claims 1-3.
Citation Information
Patent Citations
Method for recommending target object, computing device and computer storage medium
CN111523044A
Method and device for constructing federated learning model by multiple parties
CN112396189A
Model generation method and device, recommendation method and device and electronic equipment
CN115048575A