Method, device and system for training information recommendation model and information recommendation method and device
By acquiring users' historical behavior data and recommendation attribute data, the display order of information recommendation models is optimized, solving the problem of user preferences not being considered and achieving more reasonable information sorting and improved user experience.
Patent Information
- Application Number
- CN202011254218.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-11
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2040-11-11
AI Technical Summary
In existing technologies, the sorting of information recommendation lists fails to fully consider user preferences, resulting in unreasonable sorting of recommended information.
By acquiring users' historical behavior data, determining behavioral feature data, and inputting it into the information recommendation model, the display order is optimized by combining the recommendation attribute data of candidate objects, and the model is trained using a reward function to ensure that the recommendation effect matches user preferences.
The information sorting in the information recommendation list has been optimized, improving the user experience while meeting the cost and benefit requirements of the target audience, thus achieving more reasonable information recommendations.
Smart Images

Figure CN112418920B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a training method for an information recommendation model, an information recommendation method, and an apparatus. Background Technology
[0002] Currently, business platforms can recommend various types of advertisements to users to improve their life experience.
[0003] In practical applications, business platforms display a list of recommended information containing advertisements and various recommendations to users. This means ads and recommendations are presented in a mixed manner, with the position of ads typically determined by preset weight parameters. This mixed-format approach fails to consider the varying degrees of ad preference among users, such as the impact of ad quantity or ad placement on the order of recommendations, leading to unreasonable ranking of recommended information.
[0004] Therefore, how to optimize the sorting of recommended information more reasonably is an urgent problem to be solved. Summary of the Invention
[0005] This specification provides a training method, apparatus, storage medium, and electronic device for an information recommendation model, in order to partially solve the aforementioned problems existing in the prior art.
[0006] The following technical solution is adopted in this specification:
[0007] This manual provides a training method for an information recommendation model, including:
[0008] Obtain historical behavioral data of each user when browsing a recommended information list containing a defined object;
[0009] For each user, behavioral characteristic data is determined based on the historical behavior data.
[0010] The behavioral feature data and the recommendation attribute data corresponding to each candidate setting object are input into the information recommendation model to be trained to obtain each setting object to be recommended and the display order of each setting object to be recommended.
[0011] The objects to be recommended and the information to be recommended, arranged in the display order, are input into a preset information recommendation simulation system to perform simulated recommendation, so as to determine the recommendation effect characterization value of each object to be recommended under the display order.
[0012] Based on the recommendation performance characterization value, the reward value of the reward function corresponding to the information recommendation model is determined, and the information recommendation model is trained based on the reward value. The information recommendation model is used to recommend specific objects to users.
[0013] Optionally, the historical behavior data includes at least one of the following: the number of set objects that the user actually clicked and viewed among the set objects displayed in the information recommendation list, the number of set objects that the user did not click and view among the set objects displayed in the information recommendation list, and the average list length of the displayed portion of each information recommendation list when the user browses each information recommendation list.
[0014] Optionally, based on the historical behavior data, the user's behavioral characteristic data is determined, specifically including:
[0015] The user's churn rate for the information recommendation list is determined based on the number of target objects actually clicked and viewed by the user and the number of target objects not clicked and viewed by the user. The higher the number of target objects not clicked and viewed by the user in the information recommendation list, the greater the churn rate.
[0016] Optionally, based on the historical behavior data, the user's behavioral characteristic data is determined, specifically including:
[0017] The user's browsing depth is determined based on the average length of the displayed portion of each information recommendation list as the user browses them.
[0018] Optionally, the information recommendation model includes: a weight function;
[0019] The behavioral feature data and the recommendation attribute data corresponding to each candidate setting object are input into the information recommendation model to be trained to obtain each setting object to be recommended, specifically including:
[0020] The user's behavioral characteristic data is input into a preset weighting function to determine the weight parameters corresponding to the user;
[0021] Based on the weight parameters corresponding to the user and the recommendation attribute data corresponding to each candidate setting object, determine the setting objects to be recommended output by the information recommendation model to be trained.
[0022] Optionally, the information recommendation model is trained based on the reward value, specifically including:
[0023] With the goal of maximizing the reward value, the information recommendation model is trained by adjusting the model parameters contained in the information recommendation model and the weight function.
[0024] Optionally, based on the recommendation performance representation value, the reward value of the reward function corresponding to the information recommendation model is determined, specifically including:
[0025] Based on the recommendation effect characterization value, a first influence factor and a second influence factor are determined. The first influence factor is used to characterize the impact of a user's personal benefit on the set object, and the second influence factor is used to control the impact of the number of set objects recommended to the user on the user's personal benefit and the recommendation effect of the set object.
[0026] Based on the recommendation performance characterization value, the first influence factor, and the second influence factor, the reward value of the reward function corresponding to the information recommendation model is determined.
[0027] This specification provides an information recommendation method, including:
[0028] Obtain historical behavioral data of users when browsing information lists;
[0029] Based on the historical behavior data, determine the behavioral characteristic data corresponding to the user;
[0030] Obtain historical behavioral data of users when browsing information recommendation lists containing defined objects;
[0031] Based on the historical behavior data, determine the behavioral characteristic data corresponding to the user;
[0032] The behavioral feature data and the recommendation attribute data corresponding to each candidate setting object are input into a pre-trained information recommendation model to obtain each setting object to be recommended and the display order of each setting object to be recommended. The information recommendation model is trained by the above method.
[0033] According to the display order, each of the objects to be recommended is recommended to the user in a mixed manner within each piece of information to be recommended.
[0034] Optionally, the behavioral feature data and the recommendation attribute data corresponding to each candidate setting object are input into a pre-trained information recommendation model to obtain each setting object to be recommended and the order of the setting objects to be recommended, specifically including:
[0035] The behavioral feature data is input into a preset weight function to determine the weight function corresponding to the user. The weight parameters determined by the weight function are not exactly the same for different users.
[0036] Based on the weight parameters corresponding to the user and the recommendation attribute data corresponding to each candidate setting object, the information recommendation model outputs each setting object to be recommended and the order of the setting objects to be recommended.
[0037] This specification provides an apparatus for training an information recommendation model, comprising:
[0038] The acquisition module is used to obtain historical behavioral data of each user when browsing information recommendation lists containing defined objects;
[0039] The determination module is used to determine the behavioral characteristic data of each user based on the historical behavior data.
[0040] The recommendation module is used to input the behavioral feature data and the recommendation attribute data corresponding to each candidate setting object into the information recommendation model to be trained, so as to obtain each setting object to be recommended and the display order of each setting object to be recommended.
[0041] The simulation module is used to input the objects to be recommended and the information to be recommended, arranged in the display order, into a preset information recommendation simulation system to perform simulated recommendation, so as to determine the recommendation effect characterization value of each object to be recommended under the display order.
[0042] The training module is used to determine the reward value of the reward function corresponding to the information recommendation model based on the recommendation effect representation value, and to train the information recommendation model based on the reward value. The information recommendation model is used to recommend specific objects to users.
[0043] This specification provides an information recommendation device, including:
[0044] The acquisition module is used to acquire historical behavioral data of users when browsing information recommendation lists containing defined objects.
[0045] The determination module is used to determine the behavioral characteristic data corresponding to the user based on the historical behavior data;
[0046] The recommendation module is used to input the behavioral feature data and the recommendation attribute data corresponding to each candidate setting object into a pre-trained information recommendation model to obtain each setting object to be recommended and the display order of each setting object to be recommended. The information recommendation model is trained by the above method.
[0047] The push module is used to recommend each of the objects to be recommended to the user in a mixed manner within each piece of information to be recommended, according to the display order.
[0048] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the training method for the information recommendation model or the information recommendation method described above.
[0049] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements either the training method for the information recommendation model described above or the information recommendation method described above.
[0050] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:
[0051] In the training method of the information recommendation model provided in this specification, historical behavior data of each user browsing information recommendations containing specified objects is obtained, and behavioral feature data of each user is determined based on the historical behavior data. Then, the behavioral feature data and the recommendation attribute data corresponding to each candidate specified object are input into the information recommendation model to be trained to obtain each specified object to be recommended and the display order of each specified object. The specified objects to be recommended and the information to be recommended, arranged according to the display order, are input into a preset information recommendation simulation system for simulated recommendation to determine the recommendation effect representation value corresponding to each specified object. Finally, based on the recommendation effect representation value, the reward value of the reward function corresponding to the information recommendation model is determined, and the information recommendation model is trained based on the reward value. The information recommendation model is used to recommend specified objects to users.
[0052] As can be seen from the above method, this method can determine the user's preference for each set object in the information recommendation list based on the user's historical behavior data, and then train the information recommendation model. In other words, the display position of each set object in the information recommendation list is determined in conjunction with the user's preferences. Compared to existing technologies that determine the display position of each set object solely through fixed weight parameters, this avoids unreasonable information sorting in the information recommendation list, thereby optimizing the information sorting. Attached Figure Description
[0053] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0054] Figure 1 A flowchart illustrating the training method of the information recommendation model provided in the embodiments of this specification;
[0055] Figure 2A flowchart illustrating the information recommendation method provided in the embodiments of this specification;
[0056] Figure 3 This is a flowchart illustrating the information recommendation model training and recommendation information process provided in the embodiments of this specification.
[0057] Figure 4 This is a schematic diagram of the information recommendation model training device provided in the embodiments of this specification;
[0058] Figure 5 This is a schematic diagram of the information recommendation device structure provided in the embodiments of this specification;
[0059] Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this specification. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0061] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0062] In the embodiments of this specification, before making information recommendations based on users' historical behavior data, a pre-trained information recommendation model is required. Therefore, the process of training the information recommendation model will be described below. Figure 1 As shown.
[0063] Figure 1 The flowchart illustrating the training method for the information recommendation model provided in the embodiments of this specification specifically includes the following steps:
[0064] S100: Obtain historical behavioral data of each user when browsing information lists.
[0065] When users browse information recommendation lists containing specific objects, various historical behavioral data will be generated. This historical behavioral data can be used to analyze the behavioral characteristics of users browsing information recommendation lists. Based on this, the business platform can obtain the historical behavioral data of each user when browsing information recommendation lists containing specific objects. The specific objects mentioned in this specification can include advertisements, coupon redemption information, etc. For ease of explanation, the following description mainly focuses on the case where the specific object is an advertisement, illustrating the method provided in this specification.
[0066] The acquired historical behavior data may include at least one of the following: the number of specified objects that the user actually clicked to view in each specified object displayed in the information recommendation list (such as the number of ads that the user actually clicked to view in each advertisement displayed in the information recommendation list), the number of specified objects that the user did not click to view in each specified object displayed in the information recommendation list (such as the number of ads that the user did not click to view in each advertisement displayed in the information recommendation list), and the average list length of the displayed portion of each information recommendation list when the user browses each information recommendation list.
[0067] Specifically, the number of settings objects actually clicked and viewed by the user in the information recommendation list, and the number of settings objects not clicked and viewed by the user in the information recommendation list, can be calculated separately. Alternatively, the business platform can determine the other by obtaining one of the values. That is, if the number of settings objects actually clicked and viewed by the user in the information recommendation list is obtained, the number of settings objects not clicked and viewed by the user can be directly determined. Similarly, if the number of settings objects not clicked and viewed by the user in the information recommendation list is obtained, the number of settings objects actually clicked and viewed by the user can be directly determined.
[0068] In the embodiments of this specification, the information recommendation list displays multiple recommended information and designated objects. However, in reality, the user's terminal device interface often only displays a portion of the recommended information and designated objects in the information recommendation list at any given time. The remaining recommended information and designated objects require the user to perform actions such as swiping on the terminal device interface to display them. Based on this, the concept of the length of the displayed portion of the list can be introduced in the embodiments of this specification. That is, the length of the information recommendation list displayed by the terminal device when the user performs a specified operation (such as swiping) on the terminal device interface. For example, assuming a two-column information recommendation list (i.e., two information display positions in each row), if the user performs a swiping operation on the terminal device interface, causing the terminal device to successively display 40 information display positions in the information recommendation list, then the length of the displayed portion of the list is 40.
[0069] Furthermore, the business platform can calculate the length of the displayed portion of each information recommendation list when a user browses it historically, and thus determine the average length of the displayed portion of each information recommendation list when the user browses it historically.
[0070] In the embodiments of this specification, the terminal device used by the user to browse information can be a terminal device such as a mobile phone or a tablet computer. Of course, the execution entity used to obtain the information recommendation list can also be a client, application (App) installed on the terminal device, or a browser in the terminal device or client.
[0071] In addition to various defined objects, the information recommendation list also includes other information recommended to the user by the business platform. For example, search results returned to the user based on their input keywords, or links to nearby businesses proactively recommended based on the user's geographical location. Therefore, it is important to emphasize that the information recommendation list mentioned in this embodiment is obtained by arranging the various recommended information and defined objects from the business platform in a mixed-layout manner.
[0072] S102: For each user, determine the user's behavioral characteristic data based on the historical behavior data.
[0073] In the embodiments of this specification, each user has its corresponding historical behavior data. Therefore, the business platform can determine the user's behavioral characteristic data based on the historical behavior data. The user's behavioral characteristic data is used to represent some preference characteristics reflected by the user when browsing the information recommendation list for the set object. The user's behavioral characteristic data may include the set object churn rate, information browsing depth, user information click preferences, etc.
[0074] The churn rate of a user's selected profile can be determined by the number of profiles actually clicked on and viewed by the user among the profiles displayed in the information recommendation list, as well as the number of profiles not clicked on and viewed by the user among the profiles displayed in the information recommendation list. The higher the number of profiles not clicked on and viewed by the user among the profiles displayed in the information recommendation list, the higher the churn rate of the selected profile.
[0075] It should be noted that there are multiple ways to determine the churn rate of a target audience by measuring the number of target audiences actually clicked and viewed by the user in the recommended information list, and the number of target audiences that the user did not click and view. For example, suppose the business platform has historically pushed two recommended information lists to a user. Based on the number of ads (target audiences) actually clicked and viewed by the user in the recommended information lists, and the number of ads that the user did not click and view, it is determined that both recommended information lists were viewed by the user. In this case, the number of recommended information lists that the user did not view is 0, the number of recommended information lists that the user viewed is 2, and the ad churn rate (i.e., the target audience churn rate) is 0 / 2 = 0. That is to say, the user viewed all the recommended information lists pushed by the business platform, so it is considered that all the ads contained in the recommended information lists had the opportunity to be clicked and viewed by the user, and the ad churn rate is zero.
[0076] For example, suppose the business platform has pushed 10 information recommendation lists to users in the past. These 10 information recommendation lists contain a total of 100 advertisements (i.e., the target audience). Of these 100 advertisements, the user clicked and viewed 90 advertisements and did not click and view 10 advertisements. Therefore, the ad churn rate can be determined to be: 10 / 100 = 10%.
[0077] In the embodiments of this specification, the business platform can determine the user's information browsing depth, that is, the number of pieces of information that the user can view on average when browsing each information recommendation list, based on the average list length of the displayed portion of the information recommendation list when the user browses each information recommendation list.
[0078] Regarding a user's information click preferences, the business platform can determine the user's information click preferences based on the position of the information clicked by the user in the information recommendation list (including the set object and other information recommended by the business platform), the duration of the user's browsing of the clicked information, and the information category of the clicked information.
[0079] For example, by analyzing users' historical browsing history of various information recommendation lists, the business platform can determine users' information click preferences, such as whether the information clicked by the user is at the top of the list, which information category the user spends the longest time browsing, and which information category the user clicks the most.
[0080] S104: Input the behavioral feature data and the recommendation attribute data corresponding to each candidate setting object into the information recommendation model to be trained to obtain each setting object to be recommended and the display order of each setting object to be recommended.
[0081] In the embodiments of this specification, the business platform can input the determined behavioral feature data and the recommendation attribute data corresponding to each candidate setting object into the information recommendation model to be trained, to obtain each setting object to be recommended and the display order of each setting object to be recommended. The recommendation attribute data corresponding to each candidate setting object is used to characterize information such as the cost of the setting object itself and the revenue that the business platform can obtain from that setting object. For example, when the setting object is an advertisement, it may specifically include revenue per thousand impressions (CPM) (also representing the cost of advertising), gross merchandise volume (GMV), and takerate.
[0082] Based on this, the business platform can input the recommended attribute data of a specified object as its features into the information recommendation model to be trained, in order to obtain each specified object to be recommended and its display order. Each specified object to be recommended is selected from a large number of candidate specified objects. Specifically, this selection can be done by determining the ranking score of the candidate specified objects. For example, the following formula can be used to determine the ranking score of the candidate specified objects:
[0083] RankScore = f(f k1 (user_feature), f k2 (user_feature), CPM, GMV, Takerate)
[0084] Among them, f k1 f k2 It is the weight function in the information recommendation model. RankScore is used to represent the ranking score of the candidate set object, and user_feature is the user's behavioral feature data.
[0085] As can be seen from this formula, the business platform can first input user behavior feature data into a preset weight function to determine the corresponding weight parameters for that user. Then, based on the user's corresponding weight parameters and the recommendation attribute data (CPM, GMV, Takerate) of the candidate settings, the platform determines the ranking score of each candidate setting. Subsequently, based on these ranking scores, the platform selects the settings to be recommended from these candidate settings and finally outputs the settings to be recommended and their display order.
[0086] As can be seen from the above process, in practical applications, the behavioral characteristic data of different users are likely to be different. Therefore, the weight parameters determined by inputting the behavioral characteristic data of different users into the preset weight function will also be different. Consequently, the recommended settings and the display order of the recommended settings for different users are also likely to be different. In other words, the weight parameters for each user are not constant and will change depending on the optimization of the weight function or the change in the behavioral characteristic data.
[0087] It is worth emphasizing that the behavioral feature data input to the above weighting functions can be the same. For example, the formula above actually contains two weighting functions, and the behavioral feature data input to these two weighting functions is the same. However, although the behavioral feature data input to these two weighting functions can be the same, the weight parameters output by these two weighting functions can be different. In other words, these two weighting functions have different emphases in the algorithm design, which leads to the possibility that the output weight parameters will be different when the same behavioral feature data is input.
[0088] The aforementioned display order can be represented by the display position of each recommended setting object. This method not only reflects the order in which each recommended setting object is arranged in the information recommendation list subsequently shown to the user, but also reflects its actual position in the information recommendation list. For example, assuming that the display position of recommended setting object A is 2 and the display position of recommended setting object B is 8, it can be seen that recommended setting object A is displayed before recommended setting object B. Between recommended setting object A and recommended setting object B, the information recommendation list also displays 5 other pieces of information.
[0089] S106: The objects to be recommended and the information to be recommended, arranged in the display order, are input into a preset information recommendation simulation system for simulated recommendation, so as to determine the recommendation effect characterization value of each object to be recommended under the display order.
[0090] In the embodiments of this specification, the business platform can input each target object and each piece of information to be recommended into a preset information recommendation simulation system for simulated recommendation, according to the above-described display order. The simulated recommendation includes delivery simulation and revenue simulation. Delivery simulation is used to simulate the delivery of each target object and each piece of information to be recommended, in order to determine the user's browsing behavior towards each target object and each target object (e.g., by simulating the display position of the target object clicked by the user when it is displayed in this order, the number of target objects exposed, etc.). Revenue simulation determines the revenue generated by these target objects for itself and the business platform based on the simulated user browsing behavior towards each target information and each target object.
[0091] Specifically, the business platform can determine the recommendation effect representation value corresponding to each target to be recommended based on revenue simulation. The recommendation effect representation value is used to represent the revenue generated by each target to be recommended for itself and the business platform. For example, if the target is an advertisement, it can be represented by CPM, GMV, Takerate, etc.
[0092] S108: Based on the recommendation effect representation value, determine the reward value of the reward function corresponding to the information recommendation model, and train the information recommendation model based on the reward value. The information recommendation model is used to recommend specific objects to users.
[0093] During the training process of the information recommendation model, the business platform needs to consider not only the user's behavioral characteristics to recommend specific objects to the user, but also the cost of the specific objects themselves and the benefits generated by the business platform in recommending the specific objects. This can effectively ensure that the specific objects recommended by the information recommendation model can not only meet the user's personal preferences to a certain extent, but also guarantee the benefits of the specific objects themselves and the benefits of the business platform in recommending the specific objects.
[0094] Based on this, in the embodiments of this specification, the business platform can determine the reward value of the reward function corresponding to the information recommendation model based on the determined recommendation effect representation value, and train the information recommendation model according to the reward value. The reward function used can refer to the following formula:
[0095]
[0096] k1 to k6 are preset parameters, bias gmv K gmv POW gmv K res POW res These are preset parameters;
[0097] Δgmv is used to represent the difference in the total transaction amount of the set object when recommending different set objects to the user and when recommending set objects to the user in different display orders;
[0098] Δfee is used to represent the difference in commission generated when recommending different set objects to users and when recommending set objects to users in different display orders;
[0099] Δcpm is used to represent the difference in revenue obtained by the business platform from recommending different target objects to users and recommending target objects to users in different display orders;
[0100] Δres is used to represent the difference in the number of objects set each time an object is recommended;
[0101] This is used to represent the impact of user preferences on the revenue of a specified object (such as the impact of user preferences on advertising revenue). Because different users have different preferences, some users accept more specified objects, resulting in higher revenue for the business platform based on the specified objects, while some users accept fewer specified objects, resulting in lower revenue for the business platform based on the specified objects.
[0102] Furthermore, in the above formula, The first influencing factor is used to characterize the impact of an individual user's benefit to a specific target of the business platform, while This is the second influencing factor used to control the impact of the number of target users on individual users and the recommendation effect of target users. The recommendation effect of target users mentioned here can be measured in two aspects: one is the benefit the business platform gains from recommending target users, and the other is the cost of the target users themselves. In other words, through... From a business perspective, the number of objects recommended to users should be kept within a reasonable range. Too many objects may lead to a poor information browsing experience for users, while too few objects may affect the cost of the objects themselves and the revenue that the business platform can obtain from recommending objects.
[0103] It should be noted that the above-mentioned first and second impact factors are not unique in their form, and can also be expressed by other formulas, which will not be explained in detail here.
[0104] In the embodiments described in this specification, the business platform can take maximizing the reward value of the reward function as the optimization objective, and train the information recommendation model by adjusting and optimizing the model parameters contained in the aforementioned weight function. That is, through multiple rounds of iterative training, the reward value of the reward function can continuously increase and converge within a certain numerical range, thereby completing the training process of the information recommendation model.
[0105] Of course, in addition to training the information recommendation model with the maximum reward value as the optimization objective, the information recommendation model can also be trained with a preset reward value as the optimization objective by adjusting the model parameters contained in the information recommendation model and the weight function. In other words, during multiple rounds of iterative training, the reward value needs to be made to continuously approach the preset reward value. After multiple rounds of iterative training, when the reward value fluctuates around the preset reward value, it can be determined that the training of the information recommendation model is complete.
[0106] As can be seen from the above process, because the model training process not only considers user behavioral feature data, but also recommendation attribute data that can represent the cost and benefit of the selected object to a certain extent, the selected objects recommended by the trained information recommendation model to the user can not only meet the user's personal preferences, but also meet the cost and benefit of the selected object itself and the business platform in recommending the selected object. This not only optimizes the information sorting of the information recommendation list and avoids unreasonable sorting of recommended information, bringing a good user experience to the information browsing, but also effectively guarantees the benefits of the selected object provider and the business platform.
[0107] After the information recommendation model is trained, the embodiments in this specification can recommend information to users through the information recommendation model, such as... Figure 2 As shown.
[0108] Figure 2 This is a flowchart illustrating the information recommendation method provided in the embodiments of this specification.
[0109] S200: Obtain historical behavioral data of users when browsing information recommendation lists containing defined objects.
[0110] S202: Determine the behavioral characteristic data corresponding to the user based on the historical behavior data.
[0111] S204: Input the behavioral feature data and the recommendation attribute data corresponding to each candidate setting object into the pre-trained information recommendation model to obtain each setting object to be recommended and the display order of each setting object to be recommended.
[0112] S206: In accordance with the display order, recommend each of the objects to be recommended to the user in a mixed manner within each piece of information to be recommended.
[0113] In the embodiments of this specification, the business platform can obtain historical behavior data of users when the information recommendation list contains the specified objects. Based on the user's historical behavior data, it determines the user's corresponding behavioral feature data, and inputs the behavioral feature data and the recommendation attribute data corresponding to each candidate specified object into a pre-trained information recommendation model to obtain each specified object to be recommended and the display order of each specified object to be recommended. Then, according to the display order, each specified object to be recommended and each piece of information to be recommended can be recommended to the user in a mixed arrangement. The content involved in S200 to S204 is basically the same as the model training stage described above, and will not be described in detail here. The difference between S206 and the model training described above is that in the model training process, each specified object to be recommended is input into the information recommendation simulation system, while here each specified object to be recommended is recommended to the actual user.
[0114] It is important to emphasize that the weight parameters determined by the information recommendation model for the same user at different times are not exactly the same. This is because the historical behavior data generated each time a user browses information changes compared to previous historical behavior data. Therefore, the behavioral feature data of the user determined by the historical behavior data at different times are not exactly the same. By inputting the behavioral feature data at different times into the preset weight function, the weight parameters corresponding to the user will change.
[0115] Of course, the business platform can also record the determined weight parameters for each user. This way, when a user browses information again, the business platform doesn't need to re-input the user's behavioral feature data into the preset weight function to determine the weight parameters. Instead, it can call the previously determined weight parameters for that user and then, based on those weight parameters and the recommendation attribute data for each candidate setting object, determine the recommended setting objects output by the information recommendation model. This avoids calculating the user's weight parameters every time, saving network resources.
[0116] In the training process of the information recommendation model described in this specification, the business platform can use the information recommendation model to recommend each identified target object to the user, and then train the information recommendation model based on the user's behavioral data regarding these target objects. Figure 3 As shown.
[0117] Figure 3 This is a schematic diagram illustrating the process of training an information recommendation model and recommending information, as provided in an embodiment of this specification.
[0118] In the embodiments of this specification, the information recommendation model is divided into the use of an online information recommendation model and the training of an offline information recommendation model. The information recommendation model to be trained, the information recommendation simulation system, and the reward function constitute the training part of the offline information recommendation model. The business platform can determine the behavioral feature data of user A based on the historical behavioral data of user A when browsing an information recommendation list containing a set object. The behavioral feature data and the recommendation attribute data corresponding to each candidate set object are input into the information recommendation model to be trained to obtain each set object to be recommended and the display order of each set object to be recommended. Then, according to the display order, each set object to be recommended and each piece of information to be recommended are input into the information recommendation simulation system for simulated recommendation to determine the recommendation effect representation value corresponding to each set object to be recommended. Based on the recommendation effect representation value, the reward value of the reward function corresponding to the information recommendation model to be trained is determined. Based on the reward value, the information recommendation model to be trained is trained. The trained offline information recommendation model is used to update the online information recommendation model.
[0119] Then, the business platform can use the updated information recommendation model to recommend a list of information containing the specified objects to the user. After the user browses the information recommendation list, the user's historical behavior data will change. The business platform can then train the offline information recommendation model based on the user's changed historical behavior data.
[0120] The above describes one or more embodiments of the information recommendation model training method provided in this specification. Based on the same idea, this specification also provides a corresponding information recommendation model training apparatus, such as... Figure 4 As shown.
[0121] Figure 4 This is a schematic diagram of the information recommendation model training device provided in the embodiments of this specification, specifically including:
[0122] The acquisition module 400 is used to obtain historical behavior data of each user when browsing information recommendation lists containing set objects;
[0123] The determination module 402 is used to determine the behavioral characteristic data of each user based on the historical behavior data.
[0124] Recommendation module 404 is used to input the behavioral feature data and the recommendation attribute data corresponding to each candidate setting object into the information recommendation model to be trained, so as to obtain each setting object to be recommended and the display order of each setting object to be recommended.
[0125] The simulation module 406 is used to input the objects to be recommended and the information to be recommended, arranged in the display order, into a preset information recommendation simulation system to perform simulated recommendation, so as to determine the recommendation effect characterization value of each object to be recommended under the display order.
[0126] The training module 408 is used to determine the reward value of the reward function corresponding to the information recommendation model based on the recommendation effect representation value, and to train the information recommendation model based on the reward value. The information recommendation model is used to recommend specific objects to users.
[0127] Optionally, the acquisition module 400 is specifically used to acquire the historical behavior data, including at least one of the following: the number of set objects actually clicked and viewed by the user among the set objects displayed in the information recommendation list, the number of set objects not clicked and viewed by the user among the set objects displayed in the information recommendation list, and the average list length of the displayed portion of each information recommendation list when the user browses each information recommendation list.
[0128] Optionally, the determining module 402 is specifically used to determine the user's churn rate for the set objects in the information recommendation list based on the number of set objects actually clicked and viewed by the user and the number of set objects not clicked and viewed by the user in the information recommendation list. The higher the number of set objects not clicked and viewed by the user in the information recommendation list, the greater the churn rate for the set objects.
[0129] Optionally, the determining module 402 is specifically used to determine the user's information browsing depth based on the average list length of the displayed portion of each information recommendation list when the user browses each information recommendation list.
[0130] Optionally, the recommendation module 404 is specifically used to input the user's behavioral feature data into a preset weight function to determine the weight parameters corresponding to the user, and determine the recommended settings objects output by the information recommendation model to be trained based on the weight parameters corresponding to the user and the recommendation attribute data corresponding to each candidate setting object.
[0131] Optionally, the training module 408 is specifically used to train the information recommendation model by adjusting the model parameters contained in the information recommendation model and the weight function, with the goal of maximizing the reward value.
[0132] Optionally, the training module 408 is specifically used to determine a first influence factor and a second influence factor based on the recommendation effect characterization value. The first influence factor is used to characterize the impact of a user's personal benefit on a set object, and the second influence factor is used to control the impact of the number of set objects recommended to the user on the recommendation effect of the user and the set objects. Based on the recommendation effect characterization value, the first influence factor, and the second influence factor, the reward value of the reward function corresponding to the information recommendation model is determined.
[0133] Figure 5 The schematic diagram of the information recommendation device provided in the embodiments of this specification specifically includes:
[0134] The acquisition module 500 is used to acquire historical behavioral data of users when browsing information recommendation lists containing defined objects.
[0135] The determining module 502 is used to determine the behavioral characteristic data corresponding to the user based on the historical behavior data;
[0136] The recommendation module 504 is used to input the behavioral feature data and the recommendation attribute data corresponding to each candidate setting object into a pre-trained information recommendation model to obtain each setting object to be recommended and the display order of each setting object to be recommended. The information recommendation model is trained by the above-mentioned information recommendation model training method.
[0137] The push module 506 is used to recommend each of the objects to be recommended to the user in a mixed manner in each of the information to be recommended, according to the display order.
[0138] Optionally, the recommendation module 504 is specifically used to input the behavioral feature data into a preset weight function to determine the weight function corresponding to the user. The weight parameters determined by the weight function for different users are not exactly the same. Based on the weight parameters corresponding to the user and the recommendation attribute data corresponding to each candidate set object, the recommendation model outputs each set object to be recommended and the order of the set objects to be recommended.
[0139] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 The training method for the information recommendation model provided or Figure 2 Information recommendation methods are provided.
[0140] This instruction manual also provides Figure 6 The diagram shows the structure of the electronic device. Figure 6At the hardware level, the training equipment for this information recommendation model includes a processor, internal bus, network interface, memory, and non-volatile storage, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile storage into memory and then runs it to achieve the above. Figure 1 The training method for the information recommendation model described herein is the same as the information recommendation method described above. Of course, besides software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution entity of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0141] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0142] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0143] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0144] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0145] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0146] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0147] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0148] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0149] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0150] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0151] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0152] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0153] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0154] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0155] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0156] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A training method for an information recommendation model, characterized in that, The method comprises the following steps: acquiring historical behavior data of each user when browsing information recommendation lists containing setting objects; for each user, determining behavior feature data of the user according to the historical behavior data; inputting the behavior feature data and recommendation attribute data corresponding to each candidate setting object into an information recommendation model to be trained to obtain each setting object to be recommended and a display order of the setting objects to be recommended; arranging the setting objects to be recommended and each recommended information in the display order and inputting them into a preset information recommendation simulation system to simulate recommendation, so as to determine a recommendation effect representation value corresponding to each setting object to be recommended in the display order; determining a reward value of a reward function corresponding to the information recommendation model according to the recommendation effect representation value, and training the information recommendation model according to the reward value, wherein the information recommendation model is used for recommending setting objects to users; inputting the behavior feature data and recommendation attribute data corresponding to each candidate setting object into an information recommendation model to be trained to obtain each setting object to be recommended and a display order of the setting objects to be recommended, specifically comprising: inputting the behavior feature data into a preset weight function to determine a weight function corresponding to the user, wherein the weight parameters determined by different users through the weight function are not completely the same; determining each setting object to be recommended and a display order of the setting objects to be recommended output by the information recommendation model according to the weight parameter corresponding to the user and the recommendation attribute data corresponding to each candidate setting object; determining a reward value of a reward function corresponding to the information recommendation model according to the recommendation effect representation value, specifically comprising: determining a first influence factor and a second influence factor according to the recommendation effect representation value, wherein the first influence factor is used to represent the influence of the personal benefit of a user on a setting object, and the second influence factor is used to control the influence of the number of setting objects recommended to the user on the personal benefit of the user and the recommendation effect of the setting object; determining a reward value of a reward function corresponding to the information recommendation model according to the recommendation effect representation value, the first influence factor and the second influence factor.
2. The method of claim 1, wherein, The historical behavior data comprises at least one of the number of setting objects actually clicked and browsed by a user among setting objects displayed on an information recommendation list, the number of setting objects not clicked and browsed by the user among the setting objects displayed on the information recommendation list, and the average list length of the displayed part of each information recommendation list when the user browses the information recommendation list.
3. The method of claim 2, wherein, The behavior feature data of the user is determined according to the historical behavior data, specifically comprising: determining a setting object loss rate of the user for the information recommendation list according to the number of setting objects actually clicked and browsed by the user among the setting objects displayed on the information recommendation list and the number of setting objects not clicked and browsed by the user among the setting objects displayed on the information recommendation list, wherein the higher the number of setting objects not clicked and browsed by the user among the setting objects displayed on the information recommendation list, the greater the setting object loss rate.
4. The method of claim 2, wherein, According to the historical behavior data, behavior feature data of the user is determined, specifically including: According to the average list length of the displayed part of each information recommendation list when the user browses each information recommendation list, information browsing depth of the user is determined.
5. The method of claim 1, wherein, The information recommendation model includes a weight function; The behavior feature data and the recommendation attribute data corresponding to each candidate setting object are input into the information recommendation model to be trained to obtain each to-be-recommended setting object, specifically including: The behavior feature data of the user is input into the preset weight function to determine the weight parameter corresponding to the user; According to the weight parameter corresponding to the user and the recommendation attribute data corresponding to each candidate setting object, each to-be-recommended setting object output by the information recommendation model to be trained is determined.
6. The method of claim 5, wherein, According to the reward value, the information recommendation model is trained, specifically including: Taking the maximum reward value as an optimization target, the information recommendation model is trained by adjusting the model parameters contained in the information recommendation model and the weight function.
7. An information recommendation method characterized by comprising: including: Obtain historical behavior data of a user when the user browses an information recommendation list containing a setting object; According to the historical behavior data, behavior feature data corresponding to the user is determined; The behavior feature data and the recommendation attribute data corresponding to each candidate setting object are input into the information recommendation model to be trained to obtain each to-be-recommended setting object and the display order of the to-be-recommended setting objects, the information recommendation model being trained by the method of the above claims 1-6; According to the display order, the to-be-recommended setting objects are recommended to the user in a mixed arrangement manner in each to-be-recommended information.
8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-7.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the method of any one of claims 1-7.
Citation Information
Patent Citations
Content recommendation method and device, storage medium and computer equipment
CN110263244A