A training method for a prediction model, an information recommendation method and apparatus
By adjusting the predicted click-through rate of the training sample using click-through rate compensation value in the CTR prediction model, the problem of display position affecting prediction accuracy is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202111294007.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-03
AI Technical Summary
The existing CTR prediction model is affected by the display position of the recommended information during training, resulting in low accuracy of the prediction results.
By obtaining training samples containing user feature information, historical recommendation information and display environment information, input them into the prediction model to determine the predicted click-through rate, and compensate the predicted click-through rate based on the click-through rate compensation value of the display position. Finally, the model is trained based on the compensated click-through rate and actual browsing situation.
The impact of display position on the CTR prediction model is eliminated, and the prediction accuracy of the prediction model is improved.
Smart Images

Figure CN114117202B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a prediction model training method, information recommendation method, and device. Background Art
[0002] With the development of network technology, the amount of online information has shown explosive growth. Therefore, to enable users to more quickly and easily access the information they are interested in, when recommending information to users, business platforms typically use a click-through-rate (CTR) prediction model for each user. Based on the user's preferences and the information to be recommended, they predict the predicted click-through rate (CTR) of the recommended information. Information is then pushed to the user based on the predicted CTR.
[0003] When training a CTR prediction model, information historically recommended to users (i.e., information that users have already viewed) is typically used as training samples, with whether or not the user clicked on it being used as the training label. Regardless of whether the recommended information matches the user's preferences, the probability of a user clicking on the recommended information will still decrease as the recommended information's display position moves backward. Thus, the training samples of the CTR prediction model may contain recommended information whose content does not match the user's preferences. However, because the recommended information is displayed to the user in a prominent position, the user repeatedly clicks on the recommended information, reinforcing the CTR prediction model's learning of the recommended information. This results in the CTR prediction model's prediction results being less accurate.
[0004] Based on this, how to eliminate the influence of the display position of recommended information on the prediction results of the CTR prediction model in order to improve the prediction accuracy of the CTR prediction model is an urgent problem to be solved. Summary of the Invention
[0005] This specification provides a prediction model training method, information recommendation method and device to partially solve the above-mentioned problems existing in the prior art.
[0006] This manual adopts the following technical solutions:
[0007] This specification provides a prediction model training method, including:
[0008] Obtaining a training sample, wherein the training sample includes historical recommendation information recommended to a user, characteristic information of the user, and display environment information corresponding to when the historical recommendation information is browsed by the user, wherein the display environment information includes a display position on a web page when the historical recommendation information is browsed by the user;
[0009] Inputting the training samples into a prediction model so that the prediction model determines a predicted click-through rate for the historical recommendation information based on the historical recommendation information and the characteristic information of the user, and determines a click-through rate compensation value corresponding to the display position based on the display environment information, wherein the click-through rate compensation value is used to represent the probability that the recommendation information displayed at the display position is browsed by the user;
[0010] The predicted click rate of the historical recommendation information is compensated according to the click rate compensation value to obtain a compensated click rate, and the prediction model is trained according to the compensated click rate and the actual browsing situation corresponding to the historical recommendation information.
[0011] Optionally, the prediction model includes a prediction network and a position compensation network;
[0012] Inputting the training sample into a prediction model so that the prediction model determines a predicted click rate for the historical recommendation information based on the historical recommendation information and the characteristic information of the user, and determines a click rate compensation value corresponding to the display position based on the display environment information, specifically including:
[0013] Inputting the user's characteristic information and the historical recommendation information into the prediction network to obtain a predicted click rate for the historical recommendation information;
[0014] The display environment information is input into the position compensation network to obtain a click rate compensation value corresponding to the display position.
[0015] Optionally, the training sample includes a first training sample, the historical recommendation information in the first training sample is first historical recommendation information randomly recommended to the user in history, and the position compensation network includes a data processing subnetwork and a first position compensation subnetwork;
[0016] Inputting the display environment information into the position compensation network to obtain a click rate compensation value corresponding to the display position specifically includes:
[0017] The display environment information in the first training sample is input into the data processing subnetwork to obtain the display environment feature corresponding to the display position of the first historical recommendation information in the page when the user browses it, as the first display environment feature, and the first display environment feature is input into the first position compensation subnetwork to obtain the click-through rate compensation value corresponding to the display position of the first historical recommendation information in the page when the user browses it when information is recommended to the user in a random manner, as the first click-through rate compensation value.
[0018] Optionally, the training sample includes a second training sample, the historical recommendation information in the second training sample is second historical recommendation information recommended to the user based on a ranking result obtained by ranking according to predicted click-through rates in history, and the position compensation network further includes a second position compensation subnetwork;
[0019] Inputting the display environment information into the position compensation network to obtain a click rate compensation value corresponding to the display position specifically includes:
[0020] The display environment information in the second training sample is input into the data processing subnetwork to obtain the display environment feature corresponding to the display position of the second historical recommendation information in the page when the user browses it, as the second display environment feature, and the second display environment feature is input into the second position compensation subnetwork to obtain the click-through rate compensation value corresponding to the display position of the second historical recommendation information in the page when the user browses it, when recommending information to the user according to the sorting result obtained after sorting according to the predicted click-through rate, as the second click-through rate compensation value.
[0021] Optionally, the training samples include positive samples and negative samples, wherein the historical recommendation information in the positive samples is the recommendation information clicked by the user when the user browses the historical recommendation information, and the historical recommendation information in the negative samples is the training samples that the historical recommendation information is not clicked by the user when the user browses the historical recommendation information;
[0022] Compensating the predicted click rate of the historical recommendation information according to the click rate compensation value to obtain a compensated click rate, and training the prediction model according to the compensated click rate and actual browsing conditions corresponding to the historical recommendation information, specifically including:
[0023] For each positive sample in the first training samples, compensating the predicted click rate of the positive sample according to the first click rate compensation value corresponding to the display position in the positive sample, and determining a first compensated click rate of the first historical recommendation information in the positive sample;
[0024] The prediction network and the position compensation network are jointly trained according to the first compensated click-through rate corresponding to each positive sample and the predicted click-through rate corresponding to each negative sample.
[0025] Optionally, before training the prediction model based on the compensated click-through rate of the historical recommendation information and the actual browsing situation of the historical recommendation information, the method further includes:
[0026] Inputting the second display environment feature into the first position compensation subnetwork to obtain, as a third click rate compensation value, a click rate compensation value corresponding to the display position of the second historical recommendation information on a page when the user browses the page in a random manner, and determining a compensation difference between the third click rate compensation value and the second click rate compensation value;
[0027] Compensating the predicted click rate of the historical recommendation information according to the click rate compensation value to obtain a compensated click rate, and training the prediction model according to the compensated click rate and actual browsing conditions corresponding to the historical recommendation information, specifically including:
[0028] For each positive sample in the second training sample, compensating the predicted click rate of the positive sample according to the second click rate compensation value corresponding to the display position in the positive sample, and determining a second compensated click rate of the second historical recommendation information in the positive sample;
[0029] The prediction network and the position compensation network are jointly trained to minimize the first compensated click rate, the second compensated click rate, the predicted click rate corresponding to each negative sample, and the compensation difference.
[0030] Optionally, the display environment information includes: type information of the client based on which the user browses the historical recommendation information, device information of the device based on which the user browses the historical recommendation information, business scenario information corresponding to the page displaying the historical recommendation information, and at least one of the geographical location of the user when the user browses the historical recommendation information.
[0031] This specification provides a method for information recommendation, including:
[0032] Determining each candidate information to be recommended to the user, and obtaining characteristic information of the user;
[0033] For each candidate information, input the candidate information and the user's characteristic information into the prediction model, so that the prediction model determines a predicted click-through rate for the candidate information based on the candidate information and the user's characteristic information, wherein the prediction model is trained using the above-mentioned training method;
[0034] According to the predicted click rate of each candidate information, candidate information recommended to the user is selected from each candidate information as target information, and the target information is recommended to the user.
[0035] This specification provides a prediction model training device, including:
[0036] an acquisition module, configured to acquire training samples, wherein the training samples include historical recommendation information recommended to a user, characteristic information of the user, and display environment information corresponding to when the historical recommendation information was browsed by the user, wherein the display environment information includes a display position on a web page when the historical recommendation information was browsed by the user;
[0037] a prediction module, configured to input the training samples into a prediction model, so that the prediction model determines a predicted click-through rate for the historical recommendation information based on the historical recommendation information and the characteristic information of the user, and determines a click-through rate compensation value corresponding to the display position based on the display environment information, wherein the click-through rate compensation value is used to represent the probability that the recommendation information displayed at the display position will be browsed by the user;
[0038] The training module is used to compensate the predicted click rate of the historical recommendation information according to the click rate compensation value to obtain the compensated click rate, and train the prediction model according to the compensated click rate and the actual browsing situation corresponding to the historical recommendation information.
[0039] This specification provides an information recommendation device, including:
[0040] A candidate information determination module is used to determine each candidate information that needs to be recommended to the user and obtain characteristic information of the user;
[0041] a prediction module configured to input, for each candidate information, the candidate information and the user's characteristic information into the prediction model, so that the prediction model determines a predicted click-through rate for the candidate information based on the candidate information and the user's characteristic information, wherein the prediction model is trained using the above-described training method;
[0042] The recommendation module is used to select candidate information recommended to the user from each candidate information according to the predicted click rate of each candidate information as target information, and recommend the target information to the user.
[0043] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned prediction model training method or information recommendation method.
[0044] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned prediction model training method or information recommendation method is implemented.
[0045] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0046] In the prediction model training method and information display method provided in this specification, a training sample is obtained, which includes historical recommendation information recommended to the user, the user's characteristic information, and the display environment information corresponding to the historical recommendation information when the user browsed it. The display environment information includes the display position of the historical recommendation information on the page when the user browsed it. Then, the training sample is input into the prediction model so that the prediction model determines the predicted click-through rate for the historical recommendation information based on the historical recommendation information and the user's characteristic information, and determines the click-through rate compensation value corresponding to the display position based on the display environment information. The click-through rate compensation value is used to characterize the probability that the recommended information displayed at the display position is browsed by the user. Finally, the predicted click-through rate of the historical recommendation information is compensated according to the click-through rate compensation value to obtain the compensated click-through rate, and the prediction model is trained based on the compensated click-through rate and the actual browsing situation corresponding to the historical recommendation information. When in use, the candidate information that needs to be recommended to the user is determined, and the user's characteristic information is obtained. Then, for each candidate information, the candidate information and the user's characteristic information are input into the prediction model, so that the prediction model determines the predicted click-through rate for the candidate information based on the candidate information and the user's characteristic information. Finally, based on the predicted click-through rate of each candidate information, the candidate information recommended to the user is selected from each candidate information as the target information, and the target information is recommended to the user.
[0047] It can be seen from the above method that this method determines the click-through rate compensation value corresponding to the display position of the historical recommendation information, and then uses the click-through rate compensation value to compensate the predicted click-through rate corresponding to the historical recommendation information, and trains the prediction model based on the compensated predicted click-through rate and the user's actual browsing situation of the historical recommendation information. In this way, a corresponding click-through rate compensation value can be trained for each display position, and the click-through rate compensation value can be used to multiply the predicted click-through rate of the information displayed at the front by a lower weight coefficient, and at the same time, the predicted click-through rate of the information displayed at the back by a higher weight coefficient, so as to eliminate the influence of the display position on the prediction accuracy of the prediction model, thereby improving the prediction accuracy of the prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0049] Figure 1 This is a schematic diagram of the prediction model training provided in this manual;
[0050] Figure 2A schematic diagram of a prediction model training method provided in this specification;
[0051] Figure 3 A schematic diagram of a recommended method for providing information in this specification;
[0052] Figure 4 A schematic diagram of a training device for a prediction model provided in this specification;
[0053] Figure 5 A schematic diagram of a device recommended for providing information in this manual;
[0054] Figure 6 The corresponding Figure 1 or Figure 2 Schematic diagram of electronic equipment. DETAILED DESCRIPTION
[0055] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0056] In order to eliminate the influence of display position on the prediction accuracy of the click rate prediction model, this specification provides a prediction model training method and an information recommendation method. In this method, see Figure 1 The prediction model includes a prediction network and a position compensation network, and the position compensation network is composed of a data processing network, a first position compensation sub-network and a second position compensation sub-network.
[0057] The prediction network is used to determine the predicted click-through rate (CTR) for the historical recommendation information based on the historical recommendation information and user profile information. The position compensation network is used to determine the CTR compensation value corresponding to the display position of the historical recommendation information based on the display position of the historical recommendation information on the page when the historical recommendation information was viewed by the user. The CTR compensation value corresponding to the display position can reflect the probability that the recommendation information displayed at that display position will be viewed by the user.
[0058] When the historical recommendation information is the recommendation information recommended by the server to the user in a random manner, the data processing network and the first position compensation sub-network in the position compensation network are used to determine the first click rate compensation value corresponding to the display position of the first historical recommendation information based on the corresponding display environment information when the historical recommendation information recommended to the user is browsed by the user in a random manner.
[0059] When the historical recommendation information is the recommendation information recommended to the user by the server according to the sorting result obtained after sorting according to the predicted click-through rate, the data processing network and the second position compensation sub-network in the position compensation network are used to determine the second click-through rate compensation value corresponding to the display position of the first historical recommendation information based on the display environment information corresponding to the historical recommendation information recommended to the user when the information is browsed by the user.
[0060] When training a prediction model, historical recommendation information recommended to a user and the user's characteristic information can be input into a prediction network to obtain a predicted click-through rate for the historical recommendation information. Simultaneously, the display environment information corresponding to when the historical recommendation information was browsed by the user is input into a position compensation network to obtain a click-through rate compensation value corresponding to the display position. The predicted click-through rate of the historical recommendation information is then compensated based on the click-through rate compensation value to obtain a compensated click-through rate. The prediction model is then trained based on the compensated click-through rate and the actual browsing situation corresponding to the historical recommendation information (i.e., whether the user clicked on the historical recommendation information). Thus, for historically displayed historical recommendation information, if a user clicks on the historical recommendation information, the click-through rate compensation value is lower for a display position that is easily browsed by the user (display position arranged at the front), and higher for a display position that is less likely to be viewed by the user (display position arranged at the back). This reduces the impact of display position on the predicted click-through rate and improves the accuracy of the trained prediction model.
[0061] The following will describe in detail the prediction model training scheme and information recommendation scheme provided in this specification in conjunction with embodiments.
[0062] Figure 2 The following is a flowchart of a method for training a prediction model in this specification, which specifically includes the following steps:
[0063] Step S200, obtaining a training sample, wherein the training sample includes historical recommendation information recommended to the user, the user's characteristic information, and the corresponding display environment information when the historical recommendation information is browsed by the user, wherein the display environment information includes the display position of the historical recommendation information on the page when the user browses it.
[0064] The execution entity of the prediction model training method and the information recommendation method in this specification can be a server or a business platform, or a terminal device such as a desktop computer. For the sake of convenience of description, the following will only use the server as the execution entity to illustrate the prediction model training method and information recommendation method provided in this specification.
[0065] Prediction models can be applied to a variety of businesses. For example, on a news information platform, users can predict the click-through rate of recommended news information or advertisements, and based on the predicted click-through rate, recommend news information or advertisements to users. Another example is on an online shopping platform, users can predict the click-through rate of recommended products (merchants, theme events, or other user reviews), and based on the predicted click-through rate, recommend products (merchants, theme events, or other user reviews) to users. I will not list other examples one by one.
[0066] Specifically, the above-mentioned characteristic information of the user includes but is not limited to the user's historical business records and the user's attribute information. The historical business records include the business records of various businesses that the user has performed (such as the user's historical click records, the user's historical order records, the user's browsing records in the last half hour, etc.). Attribute information can refer to information that can reflect the basic personal characteristics of the user. For example, the user's attribute information may include: the user's identity document (ID), the user's age, the user's gender, the user's city, the user's place of origin, etc., which will not be explained in detail here.
[0067] Step S202: input the training sample into the prediction model so that the prediction model determines the predicted click-through rate for the historical recommendation information based on the historical recommendation information and the characteristic information of the user, and determines the click-through rate compensation value corresponding to the display position based on the display environment information. The click-through rate compensation value is used to represent the probability that the recommended information displayed at the display position is browsed by the user.
[0068] In the specific implementation, the server inputs the user's feature information and historical recommendation information into the prediction network to obtain the predicted click rate for the historical recommendation information. At the same time, the display environment information is input into the position compensation network to obtain the click rate compensation value corresponding to the display position.
[0069] In actual applications, the training samples in this specification may include a first training sample and a second training sample. The historical recommendation information in the first training sample is the first historical recommendation information randomly recommended to the user in history, and the historical recommendation information in the second training sample is the second historical recommendation information recommended to the user based on the sorting result obtained after sorting according to the predicted click-through rate in history.
[0070] When determining the click-through rate compensation value of each training sample, if the training sample is the first training sample, the server inputs the display environment information in the first training sample into the data processing subnetwork to obtain the display environment feature corresponding to the display position of the first historical recommendation information on the page when the user browses it, as the first display environment feature, and then inputs the first display environment feature into the first position compensation subnetwork to obtain the click-through rate compensation value corresponding to the display position of the first historical recommendation information on the page when the user browses it when the information is recommended to the user in a random manner, as the first click-through rate compensation value.
[0071] If the training sample is the second training sample, the server inputs the display environment information in the second training sample into the data processing subnetwork to obtain the display environment characteristics corresponding to the display position of the second historical recommendation information in the page when the user browses it, as the second display environment characteristics, and then inputs the second display environment characteristics into the second position compensation subnetwork to obtain the click rate compensation value corresponding to the display position of the second historical recommendation information in the page when the user browses it, when the information is recommended to the user according to the sorting result obtained after sorting according to the predicted click rate, as the second click rate compensation value.
[0072] Among them, the display environment information used to determine the click-through rate compensation value corresponding to the above-mentioned display position includes, in addition to the display position, the type information of the client based on which the user browses the historical recommendation information (such as APP, mini program, web page, etc.), the device information of the device based on which the user browses the historical recommendation information (such as Android model, ISO model, etc.), the business scenario information corresponding to the page displaying the historical recommendation information (such as takeout, fresh food, etc.), and at least one of the geographical location of the user when the user browses the historical recommendation information (such as the user's company address, the user's residential address, etc.).
[0073] Because the page layout displayed to users when they log in and browse through an APP, mini-program, or web page is often different, and the position and number of display locations for information displayed to users in different page layouts will also vary. Therefore, the type of client information based on which users browse historical recommendation information is an influencing factor of the display location, which in turn changes the impact of the display location on the prediction model. Similarly, the device information of the device based on which users browse historical recommendation information, the business scenario information corresponding to the page displaying historical recommendation information, etc. may all lead to different page layouts, resulting in inconsistent display locations of information displayed on the page, which in turn leads to different effects of the display location on the prediction accuracy of the prediction model. Therefore,
[0074] Regarding the geographical location of users when browsing historical recommended information, users at home often have ample time to browse information slowly, while at work, time is often less available. In this case, users may only view a small amount of information before making a decision or stopping browsing. As a result, the number of display positions of the information seen by the user will vary. Users at work are more likely to click on information that is displayed at the top.
[0075] Different users have different habits. Some tend to browse quickly, viewing two or three pages before making a decision, while others prefer to browse a large number of pages, compare multiple options, and finally make a decision. Thus, user habits also affect the number of display positions viewed by the user. Therefore, in addition to the above-mentioned display environment information, the user ID can also be used as an influencing factor in determining the click-through rate compensation value of the display position.
[0076] To sum up, the client type information based on which the user browses the historical recommendation information, the device information based on which the user browses the historical recommendation information, the business scenario information corresponding to the page displaying the historical recommendation information, the user's ID, and at least one of the geographical location of the user when the user browses the historical recommendation information can be used as input and output items of the location compensation network to determine the click-through rate compensation value corresponding to the display location, and further provide the prediction accuracy of the trained prediction model.
[0077] Step S204 , compensating the predicted click rate of the historical recommendation information according to the click rate compensation value to obtain a compensated click rate, and training the prediction model according to the compensated click rate and the actual browsing situation corresponding to the historical recommendation information.
[0078] In a specific implementation, the server can compensate the predicted click rate of each positive sample in the first training sample according to the first click rate compensation value corresponding to the display position in the positive sample, determine the first compensated click rate of the first historical recommendation information in the positive sample, and then jointly train the prediction network and the position compensation network according to the first compensated click rate corresponding to each positive sample and the predicted click rate corresponding to each negative sample.
[0079] Specifically, when training the prediction model, the loss function of the prediction model can be expressed as:
[0080] loss=loss 随机 =-(label 随机 ×log(logit) 随机 ×ipw 随机 +(1-label 随机 )×(1-log(logit) 随机 ));
[0081] Among them, label 随机 Indicates the label of the positive sample under the condition of recommending information to users in a random manner;
[0082] log(logit) 随机 It indicates the predicted click-through rate of users clicking on the historical recommended information under the condition that information is recommended to users in a random manner;
[0083] ipw 随机 Indicates the click rate compensation value corresponding to the display position of the historical recommended information on the page when the user browses it, when the information is recommended to the user in a random manner;
[0084] 1-label 随机 Indicates the label of the negative sample when recommending information to users in a random manner;
[0085] 1-log(logit) 随机 It indicates the predicted click-through rate if the user does not click on the historical recommended information when information is recommended to the user in a random manner.
[0086] Among them, the training samples disclosed above are all collected under the condition of recommending information to users in a random manner. At this time, the prediction network in the prediction model, the data processing network in the position click rate compensation network, and the first position click rate compensation network can be directly trained jointly through the above loss function to obtain the trained prediction model.
[0087] In actual business operations, training a prediction model using the above method requires a large number of training samples collected under the condition of randomly recommending information to users, which greatly impairs the user experience and affects the effectiveness of recommendations. Therefore, this specification introduces training samples collected under the condition of recommending information to users based on the ranking results obtained after sorting according to predicted click-through rates. Furthermore, during training, the second position click-through rate prediction network used under the condition of recommending information to users based on the ranking results obtained after sorting according to predicted click-through rates is trained using the training samples collected under the condition of randomly recommending information to users as a reference.
[0088] In a specific implementation, the server can also compensate the predicted click rate of each positive sample in the second training sample according to the second click rate compensation value corresponding to the display position in the positive sample, determine the second compensated click rate of the second historical recommendation information in the positive sample, and jointly train the prediction network and the position compensation network to minimize the first compensated click rate, the second compensated click rate and the compensation difference.
[0089] In addition, the server can also input the second display environment feature into the first position compensation sub-network to obtain a click-through rate compensation value corresponding to the display position of the second historical recommendation information in the page when the user browses the information in a random manner, as a third click-through rate compensation value, and determine the compensation difference between the third click-through rate compensation value and the second click-through rate compensation value, and then jointly train the prediction network and the position compensation network to minimize the first compensation click-through rate, the second compensation click-through rate and the compensation difference.
[0090] Specifically, when training the prediction model, the loss function of the prediction model can be expressed as:
[0091] loss=loss 随机 +loss 排序 +loss 补偿 =loss 随机 -(label 排序 ×log(logit) 排序 ×ipw 排序 +(1-label 排序 )×(1-log(logit) 排序 )+loss 补偿 );
[0092] Among them, loss 随机 represents the loss function corresponding to the first training sample under the condition of recommending information to users in a random manner;
[0093] label 随机Indicates the label of the positive sample when recommending information to the user based on the ranking results obtained after the predicted click-through rate;
[0094] log(logit) 随机 Indicates the predicted click-through rate of users clicking on the historical recommended information when recommending information to users based on the ranking results obtained after sorting by predicted click-through rate;
[0095] ipw 随机 Indicates the click rate compensation value corresponding to the display position of the historical recommended information on the page when the user browses it, when recommending information to the user based on the ranking results obtained after the predicted click rate ranking;
[0096] 1-label 随机 Indicates the label of the negative sample when recommending information to the user based on the ranking results obtained after the predicted click-through rate;
[0097] 1-log(logit) 随机 It indicates the predicted click rate of historical recommended information when the user does not click on it when the information is recommended to the user according to the ranking results obtained after the predicted click rate is sorted.
[0098] loss 补偿 Indicates the compensation difference between the third click rate compensation value and the second click rate compensation value.
[0099] Through the above steps, the click-through rate compensation value corresponding to the display position of the historical recommendation information is determined, and then the click-through rate compensation value is used to compensate the predicted click-through rate corresponding to the historical recommendation information. The prediction model is trained based on the compensated predicted click-through rate and the actual browsing situation of the user on the historical recommendation information. In this way, a corresponding click-through rate compensation value can be trained for each display position, and the click-through rate compensation value can be used to multiply the predicted click-through rate of the information displayed at the front by a lower weight coefficient, and at the same time, the predicted click-through rate of the information displayed at the back by a higher weight coefficient, so as to eliminate the influence of the display position on the prediction accuracy of the prediction model and improve the prediction accuracy of the prediction model.
[0100] In addition, this specification also provides a solution for information recommendation using a prediction model trained using the above-mentioned prediction model training method.
[0101] Specifically, such as Figure 3 As shown, a flowchart of a prediction model training method provided in this specification specifically includes the following steps:
[0102] Step S300: determining each candidate information to be recommended to the user, and obtaining characteristic information of the user.
[0103] In a specific implementation, if the prediction model is applied to a news information platform, the candidate information may be news or advertisements that need to be recommended to users. If the prediction model is applied to an online shopping platform, the candidate information may be products, merchants, themed events, or other user reviews that need to be recommended to users. In practice, the candidate information that needs to be recommended to users varies from business to business, and it is difficult to list them all in this specification, so I will not provide examples here.
[0104] The above-mentioned user characteristic information is characteristic information that can characterize the user, and the composition of the user characteristic information is consistent with the composition of the user characteristic information used in the prediction model training.
[0105] Step S302: For each candidate information, input the candidate information and the characteristic information of the user into the prediction model, so that the prediction model determines the predicted click rate for the candidate information based on the candidate information and the characteristic information of the user.
[0106] The prediction model is trained using the aforementioned prediction model training method. In actual applications, the server will input each candidate information and the user's characteristic information into the prediction network, so that the prediction network determines the user's predicted click-through rate for the candidate information based on the candidate information and the user's characteristic information.
[0107] It should be noted that the position compensation network described in this specification is only used in model training. In actual applications, only the prediction network of the prediction model is used to predict the user's predicted click-through rate for each candidate information. Because the impact of the information display position on the accuracy of the prediction model has been largely eliminated during the model training process, the trained prediction network is a click-through rate prediction network that eliminates the influence of display position. Therefore, the trained prediction network can be directly used to predict click-through rate and recommend information to users based on the predicted click-through rate corresponding to each candidate information.
[0108] Step S304 : selecting candidate information recommended to the user from the candidate information according to the predicted click rate of each candidate information as target information, and recommending the target information to the user.
[0109] In a specific implementation, after the server determines the predicted click-through rate corresponding to each candidate information, it can sort the candidate information in descending order according to the size of the predicted click-through rate corresponding to each candidate information, and then select the target information that needs to be recommended to the user from each candidate information based on the sorting result, and then recommend the target information to the user based on the arrangement of each display position in the page that displays the information to the user.
[0110] Furthermore, the server can also sort the candidate information in descending order based on their predicted click-through rates. Based on the sorting results, the server selects candidate information with predicted click-through rates greater than a set click-through rate threshold as target information to be recommended to the user. The server then recommends the target information to the user based on the placement of the information on the page displayed to the user. Of course, there are other methods for selecting target information, which will not be fully illustrated here.
[0111] The above is a prediction model training method and information recommendation method provided in one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding prediction model training device and information recommendation device, such as Figure 4 Or as shown in 5.
[0112] Figure 4 A schematic diagram of a training device for a prediction model provided in this specification, specifically including:
[0113] An acquisition module 400 is configured to acquire training samples, wherein the training samples include historical recommendation information recommended to a user, characteristic information of the user, and display environment information corresponding to when the historical recommendation information was browsed by the user, wherein the display environment information includes the display position of the historical recommendation information on the page when the historical recommendation information was browsed by the user;
[0114] Prediction module 401 is configured to input the training samples into a prediction model, so that the prediction model determines a predicted click-through rate for the historical recommendation information based on the historical recommendation information and the user's characteristic information, and determines a click-through rate compensation value corresponding to the display position based on the display environment information, wherein the click-through rate compensation value is used to represent the probability that the recommendation information displayed at the display position will be viewed by the user;
[0115] The training module 402 is used to compensate the predicted click rate of the historical recommendation information according to the click rate compensation value to obtain the compensated click rate, and train the prediction model according to the compensated click rate and the actual browsing situation corresponding to the historical recommendation information.
[0116] Optionally, the prediction model includes a prediction network and a position compensation network;
[0117] The prediction module 401 is specifically used to input the user's feature information and the historical recommendation information into the prediction network to obtain a predicted click-through rate for the historical recommendation information; and input the display environment information into the position compensation network to obtain a click-through rate compensation value corresponding to the display position.
[0118] Optionally, the training sample includes a first training sample, the historical recommendation information in the first training sample is first historical recommendation information randomly recommended to the user in history, and the position compensation network includes a data processing subnetwork and a first position compensation subnetwork;
[0119] The prediction module 401 is specifically used to input the display environment information in the first training sample into the data processing subnetwork to obtain the display environment feature corresponding to the display position of the first historical recommendation information on the page when the user browses it, as the first display environment feature, and input the first display environment feature into the first position compensation subnetwork to obtain the click-through rate compensation value corresponding to the display position of the first historical recommendation information on the page when the user browses it when information is recommended to the user in a random manner, as the first click-through rate compensation value.
[0120] Optionally, the training sample includes a second training sample, the historical recommendation information in the second training sample is second historical recommendation information recommended to the user based on a ranking result obtained by ranking according to predicted click-through rates in history, and the position compensation network further includes a second position compensation subnetwork;
[0121] The prediction module 401 is specifically used to input the display environment information in the second training sample into the data processing subnetwork to obtain the display environment characteristics corresponding to the display position of the second historical recommendation information on the page when the user browses it, as the second display environment characteristics, and input the second display environment characteristics into the second position compensation subnetwork to obtain the click-through rate compensation value corresponding to the display position of the second historical recommendation information on the page when the user browses it, when recommending information to the user according to the sorting result obtained after sorting according to the predicted click-through rate, as the second click-through rate compensation value.
[0122] Optionally, the training samples include positive samples and negative samples, wherein the historical recommendation information in the positive samples is the recommendation information clicked by the user when the user browses the historical recommendation information, and the historical recommendation information in the negative samples is the training samples that the historical recommendation information is not clicked by the user when the user browses the historical recommendation information;
[0123] The training module 402 is specifically used to compensate the predicted click rate of each positive sample in the first training sample according to the first click rate compensation value corresponding to the display position in the positive sample, and determine the first compensated click rate of the first historical recommendation information in the positive sample; according to the first compensated click rate corresponding to each positive sample and the predicted click rate corresponding to each negative sample, the prediction network and the position compensation network are jointly trained.
[0124] Optionally, the device further comprises:
[0125] The compensation difference determination module 403 is configured to input the second display environment feature into the first position compensation subnetwork to obtain, as a third click-through rate compensation value, a click-through rate compensation value corresponding to the display position of the second historical recommendation information on the page when the user browses the page in a random manner, and determine a compensation difference between the third click-through rate compensation value and the second click-through rate compensation value.
[0126] The training module 402 is specifically used to compensate the predicted click rate of each positive sample in the second training sample according to the second click rate compensation value corresponding to the display position in the positive sample, and determine the second compensated click rate of the second historical recommendation information in the positive sample; in order to minimize the first compensated click rate, the second compensated click rate, the predicted click rates corresponding to each negative sample and the compensation difference, the prediction network and the position compensation network are jointly trained.
[0127] Optionally, the display environment information includes: type information of the client based on which the user browses the historical recommendation information, device information of the device based on which the user browses the historical recommendation information, business scenario information corresponding to the page displaying the historical recommendation information, and at least one of the geographical location of the user when the user browses the historical recommendation information.
[0128] Figure 5 A schematic diagram of a device recommended for providing information in this manual, specifically including:
[0129] The candidate information determination module 500 is used to determine each candidate information to be recommended to the user and obtain the characteristic information of the user;
[0130] Prediction module 501 is configured to input each candidate information and the user's characteristic information into the prediction model, so that the prediction model determines a predicted click-through rate for the candidate information based on the candidate information and the user's characteristic information, wherein the prediction model is trained using the training method described above;
[0131] The recommendation module 502 is configured to select candidate information recommended to the user from each candidate information according to the predicted click rate of each candidate information as target information, and recommend the target information to the user.
[0132] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 The training method of the prediction model provided or the above Figure 2The information provided recommends the method.
[0133] This manual also provides Figure 6 The schematic structure diagram of the electronic device shown in FIG. Figure 6 As mentioned above, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The training method of the prediction model or the above Figure 2 Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0134] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0135] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0136] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0137] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0138] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0139] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0140] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0142] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0143] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0144] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0145] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0146] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0147] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0148] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0149] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A method for training a prediction model, characterized in that, it includes: Obtain training samples, where the training samples include historical recommendation information recommended to users, the user's characteristic information, and display environment information corresponding to when the historical recommendation information is browsed by the user. The display environment information includes the display position of the historical recommendation information on the page when it is browsed by the user; Input the training samples into the prediction model, so that the prediction model determines the predicted click-through rate for the historical recommendation information according to the historical recommendation information and the user's characteristic information, and determines the click-through rate compensation value corresponding to the display position according to the display environment information. The click-through rate compensation value is used to characterize the probability that the recommendation information displayed at the display position is browsed by the user; The step of determining the click-through rate compensation value corresponding to the display position according to the display environment information includes: inputting the display environment information in the training samples into the data processing sub-network to obtain the display environment characteristics corresponding to the display position of the first historical recommendation information when it is browsed by the user, and inputting the first display environment characteristics into the first position compensation sub-network to obtain the click-through rate compensation value; train a corresponding click-through rate compensation value for each display position, and multiply the predicted click-through rate of the information with a higher display position by a lower weight coefficient, and multiply the predicted click-through rate of the information with a lower display position by a higher weight coefficient through the click-through rate compensation value; Compensate the predicted click-through rate of the historical recommendation information according to the click-through rate compensation value to obtain the compensated click-through rate, and train the prediction model according to the compensated click-through rate and the actual browsing situation corresponding to the historical recommendation information.
2. The method according to claim 1, characterized in that, the prediction model includes a prediction network and a position compensation network; Inputting the training samples into the prediction model, so that the prediction model determines the predicted click-through rate for the historical recommendation information according to the historical recommendation information and the user's characteristic information, and determines the click-through rate compensation value corresponding to the display position according to the display environment information, specifically includes: Inputting the user's characteristic information and the historical recommendation information into the prediction network to obtain the predicted click-through rate for the historical recommendation information; Inputting the display environment information into the position compensation network to obtain the click-through rate compensation value corresponding to the display position.
3. The method according to claim 2, characterized in that, the training samples include first training samples, the historical recommendation information in the first training samples is the first historical recommendation information randomly recommended to users in history, and the position compensation network includes a data processing sub-network and a first position compensation sub-network; Inputting the display environment information into the position compensation network to obtain the click-through rate compensation value corresponding to the display position, specifically includes: Input the display environment information in the first training sample into the data processing sub-network to obtain the display environment feature corresponding to the display position of the first historical recommendation information when it is browsed by the user, as the first display environment feature, and input the first display environment feature into the first position compensation sub-network to obtain the click-through rate compensation value corresponding to the display position of the first historical recommendation information when it is browsed by the user in the case of randomly recommending information to the user, as the first click-through rate compensation value.
4. The method according to claim 3, wherein, the training sample includes a second training sample, the historical recommendation information in the second training sample is the second historical recommendation information recommended to the user according to the sorting result obtained by sorting according to the predicted click-through rate in history, and the position compensation network further includes a second position compensation sub-network; Inputting the display environment information into the position compensation network to obtain the click-through rate compensation value corresponding to the display position specifically includes: Input the display environment information in the second training sample into the data processing sub-network to obtain the display environment feature corresponding to the display position of the second historical recommendation information when it is browsed by the user, as the second display environment feature, and input the second display environment feature into the second position compensation sub-network to obtain the click-through rate compensation value corresponding to the display position of the second historical recommendation information when it is browsed by the user in the case of recommending information to the user according to the sorting result obtained by sorting according to the predicted click-through rate, as the second click-through rate compensation value.
5. The method according to claim 4, wherein, the training sample includes positive samples and negative samples, the historical recommendation information in the positive samples is the recommendation information that is clicked by the user when the user browses the historical recommendation information, and the historical recommendation information in the negative samples is the training sample that is not clicked by the user when the user browses the historical recommendation information; Compensating the predicted click-through rate of the historical recommendation information according to the click-through rate compensation value to obtain a compensated click-through rate, and training the prediction model according to the compensated click-through rate and the actual browsing situation corresponding to the historical recommendation information specifically includes: For each positive sample in the first training sample, compensate the predicted click-through rate of the positive sample according to the first click-through rate compensation value corresponding to the display position in the positive sample, and determine the first compensated click-through rate of the first historical recommendation information in the positive sample; Jointly train the prediction network and the position compensation network according to the first compensated click-through rate corresponding to each positive sample and the predicted click-through rate corresponding to each negative sample.
6. The method according to claim 5, wherein, before compensating the predicted click-through rate of the historical recommendation information according to the click-through rate compensation value to obtain a compensated click-through rate, and training the prediction model according to the compensated click-through rate and the actual browsing situation corresponding to the historical recommendation information, it further includes: Input the second display environment feature into the first position compensation sub-network to obtain a click-through rate compensation value corresponding to the display position on the page when the second historical recommendation information is browsed by the user under the condition of randomly recommending information to the user, as the third click-through rate compensation value, and determine the compensation difference between the third click-through rate compensation value and the second click-through rate compensation value; Compensate the predicted click-through rate of the historical recommendation information according to the click-through rate compensation value to obtain a compensated click-through rate, and train the prediction model according to the compensated click-through rate and the actual browsing situation corresponding to the historical recommendation information, which specifically includes: For each positive sample in the second training sample, compensate the predicted click-through rate of the positive sample according to the second click-through rate compensation value corresponding to the display position in the positive sample, and determine the second compensated click-through rate of the second historical recommendation information in the positive sample; Jointly train the prediction network and the position compensation network by minimizing the first compensated click-through rate, the second compensated click-through rate, the predicted click-through rates corresponding to each negative sample, and the compensation difference.
7. The method according to any one of claims 1 to 6, characterized in that, The display environment information includes at least one of: the type information of the client based on which the user browses the historical recommendation information, the device information of the device based on which the user browses the historical recommendation information, the business scenario information corresponding to the page displaying the historical recommendation information, and the geographical location where the user is located when the user browses the historical recommendation information.
8. A method for information recommendation, characterized in that, including: Determine each candidate information to be recommended to the user, and obtain the feature information of the user; For each candidate information, input the candidate information and the feature information of the user into the prediction model, so that the prediction model determines the predicted click-through rate for the candidate information according to the candidate information and the feature information of the user, and the prediction model is trained by the training method according to any one of claims 1 to 7 above; Select the candidate information to be recommended to the user from each candidate information according to the predicted click-through rate of each candidate information as the target information, and recommend the target information to the user.
9. A training device for a prediction model, characterized in that, including: An acquisition module, configured to acquire a training sample, where the training sample includes historical recommendation information recommended to the user, the feature information of the user, and the display environment information corresponding to when the historical recommendation information is browsed by the user, and the display environment information includes the display position of the historical recommendation information on the page when it is browsed by the user; A prediction module, configured to input the training samples into a prediction model, so that the prediction model determines a predicted click-through rate for the historical recommendation information according to the historical recommendation information and the user's feature information, and determines a click-through rate compensation value corresponding to the display position according to the display environment information, where the click-through rate compensation value is used to characterize the probability that the recommendation information displayed at the display position is browsed by the user; The determining the click-through rate compensation value corresponding to the display position according to the display environment information includes: inputting the display environment information in the training samples into a data processing sub-network to obtain a display environment feature corresponding to the display position where the first historical recommendation information is browsed by the user, and inputting the first display environment feature into a first position compensation sub-network to obtain a click-through rate compensation value; training a corresponding click-through rate compensation value for each display position, and multiplying a lower weight coefficient by the predicted click-through rate of the information with a higher display position, and multiplying a higher weight coefficient by the predicted click-through rate of the information with a lower display position; A training module, configured to compensate the predicted click-through rate of the historical recommendation information according to the click-through rate compensation value to obtain a compensated click-through rate, and train the prediction model according to the compensated click-through rate and the actual browsing situation corresponding to the historical recommendation information.
10. An information recommendation device, characterized in that, it includes: A candidate information determination module, configured to determine each candidate information to be recommended to the user, and obtain the user's feature information; A prediction module, configured to input, for each candidate information, the candidate information and the user's feature information into the prediction model, so that the prediction model determines a predicted click-through rate for the candidate information according to the candidate information and the user's feature information, where the prediction model is trained by the training method according to any one of the above claims 1 to 7; A recommendation module, configured to select, from each candidate information, the candidate information to be recommended to the user as the target information according to the predicted click-through rates of each candidate information, and recommend the target information to the user.
11. A computer-readable storage medium, characterized in that, the storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of the above claims 1 to 7 or 8 is implemented.
12. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the program, the method according to any one of the above claims 1 to 7 or 8 is implemented.
Citation Information
Patent Citations
Content recommendation method and device, computer readable storage medium and computer equipment
CN110263242A
Model training and information recommendation method and device
CN112966186A