Method and apparatus for processing media resources, storage medium, and electronic device

By using two training neural networks with the same model structure, the location characteristics, resource characteristics and user characteristics of media resources are processed, and the problem of low accuracy of media resource recommendation in the prior art is solved, and more accurate media resource recommendation is achieved.

CN114722268BActive Publication Date: 2025-06-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110004963.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-04
Publication Date
2025-06-20
Estimated Expiration
2041-01-04

AI Technical Summary

Technical Problem

In the prior art, the accuracy of recommending media resources based on user behavior is low, and the click-through rate prediction is mainly due to position deviation.

Method used

By using two training neural networks with the same model structure, the location characteristics, resource characteristics and user characteristics of the sample media resources are input respectively, the first predicted click rate and the second predicted click rate are obtained, and the target loss function of the second training neural network is adjusted according to the actual click results to eliminate the deviation of the media resource location information.

Benefits of technology

It improves the accuracy of media resource recommendations, can recommend media resources of interest to users more accurately, and reduces the impact of location deviation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114722268B_ABST
    Figure CN114722268B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for processing media resources, a storage medium, and an electronic device. Among them, the method includes: inputting the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into a first training neural network to obtain a first predicted click-through rate output by the first training neural network; inputting the resource feature of the sample media resource and the user feature of the sample user into a second training neural network to obtain a second predicted click-through rate output by the second training neural network; adjusting the model parameters in the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result. By using the adjusted second training neural network, the deviation of the media resource location information can be eliminated, and the purpose of accurately recommending media resources of interest to users can be achieved, thereby solving the technical problem of low accuracy in recommending media resources according to user behavior in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method and apparatus for processing media resources, a storage medium, and an electronic device. Background Art

[0002] In a recommendation scenario, most recommended products are presented to users in the form of a list, such as common e-commerce, video, news recommendations, and so on. For a user request, a sorting model scores and sorts all products in the recall candidate set, and the product set with the highest score is shown to the user. Here, there is a problem of position bias, that is, users are more likely to click on products with higher rankings, and this tendency has nothing to do with the user's true interests. If the click-through rates at different positions are statistically analyzed, it can be found that the products with higher rankings have the highest click-through rates. This phenomenon may cause users to ignore the products they are really interested in and only click on the products with higher rankings. And user behavior is an important feature for constructing the model input. Therefore, the model will also learn the user's behavior with position bias, making it impossible to rank the products that the user is most interested in at the top during scoring.

[0003] Currently, the method for eliminating position bias is a feature-based method. During offline training, position features are added, and a default feature is used to replace the position feature during online prediction. During offline training, the PAL (Position bias) model models the position feature separately and outputs a value through a shallow tower (this value can be understood as the probability of being exposed), and finally multiplies it with the pCTR (predicted click-through rate) part; during online prediction, only the pCTR part is used.

[0004] For the feature-based method, the position feature of the exposure is used during offline training, and a default value feature is used online. Therefore, the offline training auc is high, but the prediction auc is low.

[0005] The PAL model models the position feature separately and finally introduces the position information through a multiplication operation. Only the pCTR part is used during online prediction. Therefore, this model structure can weaken the influence of not having the position feature during online prediction and reduce the gap between the training auc and the prediction auc. However, modeling the position feature separately makes the position feature lack deep cross with other features, affecting the fitting ability of the model and resulting in low accuracy when recommending media resource information according to the PAL model.

[0006] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0007] An embodiment of the present invention provides a method and apparatus for processing media resources, a storage medium, and an electronic device, so as to at least solve the technical problem of low accuracy in recommending media resources according to user behavior in the prior art.

[0008] According to one aspect of the embodiments of the present invention, a method for processing media resources is provided, including: inputting the position feature of a sample media resource, the resource feature of the sample media resource, and the user feature of a sample user into a first training neural network to obtain a first predicted click-through rate output by the first training neural network, where the position feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks the sample media resource; inputting the resource feature of the sample media resource and the user feature of the sample user into a second training neural network to obtain a second predicted click-through rate output by the second training neural network, the second predicted click-through rate is used to represent the predicted probability that the sample user clicks the sample media resource, and the first training neural network and the second training neural network have the same model structure; determining a loss value of an objective loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, where the actual click result is used to represent whether the sample user actually clicks the sample media resource; adjusting model parameters in the second training neural network according to the loss value of the objective loss function.

[0009] According to another aspect of the embodiments of the present invention, there is also provided a processing device for media resources, including: a first input unit configured to input the position feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into a first trained neural network to obtain a first predicted click-through rate output by the first trained neural network, where the position feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks the sample media resource; a second input unit configured to input the resource feature of the sample media resource and the user feature of the sample user into a second trained neural network to obtain a second predicted click-through rate output by the second trained neural network, the second predicted click-through rate is used to represent the predicted probability that the sample user clicks the sample media resource, and the first trained neural network and the second trained neural network have the same model structure; a determination unit configured to determine a loss value of a target loss function of the second trained neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, where the actual click result is used to represent whether the sample user actually clicks the sample media resource; and an adjustment unit configured to adjust model parameters in the second trained neural network according to the loss value of the target loss function.

[0010] According to yet another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the above-mentioned processing method for media resources when running.

[0011] According to yet another aspect of the embodiments of the present invention, there is also provided an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to execute the above-mentioned processing method for media resources through the computer program.

[0012] In an embodiment of the present invention, by inputting the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into the first training neural network, a first predicted click-through rate output by the first training neural network is obtained. Among them, the location feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource; inputting the resource feature of the sample media resource and the user feature of the sample user into the second training neural network, a second predicted click-through rate output by the second training neural network is obtained. The second predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource. The first training neural network and the second training neural network have the same model structure; according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, the loss value of the target loss function of the second training neural network is determined. Among them, the actual click result is used to represent whether the sample user actually clicks on the sample media resource; according to the loss value of the target loss function, the model parameters in the second training neural network are adjusted, achieving the purpose of determining the loss value of the target loss function of the second training neural network according to the first predicted click-through rate output by the first training neural network, the second predicted click-through rate output by the second training neural network, and the actual click result, and adjusting the model parameters in the second training neural network according to the loss value. By using the adjusted second training neural network, the deviation of the media resource location information can be eliminated, and media resources of interest to the user can be accurately recommended to the user. That is to say, when recommending media resources to the user, the influence of the media resource location information can be eliminated, and the media resources of interest to the user can be more accurately recommended to the user, thereby solving the technical problem of low accuracy in recommending media resources according to user behavior in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0014] Figure 1 is a schematic diagram of an application environment of an optional method for processing media resources according to an embodiment of the present invention;

[0015] Figure 2 is a flowchart of an optional method for processing media resources according to an embodiment of the present invention;

[0016] Figure 3 is a schematic diagram of the sorted display of a group of media resources on a page according to an embodiment of the present invention;

[0017] Figure 4It is a schematic structural diagram of an optional first training neural network according to an embodiment of the present invention;

[0018] Figure 5 It is a schematic structural diagram of an optional second training neural network according to an embodiment of the present invention;

[0019] Figure 6 It is a schematic structural diagram of an optional third training neural network according to an embodiment of the present invention;

[0020] Figure 7 It is a flowchart of an optional click-through rate prediction model for eliminating position bias based on knowledge distillation according to an embodiment of the present invention;

[0021] Figure 8 It is a structural diagram of an optional click-through rate prediction model for eliminating position bias based on knowledge distillation according to an embodiment of the present invention;

[0022] Figure 9 It is a schematic structural diagram of an optional processing device for media resources according to an embodiment of the present invention;

[0023] Figure 10 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention. Detailed implementation manners

[0024] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] In order to better understand the embodiments provided in this application, some terms are explained as follows:

[0027] Feed stream recommendation: A type of content recommendation that aggregates information. Through the feed stream, dynamic information can be transmitted to subscribers in real time, which is an effective way for users to obtain information streams.

[0028] CTR: That is, click-through rate. In a recommendation system, usually, the content subsets retrieved are sorted according to the click-through rate, and then the content is distributed in combination with strategies.

[0029] Position bias: In a recommendation system, the attention received by each item is affected by the display position. Items with a higher position are usually more likely to be noticed by users and more likely to be clicked than items with a lower position, which may cause a bias in the model's perception of user preferences and inaccurate estimated CTR.

[0030] Knowledge distillation: Use a student network to learn the knowledge of a teacher network, that is, use the output of the student network to fit the output of the teacher network.

[0031] AUC: Randomly draw a positive and a negative sample from the positive and negative sample sets respectively. The probability that the predicted value of the positive sample is greater than that of the negative sample. It can be used to evaluate the ranking quality offline.

[0032] According to one aspect of the embodiments of the present invention, a method for processing media resources is provided. Optionally, as an alternative implementation, the above method for processing media resources can be but is not limited to being applied to an environment such as Figure 1 shown. Terminal 102, network 104, and server 106.

[0033] Server 106 inputs the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into the first training neural network to obtain the first predicted click-through rate output by the first training neural network. The location feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource. The resource feature of the sample media resource and the user feature of the sample user are input into the second training neural network to obtain the second predicted click-through rate output by the second training neural network. The second predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource. The first training neural network and the second training neural network have the same model structure. According to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, the loss value of the target loss function of the second training neural network is determined, where the actual click result is used to represent whether the sample user actually clicks on the sample media resource. According to the loss value of the target loss function, the model parameters in the second training neural network are adjusted, achieving the purpose of determining the loss value of the target loss function of the second training neural network based on the first predicted click-through rate output by the first training neural network, the second predicted click-through rate output by the second training neural network, and the actual click result, and adjusting the model parameters in the second training neural network according to the loss value. The adjusted second training neural network can eliminate the deviation of the media resource location information and accurately recommend media resources of interest to users. That is to say, when recommending media resources to users, the influence of the media resource location information can be eliminated, and the media resources of interest to users can be more accurately recommended to users, thereby solving the technical problem of low accuracy in recommending media resources according to user behavior in the prior art.

[0034] It should be noted that the above method for processing media resources may include, but is not limited to, being executed by the terminal 102, being executed by the server 106, and being jointly completed by the terminal 103 and the server 106.

[0035] Optionally, in this embodiment, the above terminal 102 may be a terminal device configured with a target client, and may include, but is not limited to, at least one of the following: mobile phone (such as Android mobile phone, iOS mobile phone, etc.), laptop computer, tablet computer, handheld computer, MID (Mobile Internet Devices), PAD, desktop computer, smart TV, etc. The target client may be a video client, instant messaging client, browser client, education client, etc. The above network may include, but is not limited to: wired network, wireless network, where the wired network includes: local area network, metropolitan area network and wide area network, and the wireless network includes: Bluetooth, WIFI and other networks for implementing wireless communication. The above server may be a single server, or a server cluster composed of multiple servers, or a cloud server. The above is only an example, and this embodiment does not make any limitation thereto.

[0036] Optionally, as an alternative implementation, as Figure 2 shown, the above method for processing media resources includes:

[0037] Step S202, input the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into the first training neural network to obtain the first predicted click-through rate output by the first training neural network, where the location feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks the sample media resource.

[0038] Step S204, input the resource feature of the sample media resource and the user feature of the sample user into the second training neural network to obtain the second predicted click-through rate output by the second training neural network, where the second predicted click-through rate is used to represent the predicted probability that the sample user clicks the sample media resource, and the first training neural network and the second training neural network have the same model structure.

[0039] Step S206, determine the loss value of the target loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, where the actual click result is used to represent whether the sample user actually clicks the sample media resource.

[0040] Step S208, adjust the model parameters in the second training neural network according to the loss value of the target loss function.

[0041] Optionally, in this embodiment, the above media resource processing method includes, but is not limited to, being used in recommendation scenarios, such as the recommendation scenario of media resources in "Look at" in WeChat, the recommendation scenario of e-commerce products, the video recommendation scenario, the news recommendation scenario, the recommendation scenario of announcement resources in games, etc.

[0042] Among them, in this embodiment, the position feature information of the above sample media resource represents the actual arrangement position of the sample media resource in a group of media resources pushed to the user. Taking a group of media resources in WeChat Look at as an example, as Figure 3 shown, it is a schematic diagram of the sorted display of a group of media resources on the page. In Figure 3 it, each media resource is in a position when displayed on the page. For example, "Alibaba's Reinforcement Learning Rearrangement Practice" is in the most prominent position on the page, indicating that this media resource has been clicked the most times among the WeChat user's friends, or the WeChat user has clicked on the resource information related to this media resource the most times. Therefore, the "Alibaba's Reinforcement Learning Rearrangement Practice" news is in the most prominent position on the page. That is to say, when a group of media resources are displayed on the page, the position of the media resource in the page can be understood as the position feature information of the media resource, and this position feature information can include, but is not limited to, being represented by coordinate information.

[0043] In this embodiment, the resource feature of the above sample media resource is used to represent the feature of the sample media resource. The resource feature can include, but is not limited to, the type of the sample media resource, such as the sample media resource being financial news, sports news, entertainment news, game resources, etc. The above user feature of the user can represent the attribute information of the user, and the user feature can include, but is not limited to, the user's gender, age, education level, work attribute, game rank, etc.

[0044] In this embodiment, the position feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user are input into the first training neural network to obtain the first predicted click-through rate output by the first training neural network. Among them, the first training neural network can include, but is not limited to, the teacher network in knowledge distillation. The first training neural network can output the first predicted click-through rate, and the first predicted click-through rate represents the predicted probability that the sample user clicks on the sample media resource. For example, it can be predicted that the probability of a user clicking on the sample media resource "Alibaba's Reinforcement Learning Rearrangement Practice" is 0.6. As Figure 4As shown in the figure, it is a schematic structural diagram of the first training neural network. The location information S1 of the sample media resource 1, the user features 1 and 2 of user A, and the resource feature of the sample media resource 1, i.e., the financial category L1, are input into the first training neural network, and the first predicted click-through rate is output. It should be noted that the user features of a user may not be limited to user features 1 and 2, but can be multiple user features in different dimensions. The more features there are, the more accurate the output first predicted click-through rate will be. The resource feature of the sample media resource 1 can be information in multiple dimensions. The more features there are, the more accurate the output first predicted click-through rate will be. The above is only an example, and this embodiment does not make any limitations in this regard.

[0045] In this embodiment, the resource feature of the sample media resource and the user features of the sample user are input into the second training neural network, and the second predicted click-through rate output by the second training neural network is obtained. The second predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource. The first training neural network and the second training neural network have the same model structure. Among them, the second training neural network may include, but is not limited to, the student network in knowledge distillation, i.e., studentnetwork. The second training neural network can output the second predicted click-through rate, which represents the predicted probability that the sample user clicks on the sample media resource. For example, it can be predicted that the probability that user clicks on the sample media resource "Ali Reinforcement Learning Rearrangement Practice" is 0.6. Such as Figure 5 As shown in the figure, it is a schematic structural diagram of the second training neural network. The user features 1 and 2 of user A and the resource feature of the sample media resource 1, i.e., the financial category L1, are input into the second training neural network, and the second predicted click-through rate is output. It should be noted that the user features of a user may not be limited to user features 1 and 2, but can be multiple user features in different dimensions. The more features there are, the more accurate the output second predicted click-through rate will be. The resource feature of the sample media resource 1 can be information in multiple dimensions. The more features there are, the more accurate the output second predicted click-through rate will be. The above is only an example, and this embodiment does not make any limitations in this regard.

[0046] Optionally, in this embodiment, according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, determining the loss value of the target loss function of the second training neural network may include: according to the first predicted click-through rate, the second predicted click-through rate, the predicted exposure rate, and the actual click result, determining the loss value of the target loss function of the second training neural network, where the predicted exposure rate is the exposure rate determined according to the location feature of the sample media resource, and the predicted exposure rate is used to represent the probability that the sample media resource is pushed to the sample user.

[0047] Among them, the location feature of the sample media resource is input into the third training neural network, and the predicted exposure rate output by the third training neural network is obtained.

[0048] As shown Figure 6 in the structural schematic diagram of the third training neural network, the position information S1 of the sample media resource 1 is input into the third training neural network, and the predicted exposure rate output by the third training neural network is obtained. Among them, the predicted exposure rate represents the probability that the sample media resource is pushed to the sample user. For example, the probability that the sample media resource "Ali Reinforcement Learning Rearrangement Practice" is pushed to the sample user is 0.8.

[0049] Optionally, in this embodiment, determining the loss value of the target loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, the predicted exposure rate, and the actual click result may include: determining a first loss value according to the first predicted click-through rate and the second predicted click-through rate; determining a second loss value according to the second predicted click-through rate, the predicted exposure rate, and the actual click result; and determining the loss value of the target loss function of the second training neural network according to the first loss value and the second loss value.

[0050] In this embodiment, the first loss value may be the loss value formed by the second predicted click-through rate fitting the first predicted click-through rate, and the second loss value may be the loss value of the second predicted click-through rate relative to the actual click result.

[0051] Optionally, in this embodiment, determining the loss value of the target loss function of the second training neural network according to the first loss value and the second loss value may include: determining the loss value of the target loss function of the second training neural network through the following formula:

[0052] L = αL (soft) +(1 - α)L (hard)

[0053] where L (soft) represents the first loss value, L (hard) represents the second loss value, L represents the loss value of the target loss function of the second training neural network, α represents a preset weight, and 0 < α < 1.

[0054] Among them, the above α can also take the value of 0 or 1, and usually it can take the value of 0.6. The above is only an example, and this embodiment does not make any limitation in this regard.

[0055] Through the embodiments provided in this application, the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user are input into the first training neural network to obtain the first predicted click-through rate output by the first training neural network, where the location feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource; the resource feature of the sample media resource and the user feature of the sample user are input into the second training neural network to obtain the second predicted click-through rate output by the second training neural network, and the second predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource. The first training neural network and the second training neural network have the same model structure; according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, the loss value of the target loss function of the second training neural network is determined, where the actual click result is used to represent whether the sample user actually clicks on the sample media resource; according to the loss value of the target loss function, the model parameters in the second training neural network are adjusted, achieving the purpose of determining the loss value of the target loss function of the second training neural network according to the first predicted click-through rate output by the first training neural network, the second predicted click-through rate output by the second training neural network, and the actual click result, and adjusting the model parameters in the second training neural network according to the loss value. The adjusted second training neural network can eliminate the bias of the media resource location information and accurately recommend media resources of interest to the user. That is to say, when recommending media resources to the user, the influence of the media resource location information can be eliminated, and the media resources of interest to the user can be recommended to the user more accurately, thereby solving the technical problem of low accuracy in recommending media resources according to user behavior in the prior art.

[0056] Optionally, in this embodiment, the adjusted second training neural network can be used for pushing media resources. That is, the influence of the location feature information of the media resource can be eliminated, a group of media resources are sorted, and the sorted group of media resources are displayed in sequence on the page, as Figure 3 shown, a group of media resources are displayed in sequence.

[0057] Optionally, in this embodiment, determining the first loss value according to the first predicted click-through rate and the second predicted click-through rate may include: determining the square of the difference between the first predicted click-through rate and the second predicted click-through rate as the first loss value; or

[0058] determining the sum or product of the second predicted click-through rate and the predicted exposure rate as the predicted output value; determining the square of the difference between the first predicted click-through rate and the predicted output value as the first loss value.

[0059] Optionally, in this embodiment, determining the second loss value according to the second predicted click-through rate, predicted exposure rate, and actual click result may include: determining the sum or product of the second predicted click-through rate and the predicted exposure rate as the predicted output value; determining the cross-entropy between the predicted output value and the actual click result as the second loss value.

[0060] Optionally, in this embodiment, adjusting the model parameters in the second training neural network according to the loss value of the target loss function may include: adjusting the model parameters in the second training neural network in the adjustment direction that reduces the loss value of the target loss function.

[0061] In this embodiment, the above loss function is used to measure the degree of inconsistency between the predicted value and the true value of the model. It is a non-negative real-valued function. The smaller the loss function, the better the robustness of the model. The process of training the model is continuous iterative calculation. In the process of iterative calculation, the optimization algorithm of gradient descent can be used to make the loss function smaller and smaller. Among them, gradient descent is an optimization algorithm that makes the loss function smaller and smaller, and can solve the model parameters of machine learning algorithms without solving, that is, the constrained optimization problem.

[0062] Optionally, in this embodiment, the above method may further include: obtaining the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user, where the resource feature of the sample media resource includes resource features in one or more dimensions, and the user feature of the sample user includes user features in one or more dimensions.

[0063] Among them, obtaining the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user may include:

[0064] Obtaining the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user, where the resource feature of the sample media resource includes the type of the sample media resource, and the user feature of the sample user includes the age of the sample user and the gender of the sample user.

[0065] Optionally, determining the loss value of the target loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result may include:

[0066] S1, inputting the first predicted click-through rate output by the first training neural network, the second predicted click-through rate output by the second training neural network, and the actual click result into the target fully connected layer module;

[0067] S2, determining the first loss value between the first predicted click-through rate and the second predicted click-through rate through the target fully connected layer module;

[0068] S3. Determine a second loss value between the second predicted click-through rate and the actual click result through the target fully connected layer module;

[0069] S4. Determine the loss value of the target loss function through the target fully connected layer module according to the first loss value and the second loss value.

[0070] In the process of prediction by the neural network, it can be divided into the training process of the neural network and the use of the trained neural network.

[0071] In this embodiment, during the training of the neural network, the fully connected layer can determine the first loss value between the first predicted click-through rate and the second predicted click-through rate according to the received first predicted click-through rate, second predicted click-through rate, and actual click result, and can also determine the second loss value between the second predicted click-through rate and the actual click result. Furthermore, in the fully connected layer, according to the first loss value and the second loss value, determine the loss value of the target loss function. Furthermore, determine the loss function according to the first loss value and the second loss value, and the training of the neural network can be ended by adjusting the loss function. It should be noted that, in order to improve the accuracy of the neural network push, during the training process, the predicted exposure rate can be used as a training parameter, that is, the predicted exposure rate can be input into the fully connected layer, and the predicted exposure rate can be used as a parameter for determining the loss value of the loss function.

[0072] In actual use, the fully connected layer can receive the second predicted click-through rate and then output the push result. It should be noted that the fully connected layer can also receive the predicted exposure rate and determine the push position of the media resource according to the predicted exposure rate and the first predicted click-through rate.

[0073] It should be noted that the fully connected layer can be the fully connected layer in the second trained neural network or outside the second trained neural network. That is to say, the fully connected layer can be a part of the neural network structure or the next layer operation structure after the neural network outputs data.

[0074] Optionally, inputting the position feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into the first trained neural network to obtain the first predicted click-through rate output by the first trained neural network may include: generating a cross feature vector of the position feature, the resource feature, and the user feature; determining the first predicted click-through rate according to the cross feature vector.

[0075] In this embodiment, a cross feature vector can be generated based on the location feature, resource feature and user feature. The cross feature can simultaneously represent the location information, resource information and user information, and can better integrate the data of each feature parameter to obtain a more accurate first predicted click rate. For example, the user feature age = 25, gender = female, the location feature is position = upper left, and the resource feature categories = entertainment news. The intersection of these three features can tell that a 25-year-old girl is very likely to like to click on the entertainment news displayed in the upper left part.

[0076] As an optional embodiment, the technical solution of the present application is described in detail below using the "Look" function in WeChat as an example: that is, taking the recommendation of a group of articles as an example, it explains how to adjust the model parameters of the second training neural network and how to recommend articles to users.

[0077] Step S1, input the position feature of the sample article, the feature of the article, and the user feature of the sample user into the first training neural network to obtain the first predicted click rate output by the first training neural network; wherein, the position feature W1 of article 1, the feature P1 of article 1, the age feature n1 of user B, and the gender feature x1 of the user are input into the first training neural network to obtain the first predicted click rate output by the first training neural network. wherein, the user feature of user B may not be limited to the age feature n1 of user B and the gender feature x1 of the user, but may be a variety of user features of different dimensions, the more features, the more accurate the output first predicted click rate, and the resource features of the sample media resource 1 may be information of multiple dimensions, the more features, the more accurate the output first predicted click rate.

[0078] Step S2, the feature P1 of article 1, the age feature n1 of user B, and the gender feature x1 of the user are input into the second training neural network, the second predicted click rate output by the second training neural network is obtained, and the second predicted click rate is output. It should be noted that the user feature of user B is not limited to the age feature n1 of user B and the gender feature x1 of the user, and can be a variety of user features of different dimensions. The more features, the more accurate the output second predicted click rate. The resource features of sample media resource 1 can be information of multiple dimensions. The more features, the more accurate the output second predicted click rate.

[0079] Step S3, determining the loss value of the target loss function of the second training neural network according to the first predicted click rate, the second predicted click rate and the actual click result, and adjusting the model parameters in the second training neural network according to the loss value of the target loss function.

[0080] As an optional embodiment, the technical solution of the present application is described in detail below by taking a game announcement in a game application as an example:

[0081] Step S1: Input the location feature of the sample announcement resource, the feature of the announcement resource, and the user feature of the sample user into the first training neural network to obtain the first predicted click-through rate output by the first training neural network. Among them, the location feature W2 of announcement resource 2, the feature P2 of announcement resource 2, the age feature n1 of user C, the gender feature x1 of user C, the game level of user C, and the damage s1 output by user C within a period of time are input into the first training neural network to obtain the first predicted click-through rate output by the first training neural network. Among them, user C may also include, but is not limited to, multiple user features in different dimensions. The more features there are, the more accurate the first predicted click-through rate output will be. The resource feature of announcement resource 2 can be information in multiple dimensions. The more features there are, the more accurate the first predicted click-through rate output will be.

[0082] Step S2: Input the feature P2 of announcement resource 2, the age feature n1 of user C, the gender feature x1 of user C, the game level of user C, and the damage s1 output by user C within a period of time into the second training neural network to obtain the second predicted click-through rate output by the second training neural network. Among them, user C may also include, but is not limited to, multiple user features in different dimensions. The more features there are, the more accurate the second predicted click-through rate output will be. The resource feature of announcement resource 2 can be information in multiple dimensions. The more features there are, the more accurate the second predicted click-through rate output will be. It should be noted that the user feature of user C can also be features in multiple different dimensions. The more features there are, the more accurate the second predicted click-through rate output will be. The resource feature of announcement resource 2 can be information in multiple dimensions. The more features there are, the more accurate the second predicted click-through rate output will be.

[0083] Step S3: Determine the loss value of the target loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, and adjust the model parameters in the second training neural network according to the loss value of the target loss function.

[0084] Among them, the adjusted second training neural network can be used as the target neural network, and this target neural network can be used for the recommendation of game announcement information in game applications, that is, to recommend the information that users are interested in to users. That is to say, the target neural network can be used to push information to users, and the content of the pushed information is the content that users are interested in, thereby improving the usage conversion rate of the information.

[0085] Optionally, the present application also provides an optional click-through rate prediction model for eliminating position bias based on knowledge distillation, as Figure 7 shown in the flowchart of the click-through rate prediction model for eliminating position bias based on knowledge distillation.

[0086] In this embodiment, the click-through rate prediction model may include but is not limited to the recommendation of information in WeChat Kanyi.

[0087] WeChat Kanyikan is a feed stream recommendation product that can recommend different content such as picture and text public accounts, videos, news, etc. The click-through rate prediction model (including the first training neural network, the second training neural network, and the third training neural network mentioned above) can estimate the click-through rate of each article, each video, and each news, and sort them according to the click-through rate, and finally recommend the content with a higher click-through rate to users.

[0088] In this embodiment, the teacher network (equivalent to the first training neural network) uses a feature-based method to take position features (location features of sample media resources) and other features (equivalent to resource features including the sample media resources and user features of sample users) as inputs of the teacher network deep recommendation model deepfm (the first training neural network model), so that the position features and other features are multi-dimensionally crossed; the student network (equivalent to the second neural network) adopts the PAL model structure, models the position features separately, and adds or multiplies the results obtained by the neural network dnn (equivalent to the third training neural network) and deepfm to obtain the final output; the cross entropy of the soft target output by the teacher network and the hard target of the true label is calculated, and linear weighting is performed to obtain the loss function of the student network

[0089] L=αL (soft) +(1-α)L (hard)

[0090] Among them, L (soft) represents the first loss value, L (hard) represents the second loss value, L represents the loss value of the target loss function of the second training neural network, α represents the preset weight, 0<α<1.

[0091] In this embodiment, through knowledge distillation, the output of the teacher network can be transmitted to the student network, so that the student network can also learn the results of the multi-dimensional intersection of position features and other features. At the same time, the student network models the position features separately, and only uses the pCTR part of the deepfm output during online prediction, which reduces the impact of not being able to get the position features online. In this embodiment, the network structure used by deepfm is an improved version of deepfm, such as Figure 8 As shown, the structural diagram of the click rate prediction model based on knowledge distillation to eliminate position bias. Figure 8As shown, the click-through rate prediction model that eliminates position bias based on knowledge distillation includes a teacher network, a student network, and a neural network. The position feature, the user feature of the user, and the resource feature of the media resource are input into the teacher network to obtain the first predicted click-through rate output by the teacher network. The position feature is input into the neural network model to obtain the predicted exposure rate output by the neural network model. The user feature of the user and the resource feature of the media resource are input into the student network to obtain the second predicted click-through rate output by the student network. According to the first predicted click-through rate and the second predicted click-through rate, a first loss value is determined; according to the first predicted click-through rate, the second predicted click-through rate, the predicted exposure rate, and the actual click result, the loss value of the target loss function of the student network is determined, and according to the loss value of the target loss function, the model parameters in the student network are adjusted.

[0092] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0093] According to another aspect of the embodiments of the present invention, there is also provided a media resource processing device for implementing the above-mentioned media resource processing method. As Figure 9 shown, the media resource processing device includes: a first input unit 91, a second input unit 93, a determination unit 95, and an adjustment unit 97.

[0094] The first input unit 91 is configured to input the position feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into the first training neural network to obtain the first predicted click-through rate output by the first training neural network, where the position feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks the sample media resource.

[0095] The second input unit 93 is configured to input the resource feature of the sample media resource and the user feature of the sample user into the second training neural network to obtain the second predicted click-through rate output by the second training neural network. The second predicted click-through rate is used to represent the predicted probability that the sample user clicks the sample media resource. The first training neural network and the second training neural network have the same model structure.

[0096] A determination unit 95, configured to determine a loss value of an objective loss function of a second training neural network according to a first predicted click-through rate, a second predicted click-through rate, and an actual click result, where the actual click result is used to indicate whether a sample user actually clicks on a sample media resource.

[0097] An adjustment unit 97, configured to adjust model parameters in the second training neural network according to the loss value of the objective loss function.

[0098] Optionally, in this embodiment, the above determination unit 95 may include: a first determination module, configured to determine a loss value of an objective loss function of a second training neural network according to a first predicted click-through rate, a second predicted click-through rate, a predicted exposure rate, and an actual click result, where the predicted exposure rate is an exposure rate determined according to a position feature of the sample media resource, and the predicted exposure rate is used to indicate the probability that the sample media resource is pushed to the sample user.

[0099] Through the embodiment provided in this application, a first input unit 91 inputs the position feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into a first training neural network, and obtains a first predicted click-through rate output by the first training neural network, where the position feature is used to indicate the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to indicate the predicted probability that the sample user clicks on the sample media resource; a second input unit 93 inputs the resource feature of the sample media resource and the user feature of the sample user into a second training neural network, and obtains a second predicted click-through rate output by the second training neural network, where the second predicted click-through rate is used to indicate the predicted probability that the sample user clicks on the sample media resource, and the first training neural network and the second training neural network have the same model structure; the determination unit 95 determines a loss value of an objective loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, where the actual click result is used to indicate whether the sample user actually clicks on the sample media resource; the adjustment unit 97 adjusts the model parameters in the second training neural network according to the loss value of the objective loss function. The purpose is achieved of determining the loss value of the objective loss function of the second training neural network according to the first predicted click-through rate output by the first training neural network, the second predicted click-through rate output by the second training neural network, and the actual click result, and adjusting the model parameters in the second training neural network according to the loss value. The adjusted second training neural network can eliminate the deviation of the media resource position information and can accurately recommend media resources of interest to the user. That is to say, when recommending media resources to the user, the influence of the media resource position information can be eliminated, and the media resources of interest to the user can be recommended to the user more accurately, thereby solving the technical problem of low accuracy in recommending media resources according to user behavior in the prior art.

[0100] Optionally, the above device may further include: a third input unit, configured to input the location feature of the sample media resource into a third training neural network, and obtain a predicted exposure rate output by the third training neural network.

[0101] Among them, the above first determination module may include: a first determination sub-module, configured to determine a first loss value according to the first predicted click-through rate and the second predicted click-through rate; a second determination sub-module, configured to determine a second loss value according to the second predicted click-through rate, the predicted exposure rate, and the actual click result; a third determination sub-module, configured to determine the loss value of the target loss function of the second training neural network according to the first loss value and the second loss value.

[0102] It should be noted that the above third determination sub-module is further configured to perform the following operations: determine the loss value of the target loss function of the second training neural network through the following formula:

[0103] L = αL (soft) +(1 - α)L (hard)

[0104] Among them, L (soft) represents the first loss value, L (hard) represents the second loss value, L represents the loss value of the target loss function of the second training neural network, α represents a preset weight, and 0 < α < 1.

[0105] The above first determination sub-module is further configured to perform the following operations: determine the square of the difference between the first predicted click-through rate and the second predicted click-through rate as the first loss value; or determine the sum or product of the second predicted click-through rate and the predicted exposure rate as the predicted output value; determine the square of the difference between the first predicted click-through rate and the predicted output value as the first loss value.

[0106] The above second determination sub-module is further configured to perform the following operations: determine the sum or product of the second predicted click-through rate and the predicted exposure rate as the predicted output value; determine the cross-entropy between the predicted output value and the actual click result as the second loss value.

[0107] Optionally, the above adjustment unit 97 may include: an adjustment module, configured to adjust the model parameters in the second training neural network in an adjustment direction that reduces the loss value of the target loss function.

[0108] Optionally, the above device may further include: an acquisition unit, configured to acquire the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user, where the resource feature of the sample media resource includes resource features in one or more dimensions, and the user feature of the sample user includes user features in one or more dimensions.

[0109] Among them, the above-mentioned acquisition unit may include: an acquisition module, configured to acquire the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user, where the resource feature of the sample media resource includes the type of the sample media resource, and the user feature of the sample user includes the age and gender of the sample user.

[0110] Optionally, the above-mentioned determination unit may further include: a first input module, configured to input the first predicted click-through rate output by the first training neural network, the second predicted click-through rate output by the second training neural network, and the actual click result into the target fully-connected layer module; a second determination module, configured to determine, through the target fully-connected layer module, a first loss value between the first predicted click-through rate and the second predicted click-through rate; a third determination module, configured to determine, through the target fully-connected layer module, a second loss value between the second predicted click-through rate and the actual click result; a fourth determination module, configured to determine, through the target fully-connected layer module, the loss value of the target loss function according to the first loss value and the second loss value.

[0111] Optionally, the above-mentioned first input unit 91 may include: a generation module, configured to generate a cross feature vector of the location feature, the resource feature, and the user feature; a fifth determination module, configured to determine the first predicted click-through rate according to the cross feature vector.

[0112] According to another aspect of the embodiments of the present invention, there is also provided an electronic device for implementing the above-mentioned method for processing media resources. The electronic device may be Figure 1 the terminal device or server shown in the figure. In this embodiment, the electronic device is taken as an example of a server for illustration. As Figure 10 shown in the figure, the electronic device includes a memory 1002 and a processor 1004. A computer program is stored in the memory 1002, and the processor 1004 is configured to execute the steps in any one of the above-mentioned method embodiments through the computer program.

[0113] Optionally, in this embodiment, the above-mentioned electronic device may be at least one network device among multiple network devices in a computer network.

[0114] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through the computer program:

[0115] S1, input the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into the first training neural network to obtain the first predicted click-through rate output by the first training neural network, where the location feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource;

[0116] S2, input the resource features of the sample media resource and the user features of the sample user into the second training neural network to obtain the second predicted click-through rate output by the second training neural network. The second predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource. The first training neural network and the second training neural network have the same model structure;

[0117] S3, determine the loss value of the objective loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, where the actual click result is used to represent whether the sample user actually clicks on the sample media resource;

[0118] S4, adjust the model parameters in the second training neural network according to the loss value of the objective loss function.

[0119] Optionally, those of ordinary skill in the art can understand that Figure 10 The structure shown is only illustrative. The electronic device can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 10 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 10 or have a different configuration from that shown in Figure 10 shown.

[0120] Among them, the memory 1002 can be used to store software programs and modules, such as the program instructions / modules corresponding to the processing method and device of the media resource in the embodiment of the present invention. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, that is, implements the above-mentioned processing method of the media resource. The memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 1002 may further include a memory remotely set relative to the processor 1004, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof. Specifically, the memory 1002 can be but is not limited to used to store information such as the first training neural network, the second training neural network, the location features of the sample media resource, the resource features of the sample media resource, the user features of the sample user, and the loss value of the objective loss function. As an example, such as Figure 10As shown in the figure, the above-mentioned memory 1002 may but is not limited to include a first input unit 91, a second input unit 93, a determination unit 95, and an adjustment unit 97 in the processing device of the above-mentioned media resources. In addition, it may also include but is not limited to other module units in the processing device of the above-mentioned media resources, which will not be elaborated in this example.

[0121] Optionally, the above-mentioned transmission device 1006 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one example, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 1006 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0122] In other embodiments, the above-mentioned terminal device or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes in the form of network communication. Among them, the nodes can form a peer-to-peer (P2P, Peer To Peer) network, and any form of computing device, such as a server, a terminal, etc., can become a node in the blockchain system by joining the peer-to-peer network.

[0123] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the processing method of the above-mentioned media resources or the processing method of the media resources provided in various optional implementation manners of the processing aspect of the media resources. Among them, the computer program is set to execute the steps in any one of the above method embodiments when running.

[0124] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be set to store a computer program for executing the following steps:

[0125] S1. Input the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into the first training neural network to obtain the first predicted click-through rate output by the first training neural network, where the location feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource;

[0126] S2. Input the resource feature of the sample media resource and the user feature of the sample user into the second training neural network to obtain the second predicted click-through rate output by the second training neural network. The second predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource. The first training neural network and the second training neural network have the same model structure;

[0127] S3. Determine the loss value of the target loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, where the actual click result is used to represent whether the sample user actually clicks on the sample media resource;

[0128] S4. Adjust the model parameters in the second training neural network according to the loss value of the target loss function.

[0129] Optionally, in this embodiment, those of ordinary skill in the art can understand that all or part of the steps in the above-mentioned various methods of the embodiment can be completed by a program instructing the relevant hardware of the terminal device. The program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0130] The serial numbers of the above-mentioned embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0131] If the integrated unit in the above-mentioned embodiment is implemented in the form of a software function unit and sold or used as an independent product, it can be stored in the above-mentioned computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in the storage medium and includes several instructions for causing one or more computer devices (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0132] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0133] In several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0134] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0136] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for processing media resources, characterized in that, Including: Inputting the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into a first training neural network to obtain a first predicted click-through rate output by the first training neural network, where the location feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource; Inputting the resource feature of the sample media resource and the user feature of the sample user into a second training neural network to obtain a second predicted click-through rate output by the second training neural network, where the second predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource, and the first training neural network and the second training neural network have the same model structure; Determining the loss value of the target loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result, where the actual click result is used to represent whether the sample user actually clicks on the sample media resource; Adjusting the model parameters in the second training neural network according to the loss value of the target loss function.

2. The method according to claim 1, characterized in that, The determining the loss value of the target loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result includes: Determining the loss value of the target loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, the predicted exposure rate, and the actual click result, where the predicted exposure rate is the exposure rate determined according to the location feature of the sample media resource, and the predicted exposure rate is used to represent the probability that the sample media resource is pushed to the sample user.

3. The method according to claim 2, characterized in that, The method further includes: Inputting the location feature of the sample media resource into a third training neural network to obtain the predicted exposure rate output by the third training neural network.

4. The method according to claim 2, characterized in that, The determining the loss value of the target loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, the predicted exposure rate, and the actual click result includes: Determining a first loss value according to the first predicted click-through rate and the second predicted click-through rate; Determining a second loss value according to the second predicted click-through rate, the predicted exposure rate, and the actual click result; Determining the loss value of the target loss function of the second training neural network according to the first loss value and the second loss value.

5. The method according to claim 4, characterized in that, The determining the loss value of the target loss function of the second training neural network according to the first loss value and the second loss value includes: Determining the loss value of the target loss function of the second training neural network through the following formula: L = αL (soft) +(1 - α)L (hard) Among them, L (soft) represents the first loss value, and L (hard) represents the second loss value. L represents the loss value of the target loss function of the second trained neural network, and α represents a preset weight, where 0 < α < 1.

6. The method according to claim 4, characterized in that, The determining the first loss value according to the first predicted click-through rate and the second predicted click-through rate includes: Determining the square of the difference between the first predicted click-through rate and the second predicted click-through rate as the first loss value; or Determine the sum or product of the second predicted click-through rate and the predicted exposure rate as the predicted output value; determine the square of the difference between the first predicted click-through rate and the predicted output value as the first loss value.

7. The method according to claim 4, characterized in that, The determining the second loss value according to the second predicted click-through rate, the predicted exposure rate, and the actual click result includes: Determine the sum or product of the second predicted click-through rate and the predicted exposure rate as the predicted output value; Determine the cross entropy between the predicted output value and the actual click result as the second loss value.

8. The method according to any one of claims 1 to 7, characterized in that, The adjusting the model parameters in the second trained neural network according to the loss value of the target loss function includes: Adjust the model parameters in the second trained neural network in the adjustment direction that reduces the loss value of the target loss function.

9. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Obtain the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user, where the resource feature of the sample media resource includes resource features in one or more dimensions, and the user feature of the sample user includes user features in one or more dimensions.

10. The method according to claim 9, characterized in that, The obtaining the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user includes: Obtain the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user, where the resource feature of the sample media resource includes the type of the sample media resource, and the user feature of the sample user includes the age of the sample user and the gender of the sample user.

11. The method according to any one of claims 1 to 7, characterized in that, The determining the loss value of the target loss function of the second trained neural network according to the first predicted click-through rate, the second predicted click-through rate, and the actual click result includes: Input the first predicted click-through rate output by the first trained neural network, the second predicted click-through rate output by the second trained neural network, and the actual click result into the target fully connected layer module; Determine the first loss value between the first predicted click-through rate and the second predicted click-through rate through the target fully connected layer module; Determine the second loss value between the second predicted click-through rate and the actual click result through the target fully connected layer module; Determine the loss value of the target loss function according to the first loss value and the second loss value through the target fully connected layer module.

12. The method according to any one of claims 1 to 7, characterized in that, The inputting the location feature of the sample media resource, the resource feature of the sample media resource, and the user feature of the sample user into the first trained neural network to obtain the first predicted click-through rate output by the first trained neural network includes: Generate a cross feature vector of the location feature, the resource feature, and the user feature; Determine the first predicted click-through rate according to the cross feature vector.

13. A processing device for media resources, characterized in that, Includes: A first input unit for inputting the location feature of a sample media resource, the resource feature of the sample media resource, and the user feature of a sample user into a first training neural network to obtain a first predicted click-through rate output by the first training neural network, where the location feature is used to represent the actual arrangement position of the sample media resource in a group of media resources pushed to the sample user, and the first predicted click-through rate is used to represent the predicted probability that the sample user clicks on the sample media resource; A second input unit for inputting the resource feature of the sample media resource and the user feature of the sample user into a second training neural network to obtain a second predicted click-through rate output by the second training neural network, the second predicted click-through rate being used to represent the predicted probability that the sample user clicks on the sample media resource, and the first training neural network and the second training neural network having the same model structure; A determination unit for determining a loss value of an objective loss function of the second training neural network according to the first predicted click-through rate, the second predicted click-through rate, and an actual click result, where the actual click result is used to represent whether the sample user actually clicks on the sample media resource; An adjustment unit for adjusting model parameters in the second training neural network according to the loss value of the objective loss function.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where the program, when running, executes the method according to any one of claims 1 to 12.

15. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 12 through the computer program.

Citation Information

Patent Citations

  • Ensemble predictor

    CN109154944A

  • Information recommendation method and device based on artificial intelligence and electronic equipment

    CN111695037A