Model training method, click rate determination method, medium, device and computing device

By adjusting the click-through rate prediction model using the static attribute data and label data of multimedia, the problem of unexposed multimedia click-through rate prediction is solved, and a more accurate and stable click-through rate prediction is achieved.

CN113902103BActive Publication Date: 2025-05-16NETEASE MEDIA TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111262649.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-28
Publication Date
2025-05-16
Estimated Expiration
2041-10-28

AI Technical Summary

Technical Problem

The prior art cannot accurately predict the click rate of unexposed multimedia.

Method used

By obtaining the static attribute data and corresponding label data of multiple training samples, the click-through rate prediction model is adjusted to achieve accurate prediction of unexposed multimedia click-through rate.

Benefits of technology

Accurate prediction of unexposed multimedia click-through rate is achieved, the application scenarios of click-through rate prediction are expanded, and the robustness and stability of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113902103B_ABST
    Figure CN113902103B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a model training method, a click rate determination method, a medium, an apparatus and a computing device. The model training method includes: determining to obtain multiple training samples, the training samples include multiple static attribute data of sample multimedia, the static attribute data represents the properties of the sample multimedia itself; obtaining label data corresponding to the training samples, the label data is used to represent the actual click rate level of the sample multimedia; inputting the training samples into a click rate prediction model to obtain a training output click rate level; adjusting the click rate prediction model according to the label data and the training output click rate level to obtain a trained click rate prediction model. The present disclosure uses the static attribute data of multimedia as the basis for predicting the click rate, which can effectively increase the application scenarios of click rate prediction, does not need to rely on the user's click behavior for modeling and analysis, and further enhances the robustness and stability of the trained click rate prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of multimedia technology. More specifically, the embodiments of the present disclosure relate to a model training method, a click-through rate determination method, a medium, an apparatus, and a computing device. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims. No description herein is admitted to be prior art by inclusion in this section.

[0003] The click-through rate of multimedia refers to the ratio of the number of times the multimedia is clicked by users to the number of times it is exposed to users. The click-through rate of multimedia can reflect the user's preference for the multimedia and the rationality of the system's recommendation of the multimedia.

[0004] The click-through rate of multimedia is a reference data for ranking multimedia to be exposed. Among them, the current methods for predicting the click-through rate of multimedia all need to combine the user's click behavior, so the click-through rate of multimedia that has not been exposed cannot be accurately predicted. Summary of the invention

[0005] In this context, embodiments of the present disclosure are intended to provide a model training method, a click-through rate determination method, a medium, an apparatus, and a computing device to solve the current problem of being unable to accurately predict the click-through rate of multimedia that has not been exposed.

[0006] In the first aspect of the embodiments of the present disclosure, a model training method is provided, comprising: obtaining multiple training samples, the training samples including multiple static attribute data of the sample multimedia, the static attribute data representing the properties of the sample multimedia itself; obtaining label data corresponding to the training samples, the label data being used to represent the actual click-through rate level of the sample multimedia; inputting the training samples into a click-through rate prediction model to obtain a training output click-through rate level; adjusting the click-through rate prediction model according to the label data and the training output click-through rate level to obtain a trained click-through rate prediction model.

[0007] In one embodiment of the present disclosure, the click-through rate prediction model is a single-classification model, the number of classes in the single-classification model is the same as the number of click-through rate level divisions, and the training samples are input into the click-through rate prediction model to obtain the training output click-through rate level, including: inputting the training samples into the click-through rate prediction model to obtain probability values ​​of different click-through rate levels; and determining the click-through rate level corresponding to the maximum probability value in the probability values ​​as the training output click-through rate level.

[0008] In another embodiment of the present disclosure, static attribute data includes: a cover image of sample multimedia and text data of the sample multimedia; inputting the training sample into a click-through rate prediction model to obtain probability values ​​of different click-through rate levels, including: inputting the cover image into a convolutional neural network to obtain a first feature vector; inputting the text data into an embedding network to obtain a second feature vector; concatenating the first feature vector and the second feature vector to obtain a concatenated feature vector; and inputting the concatenated feature vector into a fully connected layer to obtain probability values ​​of different click-through rate levels.

[0009] In another embodiment of the present disclosure, text data includes: title data, content data and associated data, the associated data is used to represent the quality of sample multimedia, the embedding network includes: a word embedding network, a pre-trained embedding network and a lookup embedding network, and the text data is input into the embedding network to obtain a second feature vector, including: inputting the title data into the word embedding network to obtain a first sub-feature vector; inputting the content data into the pre-trained embedding network to obtain a second sub-feature vector; inputting the associated data into the lookup embedding network to obtain a third sub-feature vector; and concatenating the first sub-feature vector, the second sub-feature vector and the third sub-feature vector to obtain a second feature vector.

[0010] In another embodiment of the present disclosure, multiple training samples are obtained, including: obtaining exposure logs of sample multimedia exposed within a preset historical time; preprocessing the exposure logs of the same sample multimedia to obtain non-repetitive spare data; and screening out training samples belonging to a preset sample dictionary from the spare data, the preset sample dictionary being used to represent static properties of the sample multimedia.

[0011] In another embodiment of the present disclosure, an exposure log of sample multimedia exposed within a preset historical time is obtained, including: obtaining the exposure log of sample multimedia exposed within the preset historical time at every preset time, the exposure log is generated by the multimedia in the multimedia library when it is exposed, and the multimedia in the multimedia library is updated in real time.

[0012] In another embodiment of the present disclosure, label data corresponding to a training sample is obtained, including: obtaining a current click-through rate corresponding to the training sample; determining an actual click-through rate level corresponding to the training sample based on the current click-through rate and a preset mapping rule, wherein the current click-through rate and the actual click-through rate level are positively correlated, and the actual click-through rate level is used to represent the label data.

[0013] In another embodiment of the present disclosure, the click-through rate prediction model is adjusted according to the label data and the training output click-through rate level to obtain a trained click-through rate prediction model, including: determining the cross-entropy loss function corresponding to the click-through rate prediction model according to the label data and the training output click-through rate level; adjusting the click-through rate prediction model according to the cross-entropy loss function to obtain the trained click-through rate prediction model.

[0014] In the second aspect of the embodiments of the present disclosure, a method for determining a click rate is provided, including: obtaining static attribute data of the multimedia to be detected; inputting the static attribute data of the multimedia to be detected into a click rate prediction model to obtain a target click rate level of the multimedia to be detected, wherein the click rate prediction model is obtained by training according to any model training method of the first aspect, and the target click rate level is used to represent the click rate range of the multimedia to be detected.

[0015] In one embodiment of the present disclosure, static attribute data of the multimedia to be detected is input into a click-through rate prediction model to obtain a target click-through rate level of the multimedia to be detected, including: inputting the static attribute data of the multimedia to be detected into the click-through rate prediction model to obtain probability values ​​of the multimedia to be detected at different click-through rate levels; and determining the click-through rate level corresponding to the maximum probability value in the probability values ​​as the target click-through rate level.

[0016] In one embodiment of the present disclosure, the click rate determination method further includes: using the following formula to determine the click rate score of the multimedia to be detected:

[0017]

[0018] In the above formula, s represents the click-through rate score, t represents the set maximum click-through rate level, i represents the value of the click-through rate level, and p i It represents the probability value when the click rate level is i. The click rate score is used to represent the estimated click rate of the multimedia to be detected.

[0019] In a third aspect of the embodiments of the present disclosure, a model training device is provided, comprising:

[0020] A first acquisition module is used to acquire a plurality of training samples, wherein the training samples include a plurality of static attribute data of the sample multimedia, and the static attribute data represents the properties of the sample multimedia itself;

[0021] The second acquisition module is used to acquire label data corresponding to the training sample, and the label data is used to represent the actual click rate level of the sample multimedia;

[0022] A click rate determination module is used to input the training samples into the click rate prediction model to obtain the training output click rate level;

[0023] The adjustment module is used to adjust the click rate prediction model according to the label data and the training output click rate level to obtain a trained click rate prediction model.

[0024] In one embodiment of the present disclosure, the click rate prediction model is a single classification model, the number of classifications of the single classification model is the same as the number of click rate level divisions, and the click rate determination module includes:

[0025] An input unit, used to input the training samples into the click rate prediction model to obtain probability values ​​of different click rate levels;

[0026] The determination unit is used to determine the click rate level corresponding to the maximum probability value in the probability values ​​as the training output click rate level.

[0027] In one embodiment of the present disclosure, the static attribute data includes: a cover image of the sample multimedia and text data of the sample multimedia; the input unit includes:

[0028] A first input subunit, used for inputting the cover image into the convolutional neural network to obtain a first feature vector;

[0029] A second input subunit, used for inputting text data into the embedding network to obtain a second feature vector;

[0030] A third input subunit is used to concatenate the first feature vector and the second feature vector to obtain a concatenated feature vector;

[0031] The fourth input subunit is used to input the concatenated feature vector into the fully connected layer to obtain probability values ​​of different click rate levels.

[0032] In one embodiment of the present disclosure, text data includes: title data, content data and associated data, the associated data is used to represent the quality of sample multimedia, the embedding network includes: word embedding network, pre-trained embedding network and search embedding network, and the second input sub-unit is specifically used to: input the title data into the word embedding network to obtain a first sub-feature vector; input the content data into the pre-trained embedding network to obtain a second sub-feature vector; input the associated data into the search embedding network to obtain a third sub-feature vector; and concatenate the first sub-feature vector, the second sub-feature vector and the third sub-feature vector to obtain a second feature vector.

[0033] In one embodiment of the present disclosure, the first acquisition module includes:

[0034] An acquisition unit, used to acquire exposure logs of sample multimedia that have been exposed within a preset historical time;

[0035] A processing unit, used for preprocessing the exposure logs of the same sample multimedia to obtain non-repetitive backup data;

[0036] The screening unit is used to screen out training samples belonging to a preset sample dictionary from the spare data, where the preset sample dictionary is used to represent static properties of sample multimedia.

[0037] In one embodiment of the present disclosure, the acquisition unit is specifically used to: obtain the exposure log of the sample multimedia exposed within a preset historical time at every preset time, the exposure log is generated by the multimedia in the multimedia library when it is exposed, and the multimedia in the multimedia library is updated in real time.

[0038] In one embodiment of the present disclosure, the second acquisition module is specifically used to: obtain the current click-through rate corresponding to the training sample; determine the actual click-through rate level corresponding to the training sample based on the current click-through rate and according to a preset mapping rule, the current click-through rate and the actual click-through rate level are positively correlated, and the actual click-through rate level is used to represent the label data.

[0039] In one embodiment of the present disclosure, the adjustment module is specifically used to: determine the cross entropy loss function corresponding to the click-through rate prediction model based on the label data and the training output click-through rate level; adjust the click-through rate prediction model according to the cross loss function to obtain a trained click-through rate prediction model.

[0040] In a fourth aspect of the embodiments of the present disclosure, a click rate determination device is provided, comprising:

[0041] An acquisition module, used for acquiring static attribute data of the multimedia to be detected;

[0042] The click rate determination module is used to input the static attribute data of the multimedia to be detected into a click rate prediction model to obtain a target click rate level of the multimedia to be detected. The click rate prediction model is obtained by training according to any model training device of the third aspect, and the target click rate level is used to represent the click rate range of the multimedia to be detected.

[0043] In one embodiment of the present disclosure, the click rate determination module is specifically used to: input the static attribute data of the multimedia to be detected into the click rate prediction model to obtain the probability value of the multimedia to be detected at different click rate levels; determine the click rate level corresponding to the maximum probability value in the probability value as the target click rate level.

[0044] In one embodiment of the present disclosure, the click rate determination device further includes:

[0045] The determination module is used to determine the click rate score of the multimedia to be detected by using the following formula:

[0046]

[0047] In the above formula, s represents the click-through rate score, t represents the set maximum click-through rate level, i represents the value of the click-through rate level, and p i It represents the probability value when the click rate level is i. The click rate score is used to represent the estimated click rate of the multimedia to be detected.

[0048] In a fifth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer program instructions are stored. When the computer program instructions are executed, the method of any one of the first aspect or the second aspect described above is implemented.

[0049] In a sixth aspect of the embodiments of the present disclosure, a computing device is provided, comprising: a memory and a processor; the memory is used to store program instructions; the processor is used to call the program instructions in the memory to execute a method as described in any one of the first or second aspects above.

[0050] According to the embodiment of the present disclosure, a plurality of training samples are obtained, the training samples include a plurality of static attribute data of the sample multimedia, the static attribute data represents the properties of the sample multimedia itself; label data corresponding to the training samples are obtained, the label data is used to represent the actual click rate level of the sample multimedia; the training samples are input into the click rate prediction model to obtain the training output click rate level; the click rate prediction model is adjusted according to the label data and the training output click rate level to obtain a trained click rate prediction model. On the one hand, the use of static attribute data of multimedia as the basis for predicting click rate can effectively increase the application scenarios of click rate prediction without relying on the user's click behavior for modeling and analysis. On the other hand, the use of multiple static attribute data of multimedia is the result of a combination of multiple features, and the error influence of a single feature is reduced, which can enhance the robustness and stability of the trained click rate prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:

[0052] Figure 1 The following schematically shows an application scenario diagram of the model training method according to an embodiment of the present disclosure;

[0053] Figure 2 The following is a schematic diagram showing a flow chart of steps of a model training method according to an embodiment of the present disclosure;

[0054] Figure 3 The following schematically shows a flow chart of steps of a model training method according to another embodiment of the present disclosure;

[0055] Figure 4 The structure diagram of the click rate prediction model according to an embodiment of the present disclosure is schematically shown;

[0056] Figure 5 The following is a schematic flowchart of a method for determining a click rate according to an embodiment of the present disclosure;

[0057] Figure 6 The following is a schematic flowchart showing a method for determining a click rate according to another embodiment of the present disclosure;

[0058] Figure 7 The structure diagram of a computer storage medium according to an embodiment of the present disclosure is schematically shown;

[0059] Figure 8 The structure block diagram of a model training device according to an embodiment of the present disclosure is schematically shown;

[0060] Fig. 9 The structure block diagram of a click rate determination device according to an embodiment of the present disclosure is schematically shown;

[0061] Fig.10 The structure block diagram of a computing device according to an embodiment of the present disclosure is schematically shown.

[0062] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION

[0063] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0064] Those skilled in the art will appreciate that the embodiments of the present disclosure may be implemented as a system, device, apparatus, method or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0065] According to an embodiment of the present disclosure, a model training method, medium, apparatus and computing device are proposed.

[0066] It should be understood herein that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction rather than having any limiting meaning.

[0067] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure. SUMMARY OF THE INVENTION

[0069] The inventors found that the first click-through rate prediction model uses a logistic regression method (LogisticRegression, LR), which has a simple structure and strong interpretability. Specifically, the click-through rate prediction model calculates the features of each video in the form of weighted summation, and the activation function (sigmoid) of the logistic regression has a value range of 0 to 1. The output value of the activation function is used to explain the click probability of the video. In addition, after long-term optimization and development, the click-through rate prediction model can add calculations such as squares or custom features as additional feature supplements to the original logistic regression formula to further optimize the content representation of the video.

[0070] However, the first click-through rate prediction model relies heavily on the definition of the logistic regression formula to fit the video features to the click-through rate, which is limited by prior knowledge. In addition, the logistic regression method is optimized by increasing the complexity of features. Some important cross features cannot be expressed by linear models, and these cross features need to be manually added to expand the input, which will result in the algorithm's low scalability. In addition, the difficulty of features will be greatly increased, which will lead to overfitting of sample data, making it impossible to handle complex practical application scenarios.

[0071] Among them, the second click-through rate prediction model is the Factorization Machines (FM) model. The Factorization Machines model can be expressed as the sum of the first-order cross terms and the second-order cross terms of the features, without distinguishing between static features and dynamic features of the video. The second click-through rate prediction model learns a low-dimensional embedding vector for each dimension of the feature, and then uses the inner product of the embedding vector to represent the cross coefficient. This can reduce the number of parameters of the second click-through rate prediction model, making it feasible for the model to learn cross features, and ultimately increase the practicality of the model. Corresponding each dimensional feature to an embedding vector also brings an advantage: as long as the number of non-zero values ​​of the feature in the training set is large enough, the corresponding embedding vector can be fully trained, so that the coefficient of the cross feature can be calculated through the inner product of the embedding vector, and the cross term does not require many times in the training set.

[0072] Although the second click-through rate prediction model avoids the model parameter problems caused by feature intersection and increase of features to a certain extent, it only models the second-order intersection features. However, in actual application scenarios, higher-order feature intersections are also important for the prediction of click-through rate, thus limiting the application scenarios of the second click-through rate prediction model.

[0073] In addition, in order to cope with more complex application scenarios, user behavior characteristics and video content characteristics are fully utilized. Some deep learning methods have also become a common means of estimating click-through rates. Deep learning methods often increase the depth of the network and the cross-calculation of features, so that the third click-through rate prediction model can learn high-order feature cross-calculations. Common methods include deep factorization machine (Deep FM) and deep interest network (DIN). These methods combine the feature cross-engineering that the click-through rate prediction problem focuses on with the features that the deep learning network focuses on. For example, DIN first encodes each historical behavior of the user, then calculates the attention weight vector of these historical behaviors through the local attention module, and finally multiplies the attention weight by the historical behavior encoding vector, and obtains the user's feature encoding after weighted summation; user features are spliced ​​together with other features and video features, and input into a multi-layer neural network to finally obtain the video click-through rate prediction value.

[0074] However, the third CTR prediction model needs to combine samples of user click behavior (positive samples), but these samples often only account for a small part of the total samples, which will bring great challenges to predicting positive samples. And due to the difference in the ratio of positive and negative samples, most classifiers will tend to give prediction results that are biased towards negative samples. In addition, the third CTR prediction model is a deep model, and the complexity of the deep model will also cause the model to overfit, which is reflected in the different effects of the training set and the test set.

[0075] Based on the above problems, the present disclosure provides a model training method, which can only use the static attribute data of sample multimedia to train the click-through rate prediction model, and then enable the trained click-through rate prediction model to be applied to the prediction of multimedia that has not been clicked by users, and achieve an accurate prediction result.

[0076] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure are described in detail below.

[0077] Application Scenario Overview

[0078] First reference Figure 1 , is an application scenario diagram of the model training method provided by the present disclosure. Specifically, a plurality of multimedia are stored in the multimedia library 111 of the server 11, wherein the multimedia has a click rate ranking in the multimedia library. The server 11 will search for the corresponding multimedia in the multimedia library 111, and when the user enters the page through the client 12, the multimedia will be pushed to the client 12 for display. When the multimedia is displayed, a corresponding exposure log will be generated. The server will store the exposure log of the multimedia in the log storage space 112 for calling when the server 11 trains the model.

[0079] When a user enters a page through the client 12, the page will display the title 121 of the currently playing multimedia, the multimedia 122 being played, and the titles 123 of other multimedia in sequence. In addition, the multimedia 122 being played can be the multimedia with the highest click rate, and the titles 123 of other multimedia can also be sorted according to the corresponding click rates, and when the user triggers a title 123, the corresponding multimedia can be played.

[0080] Furthermore, if the multimedia 122 being played or the multimedia corresponding to the titles 123 of other multimedia have not been exposed, it is necessary to pre-estimate the click-through rates of these multimedia, and then determine the display positions on the page according to the click-through rates.

[0081] Exemplary Methods

[0082] Combine the following Figure 1 For application scenarios, refer to Figure 2 To describe the model training method according to the exemplary embodiment of the present disclosure. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0083] Figure 2 A flow chart of the steps of a model training method provided by the present disclosure is shown, which specifically includes the following steps:

[0084] S201, obtaining multiple training samples.

[0085] The training samples include a plurality of static attribute data of the sample multimedia, and the static attribute data represent the properties of the sample multimedia itself.

[0086] In the present disclosure, sample multimedia includes content such as video, audio, picture or advertisement. The following description takes video as an example to illustrate the content of the present disclosure. Specifically, sample multimedia refers to data stored in a multimedia library and has been exposed. Among them, a large amount of multimedia data is stored in the multimedia library, part of which has been exposed, that is, displayed on the client's page and can be used as sample multimedia, and part of which has not been exposed, but any multimedia data in the multimedia library can be called by the server to form a page.

[0087] Furthermore, the sample multimedia in the multimedia library has a time limit and is updated in real time. For example, if the time limit of the sample multimedia in the multimedia library is set to 6 months, the sample multimedia stored in the multimedia library for more than 6 months will be deleted, and new multimedia will be added to the multimedia library to achieve real-time updating of the sample multimedia in the multimedia library.

[0088] In addition, the sample multimedia has static attribute data. For example, multiple items of the sample multimedia's identity information (ID), sample multimedia's title, sample multimedia's classification, points of interest, keywords, tags, publishing account name, publishing account's identity information (ID), publishing account level, quality score, quality level, timeliness, cover address (URL), logo, cover clarity, content clarity, three vulgarity scores, and clickbait logo. These static attribute data can represent the properties of the sample multimedia itself, and multiple static attribute data of a sample multimedia constitute a training sample corresponding to the sample multimedia.

[0089] S202, obtaining label data corresponding to the training sample.

[0090] The label data is used to represent the actual click rate level of the sample multimedia.

[0091] Specifically, the sample multimedia refers to multimedia data that has been exposed, and the sample multimedia has a corresponding click-through rate after being exposed. The click-through rate refers to the click-through rate (ctr) of multimedia exposure. The click-through rate represents the ratio of the number of times the multimedia is clicked by users to the number of times it is exposed to users. The click-through rate reflects the preference of the multimedia in the user dimension and the rationality of recommending it to users to a certain extent.

[0092] In addition, according to the click rate of the sample multimedia, the actual click rate level to which the sample multimedia belongs is determined, which is the label data of the corresponding training sample.

[0093] For example, if the click rate of the sample multimedia is less than 0.08, the actual click rate level is determined to be 0, and 0 is used as the label data of the training sample. If the click rate of the sample multimedia is greater than or equal to 0.08 and less than 0.13, the actual click rate level is determined to be 1, and 1 is used as the label data of the training sample. If the click rate of the sample multimedia is greater than or equal to 0.13, the actual click rate level is determined to be 2, and 2 is used as the label data of the training sample.

[0094] In the present disclosure, the training samples and label data of the current sample multimedia can be pre-stored in the memory, so when training the click-through rate prediction model, the training samples and the corresponding label data can be directly obtained to train the click-through rate prediction model.

[0095] S203, input the training samples into the click rate prediction model to obtain the training output click rate level.

[0096] Among them, the click rate prediction model is a single classification model. After the training sample is input into the click rate prediction model, a value will be output as the training output click rate level.

[0097] S204, adjusting the click rate prediction model according to the label data and the training output click rate level to obtain a trained click rate prediction model.

[0098] Among them, the label data is the true value, and the training output click rate level is the estimated value. According to the difference between the true value and the estimated value, the parameters of the click rate prediction model can be adjusted, and the adjusted click rate prediction model is adopted to continue to perform steps S203 and S204 until the adjusted click rate prediction model meets the convergence requirements or meets the number of training times. Exemplarily, the click rate prediction model can be set to be updated every day, and steps S203 and S204 are executed 10 times a day in a cycle, so as to obtain the click rate prediction model that has been trained.

[0099] In addition, since multimedia in the multimedia library will be updated in real time, and multimedia in the multimedia library will be exposed every day, generating new exposure logs, new exposure logs can be selected every day to train the click rate prediction model, and the click rate prediction model trained on the same day is used to predict the click rate of unexposed multimedia on the next day, and the click rate prediction model trained on the next day is trained based on the click rate prediction model trained on the previous day. This achieves the real-time nature of the click rate prediction model and further improves the accuracy of multimedia click rate prediction.

[0100] On the one hand, the disclosed embodiment uses multimedia static attribute data as the basis for predicting click-through rate, which can effectively increase the application scenarios of click-through rate prediction without relying on the user's click behavior for modeling and analysis. On the other hand, the use of multiple multimedia static attribute data is the result of a combination of multiple features, and the error influence of a single feature is reduced, which can enhance the robustness and stability of the trained click-through rate prediction model.

[0101] Figure 3 A flowchart of another model training method provided by the present disclosure is shown, which specifically includes the following steps:

[0102] S301, obtaining exposure logs of sample multimedia that have been exposed within a preset historical period.

[0103] Among them, at every preset time, the exposure log of the sample multimedia exposed within the preset historical time is obtained, and the exposure log is generated by the multimedia in the multimedia library when it is exposed, and the multimedia in the multimedia library is updated in real time.

[0104] Specifically, the preset historical time may be 10 days, 8 days, 5 days, or 3 days, etc. The preset time may be 24 hours or 12 hours. For example, the exposure log of the sample multimedia that has been exposed within 10 days may be obtained every 24 hours. In this way, the click-through rate prediction model is trained once every 24 hours, and the number of iterations of each training of the click-through rate prediction model is 10 cycles (epochs).

[0105] S302, pre-processing the exposure logs of the same sample multimedia to obtain non-repetitive backup data.

[0106] In the present disclosure, within a preset historical time, each sample multimedia may be exposed at different times, and each exposure will generate an exposure log. The static attribute data contained in the exposure log of the same sample multimedia is the same, and the dynamic attribute data is different. Therefore, preprocessing the exposure log of the same sample multimedia to obtain non-repeating spare data means determining that all data in one of the multiple exposure logs of the same sample multimedia is non-repeating spare data.

[0107] The sample multimedia has identity information (ID), so the exposure logs belonging to the same sample multimedia can be determined from multiple exposure logs according to the identity information.

[0108] The exposure log includes a large amount of data, including static attribute data and dynamic data, wherein the dynamic data is the user's interaction data. For example, the dynamic data includes: the identity information of the user who clicks the sample multimedia, the user's gender, age, etc.

[0109] In addition, the present disclosure can also obtain static attribute data corresponding to the sample multimedia in a preset database through the identity information of the sample multimedia. The preset database stores the corresponding relationship between the identity information of each sample multimedia and its static attribute data.

[0110] S303, screening out training samples belonging to a preset sample dictionary from the spare data.

[0111] The preset sample dictionary is used to represent the static properties of the sample multimedia. The preset sample dictionary is shown in Table 1. According to the preset sample dictionary in Table 1, the training samples belonging to the preset sample dictionary can be screened in the spare data.

[0112] Table 1

[0113] serial number Field Name Field meaning Field value attributes 1 Docid Multimedia ID String 2 Title Multimedia Title String 3 Category Multimedia Classification String 4 Poi Points of Interest String 5 Key_word Keywords String 6 Tags Label String 7 Source_name Publishing account name String 8 Source_id Publishing Account ID Integer 9 Source_lv Release account level Integer 10 Quality Multimedia Quality Score Floating point 11 Quality_lv Multimedia level Integer 12 Exp_time Multimedia Time Effect Integer 13 Cover Cover url String 14 Short_video Multimedia Logo Integer 15 Def_cover Cover clarity Integer 16 Def_video Multimedia clarity Integer 17 Vulgar Three vulgar points Floating point 18 Cheat Clickbait logo Integer

[0114] In Table 1, multimedia ID refers to the multimedia identity information. Figure 1 "How to make a delicious braised pork dish", "My hometown", "Economic layout", "Xiao Wang's life" or "Financial events" in the multimedia category. Multimedia categories include: international news, society, entertainment or life, etc. Points of interest (POI) are data describing the multimedia, which are preset systematic data. For example, for Figure 1The interest point of the video "How to make a delicious braised pork" in the video can be "food making". Keywords are data that describe the multimedia content. For example, Figure 1 The keyword of the video "How to make a delicious braised pork" in can be "braised pork". Tags are also data describing the multimedia. Figure 1 The tags of the video "How to make a delicious braised pork dish" can be "meat", "production" and "food". The publishing account name refers to the account name that produced or uploaded the multimedia. The publishing account ID refers to the identity information of the account that produced or uploaded the multimedia.

[0115] In addition, the publishing account level can be the level of the publishing account determined according to a preset rule. For example, if the publishing account has hardly published any videos, the level is 0, and if the publishing account has occasionally published a small number of videos, the level is 1. If the publishing account often publishes videos, but the video click-through rate is low, the level is 2. If the publishing account often publishes videos and the click-through rate is high, the level is 3. The multimedia quality score is a score given to the multimedia setting in advance according to the preset rules, such as 90 points, 60 points or 40 points. The multimedia level is a level given to the multimedia setting in advance according to the preset rules, such as level 1, level 2 or level 3. Multimedia timeliness refers to the length of time that the multimedia can be saved in the multimedia library, such as 6 months or 3 months. The cover url refers to the address information of the multimedia cover. For example, the multimedia identification of a multimedia video with a length of less than 30 seconds can be set to 1, indicating a short video, and the multimedia identification of a multimedia video with a length of greater than or equal to 30 seconds can be set to 2. The cover clarity includes 0, 1, 2 or 3. Among them, 3 means high clarity and 0 means very poor clarity. Multimedia clarity refers to the clarity of multimedia content, such as the clarity of video content, and can also be 0, 1, 2 or 3. Among them, 3 means high clarity, and 0 means very poor clarity. The three vulgarity points refer to the scores of three vulgarities. If 0, it means there is no three vulgar content, if it is 1, it means there is a small amount of three vulgar content, and if it is 2, there is a large amount of three vulgar content. The title party mark refers to the degree of consistency between the title and the content. If it is 0, it is completely consistent and does not belong to the title party. If it is 1, it means that the title and the content are slightly inconsistent. If it is 2, it means that the title and the content are completely different, and it is a title party video.

[0116] S304, obtaining the current click rate corresponding to the training sample.

[0117] The current click rate can also be determined from the exposure log, which contains the number of exposures and clicks of the corresponding sample multimedia. The current click rate can be determined based on the number of clicks and the number of exposures.

[0118] In addition, if the training sample is determined based on the exposure log of the sample multimedia within a preset historical time (such as 10 days), the current click rate is also determined based on the exposure log of the corresponding sample multimedia within the preset historical time (10 days).

[0119] S305: Determine the actual click rate level corresponding to the training sample according to the current click rate and the preset mapping rule.

[0120] Among them, the current click rate and the actual click rate level are positively correlated, and the actual click rate level is used to represent the label data.

[0121] Specifically, the preset mapping rule may be that if the current click rate is less than 0.08, the actual click rate level is determined to be 0, and 0 is used as the label data of the training sample. If the click rate of the sample multimedia is greater than or equal to 0.08 and less than 0.13, the actual click rate level is determined to be 1, and 1 is used as the label data of the training sample. If the click rate of the sample multimedia is greater than or equal to 0.13, the actual click rate level is determined to be 2, and 2 is used as the label data of the training sample. In addition, the preset mapping rule may be other, for example, the click rate level may be divided into any multiple levels such as 0 to 3 or 0 to 4.

[0122] S306, input the training samples into the click rate prediction model to obtain probability values ​​of different click rate levels.

[0123] Among them, the click-through rate prediction model is a single-classification model, and the number of categories in the single-classification model is the same as the number of click-through rate level divisions.

[0124] In addition, if the click rate is divided into n levels (0, 1, 2, ..., n-1), the number of single-class models is n. When the training sample is input into the click rate prediction model, the probability values ​​of different click rate levels are obtained. For example, if the click rate is divided into 3 levels (0, 1, 2), the number of single-class models is 3, and then a training sample is input into the click rate prediction model, the probability value obtained is (0.1, 0.3, 0.7), then it can be understood that the probability of the click rate level corresponding to the training sample being 0 is 0.1, the probability of being 1 is 0.3, and the probability of being 2 is 0.7.

[0125] In the present disclosure, static attribute data at least includes: cover image of sample multimedia and text data of sample multimedia. Then, inputting the training sample into the click rate prediction model to obtain probability values ​​of different click rate levels includes: inputting the cover image into the convolutional neural network to obtain the first feature vector; inputting the text data into the embedding network to obtain the second feature vector; concatenating the first feature vector and the second feature vector to obtain the concatenated feature vector; inputting the concatenated feature vector into the fully connected layer to obtain the probability values ​​of different click rate levels.

[0126] Among them, the text data includes: title data, content data and associated data, the associated data is used to represent the quality of the sample multimedia, the embedding network includes: word embedding network, pre-trained embedding network and search embedding network, and the text data is input into the embedding network to obtain a second feature vector, including: inputting the title data into the word embedding network to obtain a first sub-feature vector; inputting the content data into the pre-trained embedding network to obtain a second sub-feature vector; inputting the associated data into the search embedding network to obtain a third sub-feature vector; and concatenating the first sub-feature vector, the second sub-feature vector and the third sub-feature vector to obtain a second feature vector.

[0127] Specifically, the cover image can be obtained according to the cover URL. The title data can refer to the multimedia title in Table 1. The content data can refer to at least one of the multimedia classification, interest points, keywords, tags, and publishing account names in Table 1. The associated data can refer to at least one of the publishing account ID, publishing account level, multimedia quality score, quality level, video timeliness, multimedia identification, cover clarity, multimedia clarity, three vulgarity scores, and clickbait identification in Table 1.

[0128] The structural diagram of the click rate prediction model is as follows: Figure 4 The click rate prediction model 40 includes: an input layer 41, a convolutional neural network 42, an embedding network 43, a concatenation layer 44, a fully connected layer 45 and an output layer 46.

[0129] In the present disclosure, the input layer inputs various static attribute data. Among them, the cover image is input into the convolutional neural network, and the convolutional neural network includes: the inception v3 model and the first fully connected sublayer (fc). The inception v3 model is used to extract features of the cover image to obtain a 1024-dimensional feature vector, and then the 1024-dimensional feature vector is passed through the first fully connected sublayer to obtain a 128-dimensional first feature vector.

[0130] The word embedding network includes: word processing network (word embedding) and aggregation network (nextvlad). The title data is segmented and added through word processing (word embedding) to obtain word vectors, and then the word vectors are aggregated through the aggregation network to obtain the first sub-feature vector of 128 dimensions.

[0131] In addition, the content data is processed through a pre-trained embedding network to obtain a second sub-feature vector with a dimension of 64. Specifically, since the content data includes multimedia classification, points of interest, etc., the vocabulary size is large, such as 200,000. Therefore, the pre-trained embedding network is used to initialize the vocabulary, and each content data is processed into a 64-dimensional feature vector.

[0132] Furthermore, the associated data is attribute data that has been processed twice, so the look-up embedding network is used for vectorization to obtain a third sub-feature vector with a dimension of 4. The first sub-feature vector, the second sub-feature vector and the third sub-feature vector are concatenated in the concatenation layer (concat) to obtain the second feature vector, and then the second feature vector and the first feature vector are concatenated in the concatenation layer to obtain a concatenated feature vector. The concatenated feature vector is then input into the fully connected layer (including two layers of fc) to obtain n probability values ​​with an output dimension of 128. Among them, the n probability values ​​correspond to the probability values ​​of n different click-through rate levels.

[0133] S307: Determine the click rate level corresponding to the maximum probability value among the probability values ​​as the training output click rate level.

[0134] For example, if the three probability values ​​are (0, 0, 1), the training output click rate level is 2. If the three probability values ​​are (0.8, 0.2, 0), the training output click rate level is 0. If the three probability values ​​are (0.1, 0.7, 0.2), the training output click rate level is 1.

[0135] S308: Determine a cross entropy loss function corresponding to the click rate prediction model according to the label data and the training output click rate level.

[0136] Among them, the label data is the actual output rate level. Among them, if the label data is 1, the corresponding actual probability value is (0, 1, 0), and the training output click rate level is the above 1, then according to the above corresponding output probability value is (0.1, 0.7, 0.2), then according to the actual probability value (0, 1, 0) and the output probability value (0.1, 0.7, 0.2), the cross entropy loss function can be determined. Among them, the use of the cross entropy loss function can make the trained click rate prediction model have better click rate prediction performance.

[0137] S309, adjusting the click rate prediction model according to the cross loss function to obtain a trained click rate prediction model.

[0138] In the present disclosure, the static attribute data of multimedia is used as training samples to train the click-through rate prediction model, which can expand the application scenarios of the click-through rate prediction model. There is no need to rely on the specific click behavior of users to train the click-through rate prediction model. For example, it is only necessary to obtain the click-through rate for the sample multimedia, without obtaining specific information such as the user account, age, and preference data of the clicked sample multimedia. In addition, the basic structure of the click-through rate prediction model of the present disclosure is a deep learning network with strong generalization ability, which can ensure the consistency of the offline click-through rate prediction model's test of multimedia click-through rate with the online trained click-through rate prediction model, thereby significantly improving the user's stay time on the page and the click-through rate.

[0139] Furthermore, the present disclosure fully considers the advantages of multimedia static attribute data and maximizes the advantages of static attribute data when training the click-through rate prediction model. In addition, due to the use of static attribute data of different dimensions and types, and in the structural design of the click-through rate prediction model, different types of static attribute data are extracted with different methods for features, and finally all features are integrated for training and outputting the prediction of the click-through rate level, which greatly improves the robustness and stability of the click-through rate prediction model.

[0140] Finally, the present disclosure can timely train and update the click rate prediction model with the latest training samples, so that the click rate prediction model can be consistent with the distribution of training samples generated online, making the click rate prediction model more timely and consistent with online multimedia.

[0141] Figure 5 A flowchart of a method for determining a click rate provided by the present disclosure is shown, which specifically includes the following steps:

[0142] S501: Obtain static attribute data of the multimedia to be detected.

[0143] The multimedia to be detected may be unexposed audio or video, etc. The static attribute data of the multimedia to be detected is the data carried by the multimedia to be detected, and is the data describing the properties of the multimedia to be detected.

[0144] S502: Input static attribute data of the multimedia to be detected into a click rate prediction model to obtain a target click rate level of the multimedia to be detected.

[0145] The static attribute data may be input into the trained click rate prediction model, and the target click rate level of the multimedia may be obtained.

[0146] Among them, the click rate prediction model is based on Figure 2 or Figure 3The target click rate level is obtained by training with a model training method, and is used to represent the click rate range of the multimedia to be detected.

[0147] In the present disclosure, the click-through rate prediction model obtained through the above training is used to predict the click-through rate, and the obtained target click-through rate level is the result corresponding to the multi-dimensional features. Therefore, the error of a single static attribute data has a weak impact on the target click-through rate level, thereby improving the accuracy of predicting the target click-through rate level using the click-through rate prediction model.

[0148] Figure 6 A flowchart showing another method for determining click rate provided by the present disclosure is shown, which specifically includes the following steps:

[0149] S601, obtaining static attribute data of the multimedia to be detected.

[0150] For the specific implementation process of this step, please refer to S501 and will not be repeated here.

[0151] S602: Input static attribute data of the multimedia to be detected into a click rate prediction model to obtain probability values ​​of the multimedia to be detected at different click rate levels.

[0152] In the present disclosure, the click rate prediction model is a single classification model. Exemplarily, after the static attribute data of the multimedia to be detected is input into the click rate prediction model, the output is (0.1, 0.2, 0.7), which means that the probability of the click rate level of the multimedia to be detected being level 0 is 0.1, the probability of being level 1 is 0.2, and the probability of being level 2 is 0.7.

[0153] S603: Determine the click rate level corresponding to the maximum probability value among the probability values ​​as the target click rate level.

[0154] For example, the probability values ​​of different click rate levels output by the click rate prediction model are (0.1, 0.2, 0.7), and the target click rate level of the multimedia to be detected can be determined to be 2. Correspondingly, the click rate range of the multimedia to be detected can be determined to be greater than 0.13.

[0155] S604: Determine the click rate score of the multimedia to be detected using the following formula.

[0156]

[0157] In the above formula, s represents the click-through rate score, t represents the set maximum click-through rate level, i represents the value of the click-through rate level, and p i It represents the probability value when the click rate level is i. The click rate score is used to represent the estimated click rate of the multimedia to be detected.

[0158] For example, if the set click rate levels are 0, 1 and 2, then t is the set maximum click rate level 2. i is 0, 1 and 2 respectively. If the probability values ​​of the multimedia A to be detected at different click rate levels are (0.1, 0.2, 0.7), then p0 is 0.1, p1 is 0.2, p2 is 0.7, and the click rate score s = 2.6. If the probability values ​​of the multimedia B to be detected at different click rate levels are (0.1, 0.1, 0.8), then p0 is 0.1, p1 is 0.1, p2 is 0.8, and the click rate score s = 2.7.

[0159] The click rate score can distinguish the size of the estimated click rate values ​​of the same click rate level. For example, the click rate levels of the above-mentioned multimedia A and multimedia B to be detected are both 2. However, the click rate score of the multimedia A to be detected, 2.6, is less than the click rate score of the multimedia B to be detected, 2.7. It can be determined that the estimated click rate value of the multimedia A to be detected is less than the estimated click rate value of the multimedia B to be detected. When sorting the click rates of the multimedia to be detected, they can be sorted according to the corresponding click rate scores. For example, Figure 1 Multimedia with higher click-through rate scores are placed at the front of the page.

[0160] In the present disclosure, static attribute data of the multimedia to be detected can be used to predict the click rate of the multimedia to be detected without considering the interaction between the user and the multimedia. In addition, the click rate score can be determined according to the probability values ​​at different click rate levels output by the click rate prediction model, and then it can be applied to the online service end, expanding the application scope of the present disclosure.

[0161] Exemplary Media

[0162] After introducing the method of the exemplary embodiment of the present disclosure, next, refer to Figure 7 A computer-readable storage medium according to an exemplary embodiment of the present disclosure is described as follows.

[0163] refer to Figure 7 As shown, a program product 70 for implementing the above method according to an embodiment of the present disclosure is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto.

[0164] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0165] The readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, wherein the readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable signal medium may also be any readable medium other than a readable storage medium.

[0166] Program code for performing the disclosed operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN).

[0167] Exemplary Devices

[0168] After introducing the medium of the exemplary embodiment of the present disclosure, next, reference is made to Figure 8 The model training device according to the exemplary embodiment of the present disclosure is described as follows.

[0169] Figure 8 The structural block diagram of a model training device 80 provided by the present disclosure is shown, and the model training device 80 includes: a first acquisition module 81, a second acquisition module 82, a click rate determination module 83 and an adjustment module 84. Among them:

[0170] A first acquisition module 81 is used to acquire a plurality of training samples, wherein the training samples include a plurality of static attribute data of the sample multimedia, and the static attribute data represents the properties of the sample multimedia itself;

[0171] A second acquisition module 82 is used to acquire label data corresponding to the training sample, where the label data is used to represent the actual click rate level of the sample multimedia;

[0172] A click rate determination module 83 is used to input the training sample into the click rate prediction model to obtain the training output click rate level;

[0173] The adjustment module 84 is used to adjust the click rate prediction model according to the label data and the training output click rate level to obtain a trained click rate prediction model.

[0174] In one embodiment of the present disclosure, the click rate prediction model is a single classification model, the number of classifications of the single classification model is the same as the number of click rate level divisions, and the click rate determination module 83 includes:

[0175] An input unit, used to input the training samples into the click rate prediction model to obtain probability values ​​of different click rate levels;

[0176] The determination unit is used to determine the click rate level corresponding to the maximum probability value in the probability values ​​as the training output click rate level.

[0177] In one embodiment of the present disclosure, the static attribute data includes: a cover image of the sample multimedia and text data of the sample multimedia; the input unit includes:

[0178] A first input subunit, used for inputting the cover image into the convolutional neural network to obtain a first feature vector;

[0179] A second input subunit, used for inputting text data into the embedding network to obtain a second feature vector;

[0180] A third input subunit is used to concatenate the first feature vector and the second feature vector to obtain a concatenated feature vector;

[0181] The fourth input subunit is used to input the concatenated feature vector into the fully connected layer to obtain probability values ​​of different click rate levels.

[0182] In one embodiment of the present disclosure, text data includes: title data, content data and associated data, the associated data is used to represent the quality of sample multimedia, the embedding network includes: word embedding network, pre-trained embedding network and search embedding network, and the second input sub-unit is specifically used to: input the title data into the word embedding network to obtain a first sub-feature vector; input the content data into the pre-trained embedding network to obtain a second sub-feature vector; input the associated data into the search embedding network to obtain a third sub-feature vector; and concatenate the first sub-feature vector, the second sub-feature vector and the third sub-feature vector to obtain a second feature vector.

[0183] In one embodiment of the present disclosure, the first acquisition module 81 includes:

[0184] An acquisition unit, used to acquire exposure logs of sample multimedia that have been exposed within a preset historical time;

[0185] A processing unit, used for preprocessing the exposure logs of the same sample multimedia to obtain non-repetitive backup data;

[0186] The screening unit is used to screen out training samples belonging to a preset sample dictionary from the spare data, where the preset sample dictionary is used to represent static properties of sample multimedia.

[0187] In one embodiment of the present disclosure, the acquisition unit is specifically used to: obtain the exposure log of the sample multimedia exposed within a preset historical time at every preset time, the exposure log is generated by the multimedia in the multimedia library when it is exposed, and the multimedia in the multimedia library is updated in real time.

[0188] In one embodiment of the present disclosure, the second acquisition module 82 is specifically used to: obtain the current click-through rate corresponding to the training sample; determine the actual click-through rate level corresponding to the training sample according to the current click-through rate and the preset mapping rules, the current click-through rate and the actual click-through rate level are positively correlated, and the actual click-through rate level is used to represent the label data.

[0189] In one embodiment of the present disclosure, the adjustment module 84 is specifically used to: determine the cross entropy loss function corresponding to the click-through rate prediction model based on the label data and the training output click-through rate level; adjust the click-through rate prediction model according to the cross loss function to obtain a trained click-through rate prediction model.

[0190] The model training device provided in the present disclosure can be executed Figure 2 and / or Figure 3 The model training method shown in the figure, for specific content, please refer to the description of the above model training method, which will not be repeated here.

[0191] Next, refer to Fig. 9 The click rate determination device according to the exemplary embodiment of the present disclosure is described as follows.

[0192] Fig. 9 The structural block diagram of a click rate determination device 90 provided by the present disclosure is shown, and the click rate determination device 90 includes: an acquisition module 91 and a click rate determination module 92. Among them:

[0193] An acquisition module 91 is used to acquire static attribute data of the multimedia to be detected;

[0194] The click rate determination module 92 is used to input the static attribute data of the multimedia to be detected into the click rate prediction model to obtain the target click rate level of the multimedia to be detected. The click rate prediction model is obtained by training according to any of the above-mentioned model training devices, and the target click rate level is used to represent the click rate range of the multimedia to be detected.

[0195] In one embodiment of the present disclosure, the click rate determination module 92 is specifically used to: input the static attribute data of the multimedia to be detected into the click rate prediction model to obtain the probability value of the multimedia to be detected at different click rate levels; determine the click rate level corresponding to the maximum probability value in the probability value as the target click rate level.

[0196] In one embodiment of the present disclosure, the click rate determination device further includes:

[0197] The determination module (not shown) is used to determine the click rate score of the multimedia to be detected using the following formula:

[0198]

[0199] In the above formula, s represents the click-through rate score, t represents the set maximum click-through rate level, i represents the value of the click-through rate level, and p i It represents the probability value when the click rate level is i. The click rate score is used to represent the estimated click rate of the multimedia to be detected.

[0200] The click rate determination device provided by the present disclosure can be executed Figure 5 and / or Figure 6 The click rate determination method shown in the figure, for specific content, please refer to the description of the above-mentioned click rate determination method, which will not be repeated here.

[0201] Exemplary Computing Devices

[0202] After introducing the method, medium and apparatus of the exemplary embodiments of the present disclosure, next, reference is made to Fig.10 A computing device according to an exemplary embodiment of the present disclosure is described.

[0203] Fig.10 The computing device 100 shown is merely an example and should not bring any limitation to the functionality and scope of use of the embodiments of the present disclosure.

[0204] like Fig.10 As shown, the computing device 100 is in the form of a general computing device. The components of the computing device 100 may include but are not limited to: at least one processing unit 101, at least one storage unit 102, and a bus 103 connecting different system components (including the processing unit 101 and the storage unit 102).

[0205] The bus 103 includes a data bus, a control bus, and an address bus.

[0206] The storage unit 102 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 1021 and / or a cache memory 1022 , and may further include a readable medium in the form of a non-volatile memory, such as a read-only memory (ROM) 1023 .

[0207] The storage unit 102 may also include a program / utility 1025 having a set (at least one) of program modules 1024, such program modules 1024 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0208] The computing device 100 may also communicate with one or more external devices 104 (e.g., a keyboard, a pointing device, etc.). Such communication may be performed via an input / output (I / O) interface 105. Furthermore, the computing device 100 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 106. Fig.10 As shown, the network adapter 106 communicates with other modules of the computing device 100 via the bus 103. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the computing device 100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0209] It should be noted that although several units / modules or sub-units / modules of the model training device / click rate determination device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to an embodiment of the present disclosure, the features and functions of two or more units / modules described above may be embodied in one unit / module. Conversely, the features and functions of one unit / module described above may be further divided to be embodied by multiple units / modules.

[0210] In addition, although the operations of the disclosed method are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0211] Although the spirit and principle of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims.

Claims

1. A model training method, comprising: Acquire multiple training samples, wherein the training samples include multiple static attribute data of sample multimedia, wherein the static attribute data indicates the properties of the sample multimedia itself, wherein the static attribute data includes: a cover image of the sample multimedia and text data of the sample multimedia, wherein the text data includes: title data, content data, and associated data, wherein the associated data is used to indicate the quality of the sample multimedia; Acquire label data corresponding to the training sample, wherein the label data is used to represent an actual click rate level of the sample multimedia, wherein the actual click rate level is determined according to a current click rate corresponding to the training sample and a preset mapping rule, and the current click rate is positively correlated with the actual click rate level; Input the training samples into the click-through rate prediction model to obtain probability values ​​of different click-through rate levels; determine the click-through rate level corresponding to the maximum probability value among the probability values ​​as the training output click-through rate level, the click-through rate prediction model is a single-classification model, and the number of classifications of the single-classification model is the same as the number of the click-through rate level divisions; According to the label data and the training output click rate level, adjusting the click rate prediction model to obtain a trained click rate prediction model; The step of inputting the training sample into the click-through rate prediction model to obtain probability values ​​of different click-through rate levels includes: inputting the cover image into a convolutional neural network to obtain a first feature vector; Inputting the text data into an embedding network to obtain a second feature vector, wherein the embedding network includes: a word embedding network, a pre-trained embedding network, and a search embedding network; After concatenating the first feature vector and the second feature vector, a concatenated feature vector is obtained; Inputting the concatenated feature vector into a fully connected layer to obtain probability values ​​of different click-through rate levels; The step of inputting the text data into an embedding network to obtain a second feature vector comprises: inputting the title data into the word embedding network to obtain a first sub-feature vector; Inputting the content data into the pre-trained embedding network to obtain a second sub-feature vector; Inputting the associated data into the search embedding network to obtain a third sub-feature vector; The second feature vector is obtained by concatenating the first sub-feature vector, the second sub-feature vector and the third sub-feature vector.

2. According to the model training method of claim 1, the obtaining of multiple training samples comprises: Obtain exposure logs of sample multimedia that have been exposed within a preset historical period; Preprocess the exposure logs of the same sample multimedia to obtain non-repetitive backup data; Training samples belonging to a preset sample dictionary are screened out from the spare data, and the preset sample dictionary is used to represent static properties of the sample multimedia.

3. According to the model training method of claim 2, the step of obtaining the exposure log of sample multimedia that has been exposed within a preset historical time period comprises: At every preset time, an exposure log of sample multimedia exposed within a preset historical time is obtained, wherein the exposure log is generated when the multimedia in the multimedia library is exposed, and the multimedia in the multimedia library is updated in real time.

4. The model training method according to claim 1, further comprising: Obtain the current click rate corresponding to the training sample.

5. The model training method according to claim 1, wherein adjusting the click rate prediction model according to the label data and the training output click rate level to obtain a trained click rate prediction model comprises: Determining a cross entropy loss function corresponding to the click rate prediction model according to the label data and the training output click rate level; The click-through rate prediction model is adjusted according to the cross entropy loss function to obtain a trained click-through rate prediction model.

6. A method for determining a click rate, comprising: Obtain static attribute data of the multimedia to be detected; The static attribute data of the multimedia to be detected is input into a click rate prediction model to obtain a target click rate level of the multimedia to be detected. The click rate prediction model is obtained by training according to the model training method according to any one of claims 1 to 5. The target click rate level is used to represent the click rate range of the multimedia to be detected.

7. The click rate determination method according to claim 6, wherein the step of inputting the static attribute data of the multimedia to be detected into the click rate prediction model to obtain the target click rate level of the multimedia to be detected comprises: Inputting the static attribute data of the multimedia to be detected into the click rate prediction model to obtain the probability value of the multimedia to be detected at different click rate levels; The click rate level corresponding to the maximum probability value among the probability values ​​is determined as the target click rate level.

8. The method for determining click rate according to claim 7, further comprising: The following formula is used to determine the click rate score of the multimedia to be detected: In the above formula, s represents the click-through rate score, t represents the set maximum click-through rate level, and i represents the value of the click-through rate level. It represents the probability value when the click rate level is i. The click rate score is used to represent the estimated click rate of the multimedia to be detected.

9. A model training device, comprising: A first acquisition module is used to acquire a plurality of training samples, wherein the training samples include a plurality of static attribute data of sample multimedia, wherein the static attribute data indicates the properties of the sample multimedia itself, wherein the static attribute data includes: a cover image of the sample multimedia and text data of the sample multimedia, wherein the text data includes: title data, content data and associated data, wherein the associated data is used to indicate the quality of the sample multimedia; A second acquisition module is used to acquire label data corresponding to the training sample, wherein the label data is used to represent an actual click rate level of the sample multimedia, wherein the actual click rate level is determined according to a current click rate corresponding to the training sample and a preset mapping rule, and the current click rate is positively correlated with the actual click rate level; A click rate determination module, used for inputting the training samples into a click rate prediction model to obtain a training output click rate level; An adjustment module, used for adjusting the click rate prediction model according to the label data and the training output click rate level, so as to obtain a trained click rate prediction model; The click rate determination module includes: an input unit, which is used to input the training sample into the click rate prediction model to obtain probability values ​​of different click rate levels, wherein the click rate prediction model is a single classification model, and the number of classifications of the single classification model is the same as the number of the click rate level divisions; A determination unit, configured to determine the click rate level corresponding to the maximum probability value among the probability values ​​as the training output click rate level; The input unit includes: a first input subunit, used to input the cover image into a convolutional neural network to obtain a first feature vector; A second input subunit, used for inputting the text data into an embedding network to obtain a second feature vector, wherein the embedding network includes: a word embedding network, a pre-trained embedding network and a search embedding network; A third input subunit, configured to concatenate the first feature vector and the second feature vector to obtain a concatenated feature vector; A fourth input subunit, used for inputting the concatenated feature vector into a fully connected layer to obtain probability values ​​of different click-through rate levels; The second input subunit is specifically used to: input the title data into the word embedding network to obtain a first sub-feature vector; Inputting the content data into the pre-trained embedding network to obtain a second sub-feature vector; Inputting the associated data into the search embedding network to obtain a third sub-feature vector; The second feature vector is obtained by concatenating the first sub-feature vector, the second sub-feature vector and the third sub-feature vector.

10. The model training device according to claim 9, wherein the first acquisition module comprises: An acquisition unit, used to acquire exposure logs of sample multimedia that have been exposed within a preset historical time; A processing unit, used for preprocessing the exposure logs of the same sample multimedia to obtain non-repetitive backup data; The screening unit is used to screen out training samples belonging to a preset sample dictionary from the spare data, and the preset sample dictionary is used to represent the static properties of the sample multimedia.

11. The model training device according to claim 10, wherein the acquisition unit is specifically used for: At every preset time, an exposure log of sample multimedia exposed within a preset historical time is obtained, wherein the exposure log is generated when the multimedia in the multimedia library is exposed, and the multimedia in the multimedia library is updated in real time.

12. The model training device according to claim 9, wherein the second acquisition module is further specifically used for: Obtain the current click rate corresponding to the training sample.

13. The model training device according to claim 9, wherein the adjustment module is specifically used for: Determining a cross entropy loss function corresponding to the click rate prediction model according to the label data and the training output click rate level; The click-through rate prediction model is adjusted according to the cross entropy loss function to obtain a trained click-through rate prediction model.

14. A click rate determination device, comprising: An acquisition module, used for acquiring static attribute data of the multimedia to be detected; A click rate determination module is used to input the static attribute data of the multimedia to be detected into a click rate prediction model to obtain a target click rate level of the multimedia to be detected. The click rate prediction model is obtained by training according to the model training device according to any one of claims 9 to 13, and the target click rate level is used to represent the click rate range of the multimedia to be detected.

15. The click rate determination device according to claim 14, wherein the click rate determination module is specifically configured to: Inputting the static attribute data of the multimedia to be detected into the click rate prediction model to obtain the probability value of the multimedia to be detected at different click rate levels; The click rate level corresponding to the maximum probability value among the probability values ​​is determined as the target click rate level.

16. The click rate determination device according to claim 15, further comprising: The determination module is used to determine the click rate score of the multimedia to be detected by using the following formula: In the above formula, s represents the click-through rate score, t represents the set maximum click-through rate level, and i represents the value of the click-through rate level. It represents the probability value when the click rate level is i. The click rate score is used to represent the estimated click rate of the multimedia to be detected.

17. A computer-readable storage medium, wherein computer program instructions are stored in the computer-readable storage medium, and when the computer program instructions are executed, the method according to any one of claims 1 to 8 is implemented.

18. A computing device comprising: Memory and processor; The memory is used to store program instructions; The processor is configured to call the program instructions in the memory to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Click rate prediction model training method, recommendation method, device and electronic equipment

    CN110378434A

  • Information processing method, recommendation method and related equipment

    CN110851713A