A CTR prediction method, device, and computer-readable storage medium
By performing statistical calculations and multi-dimensional image construction on advertising log data, combining CNN and internal component algorithms to cross local features, and combining FM algorithms to cross global features, the problem of difficulty in effectively learning sparse features in the existing technology is solved, and the prediction accuracy and training efficiency of the CTR prediction model are improved.
Patent Information
- Application Number
- CN202210314233.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-03-28
AI Technical Summary
The prior art is difficult to effectively learn sparse features when training CTR prediction models, resulting in unsatisfactory prediction results.
By performing statistical calculations and multi-dimensional image construction on the advertising log data, a training data set with rich features is generated, and the local features are crossed by CNN algorithm and internal product algorithm, and the global features are crossed by FM algorithm to train the estimated model.
The prediction accuracy of the estimated model is improved, the number of parameters is reduced, the difficulty of network training is reduced, the richness of feature space is improved, and data quality and model training efficiency are improved.
Smart Images

Figure CN114880920B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Internet advertising, and in particular, to a CTR prediction method, apparatus, and computer-readable storage medium. Background Art
[0002] With the development of the Internet, Internet advertising has become an important revenue channel for Internet enterprises. As a link between users and advertisers, the advertising system brings a personalized experience to users on the one hand, and brings huge commercial value to advertisers in terms of brand expansion, product promotion, and sales increase on the other hand. Click-Through Rate (CTR) prediction in the advertising system is to model based on the historical behavior data of users on advertisements, and predict the click probability of users on advertisements according to the requested users and advertisements. As an important part of the advertising system, improving the prediction ability of the CTR model can improve the quality of the advertising system, enhance the user experience, and improve the marketing quality of advertisers, creating higher value for Internet advertising enterprises.
[0003] Currently, the most commonly used solution for CTR prediction in the industry is to directly model based on historical structured data. Among the many network structures under it, the traditional method is to learn the interaction of low-order features and high-order features from the original features. For example, the Cross layer of DCN, the Product layer of PNN, the Attention layer of AFM, etc. The above methods combined with DNN can have a good effect on learning the interaction of all features. However, in the advertising CTR scenario, there are a large number of sparse data, that is, most of the useful interactions are sparse. Therefore, it is difficult for the above methods to efficiently learn them among a large number of parameters. Therefore, directly using the traditional feature interaction method has an unsatisfactory effect.
[0004] Aiming at the problem that sparse features cannot be effectively learned when training a model in the prior art, there is currently no effective solution. Summary of the Invention
[0005] To solve the above problems, the present invention provides a CTR prediction method. By statistically calculating, labeling, and constructing a multi-dimensional portrait of advertisement log data, a training data set with rich features is obtained. When training the features, the CNN algorithm and the inner product algorithm are used to train local features, and the FM algorithm is used to train global features to obtain a prediction model, thereby strengthening the attention to sparse features and solving the problem that sparse features cannot be effectively learned when training a model in the prior art.
[0006] To achieve the above object, the present invention provides a CTR prediction method, including: obtaining advertisement log data within a preset number of days, statistically calculating the advertisement log data according to a target key value to obtain a statistical calculation result, generating a label for each target key value according to a user click behavior, and merging the statistical calculation result corresponding to each target key value with the label to obtain a first data set; constructing a multi-dimensional portrait according to the advertisement log data, and using the multi-dimensional portrait as a second data set; using the target key value as a data identifier, merging the first data set and the second data set into a third data set; performing feature engineering processing on the third data set to obtain a training data set; performing local feature crossing on the training data set by using a CNN algorithm and an inner product algorithm, and performing global feature crossing on the training data set by using an FM algorithm to train and obtain a prediction model; using the prediction model to perform CTR prediction on the data to be measured to obtain a predicted CTR result.
[0007] Further optionally, the step of performing local feature crossing on the training data set by using a CNN algorithm and an inner product algorithm, and performing global feature crossing on the training data set by using an FM algorithm to train and obtain a prediction model includes: identifying sparse features and dense features in the training data set; sequentially performing feature crossing on the sparse features through a CNN algorithm and an inner product algorithm to obtain first feature data; splicing the dense features with the vectorized sparse features to obtain second feature data; performing feature crossing on the sparse features and the dense features to obtain third feature data; training the first feature data, the second feature data and the third feature data to obtain the prediction model.
[0008] Further optionally, the step of performing feature engineering processing on the third data set includes: performing missing value processing on the third data set; and / or performing feature selection on the third data set; and / or removing outliers from the third data set; and / or performing dimensionless processing on the third data set; and / or performing data correction on the third data set.
[0009] Further optionally, the target key value is a combination of a user ID, an advertisement ID and a media ID.
[0010] Further optionally, constructing a multi-dimensional portrait according to the advertisement log data includes: extracting user feature fields from the advertisement log data, and constructing a user dimension portrait according to the user feature fields; extracting advertisement feature fields from the advertisement log data, and constructing an advertisement dimension portrait according to the advertisement feature fields; extracting media feature fields from the advertisement log data, and constructing a media dimension portrait according to the media feature fields.
[0011] On the other hand, the present invention also provides a CTR prediction device, including: a first data set generation module, configured to obtain advertisement log data within a preset number of days, perform statistical calculations on the advertisement log data according to target key values, obtain a statistical calculation result, generate tags for each target key value according to user click behaviors, and merge the statistical calculation result corresponding to each target key value with the tags to obtain a first data set; a second data set generation module, configured to construct a multi-dimensional portrait according to the advertisement log data, and use the multi-dimensional portrait as a second data set; a third data set generation module, configured to use the target key value as a data identifier to merge the first data set and the second data set into a third data set; a training data set generation module, configured to perform feature engineering processing on the third data set to obtain a training data set; a prediction model training module, configured to perform local feature crossing on the training data set by using a CNN algorithm and an inner product algorithm, perform global feature crossing on the training data set by using an FM algorithm, and train to obtain a prediction model; a prediction module, configured to use the prediction model to perform CTR prediction on the data to be measured, and obtain a predicted CTR result.
[0012] Further optionally, the prediction model training module includes: a feature recognition sub-module, configured to recognize sparse features and dense features in the training data set; a local feature crossing sub-module, configured to perform feature crossing on the sparse features successively through a CNN algorithm and an inner product algorithm to obtain first feature data; a second feature data generation sub-module, configured to splice the dense features with the vectorized sparse features to obtain second feature data; a global feature crossing sub-module, configured to perform feature crossing on the sparse features and the dense features to obtain third feature data; a model training sub-module, configured to train the first feature data, the second feature data, and the third feature data to obtain the prediction model.
[0013] Further optionally, the training data set generation module includes: a missing value processing sub-module, configured to perform missing value processing on the third data set; a feature selection sub-module, configured to perform feature selection on the third data set; an outlier removal sub-module, configured to remove outliers from the third data set; a dimensionless processing sub-module, configured to perform dimensionless processing on the third data set; a data modification sub-module, configured to modify the data of the third data set.
[0014] Further optionally, the second data set generation module includes: a user dimension portrait generation sub-module, configured to extract user feature fields from the advertisement log data and construct a user dimension portrait according to the user feature fields; an advertisement dimension portrait generation sub-module, configured to extract advertisement feature fields from the advertisement log data and construct an advertisement dimension portrait according to the advertisement feature fields; a media dimension portrait generation sub-module, configured to extract media feature fields from the advertisement log data and construct a media dimension portrait according to the media feature fields.
[0015] On the other hand, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned CTR prediction method is implemented.
[0016] The above technical solution has the following beneficial effects: The present invention performs full-scale interaction on features through FM, performs interaction on local features through CNN, and strengthens local feature interaction by using the inner product method. It combines FM and CNN, making up for the problem of unsatisfactory effects in full-scale interaction of sparse features and improving the prediction accuracy of the prediction model; through the combination of CNN and the inner product, the number of parameters is reduced and the difficulty of network training is lowered; by processing advertisement log data, the first dataset and the second dataset, the richness of the feature space is enhanced; feature engineering processing is performed on the third dataset to improve data quality, reduce redundant data, and thus improve the prediction accuracy of the model and the efficiency of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a flowchart of the CTR prediction method provided by an embodiment of the present invention;
[0019] Figure 2 is a flowchart of the prediction model training method provided by an embodiment of the present invention;
[0020] Figure 3 is a schematic diagram of the FCINN network structure provided by an embodiment of the present invention;
[0021] Figure 4 is a flowchart of the feature engineering processing provided by an embodiment of the present invention;
[0022] Figure 5 is a flowchart of the multi-dimensional portrait construction method provided by an embodiment of the present invention;
[0023] Figure 6 is a schematic diagram of the structure of the CTR prediction device provided by an embodiment of the present invention;
[0024] Figure 7 is a schematic diagram of the structure of the prediction model training module provided by an embodiment of the present invention;
[0025] Figure 8It is a schematic structural diagram of a training dataset generation module provided by an embodiment of the present invention;
[0026] Figure 9 It is a schematic structural diagram of a second dataset generation module provided by an embodiment of the present invention.
[0027] Reference numerals: 100 - First dataset generation module; 200 - Second dataset generation module; 2001 - User dimension portrait generation sub-module; 2002 - Advertisement dimension portrait generation sub-module; 2003 - Media dimension portrait generation sub-module; 300 - Third dataset generation module; 400 - Training dataset generation module; 4001 - Missing value processing sub-module; 4002 - Feature selection sub-module; 4003 - Outlier removal sub-module; 4004 - Dimensionless sub-module; 4005 - Data modification sub-module; 500 - Estimation model training module; 5001 - Feature recognition sub-module; 5002 - Local feature crossing sub-module; 5003 - Second feature data generation sub-module; 5004 - Global feature crossing sub-module; 5005 - Model training sub-module; 600 - Prediction module Detailed implementation manners
[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0029] To solve the problem that sparse features cannot be effectively learned when training a CTR estimation model, an embodiment of the present invention provides a CTR estimation method Figure 1 It is a flowchart of the CTR estimation method provided by an embodiment of the present invention. As Figure 1 shown, the method includes:
[0030] S1. Obtain advertisement log data within a preset number of days, perform statistical calculations on the advertisement log data according to target key-value pairs to obtain a statistical calculation result, generate a label for each target key-value according to the user click behavior, and merge the statistical calculation result corresponding to each target key-value with the label to obtain a first dataset;
[0031] As an optional implementation manner, the preset number of days is 30 days, 15 days, 60 days, etc. The preset number of days can be determined according to actual requirements.
[0032] Specifically, the advertisement log data includes: user ID, advertisement ID, media ID, device number, advertisement type, advertisement space attribute, order ID, advertiser ID, settlement method, advertisement channel, request timestamp, event status, event type, traffic type, mobile platform, ID of the province where the advertisement is placed, advertisement source, channel type, platform type, device brand, general advertisement identifier, width of the advertisement space, height of the advertisement space, advertisement channel stream type.
[0033] As an alternative implementation, the above-mentioned statistical calculation of the advertisement log data according to the target key-value pair includes: calculating, for the data involved under the target key-value pair, the number of exposures of the user on the previous day, the number of clicks of the user on the previous day, the click-through rate of the user on the previous day, the number of clicks per hour of the user on the previous day, the number of exposures of the advertisement on the previous day, the number of clicks of the advertisement on the previous day, and the click-through rate of the advertisement on the previous day.
[0034] As an alternative implementation, generating labels for each target key-value pair according to the user's click behavior includes: if all the data under any target key-value pair are exposures, generating the label 0 as a negative sample; if all the data under any target key-value pair are clicks, generating the label 1 as a positive sample; if the data under any target key-value pair include both clicks and exposures, retaining all the clicks and generating the label 1 as a positive sample.
[0035] As an alternative implementation, distributed computing of the advertisement log data such as calculation is performed using a Spark cluster, and development languages can use Python, Scala, etc.
[0036] As an alternative implementation, the storage of the advertisement log data is implemented using HDFS, and it can also be implemented using HBase, Hive, etc.
[0037] As an alternative implementation, using the target key-value as the data identifier, the first data set obtained by merging the label with the statistical calculation result includes: user ID, number of exposures of the user on the previous day, number of clicks of the user on the previous day, click-through rate of the user on the previous day, number of clicks per hour of the user on the previous day, advertisement ID, number of exposures of the advertisement on the previous day, number of clicks of the advertisement on the previous day, click-through rate of the advertisement on the previous day, advertisement space ID.
[0038] S2. Construct a multi-dimensional portrait based on the advertisement log data and use the multi-dimensional portrait as the second data set;
[0039] Construct a product portrait using relevant fields in the advertisement log data. In this embodiment, a multi-dimensional product portrait is constructed according to multi-dimensional relevant fields to enrich the feature space. After the multi-dimensional portrait is generated, it is stored as the second data set.
[0040] As an alternative embodiment, the multi-dimensional portrait includes: a user dimension portrait, an advertisement dimension portrait, and a media dimension portrait. On this basis, the second data set fields include:
[0041] (1) User dimension: user ID, total number of user exposures, total number of user clicks, user historical click-through rate, number of user clicks per hour in history, user device brand, user traffic type, user mobile platform, user's most recent active time, whether the user's most recent active time is a holiday, number of user clicks on holidays, number of user clicks on weekends, number of user clicks on weekdays, whether the user's continuous status is retention, whether the user's continuous status is churn, number of user clicks during the previous day's rest time, number of user clicks during historical rest times, whether the user is a new user.
[0042] (2) Advertisement dimension: advertisement ID, advertiser ID, advertisement channel, channel type, advertisement type, advertisement source, general targeting advertisement identifier, advertisement slot width, advertisement slot height, advertisement slot width-to-height ratio, advertisement channel stream type, delivery terminal type ID, platform type, settlement method, delivery province ID, historical number of advertisement exposures, historical number of advertisement clicks, historical advertisement click-through rate.
[0043] (3) Media dimension: advertisement slot ID.
[0044] S3. Using the target key value as the data identifier, merge the first data set and the second data set into a third data set;
[0045] In this embodiment, the first data set and the second data set are merged into a wide table data, that is, the third data set, according to the target key value.
[0046] The third data set includes: user ID, number of user exposures on the previous day, number of user clicks on the previous day, user click-through rate on the previous day, number of user clicks per hour on the previous day, total number of user exposures, total number of user clicks, user historical click-through rate, number of user clicks per hour in history, user device brand, user traffic type, user mobile platform, user's most recent active time, whether the user's most recent active time is a holiday, number of user clicks on holidays, number of user clicks on weekends, number of user clicks on weekdays, whether the user's continuous status is retention, whether the user's continuous status is churn, number of user clicks during the previous day's rest time, number of user clicks during historical rest times, whether the user is a new user, advertisement ID, number of advertisement exposures on the previous day, number of advertisement clicks on the previous day, advertisement click-through rate on the previous day, advertisement slot ID, advertiser ID, advertisement channel, channel type, advertisement type, advertisement source, general targeting advertisement identifier, advertisement slot width, advertisement slot height, advertisement slot width-to-height ratio, advertisement channel stream type, delivery terminal type ID, platform type, settlement method, delivery province ID, historical number of advertisement exposures, historical number of advertisement clicks, historical advertisement click-through rate.
[0047] S4. Perform feature engineering on the third dataset to obtain a training dataset;
[0048] Perform feature engineering on the third dataset to optimize any feature in the third dataset and improve data quality.
[0049] The fields of the training dataset include: user ID, advertisement ID, advertiser ID, ad slot ID, number of exposures of the user on the previous day, number of clicks of the user on the previous day, click-through rate of the user on the previous day, number of clicks per hour of the user on the previous day, whether the previous day was a holiday, weekend, or weekday, number of clicks during the user's rest time on the previous day, number of exposures of the advertisement on the previous day, number of clicks of the advertisement on the previous day, click-through rate of the advertisement on the previous day; user device, hour when the request occurred, ad slot attributes, ID of the province where the ad is placed, width of the ad slot, height of the ad slot, aspect ratio of the ad slot width to height, type of the delivery terminal, total number of user exposures, total number of user clicks, user click-through rate, user's mobile platform, whether the user's continuous status is retention or churn, advertisement source, settlement method, number of advertisement exposures, number of advertisement clicks, advertisement click-through rate.
[0050] S5. Use the CNN algorithm and the inner product algorithm to perform local feature crossing on the training dataset, and use the FM algorithm to perform global feature crossing on the training dataset to train and obtain a prediction model;
[0051] Use the CNN algorithm to perform local feature crossing on the features in the training dataset. After that, add an inner product operation to overcome the deficiency that CNN can only learn neighbor feature interactions and strengthen local feature interactions at the same time.
[0052] Use the FM algorithm to perform second-order global feature crossing on the training dataset to ensure that all features can interact.
[0053] After obtaining the local crossing result of the CNN algorithm and the inner product and the global crossing result of the FM algorithm, concatenate the two crossing results and send them into the DNN to learn low-order and high-order feature interactions to obtain a prediction model.
[0054] To prevent the loss of original information, as an optional implementation, concatenate the local crossing result, the global crossing result, and the original information and send them into the DNN for training to obtain a prediction model. The prediction model in this embodiment is called the FCINN (FM CNN Inner Product Neural Network) model.
[0055] S6. Use the prediction model to perform CTR prediction on the data to be measured to obtain the predicted CTR result.
[0056] Input the target key value into the prediction model to obtain the predicted CTR result.
[0057] In this embodiment, the use of the FCINN model is packaged into a service interface. The service framework is implemented using TF Serving, and the service deployment is implemented using Docker.
[0058] As an alternative implementation Figure 2 is a flowchart of the prediction model training method provided by the embodiments of the present invention. As Figure 2 shown, the CNN algorithm and the inner product algorithm are used to perform local feature crossing on the training data set, and the FM algorithm is used to perform global feature crossing on the training data set. The trained prediction model includes:
[0059] S501. Identify sparse features and dense features in the training data set;
[0060] Figure 3 is a schematic diagram of the FCINN network structure provided by the embodiments of the present invention. As Figure 3 shown, the FCINN network structure includes an input part, a local crossing part, an input Concat part, a global feature crossing part, and a DNN part.
[0061] Input part: Process the data input part, where spare is the sparse feature input and dense is the dense feature input. Categorical features become very sparse after one-hot encoding, resulting in extremely high computational overhead. They are reduced in dimension through the Embedding method, that is, multiplying the one-hot vector by the Embedding matrix to obtain the Embedding vector corresponding to the feature. The Embedding matrix is backpropagated along with the entire network.
[0062] As an alternative implementation, categorical features can be directly Label-encoded and then Embedded. A certain index value after Label-encoding corresponds to a certain row vector of the Embedding matrix, which is used as the Embedding vector of this feature.
[0063] S502. Sequentially pass the sparse features through the CNN algorithm and the inner product algorithm for feature crossing to obtain the first feature data;
[0064] Local feature crossing part: It includes a CNN Layer and an Inner Product Layer.
[0065] CNN Layer: In this embodiment, to address the problem of difficult sparse feature interaction learning, local feature interaction is achieved by using CNN. The core of CNN lies in its convolutional kernel Kernel, which can be understood as a small window that moves in the matrix to perform local feature crossing through convolution. Based on translational invariance, parameter sharing for the local feature crossing part can be achieved using convolution. Additionally, a max pooling mechanism is added to reduce the number of parameters required for local interaction and lower the training difficulty.
[0066] In this embodiment, the CNN Layer performs standard convolution, pooling, and fully connected operations on the features. After parameter tuning, there are a total of four convolutions in the task scenario of this embodiment. The number of convolutional kernels for each convolution is 16, 18, 22, and 24 respectively, the convolutional kernel size is 7×1, and the max pooling size is 2×1.
[0067] As an alternative implementation, if the feature dimension is extremely large, such as several thousand or several ten thousand dimensions, the convolutional part can also be replaced with the structure after CNN variants, such as Alexnet, vgg series, etc.
[0068] Inner Product Layer: To overcome the limitation that CNN can only learn neighbor feature interactions, in this embodiment, local interaction is enhanced through the inner product method after CNN. In this embodiment, the inner product is the dot product of two vectors, that is, the Embedding representations after convolution are multiplied pairwise.
[0069] As an alternative implementation, in addition to the inner product operation, local interaction can also be enhanced through the outer product, inner product + outer product methods, which are determined according to the fitting metrics in different task scenarios.
[0070] The first feature data is the result of local crossing.
[0071] S503. Concatenate the dense feature and the vectorized sparse feature to obtain the second feature data;
[0072] Input Concat part: To prevent the loss of original information after global crossing and local crossing, the dense and the spare after Embedding are connected in parallel to obtain the second feature data, so as to be jointly fed into the DNN for training together with the local crossing result and the global crossing result, ensuring that the final DNN can learn the original information.
[0073] S504. Perform feature crossing on the sparse feature and the dense feature to obtain the third feature data;
[0074] Global feature cross part: In this embodiment, FM is used for global second-order feature cross. FM learns an Embedding vector for each feature, and the inner product of the vectors is used as the weight of the cross feature.
[0075] In this embodiment, using FM for feature cross is equivalent to treating all features equally. If the influence of different feature domains needs to be considered, the global feature cross part can also use FFM, FwFM, etc. for cross between feature domains.
[0076] The third feature data is the result of the global cross.
[0077] S505. Train the first feature data, the second feature data and the third feature data to obtain a prediction model.
[0078] DNN part: In this embodiment, finally, the local cross part, the global cross part and the original input part are concatenated, and three fully connected layers are used for non-linear fitting to learn low-order features and high-order features. Finally, the forward propagation is constructed through the output layer.
[0079] As an alternative implementation, after concatenation, an attention mechanism can be added to represent the importance of the interaction features and the original features to the target. The attention network can be a simple single fully connected layer, and then mapped into attention weights by Softmax. This attention network propagates backward along with the entire network.
[0080] As an alternative implementation, both the number of DNN layers and the number of neurons are hyperparameters. In this embodiment, after debugging in this task scenario, the effect of a three-layer network is the best, so it is three fully connected layers.
[0081] Through the first concatenation layer, the original data is concatenated, which ensures that the DNN layer can learn the original information and prevent information loss; through the FM layer, the interaction of all features can be performed, and the pairwise interaction of features can be automatically learned; through the CNN layer, the interaction of local features can be performed; through the combination of CNN and inner product, while strengthening the local feature interaction, the number of parameters can be reduced and the network training difficulty can be lowered; FCINN combines FM and CNN to make up for the problem that the full interaction effect of sparse features is not ideal. Finally, non-linear fitting is performed through the DNN layer to learn low-order and high-order feature interactions.
[0082] As an alternative implementation Figure 4 is the flowchart of the feature engineering processing provided by the embodiment of the present invention. As Figure 4 shown, the feature engineering processing of the third data set includes:
[0083] S401. Process the missing values in the third data set;
[0084] In this embodiment, the processing of missing values can be divided into two steps.
[0085] The first step is to delete the columns with serious missing values based on statistical methods. For example, advertisement types, platform types, etc.
[0086] The second step is to handle the missing values of the advertisement width and height. Since the advertisement positions are adapted according to the user's device, for each user, the mode of the advertisement width and height of this user is used for filling. If the advertisement width and height of a certain user are all missing, then the mode of the advertisement width and height of all users is taken for filling.
[0087] And / or S402, perform feature selection on the third data set;
[0088] In this embodiment, feature selection is all offline calculations, that is, the online process directly uses the selected features for operations. For feature selection in this embodiment, it can be divided into two steps.
[0089] The first step is to eliminate the features with low variance based on the variance of each feature.
[0090] The second step is to pre-train a tree model based on the wide table data, use the tree model to output the importance ranking of features, and eliminate the features with low importance.
[0091] For the above-mentioned tree model, XGBOOST can be used, or GBDT, CatBoost, etc. can be used to implement it.
[0092] For the above-mentioned feature importance ranking, the number of splits of each feature of each tree can be used for ranking. The more the number of splits, the more important the feature; or the GINI value of each split node of each tree can be used for ranking. The smaller the GINI value, the more important this split feature is.
[0093] And / or S403, eliminate outliers from the third data set;
[0094] Taking the exposure and click fields as examples for outlier elimination in this embodiment: the total number of user exposures, the total number of user clicks, the number of user exposures on the previous day, the number of user clicks on the previous day, the total number of advertisement exposures, the total number of advertisement clicks, the number of advertisement exposures on the previous day, the number of advertisement clicks on the previous day. For outlier elimination in this embodiment, it can be divided into two steps.
[0095] The first step is to eliminate the fields with exposure and click being 0.
[0096] The second step is to eliminate outliers from the exposure and click fields. First, standardize the exposure and click fields, and then eliminate the outliers exceeding the upper and lower limits based on their distributions.
[0097] The above-mentioned outlier removal can be implemented using Pandas based on Boxplot, or can also be implemented using DBSCAN, LOF, etc.
[0098] and / or S404, dimensionlessize the third dataset;
[0099] In this embodiment, the dimensionlessization process can be divided into two steps.
[0100] The first step is to unify the precision. For example, unify the precision of all floating-point types to Float32.
[0101] The second step is to perform normalization processing on the total number of user exposures, the total number of user clicks, the number of user exposures on the previous day, the number of user clicks on the previous day, the total number of ad exposures, the total number of ad clicks, the number of ad exposures on the previous day, and the number of ad clicks on the previous day.
[0102] The above-mentioned normalization formula is: The normalization formula mapped to the [0, 1] interval is: where X refers to a specific column of features, X min refers to the minimum feature value in a certain feature domain, X max refers to the maximum feature value in a certain feature domain.
[0103] and / or S405, perform data correction on the third dataset.
[0104] In this embodiment, taking the click-through rate as an example for data correction: the user's historical click-through rate, the user's click-through rate on the previous day, the ad's historical click-through rate, and the ad's click-through rate on the previous day. In this embodiment, Wilson correction is adopted. Specifically, using the total number of user exposures, the total number of user exposures on the previous day, the total number of ad exposures, and the total number of ad exposures on the previous day, calculate their respective Wilson intervals, and use the lower limit of their respective Wilson intervals to correct the click-through rate to obtain the user's click-through rate after Wilson correction, the user's click-through rate on the previous day after Wilson correction, the ad's click-through rate after Wilson correction, and the ad's click-through rate on the previous day after Wilson correction.
[0105] The above-mentioned Wilson correction formula is: where p is the probability, that is, the click-through rate CTR; n is the total number of samples, which refers to the number of exposures here; z represents the statistic corresponding to a certain confidence interval. In the normal distribution, μ + z * σ will have a certain confidence level. For example, when z = 1.96, there is a 95% confidence level. The meaning of the Wilson interval is the range of the true CTR under a certain confidence level.
[0106] As a preferred implementation manner, perform missing value processing, feature selection, outlier removal, dimensionlessization, and data correction on any feature in sequence.
[0107] As an alternative implementation, the target key value is a combination of the user ID, the advertisement ID, and the media ID.
[0108] As an alternative implementation, Figure 5 is a flowchart of the multi-dimensional portrait construction method provided by the embodiments of the present invention. As Figure 5 shown, constructing a multi-dimensional portrait based on advertisement log data includes:
[0109] S201. Extract user feature fields from the advertisement log data, and construct a user-dimensional portrait according to the user feature fields;
[0110] The multi-dimensional portrait includes a user-dimensional portrait, an advertisement-dimensional portrait, and a media-dimensional portrait.
[0111] Among them, the construction of the user-dimensional portrait includes: obtaining the fields required for the construction of the user dimension, for example, the number of exposures of the user on the previous day, the number of clicks of the user on the previous day, the user device, the time period where the request is located, etc. Depict the portrait and derive the features of the user dimension according to the mapping, analysis, and statistical methods.
[0112] As an alternative implementation, the portrait fields, construction methods, and storage methods of the user dimension are shown in Table 1:
[0113]
[0114]
[0115] Table 1
[0116] In the above-mentioned fields "the number of clicks during the rest time of the user on the previous day" and "the number of clicks during the rest time of the user's history", a specific implementation example of the rest time period is: [12:00 - 14:00, 19:00 - 23:00, 0:00 - 8:00].
[0117] S202. Extract advertisement feature fields from the advertisement log data, and construct an advertisement-dimensional portrait according to the advertisement feature fields;
[0118] The construction of the advertisement-dimensional portrait includes: obtaining the fields required for the construction of the advertisement dimension, for example, the number of exposures of the advertisement on the previous day, the number of clicks of the advertisement on the previous day, the advertisement source, etc. When calculating and generating the numerical values of each field, in order to avoid the problem of data skew, the calculation method of aggregating both ends can be used to solve it. Depict the portrait and derive the features of the advertisement dimension according to the mapping, analysis, and statistical methods.
[0119] As an alternative implementation, the portrait fields, construction methods, and storage methods of the advertisement dimension are shown in Table 2:
[0120]
[0121]
[0122] Table 2
[0123] The above-mentioned two-end aggregation refers to one of the methods for solving data skew in Spark distributed computing. As an optional implementation of this embodiment, salt is added to the originally identical Keys with random numbers to make them different Keys, so that the data originally processed by one Task is spread to multiple Tasks for local aggregation, avoiding memory kill caused by excessive data in a single Task. Then, the random prefix is removed for global aggregation.
[0124] S203. Extract the media feature fields from the advertising log data, and construct a media dimension portrait according to the media feature fields.
[0125] The construction of the media dimension portrait includes: obtaining the fields required for media dimension construction, and depicting the portrait according to the mapping method: ad slot ID.
[0126] As an optional implementation, the portrait fields, construction methods, and storage methods of the media dimension are shown in Table 3:
[0127] Image field Construction method Storage example Ad slot ID Mapping {"apid":"2252"}
[0128] Table 3
[0129] As an optional implementation, update the above-mentioned user dimension portrait, advertising dimension portrait, and media dimension portrait to the user dimension table, advertising dimension table, and media dimension table respectively to complete the construction of the product portrait.
[0130] The above-mentioned product portrait storage tool can be implemented using MongoDB, or can also be implemented using HBase, Elasticsearch, etc.
[0131] The above-mentioned "mapping" refers to mapping a string or numerical value to a specified value for storage. For example, in the user traffic type, app is mapped to 1, and pc is mapped to 2; or it is stored in the way of identity mapping. For example, 278191 in the advertising ID is stored identically as 278191.
[0132] The above-mentioned "statistics" refers to counting the number of times of a user's specified behavior within a certain period of time. If a certain behavior feature has multiple feature values, they are respectively counted and summarized according to different methods.
[0133] The above-mentioned "analysis" refers to the time attribute of the user's behavior within a specified time period. Using a method defined by humans, new, discarded, or updated measures are taken according to the time attributes corresponding to different behaviors.
[0134] As an alternative implementation, an embodiment of the present invention further provides a CTR prediction device. Figure 6 It is a schematic structural diagram of the CTR prediction device provided by an embodiment of the present invention. As Figure 6 shown, the device includes:
[0135] A first data set generation module 100, configured to obtain advertisement log data within a preset number of days, perform statistical calculations on the advertisement log data according to target key-value pairs, obtain a statistical calculation result, generate labels for each target key-value according to user click behavior, and merge the corresponding statistical calculation result and label for each target key-value to obtain a first data set;
[0136] As an alternative implementation, the preset number of days is 30 days, 15 days, 60 days, etc. The preset number of days can be determined according to actual needs.
[0137] Specifically, the advertisement log data includes: user ID, advertisement ID, media ID, device number, advertisement type, advertisement space attribute, order ID, advertiser ID, settlement method, advertisement channel, request timestamp, event status, event type, traffic type, mobile platform, delivery province ID, advertisement source, channel type, platform type, device brand, general delivery advertisement identifier, advertisement space width, advertisement space height, advertisement channel stream type.
[0138] As an alternative implementation, performing statistical calculations on the advertisement log data according to target key-value pairs includes: calculating, for the data involved under the target key-value pair: the number of user exposures on the previous day, the number of user clicks on the previous day, the user click-through rate on the previous day, the number of user clicks per hour on the previous day, the number of advertisement exposures on the previous day, the number of advertisement clicks on the previous day, and the advertisement click-through rate on the previous day.
[0139] As an alternative implementation, generating labels for each target key-value according to user click behavior includes: if all the data under any target key-value pair are exposures, generating a label 0 as a negative sample; if all the data under any target key-value pair are clicks, generating a label 1 as a positive sample; if the data under any target key-value pair include both clicks and exposures, retaining all clicks and generating a label 1 as a positive sample.
[0140] As an alternative implementation, distributed calculations such as calculations on the advertisement log data are performed using a Spark cluster, and development languages can use Python, Scala, etc.
[0141] As an alternative implementation, the storage of the advertisement log data is implemented using HDFS, and can also be implemented using HBase, Hive, etc.
[0142] As an alternative implementation, taking the target key-value as the data identifier, the first data set obtained by merging the label and the statistical calculation result includes: user ID, the number of exposures of the user on the previous day, the number of clicks of the user on the previous day, the click-through rate of the user on the previous day, the number of clicks per hour of the user on the previous day, advertisement ID, the number of exposures of the advertisement on the previous day, the number of clicks of the advertisement on the previous day, the click-through rate of the advertisement on the previous day, and advertisement position ID.
[0143] The second data set generation module 200 is used to construct a multi-dimensional portrait based on the advertisement log data and use the multi-dimensional portrait as the second data set;
[0144] Use relevant fields in the advertisement log data to construct a product portrait. In this embodiment, construct a multi-dimensional product portrait based on multi-dimensional relevant fields to enrich the feature space. After the multi-dimensional portrait is generated, it is stored as the second data set.
[0145] As an alternative implementation, the multi-dimensional portrait includes: user dimension portrait, advertisement dimension portrait, and media dimension portrait. On this basis, the fields of the second data set include:
[0146] (1) User dimension: user ID, total number of user exposures, total number of user clicks, user historical click-through rate, user historical number of clicks per hour, user device brand, user traffic type, user mobile platform, user's most recent active time, whether the user's most recent active time is a holiday, number of clicks of the user on holidays, number of clicks of the user on weekends, number of clicks of the user on weekdays, whether the user's continuous status is retention, whether the user's continuous status is churn, number of clicks of the user during the rest time on the previous day, number of clicks of the user during the rest time in history, whether the user is a new user.
[0147] (2) Advertisement dimension: advertisement ID, advertiser ID, advertisement channel, channel type, advertisement type, advertisement source, general targeting advertisement identifier, width of the advertisement position, height of the advertisement position, width-to-height ratio of the advertisement position, advertisement channel stream type, ID of the delivery terminal type, platform type, settlement method, ID of the province where the advertisement is delivered, historical number of exposures of the advertisement, historical number of clicks of the advertisement, historical click-through rate of the advertisement.
[0148] (3) Media dimension: advertisement position ID.
[0149] The third data set generation module 300 is used to merge the first data set and the second data set into a third data set with the target key-value as the data identifier;
[0150] In this embodiment, the first data set and the second data set are merged into a wide table data, that is, the third data set, according to the target key-value.
[0151] The third data set includes: user ID, the number of exposures of the user on the previous day, the number of clicks of the user on the previous day, the click-through rate of the user on the previous day, the number of clicks per hour of the user on the previous day, the total number of exposures of the user, the total number of clicks of the user, the historical click-through rate of the user, the number of clicks per hour of the user's history, the brand of the user's device, the type of user traffic, the user's mobile phone platform, the time of the user's most recent activity, whether the user's most recent activity time is a holiday, the number of clicks of the user on holidays, the number of clicks of the user on weekends, the number of clicks of the user on weekdays, whether the user's continuous status is retention, whether the user's continuous status is churn, the number of clicks of the user during rest time on the previous day, the number of clicks of the user during rest time in history, whether the user is a new user, ad ID, the number of exposures of the ad on the previous day, the number of clicks of the ad on the previous day, the click-through rate of the ad on the previous day, ad slot ID, advertiser ID, ad channel, channel type, ad type, ad source, general targeting ad flag, ad slot width, ad slot height, ad slot width-to-height ratio, ad channel stream type, ID of the delivery terminal type, platform type, settlement method, ID of the province where the ad is delivered, the historical number of exposures of the ad, the historical number of clicks of the ad, the historical click-through rate of the ad.
[0152] The training data set generation module 400 is used to perform feature engineering processing on the third data set to obtain a training data set;
[0153] Perform feature engineering processing on the third data set to optimize any feature in the third data set and improve data quality.
[0154] The fields of the training data set include: user ID, ad ID, advertiser ID, ad slot ID, the number of exposures of the user on the previous day, the number of clicks of the user on the previous day, the click-through rate of the user on the previous day, the number of clicks per hour of the user on the previous day, whether the previous day is a holiday, weekend, or weekday, the number of clicks of the user during rest time on the previous day, the number of exposures of the ad on the previous day, the number of clicks of the ad on the previous day, the click-through rate of the ad on the previous day; user device, hour when the request occurs, ad slot attributes, ID of the province where the ad is delivered, ad slot width, ad slot height, ad slot width-to-height ratio, delivery terminal type, total number of exposures of the user, total number of clicks of the user, user click-through rate, user mobile phone platform, whether the user's continuous status is retention, whether the user's continuous status is churn, ad source, settlement method, number of ad exposures, number of ad clicks, ad click-through rate.
[0155] The prediction model training module 500 is used to perform local feature crossing on the training data set using the CNN algorithm and the inner product algorithm, and perform global feature crossing on the training data set using the FM algorithm to train and obtain a prediction model;
[0156] Perform local feature crossing on the features in the training data set using the CNN algorithm. After that, add an inner product operation to overcome the deficiency that CNN can only learn neighbor feature interactions and strengthen local feature interactions at the same time.
[0157] The FM algorithm is used to perform second-order cross of global features on the training dataset to ensure that all features can interact.
[0158] After obtaining the local cross results of the CNN algorithm and the inner product and the global cross results of the FM algorithm, the two cross results are connected in parallel and fed into the DNN to learn low-order and high-order feature interactions, obtaining a prediction model.
[0159] To prevent the loss of original information, as an alternative implementation, the local cross results, the global cross results, and the original information are connected in parallel and fed into the DNN for training to obtain a prediction model. The prediction model in this embodiment is called the FCINN (FM CNN Inner Product Neural Network) model.
[0160] The prediction module 600 is used to perform CTR prediction on the data to be measured using the prediction model to obtain the predicted CTR result.
[0161] By inputting the target key value into the prediction model, the predicted CTR result can be obtained.
[0162] In this embodiment, the use of the FCINN model is packaged into a service interface, the service framework is implemented using TF Serving, and the service deployment is implemented using Docker.
[0163] As an alternative implementation, Figure 7 is the structural schematic diagram of the prediction model training module provided by the embodiment of the present invention. As shown in Figure 7 shown, the prediction model training module 500 includes:
[0164] The feature recognition sub-module 5001 is used to recognize sparse features and dense features in the training dataset;
[0165] The FCINN network structure includes an input part, a local cross part, an input Concat part, a global feature cross part, and a DNN part.
[0166] Input part: Process the data input part, where spare is the sparse feature input and dense is the dense feature input. Categorical features become very sparse after one-hot encoding, resulting in extremely high computational overhead. They are dimensionally reduced through the Embedding method, that is, multiplying the one-hot vector by the Embedding matrix to obtain the Embedding vector corresponding to the feature. The Embedding matrix is backpropagated along with the entire network.
[0167] As an alternative implementation, the categorical features can be directly Label-encoded and then Embedded. A certain index value after Label-encoding corresponds to a certain row vector of the Embedding matrix, which is used as the Embedding vector of this feature.
[0168] The local feature cross sub-module 5002 is used to perform feature crossing on the sparse features through the CNN algorithm and the inner product algorithm in sequence to obtain the first feature data;
[0169] Local feature cross part: It includes a CNN Layer and an Inner Product Layer.
[0170] CNN Layer: In this embodiment, to address the problem of difficult interaction learning for sparse features, local feature interaction is achieved by using CNN. The core of CNN lies in its convolution kernel Kernel, which can be understood as a small window that moves in the matrix to perform local feature crossing through convolution. Based on translational invariance, parameter sharing for the local feature cross part can be achieved using convolution. And a max pooling mechanism is added to reduce the number of parameters required for local interaction and lower the training difficulty.
[0171] In this embodiment, the CNN Layer performs standard convolution, pooling, and fully connected operations on the features. After tuning the parameters, there are a total of four convolutions in the task scenario of this embodiment. The number of convolution kernels for each convolution is 16, 18, 22, and 24 respectively, the convolution kernel size is 7×1, and the max pooling size is 2×1.
[0172] As an alternative implementation, if the feature dimension is particularly large, such as several thousand or several ten thousand dimensions, the convolution part can also be replaced with the structure after the CNN variant, such as Alexnet, vgg series, etc.
[0173] Inner Product Layer: To overcome the deficiency that CNN can only learn neighbor feature interaction, in this embodiment, after CNN, local interaction is enhanced through the inner product method. In this embodiment, the inner product is the dot product of two vectors, that is, the Embedding representations after convolution are multiplied pairwise.
[0174] As an alternative implementation, in addition to the inner product operation, local interaction can also be enhanced through the outer product, inner product + outer product methods, which are determined according to the fitting metrics in different task scenarios.
[0175] The first feature data is the local cross result.
[0176] The second feature data generation sub-module 5003 is used to concatenate the dense features and the vectorized sparse features to obtain the second feature data;
[0177] Input Concat part: To prevent the loss of original information after global and local crosses, the sparse data after dense and Embedding are connected in parallel to obtain the second feature data, so as to be jointly fed into the DNN for training together with the local cross result and the global cross result later, ensuring that the final DNN can learn the original information.
[0178] Global feature cross sub-module 5004 is used to perform feature cross on sparse features and dense features to obtain the third feature data;
[0179] Global feature cross part: In this embodiment, the FM is used for the second-order global feature cross. The FM learns an Embedding vector for each feature, and the inner product of the vectors is used as the weight of the cross feature.
[0180] In the scenario of this embodiment, using FM for feature cross is equivalent to treating all features equally. If the influence of different feature domains is to be considered, the global feature cross part can also use FFM, FwFM, etc. for cross between feature domains.
[0181] The third feature data is the result of the global cross.
[0182] Model training sub-module 5005 is used to train the first feature data, the second feature data and the third feature data to obtain a prediction model.
[0183] DNN part: In this embodiment, finally, the local cross part, the global cross part and the original input part are concatenated, and 3 fully connected layers are used for non-linear fitting to learn low-order features and high-order features. Finally, the forward propagation is constructed through the output layer.
[0184] As an optional implementation, after Concat, an attention mechanism can be added to represent the importance of the interaction features and the original features to the target. The attention network can be a simple single fully connected layer, and then mapped into attention weights by Softmax. This attention network propagates backward together with the whole network.
[0185] As an optional implementation, the number of DNN layers and neurons are both hyperparameters. In this task scenario of this embodiment, the 3-layer network has the best effect after debugging, so it is 3 fully connected layers.
[0186] Through the first Concat layer, the original data is concatenated, ensuring that the DNN layer can learn the original information and prevent information loss; through the FM layer, all features can be interacted, and the pairwise interactions of features can be automatically learned; through the CNN layer, local feature interactions can be performed; through the combination of CNN and inner product, while strengthening the local feature interactions, the number of parameters can be reduced and the network training difficulty can be lowered; FCINN combines FM and CNN to make up for the problem of unsatisfactory all-feature interactions of sparse features, and finally, the DNN layer performs non-linear fitting to learn low-order and high-order feature interactions.
[0187] As an alternative implementation Figure 8 is a schematic structural diagram of the training dataset generation module provided by an embodiment of the present invention, as Figure 8 shown, the training dataset generation module 400 includes:
[0188] The missing value processing sub-module 4001 is used to process the missing values in the third dataset;
[0189] In this embodiment, the processing of missing values can be divided into two steps.
[0190] The first step is to delete the columns with serious missing values based on statistical methods, for example, advertisement types, platform types, etc.
[0191] The second step is to process the missing values of the advertisement width and height. Since the advertisement space will be adapted according to the user's device, for each user, the mode of the advertisement width and height of this user is used for filling. If the advertisement width and height of a certain user are all missing, the mode of the advertisement width and height of all users is taken for filling.
[0192] The feature selection sub-module 4002 is used to perform feature selection on the third dataset;
[0193] In this embodiment, feature selection is all offline calculations, that is, the online process directly uses the screened features for operations. Among them, the feature selection in this embodiment can be divided into two steps.
[0194] The first step is to eliminate the features with low variance based on the variance of each feature.
[0195] The second step is to pre-train a tree model based on the wide table data, use the tree model to output the importance ranking of features, and eliminate the features with low importance.
[0196] The above-mentioned tree model can use XGBOOST, or can be implemented using GBDT, CatBoost, etc.
[0197] For the above-mentioned sorting of feature importance, the number of splits of each feature of each tree can be used for sorting. The more splits, the more important the feature. Alternatively, the GINI value of each split node of each tree can be used for sorting. The smaller the GINI value, the more important the split feature.
[0198] The outlier removal sub-module 4003 is used to remove outliers from the third data set;
[0199] In this embodiment, the outlier removal takes the exposure and click fields as examples: the total number of user exposures, the total number of user clicks, the number of user exposures on the previous day, the number of user clicks on the previous day, the total number of ad exposures, the total number of ad clicks, the number of ad exposures on the previous day, and the number of ad clicks on the previous day. For outlier removal, this embodiment can be divided into two steps.
[0200] The first step is to remove the fields where the exposure and click are 0.
[0201] The second step is to remove the outliers from the exposure and click fields. First, standardize the exposure and click fields, and then remove the outliers that exceed the upper and lower limits based on their distribution.
[0202] For the above-mentioned outlier removal, it can be implemented using Pandas based on Boxplot, or it can also be implemented using DBSCAN, LOF, etc.
[0203] The dimensionless sub-module 4004 is used to make the third data set dimensionless;
[0204] In this embodiment, the dimensionless processing can be divided into two steps.
[0205] The first step is to unify the precision. For example, unify the precision of all floating-point types to Float32.
[0206] The second step is to perform normalization on the total number of user exposures, the total number of user clicks, the number of user exposures on the previous day, the number of user clicks on the previous day, the total number of ad exposures, the total number of ad clicks, the number of ad exposures on the previous day, and the number of ad clicks on the previous day.
[0207] The above-mentioned normalization formula is: The normalization formula for mapping to the [0,1] interval is: where X refers to a specific column feature, X min refers to the smallest feature value in a certain feature domain, X max refers to the largest feature value in a certain feature domain.
[0208] The data modification sub-module 4005 is used to correct the data of the third data set.
[0209] In this embodiment, the data correction takes the click-through rate as an example: the historical click-through rate of users, the click-through rate of users on the previous day, the historical click-through rate of advertisements, and the click-through rate of advertisements on the previous day. In this embodiment, Wilson correction is adopted. Specifically, by using the total number of user exposures, the total number of user exposures on the previous day, the total number of advertisement exposures, and the total number of advertisement exposures on the previous day, the respective Wilson intervals are calculated, and the lower limits of the respective Wilson intervals are used to correct the click-through rate, obtaining the user click-through rate after Wilson correction, the user click-through rate on the previous day after Wilson correction, the advertisement click-through rate after Wilson correction, and the advertisement click-through rate on the previous day after Wilson correction.
[0210] The above-mentioned Wilson correction formula is as follows: Among them, p is the probability, that is, the click-through rate CTR; n is the total number of samples, which refers to the number of exposures here; z represents the statistic corresponding to a certain confidence interval. In the normal distribution, μ + z * σ will have a certain confidence level. For example, when z = 1.96, there is a 95% confidence level. The meaning of the Wilson interval is the range of the true CTR under a certain confidence level.
[0211] As a preferred implementation manner, missing value processing, feature selection, outlier removal, dimensionless processing, and data correction are sequentially performed on any feature.
[0212] As an alternative implementation manner, the target key value is a combination of user ID, advertisement ID, and media ID.
[0213] As an alternative implementation manner, Figure 9 is a schematic structural diagram of the second data set generation module provided by an embodiment of the present invention, as Figure 9 shown. The second data set generation module 200 includes:
[0214] The user dimension portrait generation sub-module 2001 is used to extract user feature fields from the advertisement log data and construct a user dimension portrait according to the user feature fields;
[0215] The multi-dimensional portrait includes a user dimension portrait, an advertisement dimension portrait, and a media dimension portrait.
[0216] Among them, the construction of the user dimension portrait includes: obtaining the fields required for the construction of the user dimension, for example, the number of user exposures on the previous day, the number of user clicks on the previous day, the user device, the time period where the request is located, etc. The portrait is depicted and the features of the user dimension are derived according to the mapping, analysis, and statistical methods.
[0217] As an alternative implementation manner, the portrait fields, construction methods, and storage methods of the user dimension are shown in Table 1.
[0218] In the above-mentioned fields "number of clicks by the user during the previous day's rest time" and "number of clicks by the user during historical rest time", a specific implementation example of the rest time period is, for example: [12:00 - 14:00, 19:00 - 23:00, 0:00 - 8:00].
[0219] The advertisement dimension portrait generation sub-module 2002 is used to extract advertisement feature fields from advertisement log data and construct an advertisement dimension portrait according to the advertisement feature fields.
[0220] The construction of the advertisement dimension portrait includes: obtaining the fields required for constructing the advertisement dimension, for example, the number of exposures of the advertisement on the previous day, the number of clicks of the advertisement on the previous day, the advertisement source, etc. When calculating and generating the values of each field, to avoid the problem of data skew, the calculation method of aggregating at both ends can be used to solve it, and the portrait is depicted and the features of the advertisement dimension are derived according to the mapping, analysis, and statistical methods.
[0221] As an optional implementation method, the portrait fields, construction methods, and storage methods of the advertisement dimension are shown in Table 2.
[0222] The above-mentioned aggregation at both ends refers to one of the methods for solving data skew in Spark distributed computing. As an optional implementation method of this embodiment, random numbers are salted to the originally same Key to make it different Keys, so that the data originally processed by one Task is spread to multiple Tasks for local aggregation, avoiding memory Kill caused by excessive data of a single Task, and then the random prefix is removed for global aggregation.
[0223] The media dimension portrait generation sub-module 2003 is used to extract media feature fields from advertisement log data and construct a media dimension portrait according to the media feature fields.
[0224] The construction of the media dimension portrait includes: obtaining the fields required for constructing the media dimension and depicting the portrait according to the mapping method: advertisement position ID.
[0225] As an optional implementation method, the portrait fields, construction methods, and storage methods of the media dimension are shown in Table 3.
[0226] As an optional implementation method, the above-mentioned user dimension portrait, advertisement dimension portrait, and media dimension portrait are respectively updated to the user dimension table, advertisement dimension table, and media dimension table to complete the construction of the product portrait.
[0227] The above-mentioned product portrait storage tool can be implemented using MongoDB, or can also be implemented using HBase, Elasticserch, etc.
[0228] The "mapping" mentioned above refers to mapping a string or numerical value to a specified value for storage. For example, in user traffic types, app is mapped to 1 and pc is mapped to 2; or it is stored by means of identity mapping. For example, 278191 in the advertisement ID is stored identically as 278191.
[0229] The "statistics" mentioned above refers to counting the number of times of a user-specified behavior within a certain period. If a certain behavior feature has multiple feature values, they are respectively counted, summarized, and recorded in different ways.
[0230] The "analysis" mentioned above refers to the time attribute of user behavior within a specified time period. By using a manually defined method, new creation, discarding, or updating measures are taken according to the time attributes corresponding to different behaviors.
[0231] As an optional implementation manner, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned CTR prediction method is implemented.
[0232] The above-mentioned software is stored in the above-mentioned storage medium, and the storage medium includes but is not limited to: optical discs, floppy discs, hard disks, erasable memories, etc.
[0233] The above-mentioned technical solution has the following beneficial effects: In the embodiment of the present invention, full-scale interaction of features is performed by FM, local feature interaction is performed by CNN, and the inner product method is used to strengthen local feature interaction, integrating FM and CNN, making up for the problem of unsatisfactory effect in full-scale interaction of sparse features and improving the prediction accuracy of the prediction model; through the combination of CNN and inner product, the number of parameters is reduced and the network training difficulty is lowered; by processing advertisement log data, the first dataset and the second dataset, the richness of the feature space is enhanced; feature engineering processing is performed on the third dataset to improve data quality, reduce redundant data, and further improve the prediction accuracy of the model and the efficiency of model training.
[0234] The specific implementation manners of the above invention further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above content is only the specific implementation manners of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A CTR prediction method, characterized in that, Including: Obtain the advertisement log data within a preset number of days, perform statistical calculations on the advertisement log data according to the target key-value pair, obtain the statistical calculation result, generate tags for each target key-value according to the user click behavior, and merge the statistical calculation result corresponding to each target key-value with the tag to obtain the first data set; Construct a multi-dimensional portrait according to the advertisement log data, and use the multi-dimensional portrait as the second data set; Using the target key-value as the data identifier, merge the first data set and the second data set into a third data set; Perform feature engineering processing on the third data set to obtain a training data set; Adopt the CNN algorithm and the inner product algorithm to perform local feature crossing on the training data set, adopt the FM algorithm to perform global feature crossing on the training data set, and train to obtain a prediction model; Use the prediction model to perform CTR prediction on the data to be measured, and obtain the predicted CTR result; The step of adopting the CNN algorithm and the inner product algorithm to perform local feature crossing on the training data set, adopting the FM algorithm to perform global feature crossing on the training data set, and training to obtain a prediction model includes: Identify the sparse features and dense features in the training data set; Perform feature crossing on the sparse features in sequence through the CNN algorithm and the inner product algorithm to obtain the first feature data; Concatenate the dense features with the vectorized sparse features to obtain the second feature data; Perform feature crossing on the sparse features and the dense features to obtain the third feature data; Train the first feature data, the second feature data and the third feature data to obtain the prediction model.
2. The CTR prediction method according to claim 1, characterized in that, The step of performing feature engineering processing on the third data set includes: Perform missing value processing on the third data set; And / or perform feature selection on the third data set; And / or perform outlier removal on the third data set; And / or perform dimensionless processing on the third data set; And / or perform data correction on the third data set.
3. The CTR prediction method according to claim 1, characterized in that: The target key-value is a combination of user ID, advertisement ID and media ID.
4. The CTR prediction method according to claim 1, characterized in that ,Constructing a multi-dimensional portrait according to the advertisement log data, including: Extract the user feature fields in the advertisement log data, and construct a user dimension portrait according to the user feature fields; Extract the advertisement feature fields in the advertisement log data, and construct an advertisement dimension portrait according to the advertisement feature fields; Extract the media feature fields in the advertisement log data, and construct a media dimension portrait according to the media feature fields.
5. A CTR prediction device, characterized in that, Including: The first data set generation module is used to obtain the advertisement log data within a preset number of days, perform statistical calculations on the advertisement log data according to the target key-value pair, obtain the statistical calculation result, generate tags for each target key-value according to the user click behavior, and merge the statistical calculation result corresponding to each target key-value with the tag to obtain the first data set; The second data set generation module is used to construct a multi-dimensional portrait according to the advertisement log data, and use the multi-dimensional portrait as the second data set; The third data set generation module is used to use the target key-value as the data identifier, and merge the first data set and the second data set into a third data set; The training data set generation module is used to perform feature engineering processing on the third data set to obtain a training data set; The prediction model training module is used to perform local feature crossing on the training data set by using the CNN algorithm and the inner product algorithm, and perform global feature crossing on the training data set by using the FM algorithm to train and obtain the prediction model; The prediction module is used to use the prediction model to perform CTR prediction on the data to be measured and obtain the predicted CTR result; The prediction model training module includes: The feature recognition sub-module is used to recognize sparse features and dense features in the training data set; The local feature crossing sub-module is used to perform feature crossing on the sparse features successively through the CNN algorithm and the inner product algorithm to obtain the first feature data; The second feature data generation sub-module is used to splice the dense features with the vectorized sparse features to obtain the second feature data; The global feature crossing sub-module is used to perform feature crossing on the sparse features and dense features to obtain the third feature data; The model training sub-module is used to train the first feature data, the second feature data and the third feature data to obtain the prediction model.
6. The CTR prediction device according to claim 5, wherein, The training data set generation module includes: The missing value processing sub-module is used to process the missing values in the third data set; The feature selection sub-module is used to perform feature selection on the third data set; The outlier removal sub-module is used to remove outliers from the third data set; The dimensionless sub-module is used to perform dimensionless processing on the third data set; The data modification sub-module is used to correct the data in the third data set.
7. The CTR prediction device according to claim 5, wherein, The second data set generation module includes: The user dimension portrait generation sub-module is used to extract user feature fields from the advertisement log data and construct a user dimension portrait according to the user feature fields; The advertisement dimension portrait generation sub-module is used to extract advertisement feature fields from the advertisement log data and construct an advertisement dimension portrait according to the advertisement feature fields; The media dimension portrait generation sub-module is used to extract media feature fields from the advertisement log data and construct a media dimension portrait according to the media feature fields.
8. A computer-readable storage medium, on which a computer program is stored, wherein, When the program is executed by the processor, it implements the CTR prediction method according to any one of claims 1-4.
Citation Information
Patent Citations
Advertisement recommendation method and system based on click rate prediction model, and storage medium
CN113222647A