An Advertising Click-Through Rate Prediction Method Based on a Deep Multi-Behavior Network
Through the advertising click-through rate prediction method of deep multi-behavior network, the problem of single user behavior sequence types in the sequence model is solved. Sparse attention and two-dimensional convolutional networks are used to process different types of user behaviors. Combined with the dynamic Dropout module, the accuracy of advertising click-through rate prediction and the generalization ability of the model are improved.
Patent Information
- Application Number
- CN202310247095.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-15
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-03-15
AI Technical Summary
The existing sequence model has a single user behavior sequence type in advertising click-through rate prediction, and there is room for improvement in user behavior modeling methods, resulting in insufficient prediction accuracy.
Advertising click-through rate prediction method based on deep multi-behavior networks is adopted. By obtaining user and ad data information, sequence features, user features, advertising features and context features are extracted, and deep multi-behavior networks are used for training. Different types of user behavior sequences are processed by combining sparse attention mechanisms and two-dimensional convolutional networks, and dynamic Dropout module is introduced for interest fusion to avoid model overfitting.
It improves the accuracy of advertising click-through rate prediction and the generalization ability of the model, effectively integrates user behavior differences, avoids the model's overfitting of certain types of behaviors, and improves the user experience of the recommendation system and the fairness of advertising display.
Smart Images

Figure CN116228368B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of advertisement click-through rate prediction, and particularly relates to an advertisement click-through rate prediction method based on a deep multi-behavior network. Background Art
[0002] Nowadays, society is in an era of information explosion. The dazzling array of commodities often makes it difficult for users to make a choice. Especially when choosing commodities on mobile phones, it is particularly important to select a favorite commodity from a vast number of commodities. Therefore, predicting click-through rate (CTR) to recommend commodities to users has become an important technology.
[0003] After the existing sequence models introduce user behavior into CTR prediction, although they achieve better prediction accuracy than traditional recommendation algorithms, the types of user behavior sequences introduced by existing technologies are often relatively single. Usually, only the historical click behavior sequence of users is introduced, and there is still room for improvement in the modeling method of user behavior. Summary of the Invention
[0004] To solve the above problems existing in the prior art, the present invention proposes an advertisement click-through rate prediction method based on a deep multi-behavior network. The method includes: obtaining corresponding data information of users and advertisements, where the corresponding data information includes user and advertisement basic information data, advertisement exposure data, and user behavior log data; preprocessing the corresponding data information; extracting features of the preprocessed data information, where the features include sequence features, user features, advertisement features, and context features; inputting the data information features into a trained advertisement click-through rate prediction model based on a deep multi-behavior network to obtain an advertisement click-through rate prediction result;
[0005] The process of training the advertisement click-through rate prediction model based on a deep multi-behavior network includes:
[0006] S1: Obtaining historical advertisement click data of users and preprocessing the data; where the historical advertisement click data of users includes user behavior sequence features, advertisement features, user features, and environmental features;
[0007] S2: Inputting the user behavior sequence features, advertisement features, user features, and environmental features into a feature embedding layer to generate user behavior sequence feature vector representations, advertisement feature vector representations, user feature vector representations, and environmental feature vector representations;
[0008] S3: Inputting the user behavior sequence feature vector representations into a deep multi-behavior network to extract the behavior features of users;
[0009] S4: Inputting all the user behavior features into a multi-behavior fusion module to obtain user behavior fusion features;
[0010] S5: Input the advertisement feature vector representation, user feature vector representation, and environmental feature vector representation into the deep cross network to obtain the advertisement context fusion feature;
[0011] S6: Fuse the user behavior fusion feature and the advertisement context fusion feature to obtain the advertisement click-through rate prediction result;
[0012] S7: Calculate the loss function of the model according to the advertisement click-through rate prediction result, and use the Adam optimization algorithm to optimize the parameters of the model. When the loss function converges, the training of the model is completed.
[0013] Preferably, the user behavior sequence features include long-term post-click behavior sequences, long-term click behavior sequences, short-term click behavior sequences, and short-term exposure behavior sequences.
[0014] Preferably, the deep multi-behavior network includes a long-term post-click behavior sequence modeling module, a long-term click behavior sequence modeling module, a short-term click behavior sequence modeling module, and a short-term exposure news sequence modeling module;
[0015] The process of the long-term post-click behavior sequence modeling module processing the input data includes: inputting the long-term post-click behavior sequence and the candidate advertisement feature vector into the sparse multi-head attention layer for attention feature extraction; adding and normalizing the extracted features and the input sequence; inputting the normalized data into the fully connected layer to obtain the fusion feature; adding and normalizing the fusion feature and the input feature to obtain the user's long-term click interest representation;
[0016] The process of the long-term click behavior sequence modeling module processing the input data includes: inputting the short-term click behavior sequence into the encoder for encoding, and jointly inputting the encoded data and the candidate advertisement feature vector into the decoder for decoding to obtain the long-term post-click interest representation;
[0017] The process of the short-term exposure behavior sequence modeling module processing the input data includes: inputting the short-term sequence and the candidate advertisement feature vector into the multi-head attention layer, and adding and normalizing the output result of the multi-head attention layer and the input data; inputting the normalized result into the multi-layer two-dimensional convolutional network to obtain the short-term exposure interest representation;
[0018] The process of the short-term click behavior sequence modeling module processing the input data includes: inputting the long-term post-click behavior sequence and the candidate advertisement feature vector into the sparse multi-head attention layer for attention feature extraction; adding and normalizing the extracted features and the input sequence; inputting the normalized data into the fully connected layer to obtain the fusion feature; adding and normalizing the fusion feature and the input feature to obtain the short-term click interest representation.
[0019] Preferably, the process of the multi-behavior fusion module for fusing user behavior features includes: inputting the long-term click interest representation, long-term post-click interest representation, short-term click interest representation, and short-term exposure interest representation of the user into four Dropout layers respectively, and fusing the type embedding vectors; inputting the features of the fused type embedding vectors into a fully connected layer to generate the final fused interest vector.
[0020] Further, the monotonic function of the Dropout layer is expressed as:
[0021]
[0022] where S is the true length of the sequence, θ1, θ2 are hyperparameters for controlling monotonicity and slope, and p(S) is the resulting Dropout ratio.
[0023] Advantages of the present invention:
[0024] The present invention considers introducing multiple different user behavior sequences. For different behavior sequences, optimized transformer structures are designed. The long-term click behavior sequence modeling module, short-term click behavior sequence modeling module, short-term exposure behavior sequence modeling module, and long-term post-click behavior sequence modeling module are designed respectively by introducing different optimized transformer structures; for the long-term click behavior sequence and the long-term post-click behavior sequence, considering the influence of sequence length on model performance, a sparse attention mechanism is introduced to improve the performance of the model without sacrificing the model effect; for the short-term exposure behavior sequence, considering the large proportion of noise in the exposure data, a two-dimensional convolutional network is introduced to denoise the exposure behavior sequence representation, further improving the effect of sequence modeling; the present invention proposes an interest fusion module based on dynamic Dropout, which can capture the differences in user behavior distributions, effectively fuse user interests, and avoid the model overfitting to a certain type of behavior representation. Description of the Drawings
[0025] Figure 1 It is a flowchart of the method for predicting the click-through rate of advertisements based on the deep multi-behavior network of the present invention;
[0026] Figure 2 It is a structural diagram of the overall system framework of the present invention;
[0027] Figure 3 It is an input-output module diagram of the present invention;
[0028] Figure 4 It is a diagram of the long-term post-click behavior sequence modeling module of the present invention;
[0029] Figure 5 It is a diagram of the long-term click behavior sequence modeling module of the present invention;
[0030] Figure 6 It is a module diagram for modeling the short-term exposure behavior sequence of the present invention;
[0031] Figure 7 It is a module diagram for modeling the short-term click behavior sequence of the present invention;
[0032] Figure 8 It is a module diagram for multi-behavior fusion of the present invention. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] An advertisement click-through rate prediction method based on a deep multi-behavior network, as Figure 1 shown, the method includes: obtaining corresponding data information of users and advertisements, where the corresponding data information includes user and advertisement basic information data, advertisement exposure data, and user behavior log data; preprocessing the corresponding data information; extracting features of the preprocessed data information, where the features include sequence features, user features, advertisement features, and context features; inputting the data information features into a trained advertisement click-through rate prediction model based on a deep multi-behavior network to obtain an advertisement click-through rate prediction result.
[0035] The process of training an advertisement click-through rate prediction model based on a deep multi-behavior network includes:
[0036] S1: Obtaining historical advertisement click data of users and preprocessing the data; where the historical advertisement click data of users includes user behavior sequence features, advertisement features, user features, and environmental features;
[0037] S2: Inputting the user behavior sequence features, advertisement features, user features, and environmental features into a feature embedding layer to generate user behavior sequence feature vector representations, advertisement feature vector representations, user feature vector representations, and environmental feature vector representations;
[0038] S3: Inputting the user behavior sequence feature vector representations into a deep multi-behavior network to extract the behavior features of users;
[0039] S4: Inputting all the user behavior features into a multi-behavior fusion module to obtain user behavior fusion features;
[0040] S5: Input the advertisement feature vector representation, user feature vector representation, and environmental feature vector representation into a deep cross network to obtain the advertisement context fusion feature;
[0041] S6: Fuse the user behavior fusion feature and the advertisement context fusion feature to obtain the advertisement click-through rate prediction result;
[0042] S7: Calculate the loss function of the model according to the advertisement click-through rate prediction result, and use the Adam optimization algorithm to optimize the parameters of the model. When the loss function converges, the training of the model is completed.
[0043] A specific implementation manner of an advertisement click-through rate prediction method based on a deep multi-behavior network includes:
[0044] Step 1. Obtain user and advertisement basic information data, advertisement exposure data, and user behavior log data. Preprocess the data, and then extract user historical behavior sequence features, user features, advertisement features, and context features from the data.
[0045] Step 2. Input the above features into an advertisement click-through rate prediction model based on a deep multi-behavior network. After the calculation of the model, obtain the final advertisement click-through rate prediction result.
[0046] The preprocessing of the data includes: changing the post-click behaviors in the data to <browsing, adding to cart, liking, purchasing, etc.><1, 2, 3, 4,...> according to the mapping. Changing the continuous features in the data to discrete features through binning processing.
[0047] In this embodiment, the preprocessing of the data set is mainly divided into two steps: First, splice the user behavior sequence features. Both the advertisement exposure click samples and the user behavior records in the user behavior log are single behavior records with timestamps, while the model proposed in the present invention requires the user behavior history to be input into the model in the form of a sequence. Therefore, for the data in the advertisement exposure click samples and the user behavior log, this article aggregates them according to the user ID and behavior type. For multiple behavior records of each user under each behavior type, they are sorted in ascending order of the timestamp and organized into an ordered sequence. Based on the characteristics of this data set, the historical behaviors within 7 days are regarded as short-term behaviors, and there are no restrictions on long-term behaviors other than the total length of the sequence, so as to make full use of the limited training data as much as possible. For each sample in the advertisement exposure click samples, this article appends the long-term click, short-term click, short-term exposure, and post-click (including collection, adding to cart, and purchasing) behavior sequences of its corresponding user. Of course, when splicing, the behavior sequences should be truncated according to the timestamp of the current sample to prevent future behavior information from being introduced into the sample and causing information leakage.
[0048] In the data set preprocessing step, the training set and the test set are also divided. To make full use of the user behavior sequence data, this paper uses the date as the basis for dividing the training set and the test set. Among all 7 days of data, the first 6 days are designated as the training set, and the last 1 day is designated as the test set.
[0049] User behavior sequence features include long-term post-click behavior sequence, long-term click behavior sequence, short-term click behavior sequence, and short-term exposure behavior sequence.
[0050] The four sequence formats are designed as follows: a i represents the advertisement interacted with by the i-th user behavior.
[0051] Long-term click behavior sequence and short-term exposure behavior sequence: [a1, a2,... a L , where L represents the upper limit of the behavior sequence length.
[0052] Short-term click behavior sequence: [(a 11 , a 12 ,... a 1M ), (a 21 , a 22 ,... a 2M ),... (a N1 , a N2 ,... a NM )], where N represents the upper limit of the number of sessions, and M represents the upper limit of the number of behaviors within a single session.
[0053] Long-term post-click behavior sequence: [<a1, t1>, <a2, t2>,... <a L , t L >], where t i represents the behavior category (browse, add to cart, like, purchase, etc.) of the advertisement interacted with by the i-th user behavior.
[0054] As Figure 2 shown, the advertisement click-through rate prediction model based on the deep multi-behavior network includes: an input-output module, a feature embedding module, a multi-behavior processing module, a multi-behavior fusion module, and a deep cross-network module.
[0055] As Figure 3 shown, the input of the model: First is the user behavior sequence feature. In the DMBN, the user behavior sequence includes four types: long-term post-click behavior sequence, long-term click behavior sequence, short-term click behavior sequence, and short-term exposure behavior sequence. It also includes other features such as advertisement features, user features, and context features.
[0056] The output layer of the model receives the output of the lower-layer network and converts it into a CTR prediction value. The output layer of the DMBN model is a fully connected layer with a Sigmoid function as the activation function and an output size of 1.
[0057] In this embodiment, the deep multi-behavior network includes a long-term post-click behavior sequence modeling module, a long-term click behavior sequence modeling module, a short-term click behavior sequence modeling module, and a short-term exposure news sequence modeling module.
[0058] As Figure 4 shown, the process of the long-term post-click behavior sequence modeling module processing the input data includes: the long-term post-click behavior sequence and the candidate advertisement feature vector are input into the sparse multi-head attention layer for attention feature extraction; the extracted features are added and normalized with the input sequence; the normalized data is input into the fully connected layer to obtain the fused features; the fused features and the input features are added and normalized to obtain the user's long-term click interest representation.
[0059] Specifically, let the long-term clicked advertisement vector sequence be the position embedding vector sequence be where L is the maximum length of the user behavior sequence, and the elements in the two sequences, that is, the dimensions of the vectors E are the same. The two are added by position, and the obtained sequence is input into the upper-layer network. Let the user behavior matrix after the above transformation be S = [x1, x2,... x L , where the commodity to be predicted is t, x i ∈R D , t∈R D , and the dimension of the advertisement vector to be predicted is D.
[0060] The sequence S and the advertisement t to be input obtained above are input into the sparse multi-head attention layer, then into the addition and normalization layer, and finally into the fully connected layer. Its expression is:
[0061]
[0062] O A = Concat(head1,... head H )W O
[0063] In the formula represents the randomly initialized trainable parameter matrix, K is the dimension of the hidden vector, H is the number of attention heads, both are hyperparameters, head i is the output of the i-th attention head, O A ∈R D is the result of the transformation and will be input into the addition and normalization layer, as shown below:
[0064] O AN = LayerNorm(S + O A )
[0065] After that, it comes to the fully connected layer. This layer is essentially a multi-layer fully connected network. After transforming the input vector, it outputs to another addition and normalization layer. This part of the operation can be formally described as follows:
[0066] O LC = LayerNorm(O AN + FC(O AN ))
[0067] Among them, FC represents the fully connected network transformation.
[0068] Since there are click behaviors of users in the past six months or even a year in the long-term click behavior sequence of users, the length of its behavior sequence is relatively long. Although the traditional multi-head attention mechanism has good efficiency, it is still difficult to handle sequences with lengths of thousands or even tens of thousands. Therefore, in this paper, the multi-head attention module in the long-term click behavior sequence is changed to a sparse multi-head attention module.
[0069] As Figure 5 shown, the long-term click behavior sequence modeling module processes the input data including: inputting the short-term click behavior sequence into the encoder for encoding, and jointly inputting the encoded data and the candidate advertisement feature vector into the decoder for decoding to obtain the long-term click posterior interest representation.
[0070] Specifically, the main characteristics of the short-term click behavior are sparse data and less noise. The interest it reflects has a high degree of matching with the user's current mind and has a great impact on the final prediction. In other words, the lower limit of the effect of modeling it is relatively high, and often a simple method such as pooling can achieve seemingly good results. On the other hand, over-reliance on the user's recent click behavior is often an important factor causing the feedback loop phenomenon. Simply recommending results with high similarity to the advertisements clicked by the user recently is likely to make the user feel "the more you click, the more you are recommended", which not only affects the user experience, reduces the user's evaluation and interest in the recommended results, but also causes the Matthew effect on the advertisement side to intensify continuously, and a large number of advertisements cannot be displayed, making the recommendation system deviate from the original intention of alleviating information overload.
[0071] Similar to the long-term click behavior modeling module, before the short-term click advertisement vector sequence is input into the upper network, a position embedding vector needs to be added first. For the decoder part, the masked multi-head attention part in the original Transformer is removed here and changed to the design shown in the above figure.
[0072] As shown Figure 6 in the figure, the process of the short-term exposure behavior sequence modeling module processing the input data includes: inputting the short-term sequence and the candidate advertisement feature vector into the multi-head attention layer, and performing addition and normalization processing on the output result of the multi-head attention layer and the input data; inputting the normalization processing result into a multi-layer two-dimensional convolutional network to obtain the short-term exposure interest representation.
[0073] Specifically, the short-term exposure behavior sequence modeling module adopts a scheme that combines the multi-head attention mechanism and the convolutional neural network: the multi-head attention mechanism is responsible for extracting effective information from the sequence data and representing it as multiple vectors, and the convolutional neural network is good at extracting multi-level information from the matrix composed of multiple vectors.
[0074] Although the lower layer structure of the short-term exposure behavior modeling module seems similar to that of the long-term click behavior modeling module, due to the difference in the upper layer structure, there are still slight differences in the specific implementation. Let the short-term exposure advertisement vector sequence be the position embedding vector sequence be where L is the maximum length of the user behavior sequence, and the elements in the two sequences, that is, the vectors E i have the same dimension. The two are added by position, and the resulting sequence is input into the upper layer network. Let the user behavior matrix after the above transformation be S = [x1, x2,... x L , where the advertisement to be predicted is t, and the dimension of the advertisement vector to be predicted is D. The sequence S and the advertisement to be predicted t are input into the multi-head attention layer, and the specific transformation performed can be formally described by the following formula:
[0075]
[0076] O A = Concat(head1,... head H )W O
[0077] where, represents the randomly initialized trainable parameter matrix, K is the dimension of the hidden vector, H is the number of attention heads, both are hyperparameters, Diag is a function to create a corresponding diagonal matrix according to the vector, head i is the output of the i-th attention head, O A ∈R (L×D)As the result of the transformation, it will be input into the addition and normalization layer. Note that here it is different from the long-term click behavior sequence. The generated transformation result is a matrix rather than a vector, which is convenient for processing using a two-dimensional convolutional neural network. The result obtained after the transformation is a matrix, and then a multi-layer two-dimensional convolutional neural network further transforms it. The size of the convolutional kernel of the two-dimensional convolutional neural network and the number of layers of the convolutional neural network are both hyperparameters, which are determined by experimental attempts.
[0078] As Figure 7 shown, the process of the short-term click behavior sequence modeling module processing the input data includes: inputting the long-term post-click behavior sequence and the candidate advertisement feature vector into the sparse multi-head attention layer for attention feature extraction; adding and normalizing the extracted features with the input sequence; inputting the normalized data into the fully connected layer to obtain the fused features; adding and normalizing the fused features with the input features to obtain the short-term click interest representation.
[0079] Specifically, the main characteristics of the post-click behavior are sparse data, long time intervals between behaviors, and multiple types of behaviors. Therefore, in the design of the modeling method, it is mainly to splice multiple types of behaviors into a unified sequence and add a behavior type vector to distinguish different behavior types.
[0080] It is similar to the long-term click behavior modeling module, and extracts information related to the advertisement vector to be predicted from the sequence through the multi-head sparse attention mechanism. The difference is that the vector input into the sparse multi-head attention module for the long-term post-click behavior sequence is the sum of the post-click advertisement vector sequence, the post-click advertisement type vector sequence, and the position embedding vector in order, and the obtained sequence is input into the upper layer network. The other parts are exactly the same as the long-term click behavior modeling module and will not be elaborated here.
[0081] As Figure 8 shown, the process of the multi-behavior fusion module fusing the user behavior characteristics includes: inputting the user's long-term click interest representation, long-term post-click interest representation, short-term click interest representation, and short-term exposure interest representation into four Dropout layers respectively, and fusing the type embedding vector; inputting the features of the fused type embedding vector into the fully connected layer to generate the final fused interest vector.
[0082] Specifically, after the behavior sequence modeling, the long-term click interest representation, the long-term post-click interest representation, the short-term click interest representation, and the short-term exposure interest representation of the user can be obtained. Due to the dispersion of user behavior, the user's short-term click sequence and short-term exposure sequence may be very sparse, and the long-term click sequence and long-term post-click sequence representations may dominate, thus drowning out the user's short-term interest expression. Based on this observation, the present invention proposes an interest fusion layer based on dynamic Dropout, which can capture the differences in the user behavior distribution and effectively fuse the user interests to avoid the model from overfitting to a certain type of behavior representation.
[0083] The core idea of Dropout is to randomly discard a certain proportion of neurons during the model training process, thereby reducing the risk of model overfitting. This paper believes that the large differences in the user behavior distribution may lead to the model overfitting to a certain type of behavior representation. Therefore, an interest fusion layer based on dynamic Dropout is designed. The model takes the true lengths of the four sequences as the input of a monotonic function to obtain four Dropout ratios. The specific implementation of the monotonic function is as follows:
[0084]
[0085] In the formula, S is the true length of the sequence, θ1 and θ2 are hyperparameters that control monotonicity and slope, and p(S) is the obtained Dropout ratio.
[0086] The four obtained Dropout ratios are used as the probabilities of the true Dropout of the four behavior sequences. These probabilities are used to perform Dropout on the interest vectors of the four behavior sequences, so as to adaptively control the weights of the short-term behavior sequence expression and the long-term behavior sequence expression, realize the adaptive fusion of the four different types of behavior sequences during the training stage, and mitigate the impact of overfitting on the model.
[0087] In addition, considering that the above four modules are relatively independent, the user interest vectors generated by them may not be in the same semantic space. First, the user interest vectors need to be transformed, and this work can be completed by a fully connected layer.
[0088] Next, in order to explicitly add the type information of the four interest vectors to the model, referring to the position embedding vector in Transformer, before connecting to the fully connected layer, the interest vectors also need to be added with type embedding vectors.
[0089] After the different interest vectors are connected and added with type embedding vectors, they are then input into the fully connected layer to generate the final fused interest vector. This interest vector and the output of the deep cross network are connected to another fully connected layer, and finally the CTR prediction value is calculated.
[0090] The structure of the Deep and Cross Network of the present invention generally follows the structure of the original DCN and is optimized in some details. First, in the original DCN network, the activation functions of the deep network are all ReLU functions, while the deep network part of the present invention uses Leaky ReLU activation functions instead. Second, Batch Normalization and Dropout techniques are not applied in the original DCN, while the present invention adds a BN layer after each layer in the deep network and a Dropout layer after the last layer. In addition, in terms of input features, the present invention also selects the features accessed by the deep network and the cross network in the Deep and Cross Network respectively. According to the design idea and application experience of DCN, its deep network is mainly responsible for the generalization ability of the model, and the cross network is mainly responsible for the memory ability of the model. At the same time, due to the characteristics of the cross layer, its output width is equal to the input width and cannot be designed into a triangular structure that gradually decreases from the input layer to the output layer like a fully connected layer. Therefore, the features accessed by the deep network and the cross network should be selected to balance the prediction effect and resource consumption of the model.
[0091] The process of using the Deep and Cross Network to process the advertisement feature vector representation, user feature vector representation, and environment feature vector representation includes: The input of the model also includes other features such as advertisement features, user features, and context features. It mainly includes three categories: numerical features, single-value categorical features, and multi-value categorical features. First, there are numerical features. A typical example is the selling price of a commodity. This type of feature is generally a real number and needs to be normalized during data preprocessing, or binned according to its numerical size to convert the numerical feature into a categorical feature, thereby improving robustness and introducing a certain degree of non-linearity. Second, in the CTR prediction problem, categorical features play a very important role. Categorical features refer to features whose values are categorical codes, and the magnitude of their values often does not have numerical significance but only represents that they belong to a certain category. Categorical features can be further divided into single-value and multi-value. As the name implies, single-value features contain only one value, while multi-value features may contain multiple values. The most typical single-value categorical features are user and advertisement ID features, and the most typical multi-value categorical feature is the user's historical behavior. Multi-value categorical features are similar in form to the sequence features described above, and even in some data formats, their storage methods are the same. However, multi-value categorical features often do not distinguish the order between multiple values within the feature, while sequence features strictly distinguish the order. There are a large number of categorical features in the CTR prediction scenario, and the most typical ones are user and advertisement ID features. The value space of this type of feature is large, and the frequency of a single value appearing in the data is low. If it is encoded in a simple One-hot manner, it is easy to cause the explosion of feature dimensions. At the same time, due to the data sparsity problem, the model cannot be fully trained, thus affecting the model effect.
[0092] Therefore, after processing the original input as above, all features are converted into categorical vectors. For categorical features, a feature embedding layer is generally used to transform high-dimensional sparse categorical features into low-dimensional dense real vectors, enabling the upper-layer network to process them efficiently. The implementation principle of the feature embedding layer can be regarded as a process of looking up a table. For example, for the user ID feature, a matrix is established with the number of rows equal to the total number of values of the user ID and the number of columns equal to the dimension of the Embedding vector. In this way, each user ID can correspond to a row in the matrix, that is, each user ID can be mapped to the corresponding Embedding vector. During model training, this matrix is updated through the backpropagation algorithm, thus realizing the update of the Embedding vector during the model training process.
[0093] The LeakyReLU activation function is used in the deep network and multi-behavior fusion module, and the Swish activation function is adopted in each behavior modeling module.
[0094] The expression of the loss function of the model is:
[0095] L total = L CE + L AUC
[0096] Among them, L CE is the classical cross-entropy loss function, and L AUC is the hinge ranking loss function.
[0097] In this embodiment, the model evaluation criteria include: there are inconsistencies among the online evaluation metrics, offline evaluation metrics, and model loss function of the CTR prediction model. In the present invention, the offline experiment method is adopted to evaluate the quality of the model. Therefore, this section mainly discusses the selection of model evaluation criteria. The most commonly used model evaluation criterion in CTR prediction is AUC. The value range of AUC is [0, 1]. Its physical meaning is the probability that the predicted value of a positive sample is greater than that of a negative sample when randomly taking a positive sample and a negative sample from the dataset, that is, the probability of correct ranking. That is to say, AUC is an index representing the ranking ability of the model. For the advertising CTR prediction model, not only the order of the ranking results is required to be correct, but also the absolute value of the CTR prediction is required to be accurate. Examining the AUC index can only judge the ranking ability of the model. If we want to judge whether the absolute value of the CTR prediction is accurate, other indicators need to be examined.
[0098] For several CTR prediction samples, to calculate whether the CTR predicted value is accurate, this paper selects Logloss as the index to measure the accuracy of the absolute value of the CTR prediction. Its calculation formula is as follows:
[0099]
[0100] Among them, n is the number of samples, yi is the label of the i-th sample, is the predicted value of the i-th sample.
[0101] The above embodiments have further elaborated on the object, technical solution, and advantages of the present invention. It should be understood that the above embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting advertisement click-through rate based on a deep multi-behavior network, characterized in that, Including: Obtain the corresponding data information of users and advertisements, where the corresponding data information includes user and advertisement basic information data, advertisement exposure data, and user behavior log data; preprocess the corresponding data information. Extract the features of the preprocessed data information, where the features include sequential features, user features, advertisement features, and context features. Input the data information features into the trained advertisement click-through rate prediction model based on the deep multi-behavior network to obtain the advertisement click-through rate prediction result. The process of training the advertisement click-through rate prediction model based on the deep multi-behavior network includes: S1: Obtain the user's historical advertisement click data and preprocess the data; where the user's historical advertisement click data includes user behavior sequential features, advertisement features, user features, and environmental features. S2: Input the user behavior sequential features, advertisement features, user features, and environmental features into the feature embedding layer to generate user behavior sequential feature vector representations, advertisement feature vector representations, user feature vector representations, and environmental feature vector representations. S3: Input the user behavior sequential feature vector representation into the deep multi-behavior network to extract the user's behavior features; the deep multi-behavior network includes a long-term post-click behavior sequence modeling module, a long-term click behavior sequence modeling module, a short-term click behavior sequence modeling module, and a short-term exposure news sequence modeling module. The process of the long-term post-click behavior sequence modeling module processing the input data includes: Input the long-term post-click behavior sequence and the candidate advertisement feature vector into the sparse multi-head attention layer for attention feature extraction; add and normalize the extracted features and the input sequence; input the normalized data into the fully connected layer to obtain the fused features; add and normalize the fused features and the input features to obtain the user's long-term click interest representation. The long-term click behavior sequence modeling module processing the input data includes: Input the short-term click behavior sequence into the encoder for encoding, and input the encoded data and the candidate advertisement feature vector into the decoder for decoding to obtain the long-term post-click interest representation. The process of the short-term exposure behavior sequence modeling module processing the input data includes: Input the short-term sequence and the candidate advertisement feature vector into the multi-head attention layer, and add and normalize the output result of the multi-head attention layer and the input data; input the normalized result into the multi-layer two-dimensional convolutional network to obtain the short-term exposure interest representation. The process of the short-term click behavior sequence modeling module processing the input data includes: Input the long-term post-click behavior sequence and the candidate advertisement feature vector into the sparse multi-head attention layer for attention feature extraction; add and normalize the extracted features and the input sequence; input the normalized data into the fully connected layer to obtain the fused features; add and normalize the fused features and the input features to obtain the short-term click interest representation. S4: Input the multiple user behavior features output by the deep multi-behavior network into the multi-behavior fusion module to obtain the user behavior fusion feature. S5: Input the advertisement feature vector representation, user feature vector representation, and environmental feature vector representation into the deep cross network to obtain the advertisement context fusion feature; S6: After fusing the user behavior fusion feature and the advertisement context fusion feature, input them into the fully connected layer to obtain the advertisement click-through rate prediction result; S7: Calculate the loss function of the model according to the advertisement click-through rate prediction result, and use the Adam optimization algorithm to optimize the parameters of the model. When the loss function converges, complete the training of the model.
2. The method for predicting advertisement click-through rate based on a deep multi-behavior network according to claim 1, characterized in that, The user behavior sequence features include long-term post-click behavior sequences, long-term click behavior sequences, short-term click behavior sequences, and short-term exposure behavior sequences.
3. A method for predicting advertisement click-through rate based on a deep multi-behavior network according to claim 1, characterized in that, The preprocessing of the data includes: changing the post-click behavior in the data according to the mapping to [<a1,t1>,<a2,t2>,…,<a L ,t L >], where t L represents the behavior category of the advertisement interacted by the L-th user behavior, and a L is the advertisement interacted by the L-th user behavior; the continuous features in the data are changed to discrete features through bucketing processing.
4. A method for predicting advertisement click-through rate based on a deep multi-behavior network according to claim 1, wherein The process of the multi-behavior fusion module fusing user behavior features includes: inputting the long-term click interest representation, long-term post-click interest representation, short-term click interest representation, and short-term exposure interest representation of the user into four Dropout layers respectively, and fusing the type embedding vectors; inputting the features of the fused type embedding vectors into the fully connected layer to generate the final fused interest vector.
5. The method for predicting advertisement click-through rate based on a deep multi-behavior network according to claim 4, wherein The monotonic function of the Dropout layer is expressed as: Among them, S is the true length of the sequence, θ1 and θ2 are hyperparameters controlling monotonicity and slope, and p(S) is the obtained Dropout ratio.
6. A method for predicting advertisement click-through rate based on a deep multi-behavior network according to claim 1, characterized in that, The process of using the deep cross network to process the advertisement feature vector representation, user feature vector representation, and environmental feature vector representation includes: the deep cross network includes a deep network and a cross network; the deep network consists of multiple fully connected layers, and the cross network consists of multiple cross layers; the fully connected layer is used to obtain the deep feature information of the advertisement feature vector representation, user feature vector representation, and environmental feature vector representation, and the cross layer is used to obtain the cross feature information of the advertisement feature vector representation, user feature vector representation, and environmental feature vector representation; concatenate the feature information extracted by the deep network, the cross feature information extracted by the cross network, and the fused interest vector output by the multi-type interest vector fusion module, and then input them into the fully connected layer to obtain the final result.
7. A method for predicting advertisement click-through rate based on a deep multi-behavior network according to claim 1, characterized in that The expression of the loss function of the model is: Among them, y i is the true value, is the predicted probability value, and n is the number of samples.