A Deep Click-Through Rate Prediction Method Supporting Multiple Fine-Grained Interest Extraction
By introducing local interest activation layer, overall interest extraction layer, multi-core convolution layer and multi-head self-attention layer into the deep click-through rate estimation model, multiple fine-grained interest expressions of users are solved, and a more accurate and personalized click-through rate prediction is achieved.
Patent Information
- Application Number
- CN202210843970.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-07-18
AI Technical Summary
In the prior art, local interests are over-dominated in decision-making under the attention mechanism, and it is difficult to model different behavior patterns among users and complex behavior patterns within users.
A deep click-through rate estimation method that supports multiple fine-grained interest extraction is proposed, including local interest activation layer, overall interest extraction layer, multi-core convolutional layer and multi-head self-attention layer. Through these modules, they adaptively process user behavior sequences and extract multiple user interest representations.
It effectively solves the problem of excessive local interest under the attention mechanism, enhances the model's learning ability to express multiple interests by users, can better model complex behavior patterns between users and within users, and improves the accuracy and personalization of click-through rate estimates.
Smart Images

Figure CN115309981B_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to the technical field of information processing, and in particular to a deep click rate prediction method supporting extraction of multiple fine-grained interests. [Background technology]
[0002] In network application services such as e-commerce websites, the selection of products or advertisements and their positions on the page are determined by their click-through rate (CTR). An excellent CTR prediction model can significantly increase the CTR of displayed products or advertisements in e-commerce, thereby increasing the transaction volume of the platform. Thanks to its huge profit-enhancing ability, the CTR prediction task has become one of the most important tasks in the field of recommendation systems.
[0003] CTR prediction is to give the probability prediction of the user clicking on the candidate item given the user information, candidate item information and possible context information as input. It can be regarded as a supervised binary classification learning task or a regression learning task that outputs any value between 0 and 1. Traditional CTR prediction models such as logistic regression (LR) usually require a lot of manual feature engineering to explore the relationship between features. With the great success of deep learning in computer vision and natural language processing, researchers have proposed many models based on deep learning in recent years. Most of these models use word embedding technology to embed large-dimensional sparse vector features into low-dimensional dense vectors, and learn the implicit relationship between features through the nonlinear learning ability brought by activation functions, reducing a lot of unnecessary manual feature engineering. Furthermore, in order to further explore the complex relationship between each feature in the data, many neural network models are committed to using neural networks to achieve automatic and efficient feature interaction. These works have cleverly designed special network structures such as vector inner product and outer product to enable the model to explicitly learn high-order interactions between features.
[0004] However, the model based on feature interaction focuses on extracting information from each independent feature in the sample, ignoring the rich information contained in the user's historical behavior, which can be used as a representation of user interests after being refined. The model based on user historical behavior sequence modeling focuses on extracting useful features from the behavior sequence. They use candidate items as query items and introduce attention mechanisms to obtain user interest representations from user historical behavior sequences. Such a network structure can extract different user interest representations from the same historical sequence based on different candidate items. Intuitively, the network model using the attention mechanism can better capture multiple points of interest of users.
[0005] The model based on feature interaction and the model based on user behavior sequence each have their own advantages. However, no model solution that can well integrate the two technical routes has been proposed in the industry yet. In addition, when classic models such as Deep Interest Network (DIN) and Deep Interest Evolution Network (DIEN) use the attention mechanism to extract user interest representations from the user's historical behavior, its essence is the interaction between the candidate item and the historical items. Among them, the vector representation of the historical item similar to the candidate item will be assigned a higher weight, which plays a decisive role in the result of CTR prediction. Although such a selection result can maximize the activation of the user's local interest, it weakens the influence of other user interest points on the current click probability. Figure 1 The upper and lower parts respectively show the historical behavior sequence segments of two different users. Both user A and user B have records of clicking on a computer, but their other histories are significantly different. When the candidate item is the same digital product, the algorithm based on the attention mechanism will strengthen the influence of items such as computers, mice, and keyboards in the behavior sequence on the result, while the influence of other items is weakened. This is obviously unfair to user A because overall, user A's interest in cosmetics is much greater than that in digital products.
[0006] In addition, from Figure 1 the example, it can be seen that the user's behavior sequence is complex and changeable. The formation of the sequence is affected by many factors, which also leads to the characteristics of multi-noise and interest mutation in the sequence. In recent years, many studies have tried to divide the user's behavior sequence into different session sequences, aiming to reduce the complexity of item types in the session and lower the difficulty of the model in modeling the sequence. For example, the user's behavior is divided based on a 30-minute time interval, and the model learns and models within the divided subsequences. However, even within the same session, the user's behavior may not be homogeneous. In this way, the method based on artificial sequence division is difficult to divide the user's complex behavior into subsequences with homogeneous behavior within each sequence.
[0007] Using embedding technology to map the user's historical behavior to a low-dimensional dense vector and adding the attention mechanism to obtain the user's interest representation from the sequence representation are the main means of the current deep network click-through rate prediction model based on the user's behavior sequence. However, these methods overemphasize the role of items similar to the candidate item in the historical sequence and ignore the influence of the characteristics of different users themselves on the prediction result.
Summary of the Invention
[0008] The objective of the present invention is to solve the problems in the prior art that local interests under the attention mechanism overly dominate in decision-making, and it is difficult to model different behavioral patterns among users and complex behavioral patterns within users under manual sequence partitioning. A deep click-through rate prediction method supporting multiple fine-grained interest extraction is proposed.
[0009] To achieve the above objective, the present invention proposes a deep click-through rate prediction method supporting multiple fine-grained interest extraction, including the following steps:
[0010] S1. Use the local interest activation layer in the Deep Interest Network (DIN) to learn the local interest representation of the user;
[0011] S2. In the overall interest extraction layer, use multiple non-linear fully connected layers to expand the user's expression space and integrate it into the user behavior sequence. Then, use the key-value pair attention mechanism to learn the overall interest representation of the user. By adding the user vector representation to the overall interest representation of the user, the ability to distinguish different users in learning is improved;
[0012] S3. Use a multi-core convolutional layer to adaptively divide the long sequence into short-term behavior sequences and model these subsequences to enhance the model's learning ability for multi-interest representations of the user;
[0013] S4. Use a multi-head self-attention layer to model the user, item side, and context features, implicitly introduce second-order interaction information between features, and maintain the performance of the model in the context of scarce user behavior sequences;
[0014] S5. Use a multi-layer perceptron to predict the results of the features learned in steps S1 to S4, and output the click probability of the user for the candidate item.
[0015] Preferably, in step S1, in the local interest activation layer, the embedding vector of the candidate item is represented by v i The user historical behavior sequence B = [i1, i2, …, i maxlen is transformed into a dense matrix V B = [v i,1 , v i,2 , …, v i,maxlen through the embedding layer, where maxlen specifies the maximum length of the user behavior sequence, v i,1 is the embedding representation of the first item in the user's historical clicks, and so on. Use v i as the query item of the attention mechanism. The matrix V B is the key and value matrix of the attention mechanism. Calculate the attention scores for each historical item through the local interest activation layer, and finally obtain the local interest representation vector v of the user by weighted summing all historical item vectors with the attention scores as weights.local .
[0016] Preferably, step S2 specifically includes the following steps:
[0017] S21. In the overall interest extraction layer, the ReLU nonlinear activation function is first used to expand the embedding representation of the user representation, and the dimension of the expanded vector is consistent with the vector dimension of the item;
[0018] S22. In order to enhance the model’s recognition and memory capabilities for different users, the extended user representation v u Concatenate the vector representing each item in the user's historical behavior sequence to add user identification information;
[0019] S23. After two layers of fully connected layers, the final layer outputs the weight vector v w , whose length is maxlen;
[0020] S24. Finally, all the item vectors in the concatenated matrix are transformed according to the softmax(v w ) is weighted summed to get the user's overall interest representation v overall This process does not use candidate items to participate in the calculation. The final output vector not only contains historical sequence information, but also incorporates user identification, which can learn different overall interest representations for different users. At the same time, adding user vectors to the process of user interest calculation will also affect the learning of user vectors from the perspective of back propagation, making the final user vectors more closely related to their corresponding behavior sequence matrix. For example, using user vectors to participate in the calculation of the weight information of each historical item can make the user vector closer to the vector representation of items that appear more times in the behavior sequence in the vector space.
[0021] Preferably, in step S3, the modeling method is: the convolution kernel slides from top to bottom, and the elements in the window are multiplied and added with the corresponding convolution kernel parameters to obtain a new feature vector v map =[e1,e2,…,e maxlen ], each convolution kernel will get a corresponding feature vector after sliding through the window, and finally these feature vectors are merged into the final output v of the multi-core convolution layer through the pooling layer convs .
[0022] Preferably, step S4 specifically includes the following steps:
[0023] The embedding vector of context features, the embedding vector of item features, and the user-side features after two layers of nonlinear transformation are concatenated into a matrix Among them, m is the number of input features, d is the set embedding dimension; then, each value a in the matrix Ai,j The conversion is performed via layer normalization as follows:
[0024]
[0025] where \(i\in[1,m]\), \(j\in[1,d]\), \(a\) i,j represents the element in the \(i\)-th row and \(j\)-th column of matrix \(A\), \(\in\) Is 1E -15 , to prevent division-by-zero errors; the conversion of each row of matrix \(A\) by layer normalization is as follows:
[0026]
[0027] where \(g\) i and \(b\) i are both learnable parameters, used to prevent the destruction of the information of the original features after layer normalization; the output of layer normalization is matrix In the multi-head self-attention layer, the feature matrix is converted into three new feature matrices by three parameter matrices:
[0028]
[0029] where the parameter matrix In the multi-head self-attention layer, the feature matrices \(Q\), \(K\), and \(V\) are divided into \(h\) sub-matrices according to the second dimension, and then dot-product self-attention calculations are performed between the corresponding sub-matrices, i.e.:
[0030]
[0031] where \(i\in[1,h]\) is the \(i\)-th sub-matrix corresponding to the \(Q\), \(K\), and \(V\) matrices; the calculation results of all heads are concatenated according to the second dimension as the output of the multi-head self-attention: \(mhto = concat(head1, head2, \ldots, head h )W O , where is the parameter matrix; finally, the output after the linear transformation is connected with the original input as a residual connection to obtain the final output \(v\) of the multi-head self-attention layer mth .
[0032] Preferably, in step S5, the feature vectors extracted by the local interest activation layer, the overall interest extraction layer, the multi-core convolution layer, and the multi-head self-attention layer are concatenated in the fully connected layer as the input \(x0\) of the fully connected layer, \(x0 = Concat([v local , v overall , v convs , v mth ); the data of each layer in the fully connected layer will undergo the following transformation:
[0033] xl+1 = f(W l x l + b l ) (5)
[0034] where W l is the parameter matrix of the l-th layer, f is the activation function. The ReLU activation function is selected to provide non-linear learning ability, and the Sigmoid function is used as the activation function in the last layer of the fully connected layer. The output result is between 0 and 1, representing the probability that the model infers that the user clicks on the current candidate item.
[0035] Advantages of the present invention:
[0036] 1. A user interest extraction network that combines local and overall information is constructed. Based on the activation of local fine-grained user interests, user information is integrated into the user behavior sequence, and the user's interests are learned in terms of overall fine-grainedness, solving the problem that local interests overly dominate in decision-making under the attention mechanism and having better personalized learning ability.
[0037] 2. An adaptive multi-window user behavior sequence convolution network partitioning method is proposed. This method can perform sequence modeling on user behaviors under multiple sliding windows of different fine-grained sizes, learning different user behavior patterns, and avoiding the problem of difficult modeling of different behavior patterns between users and complex behavior patterns within users under manual sequence partitioning.
[0038] 3. A feature extraction method for a multi-head self-attention network with implicit feature interaction is proposed. A fusion feature interaction module is introduced to make up for the deficiency in exploring user, item-side, and context features, maintaining the decision-making ability of the model in scenarios where user historical behavior data is scarce.
[0039] The features and advantages of the present invention will be described in detail through embodiments in conjunction with the accompanying drawings.
BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic diagram of an example of a user behavior sequence;
[0041] Figure 2 is the overall structure diagram of the model of the present invention;
[0042] Figure 3 is a schematic diagram of the multi-core convolution network of the present invention;
[0043] Figure 4 is a schematic diagram of the multi-head self-attention network of the present invention.
DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] 1. A deep click-through rate prediction model that supports extraction of various fine-grained interests
[0045] The overall structure of the model of the present invention is as follows Figure 2 shown, including: a variety of fine-grained interest extraction modules, a feature interaction module; among them, the variety of fine-grained interest extraction modules include a local interest activation layer, an overall interest extraction layer, and a multi-core convolutional layer, and the feature interaction module uses a multi-head self-attention layer. That is: the prediction model can be divided into four modules according to the data flow direction, namely, a local interest activation layer, an overall interest extraction layer, a multi-core convolutional layer, and a multi-head self-attention layer.
[0046] The model first uses the local interest activation layer in DIN to learn the local interest representation of the user. Then, in the overall interest extraction layer, multiple non-linear fully connected layers are used to expand the user's expression space and integrate it into the user behavior sequence, and then the key-value pair attention mechanism is used to learn the overall interest representation of the user. By adding the user vector representation to the overall interest representation of the user, the ability to distinguish different users is improved. In addition, the model uses a multi-core convolutional layer to adaptively divide the long sequence into short-term behavior sequences and model these subsequences, enhancing the model's ability to learn the multi-interest representation of the user. Finally, a multi-head self-attention layer is used to model the user, item side, and context features, implicitly introducing second-order interaction information between features.
[0047] 2. Related Definitions and Sample Construction
[0048] To facilitate understanding of the data flow of the present invention, the following introduces some symbol definitions and sample construction processes used in the present invention.
[0049] 2.1 Related Definitions
[0050] Definition 1. Continuous Feature: The continuous feature in the m-th piece of data is represented as represents the set of all continuous features in the data set, where m ∈ [1, M], and M is the total number of samples in the data set. The continuous feature set contains all continuous features from the user, item side, and context information.
[0051] Definition 2. Discrete Feature: The discrete feature in the corresponding m-th piece of data is represented as represents the set of all discrete features in the data set. The discrete feature set contains all discrete features from the user, item side, and context information.
[0052] Definition 3. User Behavior Sequence: For user u, its behavior sequence is the feature B of the items that the user has interacted with in the past u =[i1, i2, …, i maxlen , where u ∈ U represents a user in the user set, i ∈ I is the representation of the item, and maxlen is the set historical sequence length.
[0053] Definition 4. Data Representation: The sample x processed through construction m can be represented using a tuple where y ∈ {0, 1} is the label value, representing whether user u has clicked on item i. 0 represents no, and 1 represents yes.
[0054] 2.2 Sample Construction
[0055] The data in the click-through rate prediction task generally comes from the operation logs of users on the portal website. This operation log records the click records of users on items at a certain moment and some context information such as time and the display position of the item on the page. In order to convert these log information into a data structure that can be processed by the model and construct the corresponding historical behavior sequence of the user, it is necessary to use a unified processing process to process the original data. The processing process of the original data set in the present invention is shown in Algorithm 1:
[0056] Algorithm 1. Sample Construction.
[0057] Input: Operation log data set.
[0058] Output: Model input training set X train and test set X test .
[0059] 1. X train = []; X test = [] / / Initialize the model input data set
[0060] 2. Normalize all continuous features
[0061] 3. Perform digital encoding on all discrete features
[0062] 4. FOR each user u ∈ U in the log data set DO
[0063] 5. B u = [] / / Initialize the user behavior sequence
[0064] 6. Sort all browsing records I of user u u by time
[0065] 7. FOR each record of user u DO
[0066] 8. B u .append(i)
[0067] 9. Randomly sample items
[0068] 10. IF the last browsing record of u THEN
[0069] 11.
[0070] 12.
[0071] 13.ELSE
[0072] 14.
[0073] 15.
[0074] 16.END IF
[0075] 17.END FOR
[0076] 18.END FOR
[0077] 19.RETURN X train ,X test
[0078] 2.3 Embedding Layer
[0079] In the data faced by the click-through rate prediction task, there are a large number of discrete features. These features are encoded into one-hot vectors through one-hot encoding. For example, after one-hot encoding the item category feature, the vector: v cate_one_hot =[o1, o2, …, o cate_num , where cate_num is the total number of item categories, o i ∈{0, 1} and If an item belongs to k categories at the same time, then there is which is called multi-hot encoding. If the value range of the discrete feature is too large, it will bring serious sparsity problems to the data. The embedding layer uses a dictionary method to save a learnable parameter vector for each value of the discrete feature, and finally retrieves the corresponding low-dimensional dense vector through the feature value as the embedding representation of the feature. For example, the item category feature after embedding encoding can be expressed as: v cate_emb =[e1, e2, …, e d , where d is the specified embedding dimension. Different embedding dimensions may have a slight impact on the generalization of the model. The embedding dimension in the present invention is set to 64.
[0080] 3. Multiple Fine-Grained Interest Extraction Modules
[0081] 3.1 Local Interest Activation Layer
[0082] As Figure 2 shown, use the local interest activation layer proposed by the DIN model to learn the local interest representation of the user. In this module, the embedding vector of the candidate item is v iIt is shown that the user historical behavior sequence B = [i1, i2, …, i maxlen is transformed into a dense matrix V B = [v i,1 , v i,2 , …, v i,maxlen through the embedding layer, where maxlen specifies the maximum length of the user behavior sequence, and v i,1 is the embedding representation of the first item in the user's historical clicks, and so on. Use v i as the query item of the attention mechanism, and the matrix V B is the key and value matrix of the attention mechanism. The attention scores are calculated for each historical item through the local interest activation layer, and finally the weighted sum of all historical item vectors is obtained with the attention scores as weights to obtain the user's local interest representation vector v local .
[0083] 3.2 Overall Interest Extraction Layer
[0084] As Figure 2 shown, in the overall interest extraction layer, first use the ReLU non-linear activation function to expand the embedding representation of the user representation, and the dimension of the expanded vector is the same as that of the item vector. To strengthen the model's recognition and memory ability for different users, concatenate the user representation v u after the expanded representation with each item representation vector in the user historical behavior sequence to add user identification information. Then, after two layers of fully connected layer's non-linear transformation, the last layer outputs the weight vector v w , whose length is maxlen. Finally, the weighted sum of all item vectors in the concatenated matrix is obtained according to the corresponding values in softmax(v w ) to obtain the user's overall interest representation v overall . This process does not use the candidate items to participate in the calculation. The finally output vector not only contains the historical sequence information, but also incorporates the user identification, and can learn different overall interest representations for different users. At the same time, adding the user vector to the process of calculating the user's overall interest will also affect the learning of the user vector from the perspective of backpropagation, making the final user vector have a closer relationship with its corresponding behavior sequence matrix. For example, using the user vector to participate in calculating the weight information of each historical item can make the user vector closer to the vector representation of the items that appear more frequently in the vector space.
[0085] 3.3 Multi-core Convolution Layer
[0086] The click behavior of users is driven by their interests. As their interests change, the click behavior sequence of users may form specific evolution characteristics. In addition to some common sequence evolution characteristics, there are also differences in the sequence evolution of different users. For example, some users have slower interest evolution and tend to visit similar items within a period of time, while other users show more drastic interest shifts or faster evolution processes. In order to model these diverse interest evolutions, a multi-core convolutional network is used to model the user's behavior sequence.
[0087] like Figure 3 As shown, the convolution kernel slides from top to bottom. At the same time, the elements in the window are multiplied and added with the corresponding convolution kernel parameters to obtain the new feature vector v map =[e1,e2,…,e maxlen ], each convolution kernel will get a corresponding feature vector after sliding through the window, and finally these feature vectors are merged into the final output v of the multi-core convolution layer through the pooling layer convs .
[0088] 4. Feature Interaction Module
[0089] In order to make full use of the information provided by the user side, item side and context features, a multi-head self-attention layer is used to extract the vector features encoded by the embedding layer. Different from the usual multi-head self-attention network, layer normalization (LN) is pre-placed to reset the data distribution to the non-saturated region of the activation function through layer normalization to alleviate the problems of gradient vanishing, gradient exploding and internal covariate shift.
[0090] like Figure 4 As shown, the embedding vector of the context feature, the embedding vector of the item feature, and the user-side feature after two layers of nonlinear transformation are concatenated into a matrix Where m is the number of input features and d is the embedding dimension set. Then, each value a in the matrix A i,j The following transformation is performed via layer normalization:
[0091]
[0092] where i∈[1,m],j∈[1,d],a i,j represents the element in the i-th row and j-th column of matrix A, ∈ is 1E -15 , to prevent division by zero errors. Layer normalization transforms each row of matrix A as follows:
[0093]
[0094] Among them, g i and bi They are all learnable parameters used to prevent the destruction of the information of the original features after layer normalization. The output of layer normalization is a matrix In the multi-head self-attention layer, the feature matrix is transformed into three new feature matrices by three parameter matrices:
[0095]
[0096] where the parameter matrix In the multi-head self-attention layer, the feature matrices Q, K, and V are divided into h sub-matrices according to the second dimension, and then dot-product self-attention calculations are performed between the corresponding sub-matrices, that is:
[0097]
[0098] where i ∈ [1, h] is the i-th sub-matrix corresponding to the Q, K, and V matrices. Finally, the calculation results of all heads are concatenated according to the second dimension as the output of the multi-head self-attention: mhto = concat(head1, head2,..., head h )W O , where is the parameter matrix. Finally, the output after the linear transformation is connected with the original input as a residual connection to obtain the final output v of the multi-head self-attention layer mth . In the experiment, h in the three-layer multi-head self-attention is set to 8, then The calculation process of the feature interaction module is shown in Algorithm 2.
[0099] Algorithm 2. Calculation of the feature interaction module.
[0100] Input: Feature matrix
[0101] Output: Feature vector mhto.
[0102] 1.
[0103] 2. FOR subscript i ∈ [0, m) DO
[0104] 3. Calculate the vector a i Mean μ i , variance σi
[0105] 4. Vector Element normalization j ∈ [0, d)
[0106] 5. Vector Element standardization
[0107] 6. END FOR
[0108] 7.
[0109] 8. Split the Q, K, and V matrices evenly into h sub - matrices along the second dimension
[0110] 9. begin = TRUE;
[0111] 10. FOR i ∈ [0, h) DO
[0112] 11.
[0113] 12. IF begin THEN
[0114] 13. h = head i
[0115] 14. begin = FALSE
[0116] 15. ELSE
[0117] 16. h = concat(h, head i )
[0118] 17. END IF
[0119] 18. END FOR
[0120] 19. mhto = hW O
[0121] 20. RETURN mhto
[0122] 5. Fully - connected layer
[0123] The feature vectors extracted by each sub - module are concatenated in this layer as the input x0 of the fully - connected layer, x0 = Concat([v local , v overall , v convs , v mth ). The data in each layer of the fully - connected layer will undergo the following transformation:
[0124] x l+1 = f(W l x l + b l ) (5)
[0125] where W lis the parameter matrix of the l-th layer, and f is the activation function. The ReLU activation function is selected to provide non-linear learning ability. The Sigmoid function is used as the activation function in the last layer of the fully connected layer, and the output result is between 0 and 1, representing the probability that the model infers the user clicks on the current candidate item.
[0126] 6. Loss Function
[0127] The negative log-likelihood function, which is common in binary classification tasks, is used as the loss function in model learning. The calculation formula of the negative log-likelihood function is shown in Equation (6):
[0128]
[0129] where y i is the label value corresponding to the sample x i , and p(x i ) is the probability that the model predicts the user clicks on a specific candidate item.
[0130] 7. Experiments and Discussions
[0131] 7.1 Dataset and Experimental Settings
[0132] Electronic (http: / / jmcauley.ucsd.edu / data / amazon / ). Amazon product data is a free dataset publicly available on the Amazon website, which contains information about Amazon website users and items, as well as users' browsing records and ratings of items. The e-commerce sub-dataset is selected, which records information of 192,403 users, 63,001 kinds of commodity information, 801 commodity categories, and 1,689,188 browsing records.
[0133] Books. It is also a sub-dataset from Amazon product data. This dataset collected 8,898,041 browsing and rating records of 603,669 users on a total of 367,983 books. The books are divided into 1,580 categories. In this paper, 1 million records are extracted from the set and made into a dataset that meets the model input requirements.
[0134] Alibaba(https: / / tianchi.aliyun.com / dataset / dataDetail?dataId=56).Ali_Display_Ad_Click is an advertising display click-through rate prediction dataset provided by Alibaba. The dataset contains randomly sampled advertising display and user click data from Taobao between May 6, 2017, and May 13, 2017. The processed dataset contains information on 460,907 users, 232,811 advertisements, 5,053 advertisement categories, and a total of 1,284,512 click records. The statistical information of the three datasets used is shown in Table 1.
[0135] Table 1 Statistical Information of Datasets
[0136]
[0137] In this invention, the neural network framework tensorflow2 is used to implement the design and construction of the model, and the contrast model and the multi-fine-grained interest network are trained on an NVIDIA TITAN RTX (24GB) graphics card. In terms of model parameter configuration, the dimension of discrete data types in all datasets in the embedding layer is set to 64. The number of neurons in the fully connected layer is set to 256, 128, 64, and 1 in sequence from the bottom layer to the end layer. Among them, the ReLU activation function is used in the first three layers, and the Sigmoid activation function is used in the last output layer. At the same time, Dropout is used during the training process to randomly turn off a part of the neurons with a probability of 0.5 to prevent overfitting. All models and data use mini-batch training of 1024, a learning rate of 0.001, and the Adam optimizer for parameter learning.
[0138] 7.2 Contrast Model
[0139] LR. Logistic regression is a classic shallow model that was applied to the click-through rate prediction scenario earlier. Its simple structure and efficient training and inference capabilities have made it widely popular in industrial deployments.
[0140] Wide&Deep uses a combination of a shallow linear structure and a deep network to simultaneously improve the model's memory and feature extraction capabilities.
[0141] DeepFM replaces the shallow sub-module of Wide&Deep with a factorization machine to enhance the model's explicit feature interaction capabilities.
[0142] DIN designs a local activation unit that adaptively assigns weights to the user's historical interaction items, thereby extracting different user interest representations for different candidate items, greatly improving the model's expressive ability.
[0143] DIEN combines the attention mechanism with the gated recurrent neural network, improving the model's sequence modeling ability while learning diverse interest representations for specific items.
[0144] Different from the usual attention mechanism, DAMIN uses the reciprocal of the Euclidean distance between vectors as an indicator of the attention score. Then, the user behavior sequence vectors after weight transformation are fed into a three-layer multi-head self-attention layer to extract multiple interest expressions of the user.
[0145] 7.3 Evaluation Metrics
[0146] AUC (Area Under Curve) refers to the area under the curve of the Receiver Operating Characteristic Curve (ROC). AUC can indicate the accurate ranking between positive and negative examples in the prediction results in binary classification tasks, so it is widely used in click-through rate prediction tasks. The calculation result of AUC ranges between 0 and 1, and the higher the score, the higher the speculation accuracy of the model. The AUC metric is used to evaluate the performance of different models.
[0147] Since the AUC result is 0.5 in the case of random model output, in order to better quantify the performance differences between different models, the present invention uses the RelaImpr metric to calculate, and its calculation formula is as follows:
[0148]
[0149] where AUC a refers to the AUC score obtained by model a, and AUC b is the AUC score obtained by model b.
[0150] 7.4 Experimental Results
[0151] Table 2 shows the experimental results of each model on three datasets. All experimental data are the means of 5 parallel experiments with standard deviations attached. Among them, the optimal experimental results corresponding to each dataset are marked in bold, and the sub-optimal results are marked with an underline.
[0152] Table 2 Experimental Results (AUC)
[0153]
[0154] As shown in the results in Table 2, the models based on deep learning have achieved a significant lead over logistic regression in terms of performance indicators, which verifies the efficient modeling ability of neural networks. In DeepFM and Wide&Deep, which both belong to the "deep and wide" structure, DeepFM has achieved a slight lead over Wide&Deep. This is because the DeepFM model uses factorization machines to enhance its feature interaction capabilities, but this advantage is not very obvious due to the limitation of the number of features in the data. In addition, it can be seen that the models based on sequence modeling, DIN, DIEN, and DAMIN, have achieved better results than the "deep and wide" structure models in terms of results, which also illustrates the importance of user behavior sequences in the task of click-through rate prediction. MFGIN achieved the best results in all three data sets. Compared with the baseline model (DIN) under the RelaImpr indicator, MFGIN improved by 2.03% on the Electronic data set, 2.27% on the Alibaba data set, and 1.82% on the Books data set. This advantage is due to the contribution of the overall interest extraction module and the feature interaction module in the model. The experimental results prove the advanced nature of MFGIN.
[0155] The above embodiments are intended to illustrate the present invention, not to limit the present invention. Any solution that is a simple transformation of the present invention belongs to the protection scope of the present invention.
Claims
1. A deep click-through rate prediction method supporting multiple fine-grained interest extraction, characterized in that: It includes the following steps: S1. Use the local interest activation layer in the deep interest network to learn the local interest representation of the user; In the local interest activation layer, the embedding vector of the candidate item is represented by v i The user's historical behavior sequence B = [i1, i2, …, i maxlen is transformed into a dense matrix V B = [v i,1 , v i,2 , …, v i,maxlen through the embedding layer. Here, maxlen specifies the maximum length of the user behavior sequence, v i,1 is the embedding representation of the first item in the user's historical clicks, and so on. v i is used as the query item of the attention mechanism. The matrix V B is the key and value matrix of the attention mechanism. The attention scores are calculated for each historical item through the local interest activation layer, and finally, the weighted sum of all historical item vectors is taken with the attention scores as weights to obtain the user's local interest representation vector v local ; S2. In the overall interest extraction layer, use multiple non-linear fully connected layers to expand the user's expression space and integrate it into the user behavior sequence, and then use the key-value pair attention mechanism to learn the overall interest representation of the user; S3. Use the multi-core convolutional layer to adaptively divide the long sequence into short-term behavior sequences and model these subsequences; S4. Use a multi-head self-attention layer to model user, item, and context features, and introduce the second-order interaction information between features in an implicit form; the embedding vector of the context feature, the embedding vector of the item feature, and the user-side features after two layers of nonlinear transformation are concatenated into a matrix Among them, m is the number of input features, d is the set embedding dimension; then, each value a in the matrix A i,j The following transformation is performed via layer normalization: where \(i\in[1,m]\), \(j\in[1,d]\), \(a\) i,j represents the element in the \(i\)-th row and \(j\)-th column of matrix \(A\), \(\in\) is \(1E\) -15 ; The layer normalization transformation for each row of matrix \(A\) is as follows: where, g i and b i are both learnable parameters; the output of layer normalization is the matrix In the multi-head self-attention layer, the feature matrix is transformed into three new feature matrices by three parameter matrices: Among them, the parameter matrix In the multi-head self-attention layer, the feature matrices Q, K, and V are respectively divided into h sub-matrices according to the second dimension, and then dot-product self-attention calculations are performed between the corresponding sub-matrices, that is: where \(i\in[1,h]\) corresponds to the \(i\)-th sub-matrix of the \(Q\), \(K\), and \(V\) matrices; the calculation results of all heads are concatenated along the second dimension as the output of the multi-head self-attention: \(mhto = concat(head1, head2, \ldots, head h )W O , where is the parameter matrix; Finally, the output after the linear transformation is connected with the original input as a residual connection to obtain the final output \(v\) of the multi-head self-attention layer mth ; S5. Use a multi-layer perceptron to predict the results of the features learned in steps S1 to S4, and output the click probability of the user on the candidate item.
2. The deep click-through rate prediction method for supporting multiple fine-grained interest extractions according to claim 1, characterized in that: Step S2 specifically includes the following steps: S21. In the overall interest extraction layer, first use the ReLU non-linear activation function to expand the embedded representation of the user representation, and the dimension of the expanded vector is the same as the vector dimension of the item; S22. Concatenate the user representation v after extended representation u with each item representation vector in the user historical behavior sequence to incorporate user identification information; S23. After the non-linear transformation through two fully connected layers, the last layer outputs the weight vector v w , whose length is maxlen; S24. Finally, perform a weighted sum of all item vectors in the splicing matrix according to the corresponding values in softmax(v w ) to obtain the overall user interest representation v overall .
3. The deep click-through rate prediction method supporting multiple fine-grained interest extractions according to claim 1, characterized in that: In step S3, the modeling method is as follows: the convolutional kernel slides from top to bottom, and at the same time, the elements in the window are multiplied and added with the corresponding convolutional kernel parameters to obtain a new feature vector v map =[e1, e2, …, e maxlen . Each convolutional kernel will obtain a corresponding feature vector after the window slides, and finally these feature vectors are merged through the pooling layer into the final output v of the multi-core convolutional layer convs .
4. The deep click-through rate prediction method supporting multiple fine-grained interest extractions according to claim 1, wherein: In step S5, the feature vectors extracted by the local interest activation layer, the overall interest extraction layer, the multi-core convolutional layer, and the multi-head self-attention layer are concatenated in the fully connected layer as the input of the fully connected layer; the data in each layer of the fully connected layer will undergo the following transformation: x l+1 = f(W l x l + b l ) (5) Among them, W l is the parameter matrix of the l-th layer, and f is the activation function. The ReLU activation function is selected to provide non-linear learning ability, and the Sigmoid function is used as the activation function in the last layer of the fully connected layer. The output result is between 0 and 1, representing the probability that the model infers that the user clicks on the current candidate item.
Citation Information
Patent Citations
Advertisement click rate estimation method based on improved Transformer
CN112381581A
Click rate prediction method based on dynamic depth attention model
CN113010774A