A Click-Through Rate Prediction Method Based on Fourier Transform
By introducing Fourier transform and attention mechanisms in click-through rate prediction, the problem of noise characteristic signal removal is solved, and more accurate click-through rate prediction and performance improvement is achieved.
Patent Information
- Application Number
- CN202211270263.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-10-18
AI Technical Summary
In click-through rate prediction, the prior art is difficult to effectively remove noise characteristic signals, resulting in model overfitting and performance degradation.
The Fourier transform-based method is used to convert the attribute features of the user's browsing information into feature embedding vectors, and process them through deep neural networks and attention mechanisms, while filtering the noise frequency signal in the frequency domain.
It effectively removes noise characteristic signals, reduces information loss, and improves the accuracy and performance of click-through rate prediction.
Smart Images

Figure CN115545211B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of digital signal processing and deep learning, and in particular to a click-through rate prediction method based on Fourier transform. Background Art
[0002] In a recommendation system, the prediction of click-through rate (CTR) is crucial, and its task is to estimate the probability that a user clicks on a recommended product. Therefore, CTR can be used to measure the quality of a recommendation algorithm. A small number of products are selected from a large number of candidate products for ranking, and then presented to users for personalized recommendation. Therefore, improving CTR is a goal of e-commerce platforms. Some traditional recommendation algorithms also play a very important role in improving CTR prediction, such as item-based collaborative filtering and matrix factorization. Now machine learning has been applied in many fields. A product has many attribute features, and when a user purchases a product, they are essentially selecting some of these features. Therefore, improving CTR prediction can be achieved through the combination of one or more features of the product. Therefore, deep learning plays a huge role in CTR prediction. Due to the large amount of product data, the input features will be very sparse and high-dimensional, causing the model to overfit. In the early stage, a logistic regression (LR) model was used for CTR prediction, but the influence of the interaction between multiple features on CTR prediction was not considered. We cannot only consider the low-order influence, and the high-order feature interaction is also an important factor affecting CTR.
[0003] Whether it is low-order or high-order feature interaction, not all feature interactions are equally useful and predictive. Interactions with useless functions may even introduce noise and have a negative impact on performance. Therefore, we need to remove these noise signals in this process. In the early stage, meaningful features were selected through manual feature engineering, but it was highly dependent on the knowledge of experts, so this was very difficult to achieve. Therefore, how to remove noise signals will be an important direction. Summary of the Invention
[0004] The purpose of the present invention is to provide a click-through rate prediction method based on Fourier transform that effectively removes noise feature signals by introducing an attention mechanism.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A click-through rate prediction method based on Fourier transform, the method includes the following steps in sequence:
[0006] (1) First, convert all the attribute features included in each piece of information browsed by the user into feature embedding vectors to obtain a combination e of feature embedding vectors;
[0007] (2) Input the combination e of the feature embedding vectors into a deep neural network, which includes multiple layers of neural networks. Each layer of the neural network outputs a hidden vector, and the last layer of the neural network outputs the first predicted value;
[0008] (3) Input each hidden vector output from the deep neural network into the attention mechanism module to obtain the second predicted value;
[0009] (4) Meanwhile, regard all the attribute features included in each piece of information browsed by the user as digital signals, directly perform Fourier transform, remove the signal of the noise frequency, and leave the effective signal to be input into the deep neural network for processing to obtain the third predicted value;
[0010] (5) Finally, add the first predicted value, the second predicted value, and the third predicted value together, and use the activation function for processing to obtain the final predicted value, so as to predict whether the user will click on this piece of information.
[0011] The specific content of step (1) is as follows: Each piece of information browsed by the user includes multiple attribute features. Combine all the attribute features together to form a feature sequence. For the i-th feature x i ∈R in this feature sequence, first convert x i into a one-hot encoded vector v i , where n i is the number of categories included in the i-th feature, that is, the dimension of the one-hot encoded vector of the i-th feature. Obtain the i-th feature embedding vector e i through formula (1):
[0012] e i =W emb,i v i (1)
[0013] Among them, e i ∈R k , is the embedding matrix of the i-th feature, and k is the dimension of the feature embedding vector;
[0014] Finally, concatenate all the feature embedding vectors together to obtain the combination e of the feature embedding vectors, as shown in the following formula:
[0015] e={e 1 ,e 2 ,e 3 …,e i ,…} (2).
[0016] Step (2) specifically refers to: using the combination e of feature embedding vectors as the input of a deep neural network, which includes multiple layers of neural networks. Each layer of the neural network outputs a hidden vector, and the first hidden vector h is calculated by formula (3). 1 :
[0017] h 1 = f(W 0 e + b 0 ) (3)
[0018] where h 1 ∈ R d , d is the output dimension, f(.) is the activation function, is the weight matrix, d 0 is the dimension of the combination e of feature embedding vectors, and b 0 ∈ R d is the offset;
[0019] Then, the second hidden vector h 2 :
[0020] h 2 = f(W 1 h 1 + b 1 ) + h 1 (4)
[0021] where h 2 ∈ R d , W 1 ∈ R d×d is the weight matrix, and b 1 ∈ R d is the offset;
[0022] And so on. Finally, the last hidden vector h l :
[0023] h l = f(W l-1 h l-1 + b l-1 ) + h l-1 (5)
[0024] where h l ∈ R d , where W l-1 ∈ R d×d is the weight matrix;
[0025] Finally, the first predicted value y H :
[0026]
[0027] Among them, W H ∈R d is the weight matrix, and b H ∈R is the offset.
[0028] Specifically, step (3) means that in the attention mechanism module, first, each hidden vector output from the deep neural network is used as input, and each hidden vector attention score a i :
[0029]
[0030] Among them, is the weight matrix, is the weight matrix, is the offset, d a is the attention layer dimension; h i is the i-th hidden vector output from the deep neural network;
[0031] Then, each hidden vector attention weight coefficient is obtained through the Softmax function:
[0032] a i ′ = Softmax(a i ) (8)
[0034] Next, each hidden vector attention weight coefficient and each hidden vector are operated to obtain the intermediate vector h a :
[0035]
[0036] Among them, h a ∈R d , and then the intermediate vector h a is input into the deep neural network to obtain the second predicted value y A :
[0037] y A = W A T h a + b A (10)
[0038] Among them, W A ∈R d is the weight matrix, and b A ∈R is the offset.
[0039] The specific content of step (4) is as follows: regarding all the attribute features included in each piece of information browsed by the user as a digital signal, which contains noise signals. The signal in the time domain is converted into a signal in the frequency domain through the discrete Fourier transform (DFT):
[0040]
[0041] The signal in the frequency domain is obtained d E is the dimension of the signal in the frequency domain. Then, the noise frequencies are filtered to obtain the signal E after removing the noise frequencies f :
[0042] F f = W f E (12)
[0043] where is the weight matrix. Then, the signal in the frequency domain is converted into a signal in the time domain through the inverse discrete Fourier transform (IDFT):
[0044]
[0045] The digital signal in the time domain is obtained d h is the dimension of the digital signal in the time domain;
[0046] Then, the third predicted value y is obtained through the deep neural network F :
[0047]
[0048] where is the weight matrix of the Fourier transform layer, and b F ∈R is the offset of the Fourier transform layer.
[0049] The specific content of step (5) is as follows: adding the first predicted value, the second predicted value, and the third predicted value, and inputting the result into the activation function sigmod to obtain the final predicted value
[0050]
[0051] It can be seen from the above technical scheme that the beneficial effects of the present invention are: first, the features contained in the text data are converted into feature vectors for processing. These features can be regarded as discrete digital signals. It is difficult to directly perform denoising on discrete digital signals. The present invention uses Fourier transform to convert time domain signals into frequency domain for processing, so that the noise frequency can be filtered out in the frequency domain and the effective signal frequency is left; second, in order to mine potential effective information, in deep learning, the latent vectors of each layer contain important information. If they are directly put into the next layer for processing, part of the information will inevitably be lost. Therefore, the present invention processes each layer of latent vectors, controls the weight of each layer of latent vectors through the attention mechanism, and then processes them through a fully connected layer, thereby reducing information loss in the deep neural network, fully mining information, and improving the user's click rate prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a flow chart of the method of the present invention;
[0053] Figure 2 It is a schematic diagram of the structure of the deep neural network in the present invention;
[0054] Figure 3 It is a comparison chart of the test results of the present invention. DETAILED DESCRIPTION
[0055] like Figure 1 As shown, a click rate prediction method based on Fourier transform includes the following steps in sequence:
[0056] (1) First, all attribute features contained in each piece of information browsed by the user are converted into feature embedding vectors to obtain a combination of feature embedding vectors e;
[0057] (2) Input the combination of feature embedding vectors e into a deep neural network. The deep neural network includes multiple layers of neural networks. Each layer of the neural network outputs a latent vector, and the last layer of the neural network outputs the first prediction value.
[0058] (3) Input each latent vector output by the deep neural network into the attention mechanism module to obtain the second prediction value;
[0059] (4) At the same time, all attribute features contained in each piece of information browsed by the user are regarded as digital signals, and Fourier transform is directly performed to remove the noise frequency signals, leaving the valid signals to be input into the deep neural network for processing to obtain the third prediction value;
[0060] (5) Finally, the first prediction value, the second prediction value, and the third prediction value are added together and processed using the activation function to obtain the final prediction value, thereby predicting whether the user will click on the information.
[0061] The specific step (1) means that each piece of information browsed by the user contains multiple attribute features. All the attribute features are combined together to form a feature sequence. For the \(i\)-th feature \(x\) in this feature sequence i ∈R, first convert \(x\) i into a one-hot encoding vector \(v\) i , where n i is the number of categories contained in the \(i\)-th feature, that is, the dimension of the one-hot encoding vector of the \(i\)-th feature. The \(i\)-th feature embedding vector \(e\) is obtained through Equation (1) i :
[0062] e i = W emb,i v i (1)
[0063] where \(e\) i ∈R k , is the embedding matrix of the \(i\)-th feature, \(k\) is the dimension of the feature embedding vector;
[0064] Finally, all the feature embedding vectors are concatenated together to obtain a combination \(e\) of feature embedding vectors, as shown in the following formula:
[0065] e = {e 1 , e 2 , e 3 …, e i ,…} (2).
[0066] The specific step (2) means that the combination \(e\) of feature embedding vectors is used as the input of a deep neural network. The deep neural network contains multiple layers of neural networks. Each layer of neural network outputs a hidden vector. The first hidden vector \(h\) is calculated by Equation (3) 1 :
[0067] h 1 = f(W 0 e + b 0 ) (3)
[0068] where \(h\) 1 ∈R d , \(d\) is the output dimension, \(f(.)\) is the activation function, is the weight matrix, \(d 0 is the dimension of the combination \(e\) of feature embedding vectors, \(b 0 ∈R d is the offset;
[0069] Then, the second hidden vector \(h\) is obtained using a residual network through Equation (4) 2 :
[0070] h 2 = f(W 1 h 1 + b 1 ) + h 1 (4)
[0071] where h 2 ∈ R d , W 1 ∈ R d×d is the weight matrix, b 1 ∈ R d is the offset;
[0072] And so on, and finally the last hidden vector h l :
[0073] h l = f(E l-1 h l-1 + b l-1 ) + h l-1 (5)
[0074] where h l ∈ R d , where W l-1 ∈ R d×d is the weight matrix;
[0075] Finally, the first predicted value y H :
[0076]
[0077] where W H ∈ R d is the weight matrix, b H ∈ R is the offset.
[0078] The specific step (3) refers to: in the attention mechanism module, first take each hidden vector output from the deep neural network as the input, and obtain the attention score a of each hidden vector through formula (7) i :
[0079]
[0080] where is the weight matrix, is the weight matrix, is the offset, d a is the dimension of the attention layer; h i is the i-th hidden vector output from the deep neural network;
[0081] Then, the attention weight coefficients of each hidden vector are obtained through the Softmax function:
[0082] a i ′ = Softmax(a i ) (8)
[0084] Next, each attention weight coefficient of the hidden vector is operated with each hidden vector to obtain an intermediate vector h a :
[0085]
[0086] where h a ∈R d Then, the intermediate vector h a is input into the deep neural network to obtain the second predicted value y A :
[0087] y A = W A T h a + b A (10)
[0088] where W A ∈R d is the weight matrix, and b A ∈R is the offset.
[0089] Step (4) specifically refers to: regarding all the attribute features included in each piece of information browsed by the user as a digital signal, and there will be noise signals in this digital signal. The signal in the time domain is converted into a signal in the frequency domain through the discrete Fourier transform DFT:
[0090]
[0091] to obtain the frequency-domain signal d E is the dimension of the frequency-domain signal. Then, the noise frequencies are filtered to obtain the signal E f without the noise frequencies:
[0092] E f = W f E (12)
[0093] where is the weight matrix. Then, the frequency-domain signal is converted into a time-domain signal through the inverse discrete Fourier transform IDFT:
[0094]
[0095] to obtain the time-domain digital signal dh is the dimension of the time-domain digital signal;
[0096] Then, a third predicted value y is obtained through a deep neural network F :
[0097]
[0098] where is the weight matrix of the Fourier transform layer, and b F ∈R is the offset of the Fourier transform layer.
[0099] Specifically, step (5) means: adding the first predicted value, the second predicted value, and the third predicted value, and inputting the result into the activation function sigmod to obtain the final predicted value
[0100]
[0101] Such as Figure 2 shown, the deep neural network consists of an input layer, an embedding layer, a hidden layer, an attention layer, a Fourier transform layer, and an output layer.
[0102] Such as Figure 3 shown, the effectiveness of the method is judged by two metrics, AUC (Area Under ROC) and Logloss (Cross Entropy). It can be seen that the present invention has the best effect on the Criteo dataset and the KKBox dataset.
[0103] In summary, the present invention studies the problem of the noise signal generated by the feature vector, searches for a more efficient and well-interpretable denoising method, converts the features into feature vectors for processing, and these feature vectors can be regarded as discrete digital signals; since it is difficult to directly denoise the discrete digital signals, the Fourier transform can convert the time-domain signal into the frequency domain for processing, so that the noise frequencies can be filtered out in the frequency domain, leaving the effective signal frequencies, and then the signal is processed subsequently. In order to mine potential effective information, in deep learning, each layer of hidden vectors contains important information. If directly put into the next layer for processing, some information will inevitably be lost. Therefore, each layer of hidden vectors is processed, the weights of each layer of hidden vectors are controlled through the attention mechanism, and then processed through a fully connected layer. The present invention has a very good effect in improving CTR prediction.
Claims
1. A click-through rate prediction method based on Fourier transform, characterized in that: This method includes the following steps in sequence: (1) First, convert all attribute features included in each piece of information browsed by the user into feature embedding vectors to obtain a combination e of feature embedding vectors; (2) Input the combination e of feature embedding vectors into a deep neural network. The deep neural network includes multiple layers of neural networks. Each layer of neural network outputs a hidden vector, and the last layer of neural network outputs a first predicted value; (3) Input each hidden vector output from the deep neural network into an attention mechanism module to obtain a second predicted value; (4) At the same time, regard all attribute features included in each piece of information browsed by the user as a digital signal, directly perform Fourier transform, remove the signal of the noise frequency therein, and leave the effective signal to be input into the deep neural network for processing to obtain a third predicted value; (5) Finally, add the first predicted value, the second predicted value, and the third predicted value, and use an activation function for processing to obtain the final predicted value, so as to predict whether the user will click on this piece of information; The specific content of step (3) is as follows: in the attention mechanism module, first, each hidden vector output from the deep neural network is used as an input, and each hidden vector attention score a is obtained through formula (7). i : a i = W i T ReLU(W a h i + b a ) (7) Among them, is the weight matrix, is the weight matrix, is the offset, d a is the attention layer dimension; h i is the i-th hidden vector output in the deep neural network; Then obtain the attention weight coefficient of each hidden vector through the Softmax function: a i ′ = Softmax(a i ) (8) Next, each hidden vector attention weight coefficient and each hidden vector are operated on to obtain an intermediate vector h a : where h a ∈R d , and then input the intermediate vector h a into the deep neural network to obtain the second predicted value y A : y A = W A T h a + b A (10) Among them, W A ∈R d is the weight matrix, and b A ∈R is the offset; The specific content of step (4) is: regard all attribute features included in each piece of information browsed by the user as a digital signal. There will be noise signals in this digital signal. Convert the signal in the time domain into a signal in the frequency domain through the discrete Fourier transform DFT: Obtain the frequency-domain signal d E is the dimension of the frequency-domain signal, then filter out the noise frequencies to obtain the signal E with the noise frequencies removed f : E f = W f E(12) Among them, is the weight matrix, and then the frequency-domain signal is converted into a time-domain signal through the inverse discrete Fourier transform IDFT: Obtain a time-domain digital signal d h is the dimension of the time-domain digital signal; Then, a third predicted value y is obtained through a deep neural network F : Among them, is the weight matrix of the Fourier transform layer, and b F ∈R is the offset of the Fourier transform layer; The specific content of the step (5) is as follows: add the first predicted value, the second predicted value, and the third predicted value, and input the result into the activation function sigmod to obtain the final predicted value 2. The click-through rate prediction method based on Fourier transform according to claim 1, characterized in that: The specific content of step (1) is as follows: Each piece of information browsed by the user contains multiple attribute features. All the attribute features are combined together to form a feature sequence. For the i-th feature x in this feature sequence i ∈R, first convert x i into a one-hot encoded vector v i , where n i is the number of categories included in the i-th feature, that is, the dimension of the one-hot encoded vector of the i-th feature. The i-th feature embedding vector e i is obtained through formula (1): e i = W emb,i v i (1) where e i ∈R k , is the embedding matrix of the i-th feature, and k is the dimension of the feature embedding vector; Finally, concatenate all the feature embedding vectors together to obtain a combination e of feature embedding vectors, as shown in the following formula: e = {e 1 , e 2 , e 3 ..., e i ,...} (2).
3. The click-through rate prediction method based on Fourier transform according to claim 1, characterized in that: The specific description of step (2) is as follows: The combination e of the feature embedding vectors is used as the input of the deep neural network. The deep neural network includes multiple layers of neural networks, and each layer of neural network outputs a hidden vector. The first hidden vector h is calculated by formula (3). 1 : h 1 = f(W 0 e + b 0 ) (3) where h 1 ∈R d , d is the output dimension, f(.) is the activation function, is the weight matrix, d 0 is the dimension of the combination e of the feature embedding vectors, b 0 ∈R d is the offset; Then, the second hidden vector h is obtained using the residual network through formula (4). 2 : h 2 = f(W 1 h 1 + b 1 ) + h 1 (4) where h 2 ∈ R d , W 1 ∈ R d×d is the weight matrix, and b 1 ∈ R d is the offset; And so on. Finally, the last hidden vector h is obtained through formula (5). l : h l = f(W l-1 h l-1 + b l-1 ) + h l-1 (5) where h l ∈R d , where W l-1 ∈R d×d is the weight matrix; Finally, the first predicted value y is output through the deep neural network H : where, W H ∈ ℝ d is the weight matrix, and b H ∈ ℝ 2 is the bias.
Citation Information
Patent Citations
Click rate prediction method based on attention mechanism
CN111538761A
Voice endpoint detection method and device based on neural network, equipment and medium
CN112489677A