A customer conversion rate prediction method for multi-behavior sparse data
By constructing customer feature information and customer-project interaction graph data, combining SeqPredict and GraphPredict models, using graph convolution network and contrast learning module, the sparse samples, delays and cold start problems of conversion rate prediction in multi-behavior sparse data are solved, significantly improving prediction accuracy and stability.
Patent Information
- Application Number
- CN202411557184.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-04
AI Technical Summary
When processing multi-behavior sparse data, existing deep learning models face problems such as sparse conversion samples, conversion delays, and cold start of new ads, which affects prediction accuracy and effectiveness.
A customer conversion rate prediction method for multi-behavior sparse data is proposed. By constructing sequence data of customer feature information and customer-project interaction graph data, combining SeqPredict model and GraphPredict model, using graph convolution network and comparison learning module, node embedding representation is optimized, false negative samples are identified and model performance is improved.
Through the multi-view method and multi-task learning framework, the accuracy of the model on customer interest modeling is improved, the impact of false negative samples is reduced, and the accuracy and stability of conversion rate prediction is significantly improved, especially in the scenario of sparse samples.
Smart Images

Figure CN119515440B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data mining and recommendation, and in particular relates to a customer conversion rate prediction method for multi-behavior sparse data. Background Art
[0002] With the rapid development of the Internet, e-commerce has become an important part of the modern economy. Globally, more and more consumers choose to shop online, and the transaction volume and number of customers of e-commerce platforms continue to rise. In this process of digital transformation, customer purchasing behavior has become an important indicator for measuring the success of e-commerce platforms. In order to increase sales and customer satisfaction, e-commerce platforms need to have a deep understanding of customer behavior and the motivations behind it, especially the various factors that affect customers from browsing products to final purchases. Conversion rate (CVR) prediction technology came into being. By analyzing and modeling customer behavior data, it helps the platform more accurately predict the possibility of customer purchases, thereby optimizing marketing strategies and customer experience.
[0003] Models based on deep neural networks have been widely used in conversion rate prediction, and existing methods can be divided into two main categories. The first category is based on the sequence data of customer feature information to model customer behavior for conversion rate prediction. The core idea of this type of method is to use the customer's historical behavior data, especially time series information, to capture the customer's behavior patterns and preferences at different time points, thereby improving the accuracy of the prediction. The second type of method focuses on the interactive relationship between customer characteristics and item characteristics (such as advertising, goods), and uses these interactive features for prediction. The core idea is to capture the key factors affecting the conversion rate by mining the complex relationship between customer and item characteristics.
[0004] However, deep learning models usually require a large amount of training data. In online advertising systems, although there may be millions to billions of ads, customers usually only click on a small number of ads and generate conversions on a smaller set. This data sparsity problem limits the predictive power of deep models. In addition, there are several key issues that further affect the accuracy and effectiveness of conversion rate predictions. The first is the sparsity of conversion samples. Too few positive samples make model training unreliable. The second is the conversion delay problem. There is usually a time delay from customers clicking on ads to actual conversions. For example, after customers click on game ads and download apps, they may not make payment until a few days later. This leads to the "false negative sample" problem in the training process, affecting the accuracy of predictions. Finally, there is the problem of cold start of new ads. Due to the lack of behavioral data, it is difficult for the system to effectively learn and deliver new ads in the early stage. Summary of the invention
[0005] The purpose of the embodiment of the present invention is to provide a customer conversion rate prediction method for multi-behavior sparse data to optimize the embedding representation of nodes, enhance the modeling ability of customer interests, and effectively identify the false negative sample problem caused by conversion delay.
[0006] In order to solve the above technical problems, the present invention provides a customer conversion rate prediction method for multi-behavior sparse data, which is performed according to the following steps:
[0007] Step S1: Data preprocessing, obtaining customer behavior data sets, constructing sequence data of customer feature information, and constructing click and conversion customer-item interaction graph data according to the customer number, item number, whether clicked, and whether conversion occurred in each record;
[0008] Step S2: defining a conversion rate prediction model CF4CVR, including constructing a conversion rate prediction model CF4CVR and a task definition, wherein the task is to predict the probability of a click and a conversion through customer-item feature information S, a customer set U, an item set V, and a multi-behavior interaction graph A1, A2;
[0009] Step S3, using the customer-project feature information S constructed in S2, the customer set U, the project set V, and the data set of the multi-behavior interaction graphs A1 and A2 to train the CF4CVR model and save the training weights;
[0010] Step S4, model loading and prediction, declare the CF4CVR model and load the training weights, input the customer into the model CF4CVR, and calculate the probability of conversion between the customer and other items.
[0011] Furthermore, the customer behavior data set of step S1 includes customer number, item number, whether clicked, whether a purchase behavior conversion occurred, item category, customer gender, job, location and age.
[0012] Furthermore, the conversion rate prediction model CF4CVR of step S2 includes a SeqPredict model and a GraphPredict model. The SeqPredict model includes an input embedding layer, a field pooling layer, and a multi-layer perceptron, and performs conversion rate prediction and click-through rate prediction through a normalization module composition; the GraphPredict model includes an input embedding layer, a neighbor-enhanced graph convolutional network, a multi-task learning module, a cascaded cross-behavior modeling module, and a comparative learning module.
[0013] Furthermore, the specific process of step S3 is as follows:
[0014] Step S31, the SeqPredict model uses the customer feature data for prediction: the embedding layer and field pooling layer of the SeqPredict model are input to map the sequence data of the customer feature information into vectors and concatenate them to represent them as x∈[N,d];
[0015] Where N represents the number of records, d represents the dimension of embedding, and x represents the feature embedding vector adopted from the feature space;
[0016] Then, the vector is normalized after passing through multiple linear layers and nonlinear activation functions to obtain the click probability of the SeqPredict model. and the probability of conversion
[0017] Step S32: The GraphPredict model uses the customer-project graph structure data for prediction. The specific process is as follows:
[0018] Using graph convolutional networks, we perform lightweight updates on nodes, aggregate high-order neighbor information of customer items in a single behavior graph, and sum and average the outputs of different layers to optimize node embedding representations. Specifically:
[0019]
[0020]
[0021] where f agg () is the aggregation function, N u,k Indicates that node u is in the behavior graph A k Neighbors below, Indicates that customer u is in Figure A k The embedding representation of the next lth layer is, Indicates that item i is in Figure A k The embedding table of the next l-1 layers, e' u,k is the final representation after average pooling of different layers, l=0 represents the 0th layer, and L represents the number of layers of the graph convolutional network;
[0022] Furthermore, considering the sparsity of customer interactions, the lack of sufficient neighbor information and customer supervision information makes it impossible to effectively learn customer preferences. In the present invention, the graph convolutional network has three layers, including input, hidden, and output layers. The embeddings of the other two layers are used as the neighbor information of the current layer embedding and as positive samples for comparative learning to strengthen the customer's supervision signal. Specifically, the embedding representations of the same node in the current layer and other layers are used as positive samples, specifically:
[0023]
[0024] Indicated in the behavior diagram Ak Next, the representation of customer u at layer l, Representation behavior diagram A k The following is the representation of item i in layer l, τ is the temperature coefficient, Indicated in the behavior diagram A k Next, the embedding of customer v in the 0th layer of the graph convolutional network is: Indicated in the behavior diagram A k The following is the embedding representation of item i in the 0th layer of the graph convolutional network. Indicated in the behavior diagram A k The following is the embedding representation of item j in the 0th layer of the graph convolutional network. Indicated in the behavior diagram A k Next, the representation of customer u at layer 0, Represents the behavior graph A through the neighbor enhancement method and the item set V k The calculated loss, Represents the neighbor enhancement method and customer set U in the behavior graph A k The calculated loss, Indicated in the behavior diagram A k The loss calculated by the neighbor enhancement method;∑ u∈U and∑ v∈U represents the traversal and summation of the customer set U, where u∈U and v∈U represent that u and v are items belonging to U, ∑ i∈V and∑ j∈V It means to traverse and sum the item set V, where i∈V and j∈V indicate that i and j are items belonging to V;
[0025] Furthermore, in order to make full use of the customer preference characteristics in the auxiliary behavior and alleviate the data sparsity in the target behavior, the customer's cascading behavior is modeled across behaviors, and the customer preference output under the current behavior graph and the customer preference under the previous behavior graph are summed as the final output under the current behavior graph:
[0026] e u,k =e u,k-1 +e′ u,k ,e i,k =e i,k-1 +e′ i,k
[0027] Among them, e u,k-1 and e i,k-1 Respectively represent node u, item i in behavior graph A k-1 The following expression, e' u,k and e' i,k Respectively represent node u, item i in behavior graph A k The output, e u,k ,e i,kFor the behavior diagram A k The final embedding representation of the customer-item;
[0028] Furthermore, due to the false negative sample problem caused by conversion delay, a contrastive learning module is designed to jointly optimize the embedding representations of customers and items learned from click behavior with the embedding representations of customers and items under conversion. Specifically:
[0029]
[0030] in, represents the embedding representation of node u,i under transformation behavior, represents the embedding representation of node u,i under the click behavior, τ represents the temperature coefficient, represents the total loss of the contrastive learning module, and represents the contrastive learning module loss of customer set U and item set V;
[0031] Then, the embedding representations under different behaviors are optimized. Specifically:
[0032]
[0033] in, Output of the GraphPredict model processing graph structure data, e i,1 represents the embedding representation of item i under behavior graph A1, The transpose of the embedding representation of customer u under behavior graph A1, e i,2 represents the embedding representation of item i under behavior graph A2, Transpose of the embedding representation of customer u under behavior graph A2;
[0034] Then, the CF4CVR model is optimized and the L2 regularization operation is performed on the parameters of the CF4CVR model, as follows:
[0035]
[0036]
[0037] Among them, k represents the behavior diagram A k Down, As positive sample pairs and negative sample pairs, represents the interaction samples that have appeared under the current behavior (interaction samples that have not appeared), σ is the sigmoid activation function, λ is the regularization coefficient, Θ represents the parameters of the model, and ∥Θ∥2 represents L2 regularization of the model parameters Θ to prevent overfitting. Represents customer u and positive sample item i in behavior graph A kThe predicted probability of interaction occurring under For customer u and negative sample item j in behavior graph A k The predicted probability of interaction occurring under represents the loss of recommendation, is the loss calculated by the neighbor enhancement method, The loss calculated for the comparison learning method, is the total loss of the GraphPredict model;
[0038] Step S33, training the model CF4CVR and maintaining the training weights.
[0039] The beneficial effects of the present invention are as follows: the present invention uses a multi-view method to comprehensively utilize customer feature sequence data and customer-project interaction information, aiming to optimize the advertising recommendation effect. The information of neighbor nodes is aggregated through GCN, and the embedded representations of the same node at different layers are regarded as positive samples, thereby reducing the impact of noise. In addition, the high-order interaction signals between customers and projects are used to deeply capture the relationship between customers and projects, thereby improving model performance. The present invention designs a cascade structure to achieve cross-behavior modeling, deeply explore the dependencies between different customer behaviors, and optimize the embedded representation under sparse conversion behavior by utilizing relatively rich click behavior embedding representations. The present invention adopts a multi-task learning framework to simultaneously predict the probability of customer interaction under different behaviors, and combines the comparative learning method to regard the embedded representations of the same customer under different behaviors as positive samples, comprehensively improving the model's ability to identify false negative samples, thereby making more accurate predictions in scenarios with sparse samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0041] Figure 1 It is a general flow chart of the method of the embodiment of the present invention. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] This embodiment includes a method for predicting customer conversion rate for multi-behavior sparse data, such as Figure 1 As shown in the figure, a dataset containing sequence data of customer feature information and graph structure data of customer-item interaction is created, and a new multi-view, multi-task and contrastive learning module is developed for conversion rate prediction to improve the accuracy of the model in modeling customer interests and predict the conversion rate more accurately.
[0044] The specific steps include:
[0045] Step S1: Data preprocessing
[0046] S11. Obtain a customer behavior data set, including customer number, item number, whether clicked, whether purchase behavior conversion occurred, item category, customer gender, job, location and age, and construct sequence data of customer feature information.
[0047] Here, conversion is an action, such as the click-to-buy behavior that occurs when a customer makes a purchase. The purchase behavior is the conversion behavior, and conversion is used when calculating the transition from one action to another.
[0048] S12. Construct click and conversion customer-item interaction graph data according to the customer number, item number, whether clicked, and whether conversion occurred in each record.
[0049] Step S2: Define the conversion rate prediction model CF4CVR
[0050] Step S21, construct a conversion rate prediction model CF4CVR, the model includes two parts: SeqPredict model and GraphPredict model, the SeqPredict model processes the feature information of customer data, and the GraphPredict model processes the customer-item interaction graph information.
[0051] The SeqPredict model consists of an input embedding layer, a field pooling layer, and a multi-layer perceptron, and is composed of normalized modules to perform conversion rate prediction and click-through rate (CTR, Click-Through Rate).
[0052] Among them, the input embedding layer maps the sequence data of customer feature information into a vector representation, and the embedding dimension of each field is 5. In the field pooling layer, the vectors of different fields are integrated through splicing operations. Subsequently, the multilayer perceptron (MLP) processes these embedded vectors, including linear layers and nonlinear activation functions. Among them, the linear layer contains two layers, the first layer performs linear transformation, and the output dimension of the second layer mapping is 128. The introduction of nonlinear activation functions provides the model with the ability to handle complex nonlinear tasks and enhances the model's expressiveness. Finally, the model uses normalization operations to predict conversion rate (CVR) and click-through rate (CTR).
[0053] Among them, the GraphPredict model includes an input embedding layer, a neighbor-enhanced graph convolutional network, a multi-task learning module, a cascade-structured cross-behavior modeling module, and a contrastive learning module.
[0054] The input embedding layer obtains the embedding vectors of customers and items involved in the click and conversion behavior interaction graph, and ensures that the dimensions of these embedding vectors are consistent with the vector dimensions output by the field pooling layer in the SeqPredict model. Next, the neighbor-enhanced graph convolutional network captures the relationship between customers and items by aggregating the information of high-order neighbor nodes to reduce the impact of irrelevant sample noise. The number of layers of the neighbor-enhanced graph convolutional network is set to 3 to better understand the complex customer-item interaction. The multi-task learning module is used to predict the probability of clicks and conversions between customers and items. The cross-behavior modeling method of the cascade structure uses the relatively rich customer intentions in click behaviors to alleviate the sparsity problem of conversion sample data. The contrastive learning method is used to treat the embedding representations of the same customer under click and conversion behaviors as positive samples, thereby alleviating the problem of false negative samples caused by conversion delays.
[0055] S22. Task Definition
[0056] Assume that the observed customer-item characteristics The sample (x, y→z) is sampled from the distribution X×Y×Z, where X represents the feature space, is the feature embedding vector sampled from the feature space, Y, Z represent the label space, y=0(1) or z=0(1) indicates that no click or conversion behavior occurs (occurs), and N represents the number of sample records; i ,y i 、z i They represent the feature embedding vector, click, and conversion labels of the i-th data in the sample record respectively; x is the feature embedding vector sampled from the feature space, y is the label indicating whether a click occurred, and z is the label indicating whether a conversion occurred.
[0057] Meanwhile, for the graph structure information of customer-project, define u as a customer and v as a project. Specifically, U = {u1, …, u m , … u |U|} and V = {v1, …, v t , … v |V|} represent the customer set and the project set respectively, where |U| and |V| represent the number of customers and projects. u |U| represents the |U|-th customer, and u m represents the m-th customer, where 1 < m < |U|. v |V| represents the |V|-th project, and v t represents the t-th project, where 1 < t < |V|. Define the behavior set Β = {b1, b2} to represent the behavior set {click, conversion}. Decompose the multi-behavior interaction graph A containing clicks and conversions into multiple single-behavior interaction graphs {A1, A2} according to the interaction type. For each behavior interaction graph A k = [a ui |U|×|V| ∈ {0, 1}, where a ui represents the value in the adjacency interaction matrix. a ui = 1 indicates that there is an interaction between customer u and project i, and a ui = 0 indicates that there is no interaction between customer u and project i. Define the bipartite graph as G = (H, ε, A), where H = U ∪ V, representing the set of customers and projects, where |B| represents the number of behaviors, and ε k represents the historical interaction under the behavior interaction graph A k . ε represents the historical interaction under all behavior interaction graphs, and k represents the k-th behavior.
[0058] The task of the present invention is to predict the probabilities of clicks and conversions through the customer-project feature information S, the customer set U, the project set V, and the multi-behavior interaction graphs A1 and A2. Specifically, the click-through rate prediction probability Pctr = p(y = 1|x), the conversion rate prediction probability Pcvr = p(z = 1|y = 1, x), and the probability of click and conversion Pctcvr = Pctr × Pcvr;
[0059]
[0060] where is the output of the SeqPredict model processing the customer feature information. Here, x represents the feature embedding vector adopted from the feature space, y is the label indicating whether a click occurs, and z is the label indicating whether a conversion occurs. is the output of the GraphPredict model processing the graph structure data, and Pctcvr is the final output of the model.
[0061] Step S3: Use the constructed dataset to train the CF4CVR model
[0062] S31, SeqPredict model uses customer feature data for prediction.
[0063] This implementation assumes that the embedding layer and field pooling layer of the input SeqPredict model maps the sequence data of customer feature information into vectors and concatenates them to represent them as x∈[N,d], where N represents the number of records, d represents the dimension of embedding, and x represents the feature embedding vector adopted from the feature space. The vector is normalized after passing through multiple linear layers and nonlinear activation functions to obtain the click probability of the SeqPredict model. and the probability of conversion
[0064] S32, GraphPredict model uses customer-project graph structure data for prediction
[0065] This implementation uses a graph convolutional network to perform lightweight updates on nodes, aggregate high-order neighbor information of customer items in a single line graph, and sum and average the outputs of different layers to optimize the node embedding representation. Specifically:
[0066]
[0067] where f agg () is the aggregation function, N u,k Indicates that node u is in the behavior graph A k Neighbors below, Indicates that customer u is in Figure A k The embedding representation of the next lth layer is, Indicates that item i is in Figure A k The embedding table of the next l-1 layers, e' u,k is the final representation after average pooling of different layers, l=0 represents the 0th layer, and L represents the number of layers of the graph convolutional network.
[0068] Considering the sparsity of customer interactions, the lack of sufficient neighbor information and customer supervision information makes it impossible to effectively learn customer preferences. In the present invention, the graph convolutional network has three layers, including input, hidden, and output layers. The embeddings of the other two layers are used as the neighbor information of the current layer embedding and as positive samples for comparative learning to strengthen the customer's supervision signal. Specifically, the embedding representations of the same node in the current layer and other layers are used as positive samples, specifically:
[0069]
[0070]
[0071] Indicated in the behavior diagram A k Next, the representation of customer u at layer l, Representation behavior diagram A k The following is the representation of item i in layer l, τ is the temperature coefficient, Indicated in the behavior diagram A k Next, the embedding of customer v in the 0th layer of the graph convolutional network is: Indicated in the behavior diagram A k The following is the embedding representation of item i in the 0th layer of the graph convolutional network. Indicated in the behavior diagram A k The following is the embedding representation of item j in the 0th layer of the graph convolutional network. Indicated in the behavior diagram A k Next, the representation of customer u at layer 0, Represents the behavior graph A through the neighbor enhancement method and the item set V k The calculated loss, Represents the neighbor enhancement method and customer set U in the behavior graph A k The calculated loss, Indicated in the behavior diagram A k The loss calculated by the neighbor enhancement method;∑ u∈U and∑ v∈U represents the traversal and summation of the customer set U, where u∈U and v∈U represent that u and v are items belonging to U, ∑ i∈V and∑ j∈V It represents the traversal and summation of the item set V, where i∈V and j∈V indicate that i and j are items belonging to V.
[0072] In addition, different behaviors will affect each other in an invisible way, so it is necessary to model customer preferences across behaviors. In order to make full use of customer preference characteristics in auxiliary behaviors and alleviate data sparsity in target behaviors, cross-behavior modeling is performed using a cascade structure. When modeling the customer's cascade behavior (for example, from click to purchase) across behaviors, the customer preference output under the current behavior graph and the customer preference under the previous behavior graph are summed as the final output under the current behavior graph:
[0073] e u,k =e u,k-1 +e′ u,k ,e i,k =e i,k-1 +e′ i,k
[0074] where e u,k-1 , e i,k-1 Respectively represent node u, item i in behavior graph A k-1 The following expression, e' u,k,e' i,k Respectively represent node u, item i in behavior graph A k The output, e u,k ,e i,k For the behavior diagram A k The final embedding representation of customer-item.
[0075] Due to the false negative sample problem caused by conversion delay, a contrastive learning module is designed. The customer interest preferences of the same customer under the construction of click behavior and conversion behavior are similar. Specifically, the embedding representation of customers and items learned from click behavior is and Embedding representation of customers and items with transformations and Perform joint optimization. It is assumed that the customer preferences constructed under auxiliary behavior and the customer preferences constructed under target behavior for the same customer should be similar, and they are regarded as positive samples.
[0076]
[0077] in represents the embedding representation of node u,i under transformation behavior, represents the embedding representation of node u,i under the click behavior, τ represents the temperature coefficient, Denotes the contrastive learning module loss of customer set U and item set V. The total loss of the contrastive learning module is
[0078] Learning customer preferences under click behaviors will affect the final customer preferences. Therefore, accurate learning of customer preferences under click behaviors is essential. Multi-task learning (MTL) jointly learns different related tasks. This implementation optimizes the embedding representation under different behaviors. Specifically:
[0079]
[0080] Output of the GraphPredict model processing graph structure data, e i,1 represents the embedding representation of item i under behavior graph A1, The transpose of the embedding representation of customer u under behavior graph A1, e i,2 represents the embedding representation of item i under behavior graph A2, Transpose of the embedding representation of customer u under behavior graph A2.
[0081] The goal is to optimize the CF4CVR model. In the graph data processing part, positive and negative sampling of interaction samples are used to calculate the BPR loss. The prediction score of samples interacting with customer u is higher than that of samples j without interaction. At the same time, in order to prevent overfitting, L2 regularization operation is performed on the model parameters, as follows:
[0082]
[0083] k represents the behavior diagram A k Down, As positive sample pairs and negative sample pairs, represents the interaction samples that have appeared under the current behavior (interaction samples that have not appeared), σ is the sigmoid activation function, λ is the regularization coefficient, Θ represents the parameters of the model, and ∥Θ∥2 represents L2 regularization of the model parameters Θ to prevent overfitting. Represents customer u and positive sample item i in behavior graph A k The predicted probability of interaction occurring under For customer u and negative sample item j in behavior graph A k The predicted probability of an interaction occurring. represents the loss of recommendation, is the loss calculated by the neighbor enhancement method, The loss calculated for the comparison learning method, is the total loss of the GraphPredict model.
[0084] S33. Train the model CF4CVR and store the training weights.
[0085] Step S4: Model loading and prediction
[0086] S41. Declare model CF4CVR and load training weights.
[0087] S42. Input the customer into the model CF4CVR and calculate the probability of conversion between the customer and other projects.
[0088] We use Ali-CCP [https: / / tianchi.aliyun.com / dataset / 408] data for experimental verification. Ali-CCP is provided by Alimama and collected from the recommendation system log of the Taobao mobile client. It contains clicks and conversion data associated with it. Taobao, as the world's largest online retail e-commerce platform, provides product recommendation services through the recommendation system to improve its user experience. Users can click on the products they are interested in when browsing the recommendation results, or further purchase the products. Therefore, the user's behavior can be abstracted into a sequence pattern: from browsing to clicking, and then to purchasing. Here, only part of the data in Ali-CCP is used for verification. The amount of data after processing is statistically shown in Table 1:
[0089] Table 1 Statistics of Ali-CCP dataset (after preprocessing)
[0090] Views Number of hits Number of purchases Number of users Number of products 8173640 379122 2409 79212 896181
[0091] In order to comprehensively evaluate the effectiveness of the proposed model, we adopted a widely used indicator AUC. AUC comprehensively considers the relationship between the true positive rate and the false positive rate. This indicator not only considers the ability to correctly classify positive samples, but also considers the ability to misclassify negative samples. Therefore, it can comprehensively evaluate the overall performance of the model. During the experiment, we used Pctr and Pctcvr in the model prediction output to calculate the corresponding CTR_AUC and CTCVR_AUC indicators and standard deviation STD to evaluate our model. The specific experiments are shown in Table 2:
[0092] Table 2 Experimental results (bold indicates the best result)
[0093] Evaluation Metrics SeqPredict SeqPredict(STD) CF4CVR CF4CVR(STD) CTR_AUC 0.6578 0.0605 0.6843(+4.028%) 0.0595 CTCVR_AUC 0.6332 0.0624 0.6486(+2.432%) 0.0674
[0094] The AUC value of CF4CVR in CTR prediction increased from 0.6578 to 0.6843, an increase of 4.028%, which shows that the method of the present invention significantly improves the distinguishing ability of the model in predicting click-through rate. At the same time, the standard deviation decreased from 0.0605 to 0.0595, indicating that the prediction results of the method of the present invention are more stable. The AUC value in CTR prediction increased from 0.6332 to 0.6486, an increase of 2.432%, which shows that the distinguishing ability of the CF4CVR model in predicting click-to-conversion rate has been improved. Although the standard deviation increased from 0.0624 to 0.0674, the overall improved AUC value shows that the model has achieved a better balance in dealing with the complexity of post-click conversion rate.
[0095] From the above evaluation results, it can be seen that the proposed method has achieved significant improvements in both CTR and CTCVR, which further verifies the effectiveness of the proposed method.
[0096] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0097] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A customer conversion rate prediction method for multi-behavior sparse data, characterized in that: The following steps are involved: Step S1: Data preprocessing, obtaining customer behavior data sets, constructing sequence data of customer feature information, and constructing click and conversion customer-item interaction graph data according to the customer number, item number, whether clicked, and whether conversion occurred in each record; Step S2: defining a conversion rate prediction model CF4CVR, including constructing a conversion rate prediction model CF4CVR and a task definition, wherein the task is to predict the probability of a click and a conversion through customer-item feature information S, a customer set U, an item set V, and a multi-behavior interaction graph A1, A2; Step S3, using the customer-project feature information S constructed in S2, the customer set U, the project set V, and the data set of the multi-behavior interaction graphs A1 and A2 to train the CF4CVR model and save the training weights; Step S4, model loading and prediction, declare the CF4CVR model and load the training weights, input the customer into the model CF4CVR, and calculate the probability of conversion between the customer and other items; The conversion rate prediction model CF4CVR includes a SeqPredict model and a GraphPredict model. The SeqPredict model includes an input embedding layer, a field pooling layer, and a multi-layer perceptron, and performs conversion rate prediction and click-through rate prediction through a normalization module. The GraphPredict model includes an input embedding layer, a neighbor-enhanced graph convolutional network, a multi-task learning module, a cascade-structured cross-behavior modeling module, and a contrastive learning module. The specific process of step S3 is as follows: Step S31, the SeqPredict model uses the customer feature data for prediction: the embedding layer and field pooling layer of the SeqPredict model are input to map the sequence data of the customer feature information into vectors and concatenate them to represent them as x∈[N,d]; Where N represents the number of records, d represents the dimension of embedding, and x represents the feature embedding vector adopted from the feature space; The vector is normalized after passing through multiple linear layers and nonlinear activation functions to obtain the click probability of the SeqPredict model and the probability of conversion ; Step S32, the GraphPredict model uses the customer-project graph structure data for prediction: S32a uses a graph convolutional network to perform lightweight updates on nodes, aggregate high-order neighbor information of customer items in a single behavior graph, and sum and average the outputs of different layers to optimize the node embedding representation. Specifically: where f agg () is the aggregation function, N u,k Indicates that node u is in the behavior graph A k Neighbors below, Indicates that customer u is in Figure A k The embedding representation of the next lth layer is, Indicates that item i is in Figure A k The embedding table of the next l-1 layers, e′ u,k is the final representation after average pooling of different layers, l=0 represents the 0th layer, and L represents the number of layers of the graph convolutional network; S32b, cross-behavior modeling is performed on the customer's cascading behavior, and the customer preference output under the current behavior graph and the customer preference under the previous behavior graph are summed as the final output under the current behavior graph: And u,k =and u,k-1 +and′ u,k ,And i,k =and i,k-1 +and′ i,k Among them, e u,k-1 and e i,k-1 Respectively represent node u, item i in behavior graph A k-1 The following expression, e′ u,k and e′ i,k Respectively represent node u, item i in behavior graph A k The output, e u,k ,e i,k For the behavior diagram A k The final embedding representation of the customer-item; S32c, jointly optimizes the embedding representations of customers and items learned from click behaviors with the embedding representations of customers and items under transformation, specifically: in, represents the embedding representation of node u,i under transformation behavior, represents the embedding representation of node u,i under the click behavior, τ represents the temperature coefficient, represents the total loss of the contrastive learning module, and represents the contrastive learning module loss of customer set U and item set V; S32d, optimizes the embedding representation under different behaviors, specifically: in, Output of the GraphPredict model processing graph structure data, ei ,1 represents the embedding representation of item i under behavior graph A1, The transpose of the embedding representation of customer u under behavior graph A1, e i,2 represents the embedding representation of item i under behavior graph A2, Transpose of the embedding representation of customer u under behavior graph A2; S32e, optimize the CF4CVR model and perform L2 regularization on the CF4CVR model parameters as follows: Among them, k represents the behavior diagram A k Down, As positive sample pairs and negative sample pairs, Indicates the interaction samples that have appeared under the current behavior. represents the interaction samples that have not appeared under the current behavior, σ is the sigmoid activation function, λ is the regularization coefficient, Θ represents the parameters of the model, and ∥Θ∥2 represents L2 regularization of the model parameters Θ to prevent overfitting. Represents customer u and positive sample item i in behavior graph A k The predicted probability of interaction occurring under For customer u and negative sample item j in behavior graph A k The predicted probability of interaction occurring under represents the loss of recommendation, is the loss calculated by the neighbor enhancement method, The loss calculated for the comparison learning method, is the total loss of the GraphPredict model; S33. Train the model CF4CVR and keep the training weights.
2. The customer conversion rate prediction method for multi-behavior sparse data according to claim 1, characterized in that: The customer behavior data set in step S1 includes customer number, item number, whether clicked, whether a purchase behavior conversion occurred, item category, customer gender, job, location and age.
3. The customer conversion rate prediction method for multi-behavior sparse data according to claim 1, characterized in that: The graph convolutional network includes input, hidden, and output layers. The embeddings of the other two layers are used as neighbor information of the current layer embedding and as positive samples for comparative learning. That is, the embedding representations of the same node in the current layer and other layers are used as positive samples. Specifically: in, Indicated in the behavior diagram A k Next, the representation of customer u at layer l, Representation behavior diagram A k The following is the representation of item i in layer l, τ is the temperature coefficient, Indicated in the behavior diagram A k Next, the embedding of customer v in the 0th layer of the graph convolutional network is: Indicated in the behavior diagram A k The following is the embedding representation of item i in the 0th layer of the graph convolutional network. Indicated in the behavior diagram A k The following is the embedding representation of item j in the 0th layer of the graph convolutional network. Indicated in the behavior diagram A k Next, the representation of customer u at layer 0, Represents the behavior graph A through the neighbor enhancement method and the item set V k The calculated loss, Represents the neighbor enhancement method and customer set U in the behavior graph A k The calculated loss, Indicated in the behavior diagram A k The loss calculated by the neighbor enhancement method;∑ u∈U and∑ v∈U represents the traversal and summation of the customer set U, where u∈U and v∈U represent that u and v are items belonging to U, ∑ i∈V and∑ j∈V It represents the traversal and summation of the item set V, where i∈V and j∈V indicate that i and j are items belonging to V.
Citation Information
Patent Citations
User portrait prediction method based on multi-source transboundary data fusion
CN114238758A
Contrast learning prediction method and system for heterogeneous knowledge graph
CN116401380A
Cited By
Sales agency strategy method for improving customer conversion rate
CN121032269A