Marketing tool output method and system based on customer data
Through the marketing tool output method based on customer data, and the use of statistical technology to perform fluctuation attribution and correlation analysis of key indicators, the problem of online ride-hailing platforms relying on experience judgment to formulate marketing strategies, achieving more efficient and accurate marketing activities.
Patent Information
- Application Number
- CN202510338680.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
AI Technical Summary
When formulating driver-side marketing activities, existing online ride-hailing platforms mainly rely on people's experience judgment, which affects the accuracy and effectiveness of marketing strategies and is difficult to adjust and optimize in a timely manner.
Using the marketing tool output method based on customer data, through statistical techniques such as hierarchical clustering, Granger causality test and Spearman rank correlation coefficient, the fluctuation attribution and correlation analysis of key indicators are recommended, reasonable marketing tools are recommended, and population and time selection is carried out.
It improves the pertinence and effectiveness of marketing activities, reduces marketing costs, improves ROI, adjusts and optimizes marketing strategies in a timely manner, and improves overall service quality and market competitiveness.
Smart Images

Figure CN120196970A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of online car-hailing travel services, specifically to the branch of driver-side marketing activities for travel platforms in the field of online car-hailing platform services, and particularly relates to a method and system for outputting marketing tools based on customer data. Background Art
[0002] The popularization of mobile intelligent terminals enables users to conveniently access the Internet and obtain online car-hailing services. This technical foundation provides a broad market space for online car-hailing platforms and also provides a technical platform for the implementation of driver-side marketing activities. Through mobile applications such as APPs and mini-programs, drivers can receive order information in real time, communicate with passengers, and improve service efficiency.
[0003] Currently, when many online car-hailing platforms formulate driver-side marketing activities, they still mainly rely on human experience judgment. With the continuous expansion of the online car-hailing market, formulating marketing strategies based on human experience judgment often requires a large amount of manual analysis and decision-making processes, which are time-consuming, laborious, and inefficient. In addition, due to the subjectivity and limitations of people, the accuracy and effectiveness of marketing strategies may be affected.
[0004] That is, due to the lack of scientific data support and quantitative evaluation means in traditional technologies, marketing strategies based on experience judgment are often difficult to accurately evaluate their effects. This makes it difficult for the platform to timely adjust and optimize marketing activities, thus affecting the overall service quality and market competitiveness.
[0005] Therefore, the present invention proposes a method and system for outputting marketing tools based on customer data. Summary of the Invention
[0006] In view of this, the present invention hopes to provide a method and system for outputting marketing tools based on customer data to solve or alleviate the technical problems existing in the prior art, that is, how to attribute actual data based on algorithmic logic, systematically solve the impacts brought by various data indicators, greatly improve the human and financial efficiency for users, and provide at least one beneficial option for this; the technical solution of the present invention is realized as follows:
[0007] In a first aspect, a method for outputting a marketing tool based on customer data:
[0008] (I) Overview:
[0009] This solution aims to attribute fluctuations to key indicators and rank indicators according to correlation through statistical techniques such as hierarchical clustering, Granger Causality Test, and Spearman's correlation coefficient for ranked data. The model will recommend reasonable marketing tools for highly correlated indicators through rule strategies, and select populations and time periods to preset the basic configuration of marketing tools. At the same time, the model will use historical data trends to summarize marketing methods, form empirical strategies, and continuously update and optimize through real-time data feedback mechanisms.
[0010] (II) Technical solution:
[0011] In order to achieve the above technical goals, this plan chooses to build a comprehensive model, and its implementation method is as follows:
[0012] 2.1 Data preprocessing:
[0013] Collect and integrate key indicator data, including user behavior data and market data, to form data set D, and perform routine preprocessing operations such as cleaning, denoising, and standardization on data set D.
[0014] 2.2 Step S1, user group division:
[0015] Apply a hierarchical clustering algorithm to the dataset D and divide users into different groups based on key indicators. i The specific sub-steps include:
[0016] 2.2.1 Step S100, initialization:
[0017] Each user in the data set D is regarded as an initial cluster, forming a cluster set C = {c1, c2, ..., c n}, where n is the number of users; and initialize the distance matrix M, where M[i][j] represents the i-th cluster c i and the jth c j The distance between.
[0018] 2.2.2 Step S101, calculate cluster distance:
[0019] The distance between each pair of clusters is calculated using Euclidean distance:
[0020] Suppose the key indicator vector of two users (i.e. clusters) u and v is u=(u1,u2,...,u k ) and v=(v1,v2,...,v k), where k is the number of key indicators. Then the Euclidean distance between users u and v is calculated as:
[0021]
[0022] u i and v i respectively represent the i-th elements in vectors u and v, k is the number of key indicators, and d(u, v) represents the Euclidean distance between users u and v.
[0023] 2.2.3 Step S102, finding the nearest cluster pair and updating the distance matrix M:
[0024] Find the smallest non-diagonal element in the distance matrix M. The two clusters corresponding to this element are the nearest cluster pair. Then merge the nearest cluster pair into a new cluster.
[0025] Update the cluster set C, remove the clusters before merging, and add the new merged cluster.
[0026] Update the distance matrix M, remove the rows and columns related to the clusters before merging, calculate the distances between the new cluster and other clusters, and add them to the distance matrix M.
[0027] 2.2.4 Step S103, iteration:
[0028] Repeat steps S101 - S102 until the stopping condition is met (such as the number of clusters reaches a preset value or the minimum value in the distance matrix exceeds a threshold). Then update the final cluster set C, C = {a1, a2,..., a n}; each cluster a i represents a user group.
[0029] 2.3 Step S2, performing index fluctuation attribution and index correlation analysis:
[0030] For each group a i , apply Granger causality test to determine the cause r ai of the index fluctuation, record and output the attribution result. Use Spearman rank correlation coefficient to measure the linear relationship between different causes r i of different groups a ai , and sort the index R according to the correlation to form the feature sequence C; its specific sub-steps include:
[0031] 2.3.1 Step S200, index fluctuation attribution:
[0032] For each group a i , select the key indicators to be analyzed, and apply Granger causality test to determine the cause r ai。
[0033] The basic idea of Granger causality test is that if indicator X is the cause of indicator Y, then the past values of X will help predict the future values of Y. The method is as follows:
[0034] Let X t and Y t be two time series (the former is the time series to be predicted or explained, and the latter is the time series that is considered to possibly have an impact on the dependent variable), and m is the maximum number of lag periods. Estimate the following autoregressive model (AR):
[0035] Y t = α0 + α1Y t-1 +…+ α m Y t-m + β1X t-1 +…+ β m X t-m + ∈ t ;
[0036] where ε t is an error term. α0 represents the intercept term of the regression model. α1,…,α m represent the regression coefficients of the lagged terms Y t-1 ,…,Y t-m of the dependent variable Y. These coefficients measure the impact of the past values of the dependent variable on the current value. β1,…,β m represent the regression coefficients of the lagged terms X t-1 ,…,X t-m of the independent variable X on the dependent variable Y t . These coefficients measure the impact of the past values of the independent variable on the current value of the dependent variable.
[0037] Then, calculate the residual sum of squares RSS1 of the above regression model, and estimate the following regression model (excluding the lagged terms of X):
[0038] Y t = γ0 + γ1Y t-1 +…+ γ m Y t-m + η t ;
[0039] where η t is another error term. γ0 represents the intercept term of the regression model. γ1,…,γ m represent the regression coefficients of the lagged terms Y t-1 ,…,Y t-m of the dependent variable Y. These coefficients measure the impact of the past values of the dependent variable on the current value.
[0040] Then, calculate the residual sum of squares RSS0 of the above regression model again, and then calculate the statistic F:
[0041]
[0042] where T is the number of samples.
[0043] Finally, according to the F statistic and the given significance level, judge whether the null hypothesis holds. If the F statistic is greater than the critical value, reject the null hypothesis, and consider that X is the Granger cause of Y, and establish a cause r ai .
[0044] 2.3.2 Step S201, Index Correlation Analysis:
[0045] Calculate the Spearman rank correlation coefficient: For each group a i , use the Spearman rank correlation coefficient ρ to measure the linear relationship between different causes r ai :
[0046]
[0047] where d is the rank difference between any two variables (group a i ), and m is the number of observations.
[0048] According to the magnitude of the Spearman rank correlation coefficient, sort the causes r i in group a ai to form a feature sequence C, where C = {r ai1 , r ai2 ,..., r aiM}, and M is the number of causes.
[0049] 2.4 Step S3, Marketing Tool Recommendation and Configuration:
[0050] Parse the feature sequence C. When the index hits the rule strategy, perform circle selection on the population and time corresponding to each index R i , and recommend the corresponding marketing tool. Its sub-steps include:
[0051] 2.4.1 Step S300, Parse the Feature Sequence C, and Initially Select Marketing Tools:
[0052] For each group a i , parse its feature sequence C = {r ai1 , r ai2 ,..., r aim}. According to the cause r ai in the feature sequence, according to the preset mapping relationship, recommend a reasonable marketing tool T ai . The mapping relationship can be expressed as: rai →T ai , where → represents the mapping operation.
[0053] 2.4.1.1 Step S3000, Read Rule Policies:
[0054] Read a preset series of rule policies S, and each rule policy corresponds to one or more specific reasons.
[0055] The rule policies are expressed as: S = {s1, s2,..., s n}, where s i represents the i-th rule policy.
[0056] 2.4.1.2 Step S3001, Match Rule Policies:
[0057] For each reason r in the feature sequence ai , the matching process can be expressed as r ai ∈ sj , where s j represents the j-th rule policy.
[0058] If the reason r ai hits a certain rule policy s j , then output the matching result. If not, execute Step S3002.
[0059] The matching result can be expressed as: (r ai , s j ), indicating that the reason r ai hits the rule policy s j .
[0060] 2.4.1.3 Step S3002, Activate Compensation Mechanism:
[0061] Calculate the similarity between the reason r ai and each rule policy s j . Select the rule policy with the highest similarity as the relatively optimal solution:
[0062] S30020, Similarity Calculation: sim(r ai , s j ), indicating the similarity between the reason r ai and the rule policy s j .
[0063] S30021, Initialize the relatively optimal solution s optimal :
[0064] s optimal = argmax sj ∈S im (r ai,sj), where S is the set of regular strategies.
[0065] S30022, with the priority assigned as priority(s j ), indicating the priority of the regular strategy s j .
[0066] S30023, obtain the relatively optimal solution s optimal :
[0067] s optima l = argmax sj ∈S unmatchedpriority (s j )
[0068] where S unmatched is the set of unhit regular strategies.
[0069] In this way, even if the reason r ai does not directly hit the preset regular strategy, we can find a relatively optimal regular strategy to deal with it through the compensation mechanism.
[0070] 2.4.2 Step S301, perform selection and basic configuration:
[0071] For each index R i , perform the selection operation to determine the corresponding population and time range.
[0072] The selection operation is: P i = f(R i ), where P i represents the selection result and f represents the selection function.
[0073] According to the selection result P i , preset the basic configuration for the marketing tool T ai . The basic configuration includes marketing content, channels, and time.
[0074] (III) Mechanism for solving technical problems:
[0075] 3.1 User group division:
[0076] Hierarchical clustering algorithm: It can divide users into different groups according to key indicators, which helps to understand user behavior and needs more accurately. Euclidean distance calculation can measure the similarity between users and is the key to the clustering process. Through continuous iteration until the stop condition is met, the stability and accuracy of the clustering result are ensured.
[0077] 3.2 Index fluctuation attribution and correlation analysis:
[0078] Granger causality test: This method can determine the reasons leading to the fluctuations of indicators, providing a favorable basis for the formulation of subsequent marketing strategies. The Spearman rank correlation coefficient measures the linear relationship between different reasons, helping to understand the interaction between various factors.
[0079] 3.3 Marketing Tool Recommendation and Configuration
[0080] Based on the reasons in the feature sequence, reasonable marketing tools are recommended, realizing the effective docking of data and marketing strategies. Even if the reasons do not directly hit the preset rule strategies, a relative optimal solution can be found through the compensation mechanism, improving the flexibility and adaptability of the system.
[0081] Based on the indicators, the target population and time range are determined, and the basic configuration is preset for the marketing tools, realizing the precise customization of marketing strategies.
[0082] In the second aspect, a marketing tool output system based on customer data:
[0083] The system includes a processor and a memory connected to the processor. Program instructions are stored in the memory. When the program instructions are executed by the processor, the processor executes the marketing tool output method as described above. And it also includes components connected to the processor:
[0084] (1) A user group division module that applies the hierarchical clustering algorithm to divide users into groups;
[0085] (2) An attribution analysis module for attributing the fluctuations of indicators and analyzing the correlation for each user group;
[0086] (3) A marketing tool recommendation and configuration module that recommends reasonable marketing tools and configures them according to the attribution results and correlation analysis.
[0087] Compared with the prior art, the beneficial effects of the present invention are:
[0088] (1) Effective attribution of indicator fluctuations: In the technical solution of the present invention, the Granger causality test is used to determine the reasons leading to the fluctuations of indicators, which helps enterprises accurately identify the key factors affecting user behavior or market trends. Through attribution analysis, enterprises can timely adjust their marketing strategies to cope with market changes or fluctuations in user needs. The Spearman rank correlation coefficient is used to measure the linear relationship between different reasons, which helps enterprises more comprehensively understand the interaction and influence between various factors. The results of the correlation analysis can provide a scientific basis for the decision-making of enterprises, avoiding blind following or judgment based on experience.
[0089] (2) Personalized marketing tool recommendation and configuration: Based on the attribution results and correlation analysis, the technical solution of the present invention can recommend reasonable marketing tools for enterprises and perform personalized configuration. This helps enterprises improve the pertinence and effectiveness of marketing activities, reduce marketing costs, and increase ROI (return on investment).
[0090] (3) Precise user group segmentation: By applying the hierarchical clustering algorithm, the technical solution of the present invention can more precisely segment user groups, enabling users within each group to have similar key indicator characteristics. This helps enterprises better understand the needs and behavior patterns of different user groups and provides strong support for formulating subsequent marketing strategies.
[0091] (4) Improve data analysis and decision-making efficiency: The technical solution of the present invention integrates multiple links such as data preprocessing, user group segmentation, indicator fluctuation attribution, and marketing tool recommendation, forming a complete data analysis and marketing strategy formulation process. This helps enterprises improve the accuracy and efficiency of data analysis, shorten the decision-making cycle, and quickly respond to market changes. Brief Description of the Drawings
[0092] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0093] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0094] Figure 2 It is a schematic diagram of the execution process of the method of the present invention;
[0095] Figure 3 It is a schematic diagram of the system composition of the present invention. Detailed Embodiments
[0096] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the drawings. Many specific details are set forth in the following description to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below;
[0097] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0098] Glossary:
[0099] (1) Key metrics: Quantifiable metrics including user activity, conversion rate, purchase frequency, average order amount, etc. These metrics can directly reflect users' purchase behaviors and consumption habits, as well as the effects of marketing activities. By conducting attribution analysis on the fluctuations of these key metrics, we can deeply understand the reasons for changes in user behavior, thus providing a basis for formulating more precise marketing strategies. The so-called quantification means normalizing the above metrics. For example, normalizing the data to an interval value between [0, 1] for analysis.
[0100] (2) Autoregressive model: Used to estimate the impact of the past values of the dependent variable Y on the current value. In the Granger causality test, this model is used as a benchmark model that does not contain the lagged terms of the independent variable X, so as to compare with the model that contains the lagged terms of X. By comparing the sum of squared residuals of the two models, we can test whether the independent variable X is the Granger cause of the dependent variable Y.
[0101] (3) Maximum lag order: The maximum number of past values of the independent variable considered. In this specific embodiment, there are actually no lagged terms of the independent variable X, only lagged terms of the dependent variable Y. Here, m still represents the maximum number of past values of Y that we consider.
[0102] Embodiment 1: Please refer to Figures 1 - 2 , this embodiment discloses a method for outputting a marketing tool based on customer data. Before implementing this solution, it is necessary to collect and integrate key metric data, including user behavior data and market data, to form a data set D, and perform conventional preprocessing operations such as cleaning, denoising, and standardization on the data set D. Then, perform the following steps S1 to S3.
[0103] In this embodiment, regarding step S1, user group division: This step aims to divide users into different groups according to key metrics through the hierarchical clustering algorithm. The hierarchical clustering algorithm is a clustering analysis method. It calculates the distances between data points and gradually merges or divides the data points to form the final clustering result. This step applies the hierarchical clustering algorithm to the data set D and divides users into different groups a i . Its specific sub-steps include:
[0104] (1) Step S100, Initialization: This step is the starting point of the hierarchical clustering algorithm. By treating each user in the dataset as an independent cluster, it provides a basis for subsequent calculations of distances between clusters and merging of clusters.
[0105] Treat each user in dataset D as an initial cluster to form a cluster set C = {c1, c2,..., c n}, where n is the number of users; and initialize the distance matrix M, where M[i][j] represents the distance between the i-th cluster c i and the j-th c j .
[0106] (2) Step S101, Calculate cluster distances: Euclidean distance is a method to measure the distance between two points in a multi-dimensional space. By calculating the Euclidean distance between the key index vectors of two users (i.e., clusters), their similarity or difference can be quantified. Calculate the distance between each pair of clusters using Euclidean distance:
[0107] Let the key index vectors of two users (i.e., clusters) u and v be u = (u1, u2,..., u k ) and v = (v1, v2,..., v k ), where k is the number of key indices. Then the Euclidean distance between users u and v is calculated as:
[0108]
[0109] u i and v i represent the i-th elements in vectors u and v respectively, k is the number of key indices, and d(u, v) represents the Euclidean distance between users u and v.
[0110] (3) Step S102, Find the nearest cluster pair and update the distance matrix M: This step is the core of the hierarchical clustering algorithm. By continuously merging the nearest cluster pairs, the number of clusters is gradually reduced until the stop condition is met. Updating the distance matrix M is to accurately calculate the distances between the new cluster and other clusters after each cluster merge.
[0111] Find the smallest non-diagonal element in the distance matrix M. The two clusters corresponding to this element are the nearest cluster pair. Then merge the nearest cluster pair into a new cluster.
[0112] Update the cluster set C, remove the clusters before merging, and add the new merged cluster.
[0113] Update the distance matrix M, remove the rows and columns related to the clusters before merging, calculate the distances between the new cluster and other clusters, and add them to M.
[0114] (4) Step S103, Iteration: This step is the end stage of the hierarchical clustering algorithm. By continuously iterating to calculate the clustering distance and merge the clustering pairs until the stopping condition is met. Each cluster in the final clustering set C represents a group of users with similar key indicators.
[0115] Repeat steps S101 - S102 until the stopping condition is met (such as the number of clusters reaches the preset value or the minimum value in the distance matrix exceeds the threshold). Then update the final clustering set C, C = {a1, a2,..., a n}; Each cluster a i represents a group of users.
[0116] It can be understood that in this step, the hierarchical clustering algorithm is used to divide users into groups, which enables marketers to more accurately understand the characteristics and needs of different user groups, thereby formulating more precise marketing strategies and improving marketing efficiency. After understanding the characteristics and needs of different user groups, enterprises can more reasonably allocate marketing resources, such as advertising budgets, promotional activities, etc., to achieve the optimal utilization of resources. Moreover, by providing personalized marketing services for different user groups, user satisfaction and loyalty can be improved, thereby enhancing the user experience.
[0117] Furthermore, the Python execution program for step S1 is as follows:
[0118] import numpy as np
[0119] def euclidean_distance(u, v):
[0120] """Calculate the Euclidean distance between two users (i.e., clusters)"""
[0121] return np.sqrt(np.sum((u - v) ** 2))
[0122] def hierarchical_clustering(D, num_clusters):
[0123] """Apply the hierarchical clustering algorithm to the dataset D and divide users into different groups"""
[0124] n = len(D) # Number of users
[0125]
[0126] # Update the distance matrix
[0127] for i in range(n):
[0128] if i != x and i != y:
[0129] M[x][i] = M[i][x] = min(M[x][i], M[y][i])
[0130] M = np.delete(M, y, axis = 0)
[0131] M = np.delete(M, y, axis = 1)
[0132] # Return the final clustering set
[0133] return list(C.values())
[0134] # Input dataset D
[0135] D = np.array([[1, 2], [2, 3], [3, 4], [8, 9], [9, 10]])
[0136] # Call the hierarchical clustering algorithm to divide users into 3 groups
[0137] clusters = hierarchical_clustering(D, 3)
[0138] print(clusters)
[0139] In the above program, the dataset D is a two-dimensional array, where each row represents a user and each column represents a key metric. num_clusters is the preset number of clusters. C is a dictionary used to store the clustering set. Initially, each user is considered as an independent cluster. M is a two-dimensional array used to store the distances between each pair of clusters. Initially, all distances are set to 0.
[0140] Use the euclidean_distance function to calculate the Euclidean distances between each pair of clusters (i.e., users) and store them in M. In each iteration, find the minimum non-diagonal element in M, and the two clusters corresponding to this element are the closest cluster pair. Merge the closest cluster pair and update C and M. Repeat this process until the number of clusters in C reaches num_clusters. Return the final clustering set C, where each cluster represents a group of users.
[0141] This program implements the hierarchical clustering algorithm, which can automatically divide users into different groups according to key indicators without manual intervention. Through cluster analysis, enterprises can gain a deeper understanding of the characteristics and needs of different user groups, providing strong support for formulating precise marketing strategies.
[0142] In this embodiment, regarding step S2, index fluctuation attribution and index correlation analysis: This step aims to apply the Granger causality test to each user group ai to determine the reason rai for the fluctuation of the key indicator and record the attribution result. Subsequently, the Spearman rank correlation coefficient is used to measure the linear relationship between different reasons rai, and the indicators are sorted according to the correlation to form the feature sequence C. That is, for each group a i , apply the Granger causality test to determine the reason r for the index fluctuation ai , record and output the attribution result. Use the Spearman rank correlation coefficient to measure the different reasons r of different groups a i and the linear relationship between them, and sort the indicator R according to the correlation to form the feature sequence C; its specific sub-steps include: ai
[0143] (1) Step S200, index fluctuation attribution:
[0144] For each group a i , select the key indicator to be analyzed, and apply the Granger causality test to determine the reason r for its fluctuation ai .
[0145] The basic idea of the Granger causality test is that if indicator X is the cause of indicator Y, then the past values of X will help predict the future values of Y. By constructing regression models with and without the lag terms of X, and calculating the sum of squared residuals of the two models, and then calculating the F statistic. According to the F statistic and the given significance level, it is judged whether X is the Granger cause of Y. This step systematically analyzes the reasons for the fluctuations of key indicators in each user group through the Granger causality test, providing a basis for subsequent correlation analysis and marketing strategy formulation.
[0146] Specifically, let X t and Y t be two time series (the former is the time series to be predicted or explained, and the latter is the time series considered to have an impact on the dependent variable), and m is the maximum number of lag periods. Estimate the following autoregressive model (AR):
[0147] Y t = α0 + α1Y t-1 + … + α m Y t-m + β1Xt-1 +…+ β m X t-m + ∈ t ;
[0148] where ε t is an error term. α0 represents the intercept term of the regression model. α1, …, α m represent the lagged terms Y t-1 , …, Y t-m of the dependent variable Y, and the regression coefficients of these terms measure the impact of the past values of the dependent variable on its current value. β1, …, β m represent the lagged terms X t-1 , …, X t-m of the independent variable X on the dependent variable Y t , and the regression coefficients of these terms measure the impact of the past values of the independent variable on the current value of the dependent variable.
[0149] Then, calculate the residual sum of squares RSS1 of the above regression model and estimate the following regression model (excluding the lagged periods of X):
[0150] Y t = γ0 + γ1Y t-1 + … + γ m Y t-m + η t ;
[0151] where η t is another error term. γ0 represents the intercept term of the regression model. γ1, …, γ m represent the lagged terms Y t-1 , …, Y t-m of the dependent variable Y, and the regression coefficients of these terms measure the impact of the past values of the dependent variable on its current value.
[0152] Then, calculate the residual sum of squares RSS0 of the above regression model again, and further calculate the statistic F:
[0153]
[0154] where T is the number of samples.
[0155] Finally, based on the F statistic and the given significance level, determine whether the null hypothesis holds. If the F statistic is greater than the critical value, reject the null hypothesis and conclude that X is the Granger cause of Y and establish a causal relationship r ai .
[0156] (2) Step S201, Index Correlation Analysis: The Spearman rank correlation coefficient is a non-parametric measure of correlation used to measure the linear relationship between two variables. By calculating the rank differences between any two reasons rai and substituting them into the formula for the Spearman rank correlation coefficient, their correlation can be obtained. In this step, by calculating the Spearman rank correlation coefficient, the linear relationship between different reasons rai is analyzed, and they are sorted according to the correlation to form the feature sequence C. This helps to identify the reasons that have the greatest impact on the behavior of the user group and provides strong support for formulating precise marketing strategies.
[0157] Specifically, calculate the Spearman rank correlation coefficient: For each group a i , use the Spearman rank correlation coefficient ρ to measure the linear relationship between different reasons r ai :
[0158]
[0159] where d is the rank difference between any two variables (group a i ), and m is the number of observations.
[0160] According to the magnitude of the Spearman rank correlation coefficient, sort the reasons r i in group a ai and form the feature sequence C, where C = {r ai1 , r ai2 ,..., r aiM}, and M is the number of reasons.
[0161] It can be understood that through Granger causality test and Spearman rank correlation coefficient analysis, the key reasons leading to the fluctuations in the behavior of the user group can be accurately identified, and more precise marketing strategies can be formulated based on these reasons. After understanding the degree of influence of different reasons on the behavior of the user group, enterprises can allocate marketing resources, such as advertising budgets, promotional activities, etc., more reasonably to achieve the optimal utilization of resources. By providing personalized marketing services targeting the key reasons for the behavior of the user group, the satisfaction and loyalty of users can be improved, thereby enhancing the user experience. Both the Granger causality test and the Spearman rank correlation coefficient analysis are methods for attributing and analyzing the correlation of actual data based on algorithm logic, with high analysis efficiency and accuracy.
[0162] Furthermore, the Python execution program for step S2 is as follows:
[0163]
[0164]
[0165]
[0166]
[0167] attribution_results, feature_sequences = step_s2(data, groups, maxlag, significance_level)
[0168] print(attribution_results)
[0169] print(feature_sequences)
[0170] In the above program:
[0171] (1) Granger causality test: The granger_causality_test function performs the Granger causality test using the grangercausalitytests method in the statsmodels library. It accepts two time series x and y, and the maximum number of lags maxlag, and returns the F-statistic and the corresponding p-value. If the p-value is less than the significance level, then x is considered the Granger cause of y.
[0172] (2) Spearman rank correlation coefficient: The spearman_correlation function
[0173] uses the spearmanr method in the scipy.stats library to calculate the Spearman rank correlation coefficient. It accepts a dataset containing two variables and returns the correlation coefficient and the corresponding p-value.
[0174] (3) Execution of step S2: The step_s2 function accepts the dataset data, the list of user groups groups, the maximum number of lags maxlag, and the significance level significance_level. It performs index fluctuation attribution and index correlation analysis for each group and returns the attribution results and the feature sequence C. The attribution results are a dictionary, and the values are a list of tuples containing the cause and the corresponding F-statistic. The feature sequence is also a dictionary, with the keys being the user groups and the values being a list of causes sorted by correlation.
[0175] In this embodiment, regarding step S3, marketing tool recommendation and configuration:
[0176] Parse the feature sequence C, and when the index hits the rule strategy, recommend reasonable marketing tools;
[0177] For each index R iPerform circle selection for the corresponding population and time, and preset the basic configuration of marketing tools.
[0178] Analyze the frequent and normal problems in historical data. Summarize the data performance as an empirical strategy. When encountering the same or similar data performance, recommend corresponding marketing tools. Its sub-steps include:
[0179] (1) Step S300, parse the feature sequence C, and initially select marketing tools:
[0180] For each group a i , parse its feature sequence C = {r ai1 , r ai2 ,..., r aim}. According to the reason r ai in the feature sequence, recommend a reasonable marketing tool T ai according to the preset mapping relationship. The mapping relationship can be expressed as: r ai → T ai , where → represents the mapping operation.
[0181] (1.1) Step S3000, read the rule strategy:
[0182] Read a series of preset rule strategies S. The rule strategies are predefined and used to guide how the system recommends marketing tools according to the reasons in the feature sequence. Each rule strategy corresponds to one or more specific reasons.
[0183] The rule strategy is expressed as: S = {s1, s2,..., s n}, where s i represents the i-th rule strategy. The system needs to load these rule strategies first for subsequent matching and recommendation.
[0184] (1.2) Step S3001, match the rule strategy: By matching the reasons in the feature sequence with the rule strategy, the system can determine which marketing tool should be recommended.
[0185] For each reason r ai in the feature sequence, the matching process can be expressed as r ai ∈ sj , where s j represents the j-th rule strategy.
[0186] If the reason r ai hits a certain rule strategy s j , then output the matching result. If not, execute step S3002. That is, the system traverses each reason, tries to match it with each rule strategy, and after finding the matching rule strategy, outputs the matching result. The matching result can be expressed as: (r ai , sj ) indicates reason r ai matches rule strategy s j .
[0187] (1.3) Step S3002, activate the compensation mechanism: The compensation mechanism is used to handle those reasons that do not directly match the rule strategy, and find the most suitable rule strategy by calculating the similarity. It includes:
[0188] Calculate reason r ai and each rule strategy s j for similarity. Select the rule strategy with the highest similarity as the relatively optimal solution:
[0189] S30020, similarity calculation: sim(r ai , s j ) indicates the similarity between reason r ai and rule strategy s j .
[0190] S30021, initialize the relatively optimal solution s optimal :
[0191] s optimal = argmax sj ∈ S im (r ai,sj ), where S is the set of rule strategies.
[0192] S30022, assign priority as priority(s j ), which indicates the priority of rule strategy s j .
[0193] S30023, obtain the relatively optimal solution s optimal :
[0194] s optima l = argmax sj ∈ S unmatchedpriority (s j )
[0195] where S unmatched is the set of rule strategies that are not matched.
[0196] In this way, even if reason r ai does not directly match the preset rule strategy, we can still find a relatively optimal rule strategy through the compensation mechanism to deal with it.
[0197] (2) Step S301, perform circle selection and basic configuration:
[0198] For each indicator R i, perform a selection operation to determine the corresponding population and time range.
[0199] The selection operation is: P i = f(R i ), where P i represents the selection result and f represents the selection function.
[0200] Based on the selection result P i , preset the basic configuration for the marketing tool T ai . The basic configuration includes marketing content, channels, and time.
[0201] It can be understood that by parsing the feature sequence C and matching the rule strategy, the system can accurately recommend appropriate marketing tools, improving the pertinence of marketing activities. The introduction of the compensation mechanism enables the system to handle the reasons that do not directly hit the rule strategy, improving the flexibility and adaptability of the system. By performing the selection operation and presetting the basic configuration, the system can efficiently configure marketing activities, reduce manual intervention, and improve human and financial efficiency. By analyzing the frequent and normal problems in historical data, the system can summarize the data performance as an empirical strategy to provide strong support for future marketing activities.
[0202] Furthermore, the Python execution program for step S3 is as follows:
[0203] # Step S3: Marketing tool recommendation and configuration
[0204] # Input description:
[0205] feature_sequences: A set of feature sequences, each element is a list containing reasons
[0206] rule_strategies: A set of rule strategies, each element is a set containing specific reasons indicators: A set of indicators, each element is an indicator and its corresponding population and time range
[0207]
[0208]
[0209] def configure_marketing_tool(marketing_tool, audience, time_range):
[0210] print(f"Configuring {marketing_tool} for {audience} during {time_range}")
[0211] # Input
[0212] feature_sequences = [["reason1", "reason2"], ["reason3"]]
[0213] rule_strategies = [["reason1"], ["reason2", "reason3"], ["reason4"]]
[0214] indicators = [{"Population": "all_users", "Time range": "2023-01-01 to 2023-01-31"}]
[0215] # Execute step S3
[0216] recommend_marketing_tools(feature_sequences, rule_strategies, indicators)
[0217] In the above program, traverse the reasons in each feature sequence. Try to match the strategies in the set of rule strategies. If a matching strategy is found, use that strategy to recommend marketing tools. If no matching strategy is found, activate the compensation mechanism.
[0218] Calculate the similarity between the reasons and each rule strategy. Select the rule strategy with the highest similarity as the basis for recommendation.
[0219] Then traverse each indicator. Perform the operation of circle selection to determine the target population and time range (simplified here to directly use the population in the indicator). Preset the basic configuration for each recommended marketing tool, including marketing content, channels, time, etc.
[0220] Example 2: Please refer to Figure 3 , on the basis of Example 1, this example further discloses a marketing tool output system based on customer data:
[0221] The system includes a processor and a memory connected to the processor. Program instructions are stored in the memory. When the program instructions are executed by the processor, the processor executes the marketing tool output method as described above. And it also includes components connected to the processor:
[0222] (1) A user group division module that uses the hierarchical clustering algorithm to divide users into groups; its sub-modules include:
[0223] (1.1) Initialization sub-module: Treat each user as an initial cluster and initialize the distance matrix M. Construct the initial cluster set and distance matrix based on the key metric data of users.
[0224] (1.2) Clustering calculation sub-module: Calculate the distance between each pair of clusters and find the closest pair of clusters for merging. Use Euclidean distance to calculate the distance between clusters and iteratively perform the cluster merging operation.
[0225] (1.3) Result output sub-module: Output the final cluster set C, where each cluster represents a user group. A clustering result report, including the key metric features and user list of each group.
[0226] (2) Attribution analysis module for attributing metric fluctuations and analyzing correlations for each user group; its sub-modules include:
[0227] (2.1) Metric fluctuation attribution sub-module: Apply Granger causality test to determine the causes leading to metric fluctuations. Construct an autoregressive model, calculate the sum of squared residuals, and perform an F-statistic test.
[0228] (2.2) Correlation analysis sub-module: Use Spearman's rank correlation coefficient to measure the linear relationship between different causes. Calculate the rank difference, calculate Spearman's rank correlation coefficient according to the formula, and perform sorting.
[0229] (3) Marketing tool recommendation and configuration module for recommending and configuring reasonable marketing tools based on the attribution results and correlation analysis; its sub-modules include:
[0230] (3.1) Rule strategy matching sub-module: Match the attribution results with the preset rule strategies. Check whether the cause hits a certain rule strategy and output the matching result.
[0231] (3.2) Compensation mechanism sub-module: When the cause does not hit any rule strategy, activate the compensation mechanism to find a relatively optimal solution. Calculate the similarity or priority between the cause and the rule strategy, and select the strategy with the highest similarity or priority as the relatively optimal solution.
[0232] (3.3) Marketing tool recommendation sub-module: Recommend reasonable marketing tools according to the matching result or the output of the compensation mechanism. A marketing tool recommendation report, including the recommended tools, target population, time range, etc.
[0233] (3.4) Execution of circle selection and configuration sub-module: Execute the circle selection operation and preset the basic configuration for the marketing tool. Determine the target population and time range according to the metrics, and configure the marketing content, channels, and time.
[0234] All of the above embodiments merely represent the implementation manners of the relevant actual applications of the present invention. The descriptions thereof are relatively specific and detailed, but should not be construed as limitations on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
[0235] For those skilled in the art, it can be further realized that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of the present invention.
Claims
1. A method for outputting marketing tools based on customer data, comprising integrating user behavior data and market data into key indicator data to form a data set D, characterized in that: The following steps are included: S1, user group division: Apply a hierarchical clustering algorithm to the data set D and divide users into different groups a according to key indicators i ; S2, perform indicator fluctuation attribution and indicator correlation analysis: for each group a i , use Granger causality test to determine the cause of the indicator fluctuations ai , record and output the attribution results; use Spearman's rank correlation coefficient to measure different groups a i Different reasons for ai The linear relationship between them is established, and the indicators R are sorted according to the correlation to form a characteristic sequence C; S3, marketing tool recommendation and configuration: parse the feature sequence C, when the indicator hits the rule strategy, for each indicator R i Circle the corresponding groups of people and time, and recommend corresponding marketing tools.
2. The marketing tool output method according to claim 1, characterized in that: In S1, the implementation method of the hierarchical clustering algorithm includes: S100, each user in the data set D is regarded as an initial cluster, forming a cluster set C = {c1, c2, ..., c n }, where n is the number of users; and initialize the distance matrix M, where M[i][j] represents the i-th cluster c i and the jth c j The distance between S101, the distance between each pair of clusters is calculated using Euclidean distance; S102, update the cluster set C, remove the cluster before the merger, and add the new cluster after the merger; update the distance matrix M, remove the rows and columns related to the cluster before the merger, calculate the distance between the new cluster and other clusters, and add them to the distance matrix M; S103, repeat S101 to S102 until the stop condition is met, and then update the final cluster set C, C = {a1, a2, ..., a n }; Each cluster a i Represents a user group.
3. The marketing tool output method according to claim 2, characterized in that: In S101, the Euclidean distance calculation method is: Among them, u i and v i denotes the i-th element in vectors u and v respectively, k is the number of key indicators, and d(u,v) denotes the Euclidean distance between users u and v.
4. The marketing tool output method according to claim 1, characterized in that: In S2, the Granger causality test method is: let X t and Y t are two time series and m is the maximum number of lags; the following autoregression model is estimated: Y t =α0+α1Y t-1 +…+a m Y t-m +β1X t-1 +…+b m X t-m +∈ t ; Among them, ε t is an error term; α0 represents the intercept term of the regression model; α1,…,α m represents the lagged term Y of the dependent variable Y t-1 ,…,Y t-m The regression coefficients of m represents the lag term X of the independent variable X t-1 ,…,X t-m For the dependent variable Y t The regression coefficient of Then, the residual sum of squares RSS1 is calculated and the following regression model is estimated: Y t =γ0+γ1Y t-1 +…+c m Y t-m +n t ; Among them, η t is another error term; γ0 represents the intercept term of the regression model; γ1,…,γ m represents the lagged term Y of the dependent variable Y t-1 ,…,Y t-m The regression coefficient of Then, the residual sum of squares RSS0 of the above regression model is calculated again, and then the statistic F is calculated: Where T is the sample size.
5. The marketing tool output method according to claim 4, characterized in that: According to the F statistic and the given significance level, determine whether the null hypothesis is valid; if the F statistic is greater than the critical value, reject the null hypothesis, and consider X to be the Granger original of Y, and establish a reason r ai .
6. The marketing tool output method according to claim 4, characterized in that: In S2, the calculation method of Spearman's rank correlation coefficient ρ is: Where d is any two variables (group a i ), m is the number of observations; According to the size of the Spearman rank correlation coefficient, i The reason ai Sort and form a characteristic sequence C, where C = {r ai1 ,r ai2 ,...,r aiM }, M is the number of causes.
7. The marketing tool output method according to claim 1, characterized in that: In S3, the feature sequence C is parsed and marketing tools are initially selected: for each group a i , analyze its characteristic sequence C = {r ai1 ,r ai2 ,...,r aim }; According to the reason r in the feature sequence ai , based on the preset mapping relationship, recommend reasonable marketing tools T ai .
8. The marketing tool output method according to claim 7, characterized in that: In S3, for each indicator R i , perform the circle selection operation to determine the corresponding population and time range; Circle operation is P i =f(R i ), where P i represents the circle selection result, and f represents the circle selection function; According to the circled result P i , a marketing tool ai Preset basic configuration; basic configuration includes marketing content, channels and time.
9. A marketing tool output system based on customer data, characterized in that: The system includes a processor and a memory connected to the processor, wherein program instructions are stored in the memory, and when the program instructions are executed by the processor, the processor executes the marketing tool output method as described in any one of claims 1-8.
10. The marketing tool output system according to claim 9, characterized in that: The processor is connected with, A user group segmentation module that uses a hierarchical clustering algorithm to segment users into groups; Attribution analysis module for attributing indicator fluctuations and performing correlation analysis on each user group; A marketing tool recommendation and configuration module that recommends and configures reasonable marketing tools based on attribution results and correlation analysis.