Customer conversion pipeline and funnel model analysis method
Patent Information
- Application Number
- CN202511981404.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-12-25
AI Technical Summary
[0003]然而,现有的策略评估方案往往无法准确评估不同策略对最终成交的实际贡献价值,进而导致营销资源的错误配置
在本申请的实施例中,通过引入倾向得分与条件平均处理效应,在观测数据中有效剥离了客户自身属性带来的选择性干扰,实现了基于反事实推断的精准归因,从而准确评估不同策略对最终成交的实际贡献价值,以便于合理配置营销资源。
Smart Images

Figure CN121937158B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to a customer conversion pipeline and funnel model analysis method. Background Technology
[0002] With the widespread application of artificial intelligence in the marketing field, AI agents and automated marketing systems can be used to reach customers at scale. In long-cycle B2B (Business to Business) or complex B2C (Business to Consumer) sales transactions, customers often go through a long decision-making chain from lead to closing, during which they are continuously influenced by various marketing strategies (such as incentive offers and regular follow-ups). Businesses urgently need to understand the actual contribution value of different strategies to the final sale.
[0003] However, existing strategy evaluation schemes often fail to accurately assess the actual contribution of different strategies to the final transaction, leading to the misallocation of marketing resources. Summary of the Invention
[0004] The embodiments of this application provide a customer conversion pipeline and funnel model analysis method, which aims to accurately evaluate the actual contribution value of different strategies to the final transaction, so as to rationally allocate marketing resources.
[0005] In a first aspect, embodiments of this application provide a customer conversion pipeline and funnel model analysis method, the customer conversion pipeline and funnel model analysis method comprising: Obtain marketing interaction logs including customer lead attribute vectors, historical marketing outreach actions, and corresponding stage flow status tags; Based on the customer lead attribute vector, determine the customer preference score for each of the historical marketing outreach actions; Based on the customer preference score and the stage transition status label, determine the conditional average processing effect of each of the historical marketing outreach actions on each of the customer lead attribute vectors; Based on the aforementioned conditional averaging effect, determine the net improvement in customer conversion rate for each of the aforementioned historical marketing outreach actions; Based on the net improvement value, generate the strategy effectiveness attribution analysis results.
[0006] In the above embodiments, by introducing propensity score and conditional averaging effects, selective interference caused by customer attributes is effectively removed from the observed data, achieving accurate attribution based on counterfactual inference. This provides a rigorous mathematical basis for the refined allocation and automated scheduling of marketing resources, and improves the scientific nature and conversion efficiency of marketing funnel management.
[0007] In one embodiment, obtaining marketing interaction logs, including customer lead attribute vectors, historical marketing outreach actions, and corresponding stage transition status tags, includes: Acquire customer conversion history data, which includes at least one of the following: AI agent interaction data, enterprise profile static data, and customer acquisition channel traceability data; Based on the customer conversion history data, a four-dimensional data cube is generated, which includes the lead dimension, marketing stage dimension, time dimension and strategy dimension. The marketing interaction log is obtained by slicing the four-dimensional data cube.
[0008] In the above embodiments, by constructing a four-dimensional data cube, the original business data from multiple sources and heterogeneity is reconstructed into a standardized multi-dimensional tensor model. By utilizing the structured characteristics of the cube and the slicing technique, high-quality training samples with feature alignment and temporal consistency can be extracted quickly and in batches, which greatly improves the data retrieval efficiency.
[0009] In one embodiment, after generating the four-dimensional data cube based on the customer conversion history data, the method further includes: Based on the four-dimensional data cube, determine the stage inventory saturation, inflow-outflow balance, average conversion time deviation, and strategy activity of each marketing stage in the preset marketing funnel. The health of the preset marketing funnel is obtained by weighting and summing the stage inventory saturation, the inflow and outflow balance, the average conversion time deviation, and the strategy activity.
[0010] In the above embodiments, by leveraging the structured advantages of the four-dimensional data cube, four core indicators reflecting the operational status of the sales pipeline are quickly calculated, and a comprehensive funnel health value is generated through a weighted fusion algorithm, thereby sensitively capturing abnormal states such as blockages, idleness, and inefficiency in the marketing process.
[0011] In one embodiment, after obtaining the funnel health of the preset marketing funnel, the method further includes: Determine whether the health status of the funnel is less than a preset benchmark health threshold; If the funnel health is less than the baseline health threshold, determine the contribution weight of each marketing stage to the decline in customer conversion rate; Based on the contribution weights, the target marketing stage among the multiple marketing stages is determined and designated as the bottleneck stage. For the bottleneck stage, corresponding bottleneck analysis data is generated and output.
[0012] In the above embodiments, through automated threshold determination and attribution weight calculation, the core bottleneck links that lead to the decline in overall efficiency are accurately extracted from the complex chain conversion process, and multi-dimensional diagnostic data is provided, enabling managers to quickly concentrate resources to solve key bottlenecks, thereby maximizing the recovery of potential conversion losses.
[0013] In one embodiment, after generating the strategy effectiveness attribution analysis results based on the net improvement value, the method further includes: Determine the computational cost per instance of AI agent retrieval for each of the aforementioned historical marketing outreach actions; Based on the strategy performance attribution analysis results and the computing power cost of a single AI agent call, an operations research optimization model is constructed. The operations research optimization model takes maximizing the expected total revenue as the objective function and uses the upper limit of the number of concurrent threads of the AI agent and the preset maximum allowed frequency of customer disturbance as constraints. Solving the operations research optimization model yields AI agent task instructions, which include target marketing outreach actions to be performed by the AI agent.
[0014] In the above embodiments, by constructing a constrained operations research optimization model, the problem of formulating marketing strategies is transformed from experience-based judgment into a mathematical programming problem. Under the premise of fully considering the system's computing power bottleneck and customer experience constraints, the accurate strategy attribution results (net improvement value) are used to guide decision-making. This enables the automatic search for the optimal strategy combination that can bring maximum economic benefits under limited computing resources and reach opportunities, thereby improving the input-output ratio of marketing activities.
[0015] In one embodiment, determining the conditional average processing effect of each of the historical marketing outreach actions on each of the customer lead attribute vectors based on the customer propensity score and the stage transition status label includes: By using the customer lead attribute vector to perform regression prediction on the stage flow status label, a theoretical customer conversion probability value is obtained. Based on the theoretical customer conversion probability value, the customer conversion deviation value is determined; Using the customer preference score, the action deviation value of the corresponding historical marketing outreach action is obtained; A causal forest model that fits the correlation between the customer conversion deviation value and the action deviation value; The conditional average treatment effect was determined using the causal forest model.
[0016] In the above embodiments, the dual orthogonalization concept combined with the causal forest algorithm is used to eliminate the confounding effects of customer characteristics on conversion results and action selection, effectively solving the common selection bias problem in observation data. It can accurately identify the heterogeneous effects of marketing strategies in different customer segments, providing a statistically rigorous quantitative basis for differentiated marketing decisions.
[0017] In one embodiment, determining the conditionally averaged treatment effect using the causal forest model includes: The stage transition state labels are mapped to the discrete state space of a semi-Markov model. Using the semi-Markov model, a time series analysis is performed on the customer lead attribute vector to determine the dwell time distribution function and state transition probability matrix for each discrete state. Based on the dwell time distribution function and the state transition probability matrix, determine the sample survival probability value of the customer lead attribute vector within the current observation time window; Using the sample survival probability value, determine the sample correction weights for the customer lead attribute vector; Substituting the sample correction weights into the causal forest model, a weighted fit is performed on the association between the customer conversion deviation value and the customer propensity score to obtain the conditional average treatment effect.
[0018] In the above embodiments, a semi-Markov model is introduced to model the temporal characteristics of the marketing process, calculate the sample survival probability value reflecting the degree of data censoring, and generate the sample survival probability value based on this and inject it into the causal forest model. This effectively solves the sample bias problem caused by the truncation of long-term sales leads under a limited observation window, thereby improving the accuracy and robustness of the model in evaluating the causal effects of marketing strategies in complex, long-chain B2B sales scenarios.
[0019] In one embodiment, determining the customer propensity score for each of the historical marketing outreach actions based on the customer lead attribute vector includes: Obtain a gradient boosting decision tree classification model, wherein each of the historical marketing outreach actions is used as the classification label of the gradient boosting decision tree classification model; Using the gradient boosting decision tree classification model, classification prediction is performed based on the customer lead attribute vector to obtain the predicted probability value for each of the classification labels; The predicted probability value is truncated within a numerical range to obtain the customer preference score.
[0020] In the above embodiments, the powerful nonlinear fitting capability of gradient boosting decision trees is utilized to accurately extract the distribution pattern of marketing actions from high-dimensional sparse customer features. The numerical truncation technique solves the probability boundary stability problem in causal inference, ensuring that the generated propensity score can not only truly reflect the selection bias of historical data, but also meet the strict requirements of numerical stability for subsequent causal effect estimation, thus providing a reliable mathematical basis for eliminating confounding factors.
[0021] In one embodiment, determining the net improvement in customer conversion rate for each of the historical marketing outreach actions based on the conditional averaging effect includes: Determine the expected value of the conditional average treatment effect and the variance of the effect estimate for each of the aforementioned historical marketing outreach actions; Based on the variance of the effect estimate, the statistical error boundary value is determined; Based on the expected value of the effect and the statistical error boundary value, the net improvement value of each of the historical marketing outreach actions on customer conversion rate is determined.
[0022] In the above embodiments, when evaluating the effectiveness of marketing actions, a statistical error boundary correction mechanism based on variance estimation is introduced. By calculating the lower bound of the confidence interval as the net improvement value, not only is the expected return of the strategy taken into account, but the uncertainty risk of the prediction is also rigorously quantified, so that the system can automatically filter out those "pseudo-efficient" strategies that seem to have high returns but actually have insufficient samples and low credibility.
[0023] In one embodiment, generating strategy effectiveness attribution analysis results based on the net improvement value includes: Based on the net increase value, determine the strategic marginal contribution value of each of the historical marketing outreach actions; Based on the marginal contribution value of the strategy, determine the strategy return rate indicator; Based on the marginal contribution value of the strategy and the return rate of the strategy, the attribution analysis results of the strategy effectiveness are generated.
[0024] In the above embodiments, by calculating the marginal contribution value of the strategy and the strategy return rate, the technical indicators (conversion rate improvement) of marketing outreach actions are transformed into financial indicators (revenue increase and return on investment). This dual-dimensional attribution analysis can not only identify the key actions that contribute the most to the total revenue, but also identify the most efficient low-cost leverage actions, providing precise quantitative decision support for enterprises to find the optimal balance between "increasing revenue" and "controlling costs".
[0025] Secondly, embodiments of this application provide a customer conversion pipeline and funnel model analysis system, which is used to perform the customer conversion pipeline and funnel model analysis method as described in any of the preceding claims.
[0026] The beneficial effects of the embodiments of this application are as follows: In the embodiments of this application, by introducing propensity score and conditional averaging effects, selective interference caused by customer attributes is effectively removed from the observed data, achieving accurate attribution based on counterfactual inference, thereby accurately assessing the actual contribution value of different strategies to the final transaction, so as to rationally allocate marketing resources. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a schematic flowchart of an embodiment of the customer conversion pipeline and funnel model analysis method provided in this application. Detailed Implementation
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. In addition, in the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0030] In a first aspect, embodiments of this application provide a customer conversion pipeline and funnel model analysis method, the execution subject of which is a customer conversion pipeline and funnel model analysis system (hereinafter referred to as "the system").
[0031] Specifically, refer to Figure 1 Customer conversion pipeline and funnel model analysis methods may include: S101. Obtain marketing interaction logs including customer lead attribute vectors, historical marketing outreach actions, and corresponding stage flow status tags.
[0032] In this application embodiment, marketing interaction logs refer to raw time-series data recorded in a customer relationship management (CRM) system or marketing automation platform, which contains structured information in multiple dimensions.
[0033] A customer lead attribute vector is a high-dimensional numerical array used to digitally represent the characteristics of a single potential customer. It covers static attributes (such as company size, industry, and registered capital) and dynamic behavioral attributes (such as historical browsing frequency, average dwell time, and active time period).
[0034] Historical marketing outreach actions refer to the specific interventions that a company has implemented for customers in the course of its past business. As a treatment variable in the causal inference model, its value can be a discrete action type number or a sparse vector that has undergone one-hot encoding.
[0035] The stage transition status label refers to the binary result or probability identifier of whether a customer has successfully entered the next marketing stage within a preset time window after receiving a certain historical marketing outreach action, and serves as the outcome variable in causal inference.
[0036] S102. Based on the customer lead attribute vector, determine the customer preference score for each historical marketing outreach action.
[0037] In this embodiment, the customer propensity score refers to the conditional probability that a customer will accept a specific historical marketing outreach given a customer lead attribute vector. This metric is statistically used to quantify selection bias in a sample, specifically addressing the interference of confounding factors such as "high-quality customers are often prioritized for high-quality resources" on attribution results.
[0038] S103. Based on customer preference scores and stage transition status labels, determine the conditional average processing effect of each historical marketing outreach action on each customer lead attribute vector.
[0039] In the embodiments of this application, the Conditional Average Treatment Effect (CATE) refers to the difference between the expected gain obtained by applying a specific historical marketing outreach action to an individual with specific characteristics (i.e., a specific customer lead attribute vector) through counterfactual inference and the expected result of not applying the action.
[0040] In some embodiments of this application, the conditional average treatment effect can be determined using a Double Machine Learning (DML) framework or a Causal Forest algorithm. Specifically, the system uses customer preference scores to perform inverse probability weighting (IPW) or orthogonalization on the samples, eliminating the influence of customer lead attribute vectors on the selection of historical marketing outreach actions, as well as their direct influence on stage transition status labels (i.e., the main effect). On this residual space, the local correlation between the residuals of historical marketing outreach actions and stage transition status labels is fitted. This non-parametric estimation method can accurately capture the heterogeneity sensitivity of different customer groups to the same strategy, i.e., identify which customer lead attribute vectors are most effective for the outreach action.
[0041] S104. Based on the conditional averaging effect, determine the net increase in customer conversion rate for each historical marketing outreach action.
[0042] In the embodiments of this application, the net increase in customer conversion rate refers to the additional conversion probability increment that is entirely attributable to the historical marketing outreach action after mathematical calculations have stripped away natural conversion factors.
[0043] In some embodiments of this application, the determination of the net improvement value may be combined with confidence interval estimation. Since the conditional averaging effect is an estimate based on a statistical model, the system calculates the variance or standard error of the effect. The conditional averaging effect is considered a valid net improvement value only if its lower bound is greater than zero and passes a statistical significance test. Furthermore, the system maps the conditional averaging effect to specific business metrics, such as converting probability differences into expected percentage increases in conversion rates, thereby transforming abstract algorithm parameters into measurable business performance indicators.
[0044] S105. Based on the net improvement value, generate the strategy effectiveness attribution analysis results.
[0045] In this application embodiment, the strategy effectiveness attribution analysis result refers to a structured report or data instruction generated based on the net improvement value, used to guide the subsequent allocation of marketing resources.
[0046] In some embodiments of this application, generating strategy performance attribution analysis results includes sorting all historical marketing outreach actions in descending order of net improvement value and selecting the strategy combination that performs best for the current customer lead attribute vector. The system generates a visual performance comparison chart showing the true contribution of different strategies after eliminating selection bias. Furthermore, this result can be encoded as input parameters for an automated scheduling system to update the execution strategy library of the intelligent agent, ensuring that in future marketing campaigns, outreach actions with the highest net improvement value are prioritized for customers with similar attributes.
[0047] As can be seen, by introducing the propensity score and conditional averaging effects, the embodiments of this application effectively remove the selective interference caused by customer attributes in the observed data, realize accurate attribution based on counterfactual inference, provide a rigorous mathematical basis for the refined allocation and automated scheduling of marketing resources, and improve the scientific nature and conversion efficiency of marketing funnel management.
[0048] In some embodiments of this application, obtaining marketing interaction logs including customer lead attribute vectors, historical marketing outreach actions, and corresponding stage transition status tags includes: S201. Obtain customer conversion history data, which includes at least one of the following: AI agent interaction data, enterprise profile static data, and customer acquisition channel traceability data.
[0049] In this embodiment, customer conversion history data refers to a collection of raw records scattered across different business systems of an enterprise. AI agent interaction data refers to the complete records generated when an AI agent communicates with a customer, encompassing text transcribed by Automatic Speech Recognition (ASR), semantic intent extracted by Natural Language Processing (NLP), call duration, and sentiment polarity scores. Enterprise profile static data refers to customer background information supplemented through external business databases or third-party application programming interfaces (APIs), including but not limited to registered capital, business scope, financing rounds, and core decision-making structure. Customer acquisition channel tracing data refers to link tracking parameters recording the customer's initial contact point, such as Urchin Tracking Module (UTM) parameters, advertising campaign identifiers, and source landing page addresses.
[0050] S202. Based on customer conversion history data, generate a four-dimensional data cube, which includes the lead dimension, marketing stage dimension, time dimension, and strategy dimension.
[0051] In this embodiment, the four-dimensional data cube refers to a multi-dimensional data model constructed based on Online Analytical Processing (OLAP) technology, which reconstructs a dispersed planar table structure into a high-dimensional tensor structure. The lead dimension represents a unique customer identifier and its set of attribute characteristics, constituting the cube's entity axis; the marketing stage dimension represents standardized marketing funnel process states (such as initial contact, needs confirmation, solution demonstration, business negotiation, and contract signing), constituting the cube's state axis; the time dimension represents a time-series axis scaled in discrete time steps (such as days or weeks), used to record the moments of state changes; and the strategy dimension represents the set of specific intervention measures taken by the enterprise, constituting the cube's operational axis.
[0052] S203. Perform a slicing operation on the four-dimensional data cube to obtain the marketing interaction log.
[0053] In this embodiment, slicing refers to the process of fixing the value range of one or more dimensions in a multidimensional data model to extract a specific subset of data. In this process, slicing is not merely a simple query; it also includes flattening a high-dimensional tensor into a two-dimensional table structure suitable for input to a machine learning model.
[0054] In some embodiments of this application, the system, based on analysis requirements, fixes the time dimension (selecting a specific historical period) and the marketing stage dimension (selecting the funnel stage to be analyzed), vertically segmenting a data plane containing lead characteristics and strategic behaviors from a four-dimensional data cube. Subsequently, the system formats and transforms the sliced data, mapping the values in the tensor to customer lead attribute vectors, historical marketing outreach actions, and corresponding stage transition status labels in the marketing interaction log. This cube-based extraction method ensures that the extracted sample data is strictly aligned in the time series and automatically associates with the context state, avoiding the performance bottlenecks and logical errors caused by the join queries of traditional Structured Query Language (SQL).
[0055] As can be seen, the embodiments of this application construct a four-dimensional data cube to reconstruct the original business data from multiple sources and heterogeneity into a standardized multi-dimensional tensor model. By utilizing the structured characteristics of the cube and the slicing technique, high-quality training samples with feature alignment and temporal consistency can be extracted quickly and in batches, which greatly improves the data retrieval efficiency.
[0056] In some embodiments of this application, after generating a four-dimensional data cube based on customer conversion history data, the method further includes: S301. Based on the four-dimensional data cube, determine the stage inventory saturation, inflow and outflow balance, average conversion time deviation, and strategy activity of each marketing stage in the preset marketing funnel.
[0057] In this application embodiment, the stage inventory saturation refers to the ratio between the number of customer leads currently in a specific marketing stage and the preset optimal carrying capacity of that stage, used to measure whether there is resource backlog or idleness in that stage. Inflow-outflow balance refers to the ratio between the number of new leads entering that marketing stage and the number of leads leaving that stage (including those transferred to the next stage or lost) within a specific observation period, used to determine the smoothness of the flow channel. Average conversion time deviation refers to the difference between the actual average dwell time of current customer leads in that marketing stage and the preset baseline conversion period; if this value is positive and large, it indicates a significant conversion lag. Strategy activity refers to the frequency density of marketing outreach actions actually triggered for customer leads within that marketing stage, used to evaluate the intensity of the company's intervention measures.
[0058] In some embodiments of this application, the system calculates the aforementioned metrics by performing dicing and drill-down operations on a four-dimensional data cube. Specifically, for stage inventory saturation, the system counts the current cross-sectional count of the cube in the "marketing stage dimension"; for inflow-outflow balance, the system calculates the difference in state changes between adjacent time steps in the "time dimension"; for average conversion time deviation, the system calculates the state retention time of each lead ID based on the "time dimension" and takes the average; for strategy activity, the system counts the proportion of non-empty operation records in the "strategy dimension". Through this multi-dimensional data processing, the abstract sales status is deconstructed into quantifiable physical metrics.
[0059] S302. The stage inventory saturation, inflow-outflow balance, average conversion time deviation, and strategy activity are weighted and summed to obtain the funnel health of the preset marketing funnel.
[0060] In this embodiment, funnel health refers to a normalized scalar value used to comprehensively characterize the operational status of the entire sales pipeline. Weighted summation refers to assigning different weight coefficients to the four indicators based on the emphasis of the business scenario, and then performing a linear combination calculation.
[0061] As can be seen, this application embodiment utilizes the structured advantages of a four-dimensional data cube to quickly calculate four core indicators reflecting the operational status of the sales pipeline, and generates a comprehensive funnel health value through a weighted fusion algorithm, thereby sensitively capturing abnormal states such as blockages, idleness, and inefficiency in the marketing process.
[0062] In some embodiments of this application, after obtaining the funnel health of the preset marketing funnel, the method further includes: S401. Determine whether the funnel health status is less than the preset baseline health threshold.
[0063] In this embodiment, the baseline health threshold refers to the minimum health standard that must be achieved to maintain normal business operations, determined based on the statistical distribution of the company's historical long-term operational data. This threshold can be a fixed scalar value (e.g., set based on the historical mean minus twice the standard deviation) or a time-varying parameter that is dynamically adjusted according to the peak and off-peak seasons of business.
[0064] In some embodiments of this application, the system periodically reads the latest funnel health status through a real-time monitoring module and compares it with a baseline health threshold configured in the rule engine. If the comparison result indicates that the current funnel health status is lower than the threshold, the system is triggered to enter an anomaly diagnosis mode; otherwise, the system continues to maintain normal monitoring.
[0065] S402. If the funnel health is less than the baseline health threshold, determine the contribution weight of each marketing stage to the decline in customer conversion rate.
[0066] In this application embodiment, contribution weight refers to a numerical indicator used to quantify the degree of responsibility that each marketing stage bears for the deterioration of the overall funnel health.
[0067] In some embodiments of this application, contribution weights are determined using an attribution algorithm based on Shapley Value or a gradient-based sensitivity analysis method. The system simulation restores the conversion metrics of each marketing stage to a baseline level, calculating the theoretical recovery rate of the overall funnel health. A larger recovery rate indicates a more significant impact of that marketing stage on the current decline in health, and its corresponding contribution weight is also higher. This calculation method comprehensively considers the conversion rate decline at each stage and the traffic volume carried by that stage, avoiding misjudgments caused by focusing only on the conversion rate decline while ignoring traffic weights.
[0068] S403. Based on contribution weights, identify the target marketing stage among multiple marketing stages and designate it as the bottleneck stage.
[0069] In this embodiment, the target marketing stage refers to a specific stage among numerous marketing stages that is identified as primarily causing performance degradation. The bottleneck stage refers to the specific state identifier of this target marketing stage within the system, indicating that it is a limiting factor in the current business flow.
[0070] In some embodiments of this application, the system sorts the calculated contribution weights of each marketing stage in descending order and selects the top few marketing stages, either ranking first or having a cumulative weight exceeding a preset proportion (e.g., 80%), as target marketing stages. The system marks these stages as bottleneck stages and adds anomaly tags to the corresponding dimensions of the four-dimensional data cube for subsequent deeper-dimensional drill-down analysis.
[0071] S404. For the bottleneck stage, generate and output the corresponding bottleneck analysis data.
[0072] In this embodiment of the application, bottleneck analysis data refers to detailed diagnostic report data about the bottleneck stage, which includes not only the abnormal indicators themselves, but also potential related dimension information that leads to the abnormality.
[0073] In some embodiments of this application, the system performs drill-down operations within a four-dimensional data cube to analyze the performance of metrics across different strategy dimensions and lead source dimensions within that stage. This generates a structured data package containing information on the cause of the anomaly (e.g., a sudden drop in lead quality from a specific channel or the failure of a certain strategy), a comparison curve with historical data from the same period, and suggested optimization directions. This data package is then pushed to a front-end visualization interface for rendering or sent as an alarm message to business management personnel.
[0074] As can be seen, the embodiments of this application accurately identify the core bottleneck links that lead to the decline in overall efficiency from the complex chain conversion process through automated threshold determination and attribution weight calculation, and provide multi-dimensional diagnostic data, enabling managers to quickly concentrate resources to solve key bottlenecks, thereby maximizing the recovery of potential conversion losses.
[0075] In some embodiments of this application, after generating the strategy performance attribution analysis results based on the net improvement value, the method further includes: S501, Determine the computing power cost of each AI agent's single retrieval for each historical marketing outreach action.
[0076] In this embodiment, the computational cost of a single AI agent call refers to the computational resource metric required to execute a complete marketing interaction by calling the AI agent. This cost is typically expressed as the time (in milliseconds) occupied by the graphics processing unit (GPU), the number of tokens consumed by calling the application programming interface, or the billing amount of cloud computing services.
[0077] In some embodiments of this application, the system maintains a pre-set computing power cost comparison table, which records the average resource consumption value corresponding to different types of historical marketing outreach actions (such as voice outbound calls and email composition). For complex actions involving real-time generation, the system monitors the memory usage and processing time of the agent when performing inference tasks in real time, and dynamically calculates the current single call computing power cost to ensure the real-time nature and accuracy of cost estimation.
[0078] S502. Based on the strategy performance attribution analysis results and the computing power cost of a single AI agent call, construct an operations research optimization model. The operations research optimization model takes maximizing the expected total revenue as the objective function and uses the upper limit of the number of concurrent threads of the AI agent and the preset maximum allowed frequency of customer disturbance as constraints.
[0079] In this application embodiment, the operations research optimization model refers to a decision support model built based on mathematical programming theory, typically in the form of Integer Linear Programming (ILP) or mixed integer programming. The expected total revenue refers to the sum of the product of the expected new conversion probability calculated based on the net increase value and the customer lifetime value (LTV) after executing the selected strategy combination for the target customer group. The upper limit of the number of concurrent threads for the artificial intelligence agent refers to the maximum number of AI instances that can run simultaneously, representing the instantaneous throughput capacity boundary of the system. The preset maximum allowed customer disturbance frequency is a hard threshold set per unit time to ensure user experience and prevent excessive marketing from causing customer aversion or blocking.
[0080] In some embodiments of this application, the process of constructing an operations research optimization model is a process of transforming a business problem into a mathematical formula. The system defines a set of binary decision variables X_ij, representing whether to perform the j-th marketing outreach action on the i-th customer. The objective function is constructed as maximizing SUM(X_ij × Net Improvement Value_ij × Customer Value_i), where SUM represents a summation operation. Constraints are formalized as a set of inequalities: First, the concurrent demand after discounting the computational cost of all selected actions must not exceed the total number of system threads; second, for any single customer i, SUM(X_ij) must not exceed a set frequency threshold. Furthermore, the model can also incorporate a total budget constraint, i.e., the sum of the computational cost of all selected actions multiplied by the unit price must not exceed a preset marketing budget.
[0081] S503. Solve the operations research optimization model to obtain the AI agent task instructions, which include the target marketing outreach actions to be executed by the AI agent.
[0082] In this embodiment, solving the operations research optimization model refers to the process of using a mathematical solver to search for a combination of decision variables that satisfies all constraints and maximizes the objective function value. Artificial intelligence agent task instructions refer to control instructions generated by the system based on the optimal solution, which are machine code or JavaScript Object Notation (JSON) format that can be directly recognized by the execution engine. These instructions explicitly specify which clients, when, and through which channels specific operations should be performed.
[0083] In some embodiments of this application, the system invokes a high-performance commercial or open-source solver (such as a branch and bound algorithm solver) to solve the aforementioned integer programming problem precisely. For large-scale data scenarios, if the precise solution takes too long, the system can switch to a heuristic algorithm (such as a greedy algorithm or a genetic algorithm) to obtain an approximate optimal solution. After the solution is completed, the system maps actions with a decision variable of 1 (i.e., selected) to specific target marketing outreach actions, encapsulates them into task queues, and distributes them to various AI agent instances for execution.
[0084] As can be seen, the embodiments of this application transform the problem of formulating marketing strategies from empirical judgment to mathematical programming problem by constructing a constrained operations research optimization model. Under the premise of fully considering the system's computing power bottleneck and customer experience constraints, the accurate strategy attribution results (net improvement value) guide the decision-making, and realize the automatic search for the optimal strategy combination that can bring maximum economic benefits under limited computing resources and reach opportunities, so as to improve the input-output ratio of marketing activities.
[0085] In some embodiments of this application, the conditional average processing effect of each historical marketing outreach action on each customer lead attribute vector is determined based on customer preference scores and stage transition status labels, including: S601. Use customer lead attribute vectors to perform regression prediction on stage flow status labels to obtain theoretical customer conversion probability values.
[0086] In this embodiment of the application, the theoretical customer conversion probability value refers to the baseline probability that a customer can naturally move to the next sales stage based solely on their own attribute characteristics (without considering what kind of marketing outreach they actually received).
[0087] In some embodiments of this application, the system employs machine learning regression algorithms, such as Random Forest Regressor, eXtreme Gradient Boosting (XGBoost), or deep neural networks, to construct a baseline prediction model. The model's input is a high-dimensional vector of customer lead attributes, and its output is a continuous probability value in the interval [0, 1]. By training on all historical data, the model can learn general principles such as "larger company size leads to higher conversion rates," and the resulting theoretical customer conversion probability value represents the conversion potential determined by the customer's "natural endowment."
[0088] S602. Determine the customer conversion deviation value based on the theoretical customer conversion probability value.
[0089] In this embodiment of the application, the customer conversion deviation value refers to the difference between the actual observed conversion result and the theoretical baseline predicted by the model.
[0090] In some embodiments of this application, the system subtracts the theoretical customer conversion probability value obtained in step S601 from the actual stage transition status label (usually a binary value of 0 or 1) recorded in the marketing interaction log to obtain the customer conversion deviation value. Statistically, this step achieves orthogonalization of the outcome variable, that is, it eliminates the variance directly explained by customer characteristics, and the remaining residual mainly includes the processing effect caused by marketing outreach actions and random noise.
[0091] S603. Utilize customer preference scores to obtain the action deviation values for corresponding historical marketing outreach actions.
[0092] In the embodiments of this application, the action deviation value refers to the difference between the actual marketing outreach action and the theoretical action probability predicted based on the propensity score.
[0093] In some embodiments of this application, the system extracts the actual values of historical marketing outreach actions from marketing interaction logs (e.g., 1 if a specific action was performed, 0 if not), and subtracts the customer propensity score calculated in the preceding steps. Since the customer propensity score represents the probability of performing the action given certain characteristics, the residual (action bias value) obtained by subtracting the two represents the "randomness" portion of the outreach action that cannot be explained by customer characteristics. This process also involves orthogonalizing the processing variables, aiming to sever the correlation between the processing variables and covariates.
[0094] S604, a causal forest model that fits the relationship between customer conversion deviation value and action deviation value.
[0095] In this embodiment, the causal forest model is a non-parametric causal inference algorithm based on generalized random forest, designed to estimate the effects of heterogeneity treatment.
[0096] In some embodiments of this application, the system does not directly fit the original variables, but instead constructs a regression model with action deviation values as independent variables and customer conversion deviation values as dependent variables. More specifically, the system constructs a random forest composed of multiple causal trees. When each causal tree splits a node, its optimization objective is no longer to minimize prediction error (such as mean squared error (MSE)), but to maximize the heterogeneity of the treatment effect estimate, that is, to find the splitting feature that can best distinguish subgroups with different sensitivities to marketing actions. In this way, the model can capture the true causal relationship between actions and conversions in the residual space.
[0097] In a supplementary embodiment of this application, for the process of constructing split nodes in a single tree of the causal forest in S604, the system adopts a gradient-based approximation criterion to find the optimal split point. For any parent node P in the tree, the system attempts to split it into a left child node L and a right child node R. The system traverses all feature dimensions and their corresponding splitting thresholds, and for each candidate splitting scheme, calculates the following objective function value Delta: Delta=(n_L×n_R / n_P^2)×(tau_L-tau_R)^2 Where n_P, n_L, and n_R represent the number of samples in the parent node, left child node, and right child node, respectively; tau_L and tau_R represent the average treatment effect estimates in the left and right child nodes, respectively; and ^2 represents the square operation.
[0098] To calculate tau_L and tau_R, the system does not directly use the final effect value. Instead, it uses the "honesty" principle to divide the training data into a split set and an estimation set. During the splitting phase, the pseudo-residual rho for each sample under the loss function is calculated using the idea of gradient boosting trees. Therefore, tau_L is calculated as follows: tau_L=(1 / n_L)×SUM(rho_i) That is, tau_L equals the sum of the pseudo residuals rho_i of all samples i within the left child node L divided by the number of samples n_L. The system selects the splitting feature and threshold that maximizes the objective function value Delta as the optimal splitting strategy for that node. This step essentially involves finding the feature boundary that can best distinguish between two subgroups that "respond significantly differently to marketing actions".
[0099] S605. Using the causal forest model, determine the effect of conditional average treatment.
[0100] In some embodiments of this application, for any given new customer lead attribute vector, it is input into a trained causal forest model. Each tree in the model assigns the sample to a specific leaf node. Within each leaf node, the system estimates the local treatment effect of that leaf node using locally weighted regression or simple ratio calculation (i.e., the ratio between the covariance of the conversion bias value and the action bias value of the sample within that leaf node, and the variance of the conversion bias value and the action bias value of the sample within that leaf node). Finally, the system averages the estimation results of all trees in the forest for that sample to obtain the conditional average treatment effect under that specific customer characteristic. This value accurately reflects the causal gain of implementing a marketing action compared to not implementing one under that specific customer attribute condition.
[0101] As can be seen, the embodiments of this application adopt the idea of dual orthogonalization combined with the causal forest algorithm to eliminate the confounding effects of customer characteristics on conversion results and action selection, effectively solving the common selection bias problem in observation data, and accurately identifying the heterogeneous effects of marketing strategies in different customer segments, providing a statistically rigorous quantitative basis for differentiated marketing decisions.
[0102] In some embodiments of this application, a causal forest model is used to determine the conditional average treatment effect, including: S701. Map the stage transition state labels to the discrete state space of a semi-Markov model.
[0103] In this embodiment, the semi-Markov model (SMM) is a stochastic process model. Its difference from a regular Markov chain lies in the fact that the dwell time of the system before entering the next state does not follow a fixed geometric or exponential distribution, but can follow any probability distribution. The discrete state space refers to a set consisting of a finite number of mutually exclusive and complete sales stages, such as {"potential customer", "intent", "opportunity", "winner", "loser"}.
[0104] In some embodiments of this application, the system traverses the marketing interaction logs, takes the stage transition status label in each record as an observation value, arranges them in chronological order, thereby constructing the customer's state trajectory in the funnel, and defines these states as SMM state nodes.
[0105] S702. Using a semi-Markov model, perform time series analysis on the customer lead attribute vector to determine the dwell time distribution function and state transition probability matrix for each discrete state.
[0106] In this embodiment, the dwell time distribution function describes the probabilistic pattern of the length of time a customer stays in a specific state, typically fitted using a Weibull distribution or a log-normal distribution. The state transition probability matrix describes the probability that a customer will transition from the current state to any other arbitrary state.
[0107] In some embodiments of this application, the system employs Maximum Likelihood Estimation (MLE) to parametrically estimate the model parameters of the SMM by combining customer lead attribute vectors as covariates. Specifically, the system assumes that the transition probability and dwell time parameter are functions of customer attributes, and calculates the specific values of the dwell time distribution function parameters and the state transition probability matrix for a specific attribute vector through regression analysis.
[0108] In a supplementary embodiment of this application, for estimating the dwell time distribution function, the system employs Weibull regression within the Accelerated Failure Time (AFT) framework. Specifically, it is assumed that the dwell time T_i in state i follows a Weibull distribution, and its probability density function is determined by the shape parameter k_i and the scale parameter lambda_i. The system establishes the following log-linear regression relationship to map the customer lead attribute vector X to the scale parameter: ln(T_i)=beta_0+beta_1×X_1+...+beta_m×X_m+sigma×W Where ln represents the natural logarithm, beta is the vector of regression coefficients to be estimated, X_1 to X_m are the dimensional components of the customer lead attribute vector, sigma is the scaling factor, and W is the random error term following an extreme value distribution. This means that the scaling parameter lambda_i is related to the attribute vector X as follows: lambda_i(X)=exp(-(beta×X) / sigma) The system uses a multinomial logistic regression model to model the state transition probability matrix P_ij(X). The probability of transitioning from the current state i to the next state j is expressed as: P_ij(X)=exp(gamma_ij×X) / SUM(exp(gamma_iz×X)) Where gamma_ij is the feature weight vector of the corresponding transition path, exp represents the exponential function, and SUM in the denominator is the summation of all possible next states z in the set S_next. The system uses the maximum likelihood method to jointly optimize the parameters of the two regression models mentioned above, thereby obtaining the personalized residence distribution and transition matrix for a specific customer X.
[0109] S703. Based on the dwell time distribution function and the state transition probability matrix, determine the sample survival probability value of the customer lead attribute vector within the current observation time window.
[0110] In this embodiment, the sample survival probability value refers to the probability that, at the end of a given observation time window, the customer lead has not yet been ultimately converted (e.g., it has neither been completely lost nor ultimately closed, but remains in the intermediate state of the funnel). This corresponds to the survival function value of right-censored data in survival analysis.
[0111] In some embodiments of this application, the system extrapolates the time path of the customer from entering the system to the current moment based on the SMM model. By integrating the dwell time distribution function within the observation window, and combining this with state transition logic, the probability that the sample remains "alive" at the observation cutoff time is calculated, yielding the sample survival probability value. This survival probability value reflects the completeness of the observation data; a higher survival probability means that the final result of the sample is more uncertain (i.e., the greater the possibility of truncation).
[0112] In a supplementary embodiment of this application, the system executes the following specific numerical processing logic for the calculation process in S703: First, the system determines the observed duration T_obs_i of the current sample i, that is, the time span from the moment the sample enters the current state to the data acquisition deadline (or the end of the observation window). Second, the system calls the Weibull distribution parameters for sample i obtained from the regression in step S702: the shape parameter k_i and the scale parameter lambda_i. The system uses the Weibull distribution's survival function formula to calculate the sample survival probability value S_prob_i: S_prob_i=exp(-(T_obs_i / lambda_i)^k_i) Here, exp represents an exponential function with the natural constant e as its base, and the symbol ^ represents exponentiation. The physical meaning of S_prob_i is: assuming the customer perfectly follows historical statistical patterns, the probability that their dwell time in the current state exceeds T_obs_i. Furthermore, if considering state transition probabilities, and the goal is to calculate the survival probability of "no specific positive transition (such as a winning trade)", the system needs to incorporate the state transition probability matrix M. Let p_win_i be the probability that state i directly jumps to the target transition state (winning trade), then the corrected sample survival probability value is: S_final_i=S_prob_i+(1-S_prob_i)×(1-p_win_i) The formula indicates that a sample "survival" (not transformed) includes two situations: one is that the residence time has not ended (S_prob_i), and the other is that the residence time has ended but the sample has been transferred to other states that are not winning orders (the second part of the item).
[0113] S704. Use the sample survival probability value to determine the sample correction weights for the customer lead attribute vector.
[0114] In the embodiments of this application, the sample correction weight is a weighting coefficient used to correct statistical bias caused by data censoring.
[0115] In some embodiments of this application, the system calculates the reciprocal of the sample survival probability value as the basic weight. To avoid weight explosion caused by extremely small probability values, the weights are often modified or truncated. For samples that have undergone final state transformation (non-censored), their weight is assigned as the reciprocal of the sample survival probability value; for censored samples, their weight is set to 0 or adjusted according to the specific algorithm. The physical meaning of this weight is to use the existing complete observation samples to represent those samples with similar characteristics but unfortunately not fully observed, thereby restoring the representativeness of the data.
[0116] S705. Substitute the sample correction weights into the causal forest model to perform a weighted fitting of the relationship between customer conversion deviation value and customer preference score, and obtain the conditional average treatment effect.
[0117] In some embodiments of this application, when training each tree in the causal forest, each sample is no longer treated equally; instead, the sample correction weight is used as an importance factor input into the loss function. During split node selection and leaf node estimation, the model tends to fit samples with larger weights. This means that the model can automatically correct the bias caused by the limited observation period, resulting in overexpression of "fast-transformation" samples and neglect of "slow-transformation" samples, thus obtaining a more unbiased conditional averaging effect after time-series truncation correction.
[0118] As can be seen, the embodiments of this application introduce a semi-Markov model to model the temporal characteristics of the marketing process, calculate the sample survival probability value reflecting the degree of data censoring, and generate the sample survival probability value based on this and inject it into the causal forest model. This effectively solves the sample bias problem caused by the truncation of long-term sales leads under a limited observation window, thereby improving the accuracy and robustness of the model in evaluating the causal effects of marketing strategies in complex, long-chain B2B sales scenarios.
[0119] In some embodiments of this application, customer preference scores for each historical marketing outreach action are determined based on customer lead attribute vectors, including: S801. Obtain the gradient boosting decision tree classification model, wherein each historical marketing outreach action is used as the classification label of the gradient boosting decision tree classification model.
[0120] In this embodiment, the Gradient Boosting Decision Tree (GBDT) classification model is an iterative decision tree ensemble algorithm that continuously constructs new decision trees to fit the residuals of the previous iteration, thereby gradually improving the model's classification accuracy. The classification label refers to the predicted target of the model's output layer. In this scenario, each discrete historical marketing outreach action (e.g., "sending coupons," "phone follow-up," "invitation to try") corresponds to a category label in the model.
[0121] In some embodiments of this application, the system employs high-performance frameworks such as XGBoost to construct the gradient boosting decision tree classification model. Training data is derived from historical marketing logs, where the input feature (X) is a customer lead attribute vector, and the target variable (Y) is the actual marketing action performed on the customer. If multiple mutually exclusive marketing outreach actions exist, a multi-classification model is constructed; if the actions are not mutually exclusive (i.e., the same customer can receive multiple actions simultaneously), a multi-label model or multiple binary classification models are constructed. By learning the feature distribution in historical data, this model can capture the statistical regularity of customers with specific attributes being subjected to specific marketing actions in past manual decision-making or rule-based systems.
[0122] In a supplementary embodiment of this application, to ensure that the generated predicted probability values accurately reflect the true distribution rather than being overwhelmed by the majority class (i.e., "no action" or "0 label"), the system introduces a "positive and negative sample weight balancing" mechanism when constructing the XGBoost classifier. Specifically, the system first counts the number of samples N_pos with a classification label of 1 (accepted action) and the number of samples N_neg with a classification label of 0 (not accepted action) in the training data. When defining the model's objective function, the system sets a scaling factor scale_pos_weight: scale_pos_weight=N_neg / N_pos The system uses weighted binary cross-entropy as the optimized loss function L: L=-SUM[scale_pos_weight×y_i×log(p_i)+(1-y_i)×log(1-p_i)] Where y_i is the true label (0 or 1), and p_i is the model's predicted probability. By introducing this coefficient, the model increases the penalty for the minority class (samples that actually received the marketing action) during the iteration process, forcing the model to "pay attention" to those sparse marketing events, ensuring that the propensity score does not generally collapse into a value range close to 0, thus guaranteeing the effectiveness of the numerical range truncation operation in S803.
[0123] S802. Using a gradient boosting decision tree classification model, classification prediction is performed based on customer lead attribute vectors to obtain the predicted probability value for each classification label.
[0124] In the embodiments of this application, the predicted probability value refers to the confidence level value output by the model for each possible marketing outreach action, and its value is usually between 0 and 1. In the field of causal inference, this probability value corresponds to the propensity score, which represents the conditional probability of accepting a certain treatment (marketing action) given the covariate (customer attributes).
[0125] In some embodiments of this application, the system feeds the customer lead attribute vector to be analyzed as input data into a trained GBDT model. The model output layer is normalized using a Softmax function (for multi-class classification) or a Sigmoid function (for binary classification), thereby outputting a probability vector. For example, the output vector [0.1, 0.7, 0.2] indicates that the customer has a 10% probability of being "emailed," a 70% probability of being "called," and a 20% probability of being "ignored." This step accurately quantifies the degree of preference for assigning a customer to a particular group solely based on their characteristics.
[0126] S803. Perform a numerical interval truncation operation on the predicted probability value to obtain the customer preference score.
[0127] In this embodiment of the application, the numerical range truncation operation refers to the process of forcibly adjusting the probability value that exceeds the preset numerical range to the boundary of the range.
[0128] In some embodiments of this application, the system sets a very small positive number ε (e.g., 0.01 or 0.05) as a truncation threshold. For a predicted probability value p, if p < ε, then p is forced to = ε; if p > 1 - ε, then p is forced to = 1 - ε. The final processed value is the customer preference score. The purpose of this operation is to satisfy the overlap assumption or positive value assumption in causal inference. Without truncation, extreme probability values (close to 0 or 1) would lead to extremely unstable values or variance explosion in subsequent calculations, thereby compromising the robustness of causal effect estimation.
[0129] As can be seen, the embodiments of this application utilize the powerful nonlinear fitting capability of gradient boosting decision trees to accurately extract the distribution pattern of marketing actions from high-dimensional sparse customer features. Furthermore, the numerical truncation technique solves the probability boundary stability problem in causal inference, ensuring that the generated propensity score can not only truly reflect the selective bias of historical data but also meet the strict requirements of numerical stability for subsequent causal effect estimation, thus providing a reliable mathematical basis for eliminating confounding factors.
[0130] In some embodiments of this application, the net improvement in customer conversion rate for each historical marketing outreach action is determined based on the conditional averaging effect, including: S901. Determine the expected value of the conditional average treatment effect and the variance of the effect estimate for each historical marketing outreach action.
[0131] In this embodiment, the expected value of the effect refers to the predicted mean of the conditional average treatment effect output by the causal forest model for a specific customer lead attribute vector. It represents the expected increase in conversion rate statistically significant when the marketing outreach action is implemented compared to not implementing the action. The variance of the effect estimate is a measure of the uncertainty or dispersion of the model's effect estimation result.
[0132] In some embodiments of this application, the system utilizes the infinitesimal jackknife or bootstrap method to estimate the variance of the prediction results of the causal forest. Specifically, each tree in the causal forest model produces an estimate, and the expected effect is a weighted average of all tree estimates; the variance of the effect estimate is calculated based on the differences in the prediction results among these trees and the sample weights. A larger variance value indicates that the model is less confident in its judgment of the causal effect of the customer's action, and may be affected by sample sparsity or feature noise.
[0133] S902. Based on the variance of the effect estimate, determine the statistical error boundary value.
[0134] In the embodiments of this application, the statistical error boundary value refers to a safety margin set to deduct predicted values in order to ensure the robustness of decision-making. It typically corresponds to the lower bound deviation of the confidence interval.
[0135] In some embodiments of this application, the system queries a standard normal distribution table to obtain the corresponding critical value (such as the Z-score, 1.96 or 1.645) based on a preset confidence level (e.g., 95% or 90%). Subsequently, the system multiplies this critical value by the square root of the effect estimation variance to obtain the statistical error boundary value. This step introduces conservative decision-making logic, considering not only "how much improvement can be achieved on average" but also "how much improvement can be achieved in the worst-case scenario."
[0136] S903. Based on the expected value of the effect and the statistical error boundary value, determine the net improvement value of each historical marketing outreach action on the customer conversion rate.
[0137] In some embodiments of this application, the system subtracts the statistical error boundary value from the expected effect value to obtain the net boost value. That is, the lower bound of the confidence interval is used as the final indicator. If the calculated result is negative, it is set to 0 or retained as negative to indicate that the action may have a negative effect. The system uses this net boost value to rank all available marketing actions. Compared to directly using the mean, the Lower Confidence Bound Algorithm (LCB) can effectively penalize strategies with high variance and uncertainty, and prioritize "safe" strategies that have both positive gain and high certainty, thereby achieving a balance between exploration and exploitation and avoiding decision-making risks caused by model overfitting or insufficient data.
[0138] As can be seen, in evaluating the effectiveness of marketing actions, this application's embodiments introduce a statistical error boundary correction mechanism based on variance estimation. By calculating the lower bound of the confidence interval as the net improvement value, it not only considers the expected returns of the strategy but also rigorously quantifies the uncertainty risk of the prediction, enabling the system to automatically filter out those "pseudo-efficient" strategies that seem to have high returns but actually have insufficient samples and low credibility.
[0139] In some embodiments of this application, strategy performance attribution analysis results are generated based on the net improvement value, including: S1001. Based on the net increase value, determine the strategic marginal contribution value of each historical marketing outreach action.
[0140] In this embodiment, the strategy marginal contribution value refers to the expected economic increment that can be brought about by a single execution of a specific marketing outreach action, excluding interference from other factors. This value not only reflects the improvement in conversion rate but also incorporates the potential value of the customer.
[0141] In some embodiments of this application, the system first obtains the estimated lifetime value of the customer lead or the contract amount of the current sales opportunity. Then, the system multiplies the net increase in value by the customer's value to obtain the strategy's marginal contribution value. For example, if an action increases the conversion rate by 5%, and the customer's potential value is 1 million yuan, then the marginal contribution value of that action is 50,000 yuan. This metric intuitively demonstrates the absolute contribution of the action to the revenue target.
[0142] S1002. Determine the strategy return rate indicator based on the strategy's marginal contribution value.
[0143] In this embodiment of the application, the strategy return rate index refers to the ratio of economic output obtained per unit cost input, which is used to measure the economic efficiency of the strategy.
[0144] In some embodiments of this application, the system retrieves the execution cost data of the historical marketing outreach actions. This cost may include fixed costs (such as R&D expenses allocated to developing the strategy) and variable costs (such as API call fees, SMS channel fees, and sales staff hourly costs). The system uses the formula: (Strategy Marginal Contribution Value - Execution Cost) / Execution Cost to calculate the strategy return on investment (ROI). If the result is greater than 0, it indicates that the strategy is profitable; if it is less than 0, it indicates a loss. This indicator makes strategies with different cost levels (such as expensive offline visits versus inexpensive email pushes) comparable.
[0145] S1003. Generate strategy performance attribution analysis results based on strategy marginal contribution value and strategy return rate indicators.
[0146] In this application embodiment, the strategy effectiveness attribution analysis result refers to a multi-dimensional comprehensive evaluation report or data structure that clearly indicates which strategies are "star strategies" with high contribution and high return, and which are "elimination strategies" with low contribution and low return.
[0147] In some embodiments of this application, the system constructs a two-dimensional quadrant matrix (a variant of the Boston Matrix), with the horizontal axis representing the strategy return rate and the vertical axis representing the strategy marginal contribution value. The system maps each action to this matrix and encapsulates this distribution and the specific numerical list as the strategy effectiveness attribution analysis result. Furthermore, this analysis result also includes sensitivity labels for different customer segments, such as indicating that "for customer segment A, the telephone follow-up strategy has the highest marginal contribution value." This result will directly serve as input parameters for subsequent operations optimization models or as a basis for management to adjust marketing budget allocations.
[0148] As can be seen, this application embodiment transforms the technical indicators (conversion rate improvement) of marketing outreach actions into financial indicators (revenue increase and return on investment) by calculating the marginal contribution value of the strategy and the strategy return rate indicator. This dual-dimensional attribution analysis can not only identify the key actions that contribute the most to the total revenue, but also identify the most efficient low-cost leverage actions, providing precise quantitative decision support for enterprises to find the optimal balance between "increasing revenue" and "controlling costs".
[0149] In some embodiments of this application, determining the contribution weight of each marketing stage to the decline in customer conversion rate includes: S1101. Obtain the feature set of customers who have not completed conversion and have fallen out during the marketing stage, and construct a feature subset space for the feature set of the marketing stage.
[0150] In this embodiment, step S1101 aims to construct a combinatorial space for game-theoretic attribution analysis. Unconverted churned customers refer to samples whose final state is marked as "loss" or "churn" within the observation period. The marketing stage feature set refers to a vector consisting of all marketing stages experienced by the customer before churn and their corresponding interaction attributes (such as stage dwell time and number of interactions within the stage).
[0151] In some embodiments of this application, the system employs permutation or Monte Carlo sampling methods to construct a feature subset space. This space contains all possible proper subsets (coalitions) of the feature set for that marketing stage. For example, if a customer has experienced stages A, B, and C, the subset space contains combinations such as {A}, {B}, {A,B}, and {A,C}. Each subset represents a hypothetical, "incomplete" marketing journey scenario.
[0152] S1102. Using the causal forest model, calculate the counterfactual transformation probability prediction value of each feature subset in the feature subset space.
[0153] In some embodiments of this application, the system feeds each feature subset generated in step S1101 as an input variable into a pre-trained causal forest model, while marginalizing or assigning baseline values to other stage features not in that subset, thereby calculating a counterfactual conversion probability prediction value for the marketing stage combination that retains only that subset. This prediction value reflects the theoretical conversion probability of the customer when assuming the absence of certain subsequent or intermediate stages.
[0154] S1103. For the target marketing stage, determine the marginal negative contribution value of the target marketing stage based on the difference between the counterfactual conversion probability prediction values between the first feature subset containing the target marketing stage and the second feature subset not containing the target marketing stage.
[0155] The target marketing stage refers to the specific sales segment whose "disruptive power" is currently being assessed. In some embodiments of this application, the system traverses the feature subset space, searching for all subset pairs that differ only in that "target marketing stage" (i.e., first feature subset = second feature subset ∪ {target marketing stage}). The system calculates the difference in the counterfactual conversion probability predictions for each subset pair. If the predicted conversion probability decreases significantly after adding the target marketing stage, the difference is negative, indicating that the stage has a negative effect. The system performs a weighted average of the differences between all subset pairs (the weights are determined by the number of combinations of subset sizes) to obtain the marginal negative contribution value of the target marketing stage. This calculation process follows the axiomatic definition of the Shapley value, achieving a fair allocation of causal effects.
[0156] S1104. Normalize the marginal negative contribution values of all marketing stages to obtain the contribution weight of each marketing stage to the decline in customer conversion rate.
[0157] In some embodiments of this application, the system extracts marginal negative contribution values (i.e., only focusing on stages that have side effects). The system calculates the absolute values of these negative values and sums them as the denominator. Subsequently, the absolute value of the marginal negative contribution value of each marketing stage is divided by this denominator to obtain a normalized percentage value, which is the contribution weight. The larger this weight value, the more significant it is, indicating that the marketing stage is the "primary responsible stage" for customer churn. For example, the calculation results show that the contribution weight of the "business negotiation stage" is 60%, indicating that even if the churn occurs in a later stage, it means that the negative experience in this stage implicitly led to the final failure.
[0158] As can be seen, the embodiments of this application use the causal forest model as the value function in game theory. By calculating the causal Shapley value, the marginal responsibility of each marketing stage for conversion failure is quantified. It can penetrate the complex nonlinear marketing chain and accurately locate those seemingly normal but actually secretly reducing the conversion probability "hidden killer" stages, providing in-depth attribution insights for optimizing the sales process.
[0159] Secondly, embodiments of this application provide a customer conversion pipeline and funnel model analysis system, which is used to perform the customer conversion pipeline and funnel model analysis method as described in any of the above embodiments.
[0160] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A customer conversion pipeline and funnel model analysis method, characterized in that, The customer conversion pipeline and funnel model analysis method includes: Obtain marketing interaction logs including customer lead attribute vectors, historical marketing outreach actions, and corresponding stage flow status tags; Based on the customer lead attribute vector, determine the customer preference score for each of the historical marketing outreach actions; Based on the customer preference score and the stage transition status label, determine the conditional average processing effect of each of the historical marketing outreach actions on each of the customer lead attribute vectors; Based on the aforementioned conditional averaging effect, determine the net improvement in customer conversion rate for each of the aforementioned historical marketing outreach actions; Based on the net improvement value, generate the strategy effectiveness attribution analysis results; The step of determining the conditional average treatment effect of each historical marketing outreach action on each customer lead attribute vector based on the customer propensity score and the stage transition status label includes: performing regression prediction on the stage transition status label using the customer lead attribute vector to obtain a theoretical customer conversion probability value; determining a customer conversion deviation value based on the theoretical customer conversion probability value; obtaining an action deviation value for the corresponding historical marketing outreach action using the customer propensity score; fitting a causal forest model to the correlation between the customer conversion deviation value and the action deviation value; and determining the conditional average treatment effect using the causal forest model. The step of determining the conditional average treatment effect using the causal forest model includes: mapping the stage transition state labels to the discrete state space of a semi-Markov model; using the semi-Markov model to perform time-series analysis on the customer lead attribute vector to determine the dwell time distribution function and state transition probability matrix for each discrete state; based on the dwell time distribution function and the state transition probability matrix, determining the sample survival probability value of the customer lead attribute vector within the current observation time window; using the sample survival probability value to determine the sample correction weight for the customer lead attribute vector; and substituting the sample correction weight into the causal forest model to perform a weighted fitting of the correlation between the customer conversion deviation value and the customer propensity score to obtain the conditional average treatment effect. The step of determining the customer preference score for each of the historical marketing outreach actions based on the customer lead attribute vector includes: obtaining a gradient boosting decision tree classification model, wherein each of the historical marketing outreach actions is used as a classification label for the gradient boosting decision tree classification model; using the gradient boosting decision tree classification model to perform classification prediction based on the customer lead attribute vector to obtain a predicted probability value for each of the classification labels; and performing a numerical interval truncation operation on the predicted probability value to obtain the customer preference score. The step of determining the net improvement in customer conversion rate for each of the historical marketing outreach actions based on the conditional average treatment effect includes: determining the expected value of the conditional average treatment effect and the variance of the effect estimate for each of the historical marketing outreach actions; determining the statistical error boundary value based on the variance of the effect estimate; and determining the net improvement in customer conversion rate for each of the historical marketing outreach actions based on the expected value of the effect and the statistical error boundary value.
2. The customer conversion pipeline and funnel model analysis method as described in claim 1, characterized in that, The acquisition of marketing interaction logs, including customer lead attribute vectors, historical marketing outreach actions, and corresponding stage transition status tags, includes: Acquire customer conversion history data, which includes at least one of the following: AI agent interaction data, enterprise profile static data, and customer acquisition channel traceability data; Based on the customer conversion history data, a four-dimensional data cube is generated, which includes the lead dimension, marketing stage dimension, time dimension and strategy dimension. The marketing interaction log is obtained by slicing the four-dimensional data cube.
3. The customer conversion pipeline and funnel model analysis method as described in claim 2, characterized in that, After generating the four-dimensional data cube based on the customer conversion history data, the process also includes: Based on the four-dimensional data cube, determine the stage inventory saturation, inflow-outflow balance, average conversion time deviation, and strategy activity of each marketing stage in the preset marketing funnel. The health of the preset marketing funnel is obtained by weighting and summing the stage inventory saturation, the inflow and outflow balance, the average conversion time deviation, and the strategy activity.
4. The customer conversion pipeline and funnel model analysis method as described in claim 3, characterized in that, After obtaining the funnel health status of the preset marketing funnel, the process further includes: Determine whether the health status of the funnel is less than a preset benchmark health threshold; If the funnel health is less than the baseline health threshold, determine the contribution weight of each marketing stage to the decline in customer conversion rate; Based on the contribution weights, the target marketing stage among the multiple marketing stages is determined and designated as the bottleneck stage. For the bottleneck stage, corresponding bottleneck analysis data is generated and output.
5. The customer conversion pipeline and funnel model analysis method as described in claim 1, characterized in that, After generating the strategy effectiveness attribution analysis results based on the net improvement value, the method further includes: Determine the computational cost per instance of AI agent retrieval for each of the aforementioned historical marketing outreach actions; Based on the strategy performance attribution analysis results and the computing power cost of a single AI agent call, an operations research optimization model is constructed. The operations research optimization model takes maximizing the expected total revenue as the objective function and uses the upper limit of the number of concurrent threads of the AI agent and the preset maximum allowed frequency of customer disturbance as constraints. Solving the operations research optimization model yields AI agent task instructions, which include target marketing outreach actions to be performed by the AI agent.
6. The customer conversion pipeline and funnel model analysis method as described in claim 1, characterized in that, The step of generating strategy effectiveness attribution analysis results based on the net improvement value includes: Based on the net increase value, determine the strategic marginal contribution value of each of the historical marketing outreach actions; Based on the marginal contribution value of the strategy, determine the strategy return rate indicator; Based on the marginal contribution value of the strategy and the return rate of the strategy, the attribution analysis results of the strategy effectiveness are generated.
Citation Information
Patent Citations
Model training method and device, effect prediction method and device, medium and equipment
CN119227802A
Marketing strategy acquisition method and device, electronic equipment and storage medium
CN119444297A