Loss customer prediction method and device

By constructing a customer churn model and using time-dependent covariates in historical data, the problem of ignoring the time dimension in the existing technology is solved, and accurate prediction of customer churn time and refinement of customer operations is achieved.

CN119989133APending Publication Date: 2025-05-13INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311596143.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has shortcomings in predicting customer churn time and refining customer operations, especially the life cycle change patterns in the time dimension are ignored, resulting in poor estimation results.

Method used

By constructing a customer churn model, using customer feature data and time-dependent covariate data in historical churn data for training, churn probability distribution data is generated to predict customer churn time, and customer strategy is determined based on the churn contribution value.

Benefits of technology

It realizes accurate prediction of customer churn time, refines customer operations, and improves operational efficiency and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989133A_ABST
    Figure CN119989133A_ABST
Patent Text Reader

Abstract

The invention provides a lost customer prediction method and device, relates to the technical field of artificial intelligence, and can be applied to the technical field of finance or other technical fields. The lost customer prediction method comprises the following steps: inputting customer feature data into a customer loss model to obtain loss probability distribution data; obtaining a loss prediction result according to the loss probability distribution data; wherein the customer churn model is obtained through training according to historical churn data, and the historical churn data comprises historical customer feature data and time dependency covariable data generated according to the historical customer feature data. The customer loss time can be accurately predicted, and customer operation is refined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for predicting lost customers. Background Art

[0002] With the continuous development of mobile applications and tracking technology, banking business has gradually shifted from offline to online. It is of great significance to collect customer behavior information to predict changes in customer life cycle, promote activation in a timely manner, conduct precision marketing, and improve service quality. However, there are still many shortcomings:

[0003] Currently, predictions for customer churn and retention are mostly based on machine learning models, but most models are doing classification, that is, judging whether a user will undergo a life cycle change, and the time factor is often ignored. The law of customer life cycle changes in the time dimension is of great significance for the analysis of customer operational health. Most of the covariates of customer behavior analysis based on survival analysis do not change over time, and are based on linear relationships, which cannot accurately describe the actual scenario, resulting in poor estimation results. In addition, the time for customer life cycle transitions is often long, and the transition of the life cycle may not be observed within a period of time. Compared with most machine learning models, survival analysis adds a time dimension, which greatly limits the sample size under the same computing power and has high requirements for sample quality. Summary of the invention

[0004] The main purpose of the embodiments of the present invention is to provide a method and device for predicting lost customers, so as to accurately predict the time of customer loss and refine customer operations.

[0005] In order to achieve the above object, an embodiment of the present invention provides a method for predicting lost customers, comprising:

[0006] Input customer feature data into the customer churn model to obtain churn probability distribution data;

[0007] Obtaining a churn prediction result according to the churn probability distribution data;

[0008] The customer churn model is trained based on historical churn data, and the historical churn data includes historical customer feature data and time-dependent covariate data generated based on the historical customer feature data.

[0009] In one embodiment, the step of creating a customer churn model includes:

[0010] Inputting the training data in the historical churn data into the initial customer churn model to obtain a churn probability prediction value;

[0011] Determine a loss function according to the actual churn status in the historical churn data and the predicted value of the churn probability;

[0012] Adjusting the initial customer churn model parameters according to the loss function until the loss function converges;

[0013] The initial customer churn model is verified according to the verification data in the historical churn data, and the initial customer churn model is determined as the customer churn model according to the verification result.

[0014] In one embodiment, generating time-dependent covariate data according to the historical customer characteristic data includes:

[0015] Conduct a proportional risk test on the historical customer characteristic data, and determine the customer time-related data based on the test results;

[0016] Processing the client time association data in segments;

[0017] The time-dependent covariate data is determined according to the segmented historical customer feature data and the corresponding time.

[0018] In one embodiment, it also includes:

[0019] Performing censoring processing on the historical customer characteristic data to obtain censored data;

[0020] The historical churn data is constructed based on the censored data, the historical customer characteristic data and the time-dependent covariate data.

[0021] In one embodiment, the lost customer prediction method further includes:

[0022] The actual churn status corresponding to the historical customer feature data is determined according to the censoring status of the censoring processed data.

[0023] In one embodiment, obtaining a churn prediction result according to the churn probability distribution data includes:

[0024] Obtaining customer churn time according to the churn probability distribution data;

[0025] The churn prediction result is obtained according to the customer churn time.

[0026] In one embodiment, it also includes:

[0027] Determine the churn contribution value of each feature in the customer feature data according to the churn probability distribution data;

[0028] A customer strategy is determined according to the churn contribution value.

[0029] The embodiment of the present invention further provides a lost customer prediction device, comprising:

[0030] A churn probability distribution data module is used to input customer feature data into a customer churn model to obtain churn probability distribution data;

[0031] A churn prediction result module, used to obtain a churn prediction result according to the churn probability distribution data;

[0032] The customer churn model is trained based on historical churn data, and the historical churn data includes historical customer feature data and time-dependent covariate data generated based on the historical customer feature data.

[0033] In one embodiment, it also includes:

[0034] A churn probability prediction value module, used to input the training data in the historical churn data into the customer churn initial model to obtain a churn probability prediction value;

[0035] A loss function module, used to determine a loss function according to the actual churn status in the historical churn data and the predicted value of the churn probability;

[0036] An iteration module, used for adjusting the parameters of the initial customer churn model according to the loss function until the loss function converges;

[0037] The customer churn model module is used to verify the initial customer churn model according to the verification data in the historical churn data, and determine the initial customer churn model as the customer churn model according to the verification result.

[0038] In one embodiment, it also includes:

[0039] A customer time association data module is used to perform a proportional risk test on the historical customer characteristic data and determine the customer time association data according to the test result;

[0040] A segment processing module, used for segment processing of the client time association data;

[0041] The time-dependent covariate data module is used to determine the time-dependent covariate data according to the segmented historical customer feature data and the corresponding time.

[0042] In one embodiment, it also includes:

[0043] A deletion module, used for performing deletion processing on the historical customer characteristic data to obtain deleted data;

[0044] The historical churn data module is used to construct the historical churn data according to the censored data, the historical customer feature data and the time-dependent covariate data.

[0045] In one embodiment, it also includes:

[0046] The actual churn status module is used to determine the actual churn status corresponding to the historical customer feature data according to the deletion status of the deletion processing data.

[0047] In one embodiment, the churn prediction result module includes:

[0048] A customer churn time unit, used to obtain the customer churn time according to the churn probability distribution data;

[0049] The churn prediction result unit is used to obtain the churn prediction result according to the customer churn time.

[0050] In one embodiment, it also includes:

[0051] A churn contribution value module, used to determine the churn contribution value of each feature in the customer feature data according to the churn probability distribution data;

[0052] A customer strategy module is used to determine a customer strategy according to the churn contribution value.

[0053] An embodiment of the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the lost customer prediction method when executing the computer program.

[0054] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the lost customer prediction method are implemented.

[0055] An embodiment of the present invention further provides a computer program product, including a computer program / instruction, which implements the steps of the lost customer prediction method when the computer program / instruction is executed by a processor.

[0056] The churn customer prediction method and device of the embodiment of the present invention obtains a customer churn model based on historical customer feature data and corresponding generated time-dependent covariate data training, and inputs the customer feature data into the customer churn model to obtain churn probability distribution data to further obtain churn prediction results, which can accurately predict customer churn time and refine customer operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0058] Figure 1 is a flow chart of a lost customer prediction method according to an embodiment of the present invention;

[0059] Figure 2 is a flow chart of a lost customer prediction method in another embodiment of the present invention;

[0060] Figure 3 is a flow chart of generating time-dependent covariate data in an embodiment of the present invention;

[0061] Figure 4 is a flow chart of constructing historical loss data in an embodiment of the present invention;

[0062] Figure 5 is a flow chart of creating a customer churn model in an embodiment of the present invention;

[0063] Figure 6 is a flow chart of S102 in an embodiment of the present invention;

[0064] Figure 7 is a flow chart of determining a customer strategy in an embodiment of the present invention;

[0065] Figure 8 is a schematic diagram of different censoring events in an embodiment of the present invention;

[0066] Fig. 9 is a structural block diagram of a lost customer prediction device in an embodiment of the present invention;

[0067] Fig.10 A schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. DETAILED DESCRIPTION

[0068] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0069] Those skilled in the art will appreciate that the embodiments of the present invention may be implemented as a system, device, apparatus, method or computer program product. Therefore, the present disclosure may be implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0070] The acquisition, storage, use, and processing of data in the technical solution of the present invention are in compliance with the relevant provisions of national laws and regulations. The user information in the embodiments of this application is obtained through legal and compliant channels, and the acquisition, storage, use, and processing of user information are authorized and agreed by the customer.

[0071] Based on the bank customer behavior data and life cycle transition (taking user retention as an example), the present invention uses deep learning methods to estimate the parameters of the risk proportional survival analysis model under the influence of nonlinear covariates, and uses complex deletion mechanisms to reduce complexity and improve computational efficiency. It predicts the customer's retention probability at each time point, and gives the optimal operation strategy based on the reasons for competition failure, thereby conducting refined customer operations.

[0072] Survival function and risk function are two basic functions in survival analysis. The survival function is represented by S(t) = Pr(T>t), which indicates the possibility of customer loss after time t. The risk function is a risk measure at time t, corresponding to the probability that a customer has not lost before time t but loses at time t, so the risk function λ(t) is defined as:

[0073]

[0074] The customer churn probability under the Cox proportional hazard model is not only a function of time, but also affected by multiple factors X, such as age, gender, deposit balance, financial management balance, etc., which are called covariates (customer characteristic data). The Cox proportional hazard model assumes that the customer churn probability function is in the form of:

[0075] λ(t|x)=λ0(t)·e h(x) ;

[0076] Where λ0(t) is the probability distribution of customer churn when it is not affected by other factors X h(x) = β T x describes the relationship between other factors x, and β is a parameter that reflects the degree of influence of each factor on the probability of customer churn.

[0077] However, the Cox proportional hazard model assumes that the covariates are not affected by time and are linear, which is too simple to fit real-world data sets. Based on this, the present invention broadens the scope of h(x) to nonlinear time dependence:

[0078]

[0079] Among them, θ is a parameter that describes the relationship between the factors X that change over time and their effect on the probability of customer churn.

[0080] For the internal time-dependent covariate, that is, a certain covariate x is affected by time, and the neural network parameter θ is not affected by time, x can be segmented to split the different stage characteristics of a customer into the characteristics of multiple customers. For the external time-dependent covariate, that is, the covariate x is not affected by time, and the neural network parameter θ of the customer churn model is affected by time, the same layered processing can be performed, or the time-dependent covariate z can be introduced, and θ(t) = a + blog(t). Then, in the neural network training, the linear relationship of the i-th intermediate layer covariate can be transformed into Among them, z = log(t)x is a new interaction term. Next, we can build a nonlinear Cox model including z. The survival function and probability density function of customer retention are:

[0081]

[0082]

[0083] Among them, S0(t) and f0(t) are the probability density functions of the customer's retention after time t and the customer's loss at time t, respectively, without being affected by other factors except time.

[0084] The present invention assumes that the final loss of each customer may be displayed in various forms, such as deposit withdrawal, financial redemption, credit card cancellation, debit card cancellation, etc., and such events are recorded as C = {c1, c2, ..., c p}, we can calculate the probability p(t,c|x,z) of each customer churn due to event c at time t.

[0085] The present invention is described in detail below with reference to the accompanying drawings.

[0086] Figure 1 It is a flow chart of the lost customer prediction method in an embodiment of the present invention. Figure 2 FIG. 1 is a flow chart of a method for predicting lost customers in another embodiment of the present invention. Figure 1-Figure 2 As shown, the methods for predicting lost customers include:

[0087] S101: Input customer feature data into a customer churn model to obtain churn probability distribution data.

[0088] S102: Obtaining a churn prediction result according to the churn probability distribution data.

[0089] The customer churn model is trained based on historical churn data, and the historical churn data includes historical customer feature data and time-dependent covariate data generated based on the historical customer feature data.

[0090] Figure 3 FIG. 1 is a flow chart of generating time-dependent covariate data in an embodiment of the present invention. Figure 3As shown, generating time-dependent covariate data according to the historical customer feature data includes:

[0091] S201: Perform a proportional risk test on the historical customer characteristic data, and determine customer time-correlation data according to the test result.

[0092] The customer retention rate of a bank is affected by many factors, and the selection of covariates should include internal and external factors. Internal factors should fully reflect the characteristics of customers, and external factors should fully reflect the characteristics of user behavior. Data should be easy to collect and have strong interpretability, so that the obtained indicators are complete, scientific, and practical. Based on the above requirements, the data used in this invention is based on the change in the probability of customer retention of mobile banking. Each historical churn data (0) =(X (0) ,T,δ,C) is composed of covariates (historical customer feature data)X (0) , churn time T, churn status δ and churn reason C, among which historical customer feature data Includes demographic indicators and customer behavior indicators of q customers.

[0093] Demographic variables: describe some basic information of customers, including gender, age, education, marital status, total deposits, total loans, total financial management, etc.

[0094] Customer behavior indicators: including number of logins, login duration, number of clicks, browsing length, number of likes, number of favorites, purchase frequency, last purchase time, whether to purchase products, whether to open a credit card, etc.

[0095] The churn time T is the time point when the customer churns, the churn status δ is whether the customer churns, and the churn reason C can be classified into account cancellation, financial redemption, deposit withdrawal, etc. according to the business classification.

[0096] Set the expected observation end time T0, the maximum observed customer churn number m (the ideal number of samples to meet the training data volume), the minimum observed customer churn number k (the minimum acceptable number of samples when there is data loss), and take n groups of churn customer samples Η (0) , it should be ensured that at least m customers have been lost at the estimated end time of observation T0.

[0097] Before executing S201, the historical customer feature data needs to be preprocessed:

[0098] 1. Fill or interpolate missing values ​​in historical customer feature data. To ensure the consistency of data changes in the time dimension, samples with serious data missing after a certain point in time can be censored.

[0099] 2. Standardize the continuous features in the historical customer feature data to ensure that they have similar scales and distributions. Common standardization methods include mean normalization and standard deviation normalization.

[0100] 3. Perform one-hot encoding or embedding encoding on the categorical features in the historical customer feature data to convert them into numerical representations that can be used by deep learning models.

[0101] 4. Bin the long-tail or unbalanced data in historical customer feature data.

[0102] S202: Segment processing is performed on the client timing data.

[0103] S203: Determine the time-dependent covariate data according to the segmented historical customer feature data and the corresponding time.

[0104] In the specific implementation, the proportional risk (PH) hypothesis test is performed on the historical customer feature data. When the test result is the customer time-correlated data, the customer time-correlated data is segmented. Then the historical customer feature data X is determined. (0) The corresponding time-dependent covariate data z=log(t)x.

[0105] Figure 4 is a flow chart of constructing historical loss data in an embodiment of the present invention. Figure 4 As shown, the lost customer prediction method also includes:

[0106] S301: Perform censoring processing on the historical customer characteristic data to obtain censored data.

[0107] Figure 8 Schematic diagram of different deletion events in the embodiment of the present invention. Figure 8 As shown, the data required by the present invention consists of four elements: covariate data X, churn cause C, event time T and event indicator δ. If a churn event is observed, the time interval T corresponds to the time from data collection to the occurrence of the churn event, and the event indicator is δ = 1. If a churn event is not observed, the time interval T corresponds to the time from data collection to the last record of customer behavior information (e.g., the end of the trial period), and the event indicator is δ = 0. The customer is considered to be right-censored, which needs to be specially considered in the modeling process.

[0108] Customer life cycle conversion often takes a long time. When collecting data, if a long time period is set, the amount of data will be too large to be trained. If a short period is set, it will easily lead to a small number of lost customer samples. In addition, the quality of customer behavior data is often difficult to control, and data loss or errors often occur. Survival analysis has high requirements for covariate data within the time period. If a large amount of covariate data is lost during the test period, such as data quality problems, the customer sample can be deleted to maximize the use of the full sample. Actively deleting some samples during the test can also improve computing performance.

[0109] Generalized stepwise mixed deletion not only controls the test period within an appropriate range, but also ensures a certain amount of data. Its mechanism is as follows: set the expected end time of the observation period T0, the total number of customer samples n, the expected sample size with complete data m, the minimum acceptable sample size k, and record each loss time as t i:m:n (i=1,2,…,m), and autonomously at time point τ i Censored R i samples until time T*=max{t k:m:n ,min{t m:m:n ,T0}} The experiment stops, so according to t k:m:n , t m:m:n Depending on the size of T0, three different situations may occur:

[0110] Event 1:t 1:m:n ,t 2:m:n ,…,t k:m:n , when T0<t k:m:n <t m:m:n ;

[0111] Event 2:t 1:m:n ,…,t k:m:n ,…,t d:m:n , when t k:m:n <T0<t m:m:n ;

[0112] Event 3:t 1:m:n ,…,t k:m:n ,…,t m:m:n , when t k:m:n <t m:m:n <T0.

[0113] Generalized stepwise mixed censoring can also achieve most of the single-line censoring effects such as type II censoring, stepwise censoring, and adaptive stepwise censoring by adjusting parameters.

[0114] Record the number of missing and the time point of missing (R, T S). When the actual sample number m*≥k, the observation data collection and data organization are completed according to the generalized mixed deletion design scheme; when m*<k, the time interval can be extended to meet the minimum sample number of lost customers, or the sample can be increased and the above steps can be repeated to meet event 2 and event 3.

[0115] S302: Constructing the historical churn data according to the censored data, the historical customer characteristic data and the time-dependent covariate data.

[0116] Among them, the historical loss data after the above processing is Η (1) =(X (1) ,T (1) ,Z,δ,C,R).

[0117] Figure 5 is a flow chart of creating a customer churn model in an embodiment of the present invention. Figure 5 As shown, the steps to create a customer churn model include:

[0118] S401: Input the training data in the historical churn data into the initial customer churn model to obtain a churn probability prediction value.

[0119] In the specific implementation, the historical churn data is randomly divided into r groups for model training and cross-validation. The initial customer churn model is a semi-parametric model, in which the baseline probability density function f0(t) is non-parametric and h θ The goal of the present invention is to train the network to learn the predicted value of churn probability p(t,c|x,z), that is, the estimation of the joint probability distribution of churn time and churn risk event.

[0120] DeepHit is a multi-task network that consists of a shared subnetwork and p churn-causing specific subnetworks. DeepHit first uses a single softmax layer as the output layer of DeepHit to ensure that the network learns the joint distribution of p competing events rather than the marginal distribution of each event. Secondly, a residual connection is maintained from the input covariate to the input of each churn-specific subnetwork. The deep learning neural network is used for C = {c1, c2, …, c p}Shared subnetworks and churn reasons c k The specificity network is composed of L S and The shared subnetwork takes the covariate x and the time factor z as input and outputs a vector f S (x,z), which expresses the common part of p churn events. Each churn-specific subnetwork L C,K Let q=(f S (x,z),x,z) as input and produce output It corresponds to a specific cause c k The summary of these outputs is a joint probability distribution over first hit events and times, and the churn cause-specific subnetworks learn in parallel the marginal distribution of the time to churn for each cause.

[0121] The output of the neural network is an estimate of the probability distribution That is, given a covariate x and a time covariate factor z, the output p c,t is an estimate of the probability that a customer will churn at time t in the form of event c Cumulative incidence function, that is, the probability of a customer losing in the form of event c before time t is

[0122]

[0123] However, the true cumulative distribution function represented by this formula is unknown and can be estimated using the following formula:

[0124]

[0125] S402: Determine a loss function according to the actual churn status in the historical churn data and the churn probability prediction value.

[0126] In one embodiment, the lost customer prediction method further includes: determining the actual lost status corresponding to the historical customer feature data according to the censoring status of the censoring processed data.

[0127] Among them, when the deletion status is deleted, the actual churn status of the customer is 0, indicating that the customer has not been lost; when the deletion status is complete, the actual churn status of the customer is 1, indicating that the customer has been lost.

[0128] In order to train the deep learning network, the present invention designs a loss function L that can perform specific processing on censored data. The loss function is the sum of two terms L=L1+L2, where L1 is the log-likelihood function of the joint distribution of the first hit time and the event; L2 is a loss function that combines specific churn causes, and considering p churn causes, it is modified to generalized mixed censoring. For non-censored customers, it can capture the churn events and the time when they occur; for censored customers, it records the time when the customer is censored, thereby providing information that the customer can be retained until the censoring time.

[0129] According to the setting of the missing mode, the log-likelihood function L1 of the joint distribution of hit time and event without constant term is as follows:

[0130]

[0131] Among them, δ iis the actual churn status, which is 1 when the customer finally churns, and 0 when no customer churn is observed or the customer is deleted. m* is the actual sample size in the three cases. is the estimated value of customer churn probability (predicted value of churn probability) obtained by neural network training, is the customer churn distribution function value obtained based on the predicted value of churn probability.

[0132] Since the churn reason and churn probability are independent of each other, the present invention uses a ranking loss function that conforms to the coordination concept. The risk of customers who churn at time t should be higher than that of customers who remain at time t. c,i,j =1(c i =c,t i <t j ) is used by customers at different time points (t i ,t j ) experience event c, when t i <t j When , it is 1, otherwise it is 0, then the loss function part L2 that combines the specific loss cause is:

[0133]

[0134] Among them, α c is a hyperparameter, which can also be directly specified by business personnel, representing the loss preference for different churn events c. c When it is a constant, it means that the risk ranking loss of each event is the same. G(x,y) is a convex loss function. σ is a hyperparameter, and its appropriate value can be selected through multiple model training. Incorporating L2 into the total loss function will penalize the incorrect sorting of each churn cause event, so minimizing the total loss will encourage the correct sorting, thereby obtaining the cause of each customer churn.

[0135] S403: Adjusting the initial customer churn model parameters according to the loss function until the loss function converges.

[0136] S404: verifying the initial customer churn model according to the verification data in the historical churn data, and determining the initial customer churn model as the customer churn model according to the verification result.

[0137] In the specific implementation, cross-validation is adopted, and one of the r groups of historical churn data is used as validation data, and the others are used as training data. The validation data is brought into the model obtained by the training data to calculate the error. After the validation is completed, the next group of data is selected as validation data in turn until r groups of models and corresponding validation results are obtained. The model with the best generalization ability in the r groups is used as the final customer churn model.

[0138] Figure 6 is a flow chart of S102 in an embodiment of the present invention. Figure 6 As shown, S102 includes:

[0139] S501: Obtaining customer churn time according to the churn probability distribution data.

[0140] In the specific implementation, the corresponding customer feature data of each customer is brought into the model, and the distribution of the customer's churn time due to event c can be obtained. According to the churn probability distribution data p c,t =0 to predict the customer's churn time.

[0141] S502: Obtain the churn prediction result according to the customer churn time.

[0142] In one embodiment, a warning line can be set, and when the customer churn time is lower than the warning value, the operation personnel are notified to take corresponding measures according to the cause of churn. For example, when it is predicted that a customer will churn due to event c in 30 days, the churn contribution value can be analyzed, and an operation strategy can be formulated according to the above method to postpone the customer churn time.

[0143] Figure 7 FIG. 1 is a flow chart of determining a customer strategy in an embodiment of the present invention. Figure 7 As shown, the lost customer prediction method also includes:

[0144] S601: Determine the churn contribution value of each feature in the customer feature data according to the churn probability distribution data.

[0145] S602: Determine a customer strategy according to the churn contribution value.

[0146] The DeepHit model used in the present invention is a black box model. The relationship between customer feature data X and the probability of customer churn is nonlinear, and the degree of influence of customer feature data on customer churn is exactly what the operation personnel are concerned about. Therefore, after the training is completed, the present invention uses the SHAP value (churn contribution value) to calculate the marginal contribution of the feature to the model output to assist in operational decision-making. The SHAP value after model analysis is additive.

[0147]

[0148] in, is the i-th sample due to event c i SHAP value of the jth covariate at churn, is the loss probability distribution data, is the probability density SHAP benchmark value.

[0149] Thus, we can get the influence of each customer feature data on the probability of each customer churn due to event c. By averaging the churn contribution values ​​of individual customer feature data under event c for all customers, we can get the influence of each customer feature data on the probability of churn due to event c. When the churn contribution value is positive, the customer characteristic data may lead to an increase in the probability of customer churn. When the churn contribution value is negative, the customer characteristic data may lead to a decrease in the probability of customer churn. The larger the absolute value of the churn contribution value, the greater the impact.

[0150] The present invention can formulate targeted operation strategies based on the effects of different covariates described by the churn contribution value on the churn of different reasons c. For example, the login time of mobile banking is the jth covariate, and its churn contribution value f(x ij )=-0.2, indicating that the increase in the login time of mobile banking will reduce the probability of customer i churn. If the average churn contribution value of the login time of all customers’ mobile banking is calculated, f c (x j )=-0.1, which means that the increase of mobile banking login time will generally reduce the probability of customer churn. It is recommended to formulate corresponding operation strategies to increase the customer mobile banking login time to improve customer retention.

[0151] The present invention also analyzes the impact of each characteristic data on different types of customers, explores the characteristics of high-quality customers with late churn and long life cycle, allocates marketing resources, and optimizes customer structure. For example, if it is found that under the same other conditions, customers with larger deposit balances are predicted to eventually churn later than customers with smaller deposit balances due to account closure, then customers with larger deposit balances are considered to be high-quality customers, and a multi-party platform attraction method can be adopted to explore more high-quality customer groups that may have this characteristic, thereby improving the overall retention time of customers in the bank and creating more value.

[0152] In summary, the lost customer prediction method of the embodiment of the present invention constructs a nonlinear time-dependent proportional risk model under generalized stepwise mixed censoring, designs a loss function using its likelihood function and competitive churn risk, and uses the deep learning network DeepHit for training, and finally explains the impact of the feature data through the SHAP value. Compared with right censoring, generalized stepwise mixed censoring allows data segmentation, low-quality data exit and other operations during the test cycle, while retaining sample information to the greatest extent, making model training more flexible and efficient. Ultimately, the customer's retention probability at each time point, the churn risk of each method, and the degree of influence of the feature data are obtained, which can effectively evaluate the health of customer operations and make targeted recovery measures for customers who may churn, effectively improve operational efficiency, and enhance service quality.

[0153] Based on the same inventive concept, an embodiment of the present invention further provides a lost customer prediction device. Since the principle of solving the problem by the device is similar to that of the lost customer prediction method, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0154] Fig. 9 is a structural block diagram of a lost customer prediction device in an embodiment of the present invention. Fig. 9 As shown, the lost customer prediction device includes:

[0155] A churn probability distribution data module is used to input customer feature data into a customer churn model to obtain churn probability distribution data;

[0156] A churn prediction result module, used to obtain a churn prediction result according to the churn probability distribution data;

[0157] The customer churn model is trained based on historical churn data, and the historical churn data includes historical customer feature data and time-dependent covariate data generated based on the historical customer feature data.

[0158] In one embodiment, it also includes:

[0159] A churn probability prediction value module, used to input the training data in the historical churn data into the customer churn initial model to obtain a churn probability prediction value;

[0160] A loss function module, used to determine a loss function according to the actual churn status in the historical churn data and the predicted value of the churn probability;

[0161] An iteration module, used for adjusting the parameters of the initial customer churn model according to the loss function until the loss function converges;

[0162] The customer churn model module is used to verify the initial customer churn model according to the verification data in the historical churn data, and determine the initial customer churn model as the customer churn model according to the verification result.

[0163] In one embodiment, it also includes:

[0164] A customer time association data module is used to perform a proportional risk test on the historical customer characteristic data and determine the customer time association data according to the test result;

[0165] A segment processing module, used for segment processing of the client time association data;

[0166] The time-dependent covariate data module is used to determine the time-dependent covariate data according to the segmented historical customer feature data and the corresponding time.

[0167] In one embodiment, it also includes:

[0168] A deletion module, used for performing deletion processing on the historical customer characteristic data to obtain deleted data;

[0169] The historical churn data module is used to construct the historical churn data according to the censored data, the historical customer feature data and the time-dependent covariate data.

[0170] In one embodiment, it also includes:

[0171] The actual churn status module is used to determine the actual churn status corresponding to the historical customer feature data according to the deletion status of the deletion processing data.

[0172] In one embodiment, the churn prediction result module includes:

[0173] A customer churn time unit, used to obtain the customer churn time according to the churn probability distribution data;

[0174] The churn prediction result unit is used to obtain the churn prediction result according to the customer churn time.

[0175] In one embodiment, it also includes:

[0176] A churn contribution value module, used to determine the churn contribution value of each feature in the customer feature data according to the churn probability distribution data;

[0177] A customer strategy module is used to determine a customer strategy according to the churn contribution value.

[0178] In summary, the churn customer prediction device of the embodiment of the present invention obtains a customer churn model based on historical customer feature data and correspondingly generated time-dependent covariate data training, and inputs the customer feature data into the customer churn model to obtain churn probability distribution data, so as to further obtain churn prediction results, which can accurately predict customer churn time and refine customer operations.

[0179] Fig.10 FIG. 9 is a schematic block diagram of the system structure of the electronic device 9600 according to an embodiment of the present application. Fig.10 As shown, the electronic device 9600 may include a central processor 9100 and a memory 9140; the memory 9140 is coupled to the central processor 9100. It is worth noting that Fig.10 is exemplary; other types of structures may also be used to supplement or replace this structure to implement telecommunication functions or other functions.

[0180] In one embodiment, the lost customer prediction method function may be integrated into the central processing unit 9100. The central processing unit 9100 may be configured to perform the following control:

[0181] Input customer feature data into the customer churn model to obtain churn probability distribution data;

[0182] Obtaining a churn prediction result according to the churn probability distribution data;

[0183] The customer churn model is trained based on historical churn data, and the historical churn data includes historical customer feature data and time-dependent covariate data generated based on the historical customer feature data.

[0184] From the above description, it can be seen that the churn customer prediction method provided in the present application obtains a customer churn model based on historical customer feature data and correspondingly generated time-dependent covariate data training, and inputs the customer feature data into the customer churn model to obtain churn probability distribution data, so as to further obtain churn prediction results, which can accurately predict customer churn time and refine customer operations.

[0185] In another embodiment, the churned customer prediction device may be configured separately from the central processor 9100. For example, the churned customer prediction device may be configured as a chip connected to the central processor 9100, and the functions of the churned customer prediction method may be implemented under the control of the central processor.

[0186] like Fig.10 As shown, the electronic device 9600 may also include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily have to include Fig.10 In addition, the electronic device 9600 may also include Fig.10 For components not shown, reference may be made to the prior art.

[0187] like Fig.10 As shown, the central processing unit 9100 is sometimes also referred to as a controller or an operation control, and may include a microprocessor or other processor device and / or logic device. The central processing unit 9100 receives input and controls the operation of various components of the electronic device 9600.

[0188] The memory 9140 may be, for example, one or more of a cache, a flash memory, a hard drive, a removable medium, a volatile memory, a non-volatile memory or other suitable devices. The above-mentioned information related to the failure may be stored, and a program for executing the relevant information may also be stored. The CPU 9100 may execute the program stored in the memory 9140 to implement information storage or processing, etc.

[0189] The input unit 9120 provides input to the central processing unit 9100. The input unit 9120 is, for example, a key or a touch input device. The power supply 9170 is used to provide power to the electronic device 9600. The display 9160 is used to display display objects such as images and texts. The display may be, for example, an LCD display, but is not limited thereto.

[0190] The memory 9140 may be a solid-state memory, such as a read-only memory (ROM), a random access memory (RAM), a SIM card, etc. It may also be a memory that saves information even when the power is off, can be selectively erased, and is provided with more data, examples of which are sometimes referred to as EPROMs, etc. The memory 9140 may also be some other type of device. The memory 9140 includes a buffer 9141 (sometimes referred to as a buffer memory). The memory 9140 may include an application / function storage unit 9142, which is used to store application programs and function programs or processes for executing the operation of the electronic device 9600 through the central processor 9100.

[0191] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for communication functions of the electronic device and / or for executing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0192] The communication module 9110 is a transmitter / receiver 9110 that sends and receives signals via an antenna 9111. The communication module (transmitter / receiver) 9110 is coupled to the central processor 9100 to provide input signals and receive output signals, which may be the same as the case of a conventional mobile communication terminal.

[0193] Based on different communication technologies, multiple communication modules 9110 may be provided in the same electronic device, such as a cellular network module, a Bluetooth module and / or a wireless LAN module, etc. The communication module (transmitter / receiver) 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide an audio output via the speaker 9131 and receive an audio input from the microphone 9132, thereby realizing a common telecommunication function. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. In addition, the audio processor 9130 is also coupled to the central processor 9100, so that recording can be performed on the local machine through the microphone 9132, and the sound stored on the local machine can be played through the speaker 9131.

[0194] The embodiment of the present invention also provides a computer-readable storage medium capable of implementing all the steps of the lost customer prediction method in the above embodiment, where the execution subject is a server or a client. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, all the steps of the lost customer prediction method in the above embodiment are implemented. For example, when the processor executes the computer program, the following steps are implemented:

[0195] Input customer feature data into the customer churn model to obtain churn probability distribution data;

[0196] Obtaining a churn prediction result according to the churn probability distribution data;

[0197] The customer churn model is trained based on historical churn data, and the historical churn data includes historical customer feature data and time-dependent covariate data generated based on the historical customer feature data.

[0198] In summary, the computer-readable storage medium of the embodiment of the present invention obtains a customer churn model based on historical customer feature data and correspondingly generated time-dependent covariate data, and inputs the customer feature data into the customer churn model to obtain churn probability distribution data, so as to further obtain churn prediction results, which can accurately predict customer churn time and refine customer operations.

[0199] The embodiment of the present invention also provides a computer program product capable of implementing all the steps of the lost customer prediction method in the above embodiment, where the execution subject is a server or a client. The computer program product includes a computer program / instruction. When the computer program / instruction is executed by a processor, all the steps of the lost customer prediction method in the above embodiment are implemented. For example, when the processor executes the computer program, the following steps are implemented:

[0200] Input customer feature data into the customer churn model to obtain churn probability distribution data;

[0201] Obtaining a churn prediction result according to the churn probability distribution data;

[0202] The customer churn model is trained based on historical churn data, and the historical churn data includes historical customer feature data and time-dependent covariate data generated based on the historical customer feature data.

[0203] In summary, the computer program product of the embodiment of the present invention obtains a customer churn model based on historical customer feature data and correspondingly generated time-dependent covariate data, and inputs the customer feature data into the customer churn model to obtain churn probability distribution data, so as to further obtain churn prediction results, which can accurately predict customer churn time and refine customer operations.

[0204] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the hardware + program embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0205] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0206] Although the present application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative labor. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When the actual device or client product is executed, it can be executed in the order of the method shown in the embodiments or the drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment).

[0207] Although the present specification embodiment provides the method operation steps as described in the embodiment or flow chart, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiment is only one way in the order of execution of many steps, and does not represent a unique execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel (such as a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment) according to the method shown in the embodiment or the accompanying drawings. The term "include", "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements not only includes those elements, but also includes other elements not clearly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements.

[0208] For the convenience of description, the above devices are described in various modules according to their functions. Of course, when implementing the embodiments of this specification, the functions of each module can be implemented in the same or more software and / or hardware, or the module implementing the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0209] Those skilled in the art also know that, in addition to implementing the controller in a purely computer-readable program code, the controller can be made to implement the same function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered as a hardware component, and the devices for implementing various functions included therein can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules for implementing the method and structures within the hardware component.

[0210] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0211] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0212] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0213] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0214] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0215] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0216] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, the embodiments of this specification may take the form of complete hardware embodiments, complete software embodiments or embodiments combining software and hardware. Moreover, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0217] The various embodiments in this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0218] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiments of this specification. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0219] The above is only an example of the embodiment of the present specification and is not intended to limit the embodiment of the present specification. For those skilled in the art, the embodiment of the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiment of the present specification shall be included in the scope of the claims of the embodiment of the present specification.

Claims

1. A method for predicting lost customers, characterized in that: include: Input customer feature data into the customer churn model to obtain churn probability distribution data; Obtaining a churn prediction result according to the churn probability distribution data; The customer churn model is trained based on historical churn data, and the historical churn data includes historical customer feature data and time-dependent covariate data generated based on the historical customer feature data.

2. The method for predicting lost customers according to claim 1, characterized in that: The steps to create a customer churn model include: Inputting the training data in the historical churn data into the initial customer churn model to obtain a churn probability prediction value; Determine a loss function according to the actual churn status in the historical churn data and the predicted value of the churn probability; Adjusting the initial customer churn model parameters according to the loss function until the loss function converges; The initial customer churn model is verified according to the verification data in the historical churn data, and the initial customer churn model is determined as the customer churn model according to the verification result.

3. The method for predicting lost customers according to claim 1, characterized in that: Generating time-dependent covariate data according to the historical customer characteristic data includes: Conduct a proportional risk test on the historical customer characteristic data, and determine the customer time-related data based on the test results; Processing the client time association data in segments; The time-dependent covariate data is determined according to the segmented historical customer feature data and the corresponding time.

4. The method for predicting lost customers according to claim 2, characterized in that: Also includes: Performing censoring processing on the historical customer characteristic data to obtain censored data; The historical churn data is constructed based on the censored data, the historical customer characteristic data and the time-dependent covariate data.

5. The method for predicting lost customers according to claim 4, characterized in that: The lost customer prediction method further includes: The actual churn status corresponding to the historical customer feature data is determined according to the censoring status of the censoring processed data.

6. The method for predicting lost customers according to claim 1, characterized in that: The churn prediction results obtained according to the churn probability distribution data include: Obtaining customer churn time according to the churn probability distribution data; The churn prediction result is obtained according to the customer churn time.

7. The method for predicting lost customers according to claim 1, characterized in that: Also includes: Determine the churn contribution value of each feature in the customer feature data according to the churn probability distribution data; A customer strategy is determined according to the churn contribution value.

8. A lost customer prediction device, characterized in that: include: A churn probability distribution data module is used to input customer feature data into a customer churn model to obtain churn probability distribution data; A churn prediction result module, used to obtain a churn prediction result according to the churn probability distribution data; The customer churn model is trained based on historical churn data, and the historical churn data includes historical customer feature data and time-dependent covariate data generated based on the historical customer feature data.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the lost customer prediction method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the lost customer prediction method according to any one of claims 1 to 7 are implemented.