Telecom Customer Churn Prediction Method, Device, Equipment and Storage Medium Based on XGBoost Algorithm

Through the telecom customer churn prediction method based on the XGBoost algorithm, combined with data exploration, preprocessing and feature engineering, the problem that telecom service providers cannot predict the high churn risk is solved, and the advance prediction of customer churn risk and the push of retention solutions is realized, which reduces the customer churn rate and improves service quality.

CN117522461BActive Publication Date: 2025-05-27CHINA COMM SERVICE APPL & SOLUTION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311656170.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-05-27
Estimated Expiration
2043-12-05

AI Technical Summary

Technical Problem

Existing telecommunications service providers cannot effectively predict which customers face high churn risk, resulting in a high customer churn rate.

Method used

The telecom customer churn prediction method based on the XGBoost algorithm is adopted. By obtaining the customer sample data and customer portrait tags of the sample customers, data exploration analysis, data preprocessing, feature engineering analysis and model training are carried out to obtain the lost telecom customer portrait model. Then import the data of the target customer into the model, output the confidence of the target customer divided into different customer portrait labels, and push the corresponding retention plan based on the confidence.

Benefits of technology

It realizes the advance prediction of the risk of churn of telecommunications customers, reduces the customer churn rate by pushing preset retention plans, provides high service quality, and facilitates practical application and promotion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117522461B_ABST
    Figure CN117522461B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, equipment and storage medium for predicting telecom customer churn based on the XGBoost algorithm, which relates to the field of telecom operation technologies. The method first performs exploratory data analysis processing, data preprocessing, feature engineering analysis processing and model training of a regression classification model based on the XGBoost algorithm on the customer sample data and customer portrait labels of multiple sample customers, and a churn telecom customer portrait model can be obtained. Then, the data to be measured of the target customer is imported into the portrait model, and the confidence levels of dividing the target customer into different customer portrait labels can be obtained. Finally, when the confidence level of dividing the target customer into a certain customer portrait label exceeds the preset confidence threshold, a preset retention plan corresponding to the certain customer portrait label is pushed to the current telecom service provider of the target customer. In this way, customers who are about to churn can be predicted in advance, and by pushing the preset retention plan, the customer churn rate can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of telecommunications operation, and particularly relates to a method, device, equipment and storage medium for predicting telecom customer churn based on the XGBoost algorithm. Background Art

[0002] In the era of big data, due to the excessive and miscellaneous information, it is very difficult for enterprise managers or individuals to accurately obtain the information they want. Especially, enterprise managers urgently need to extract important information from such a huge amount of information for enterprise operation decision-making and quickly respond to market changes. Therefore, a series of machine learning algorithms have emerged. The machine learning XGBoost algorithm model has the characteristics of high efficiency, flexibility and portability, and has been widely used in the fields of data mining, recommendation systems, etc.

[0003] In recent years, in the telecommunications industry, with the development of business homogenization, the market competition has become white-hot. Telecom customers can choose from various service providers and switch from one service provider to another. Therefore, the annual churn rate of telecom services is relatively high. In order to reduce customer churn, telecom companies need to timely predict which customers are at high risk of churn. Only in this way can enterprises timely implement effective customer retention programs. Only by solving the problem of customer churn can enterprises maintain their market position and develop and grow.

[0004] Therefore, how to combine portrait modeling and the XGBoost algorithm to achieve telecom customer churn prediction in order to propose a recommended solution for telecom products or operations is an urgent research topic for those skilled in the art. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, device, computer equipment and computer-readable storage medium for predicting telecom customer churn based on the XGBoost algorithm, so as to solve the problem that existing telecom service providers cannot predict which customers are at high risk of churn.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] In a first aspect, a method for predicting telecom customer churn based on the XGBoost algorithm is provided, including:

[0008] Obtain customer sample data and customer portrait labels of multiple sample customers, where the sample customers refer to the telecom customers who have churned, and the customer sample data includes static information of the corresponding customers and dynamic information in the most recent unit period before churn;

[0009] Perform exploratory data analysis processing on the customer sample data and the customer portrait labels of the multiple sample customers to obtain an exploratory data analysis result;

[0010] According to the data exploratory analysis results, data preprocessing is performed on the customer sample data of the multiple sample customers to obtain multiple standardized customer sample data that correspond one-to-one to multiple customer portrait labels and have clean and continuous data characteristics;

[0011] Performing feature engineering analysis on the plurality of standardized customer sample data to select and obtain a plurality of first data feature sets that correspond one-to-one to the plurality of customer portrait labels and are suitable for model input;

[0012] The plurality of first data feature sets are used as model input items, and the plurality of customer portrait labels are used as model output items, and are imported into a regression classification model based on an XGBoost algorithm for model training to obtain a lost telecommunications customer portrait model;

[0013] Acquire the static information of the target customer and the dynamic information in the most recent unit period;

[0014] Converting the static information and the dynamic information of the target customer into standardized customer data with clean and continuous characteristics;

[0015] Selecting a second data feature set from the standardized customer data, the data type of which is consistent with the first data feature set;

[0016] Importing the second data feature set into the lost telecommunications customer portrait model, and outputting the confidence of classifying the target customers into different customer portrait labels;

[0017] When the confidence level of classifying the target customer as a certain customer portrait label exceeds a preset confidence threshold, a preset retention plan corresponding to the certain customer portrait label is pushed to the current telecommunications service provider of the target customer.

[0018] Based on the above invention content, a new solution for realizing telecommunication customer churn prediction by combining portrait modeling and XGBoost algorithm is provided, that is, firstly, customer sample data and customer portrait labels of multiple sample customers are subjected to data exploratory analysis and processing, data preprocessing, feature engineering analysis and processing, and model training of a regression classification model based on the XGBoost algorithm in sequence, so as to obtain a portrait model of churned telecommunication customers, and then the target customer's test data is imported into the portrait model, so as to obtain the confidence of dividing the target customer into different customer portrait labels, and finally, when the confidence of dividing the target customer into a certain customer portrait label exceeds a preset confidence threshold, a preset retention plan corresponding to the certain customer portrait label is pushed to the target customer's current telecommunication service provider, so as to predict some customers who are about to churn in advance, and by pushing the preset retention plan, the customer churn rate can be reduced, high service quality can be provided, and practical application and promotion can be facilitated.

[0019] In a possible design, exploratory data analysis processing is performed on the customer sample data of the multiple sample customers and the customer portrait tags, and an exploratory data analysis result is obtained, including:

[0020] Perform an overview data profile viewing process on the customer sample data of the multiple sample customers and the customer portrait tags to determine whether there is abnormal feature data that needs to be deleted.

[0021] In a possible design, exploratory data analysis processing is performed on the customer sample data of the multiple sample customers and the customer portrait tags, and an exploratory data analysis result is obtained, including:

[0022] Perform a label distribution viewing process on the customer portrait tags of the multiple sample customers to determine whether there is customer sample data and customer portrait tags corresponding to outlier customers that need to be deleted.

[0023] In a possible design, perform a feature distribution viewing process on the customer sample data of the multiple sample customers to determine whether various situations are covered.

[0024] In a possible design, perform a correlation viewing process between features of the customer sample data of the multiple sample customers to determine whether there is redundant feature data that needs to be deleted;

[0025] And / or, perform a correlation viewing process between features and labels on the customer sample data of the multiple sample customers and the customer portrait tags to determine whether there is feature data with low correlation that needs to be deleted.

[0026] In a possible design, the data preprocessing includes removing features that are not important for data mining, filling in missing data values, processing and filtering abnormal data, eliminating redundant data through data deduplication, uniformly converting data with inconsistent units, reducing the dimension of high-dimensional features, numericalizing categorical features, and / or making the data dimensionless.

[0027] In a possible design, the feature engineering analysis processing includes feature selection through filtering methods, wrapper methods, and / or embedding methods.

[0028] In a second aspect, a telecommunications customer churn prediction device based on the XGBoost algorithm is provided, including a customer data acquisition module, an exploratory data analysis module, a data preprocessing module, a feature engineering analysis module, a model training module, a customer data conversion module, a feature data selection module, a model application module, and a retention plan recommendation module;

[0029] The customer data acquisition module is used to acquire customer sample data and customer portrait labels of multiple sample customers, where the sample customers refer to the lost telecom customers, and the customer sample data includes the static information of the corresponding customers and the dynamic information in the most recent unit period before the loss;

[0030] The data exploratory analysis module is communicatively connected to the customer data acquisition module and is used to perform data exploratory analysis processing on the customer sample data and the customer portrait labels of the multiple sample customers to obtain a data exploratory analysis result;

[0031] The data preprocessing module is communicatively connected to the data exploratory analysis module and is used to perform data preprocessing on the customer sample data of the multiple sample customers according to the data exploratory analysis result to obtain multiple standardized customer sample data that are in one-to-one correspondence with multiple customer portrait labels and have the characteristics of clean and continuous data;

[0032] The feature engineering analysis module is communicatively connected to the data preprocessing module and is used to perform feature engineering analysis processing on the multiple standardized customer sample data to select multiple first data feature sets that are in one-to-one correspondence with the multiple customer portrait labels and are suitable for model input;

[0033] The model training module is communicatively connected to the feature engineering analysis module and is used to use the multiple first data feature sets as model input items and the multiple customer portrait labels as model output items, and import them into a regression classification model based on the XGBoost algorithm for model training to obtain a lost telecom customer portrait model;

[0034] The customer data acquisition module is also used to acquire the static information of the target customer and the dynamic information in the most recent unit period;

[0035] The customer data conversion module is communicatively connected to the customer data acquisition module and is used to convert the static information and the dynamic information of the target customer into standardized customer data with the characteristics of clean and continuous data;

[0036] The feature data selection module is communicatively connected to the customer data conversion module and is used to select a second data feature set from the standardized customer data whose data type is consistent with the first data feature set;

[0037] The model application module is communicatively connected to the model training module and the feature data selection module respectively, and is used to import the second data feature set into the lost telecom customer portrait model and output the confidence levels for classifying the target customer into different customer portrait labels;

[0038] The retention plan recommendation module is communicatively connected to the model application module and is configured to, when the confidence level of classifying the target customer into a certain customer portrait label exceeds a preset confidence threshold, push a preset retention plan corresponding to the certain customer portrait label to the current telecommunications service provider of the target customer.

[0039] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the telecommunications customer churn prediction method as described in the first aspect or any possible design in the first aspect.

[0040] In a fourth aspect, the present invention provides a computer-readable storage medium, on which instructions are stored. When the instructions are run on a computer, the telecommunications customer churn prediction method as described in the first aspect or any possible design in the first aspect is executed.

[0041] In a fifth aspect, the present invention provides a computer program product containing instructions. When the instructions are run on a computer, the computer is made to execute the telecommunications customer churn prediction method as described in the first aspect or any possible design in the first aspect.

[0042] Beneficial effects of the above solution:

[0043] (1) The present invention creatively provides a new solution for realizing telecommunications customer churn prediction by combining portrait modeling and the XGBoost algorithm. That is, first, exploratory data analysis processing, data preprocessing, feature engineering analysis processing, and model training of a regression classification model based on the XGBoost algorithm are sequentially performed on the customer sample data and customer portrait labels of multiple sample customers to obtain a churn telecommunications customer portrait model. Then, the test data of the target customer is imported into this portrait model to obtain the confidence levels of classifying the target customer into different customer portrait labels. Finally, when the confidence level of classifying the target customer into a certain customer portrait label exceeds a preset confidence threshold, a preset retention plan corresponding to the certain customer portrait label is pushed to the current telecommunications service provider of the target customer. In this way, some customers who are about to churn can be predicted in advance, and by pushing the preset retention plan, the customer churn rate can be reduced, the service quality can be improved, and it is convenient for practical application and promotion. Description of the Drawings

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0045] Figure 1 It is a schematic flowchart of the telecom customer churn prediction method based on the XGBoost algorithm provided by the embodiments of this application.

[0046] Figure 2 It is a schematic structural diagram of the telecom customer churn prediction device based on the XGBoost algorithm provided by the embodiments of this application.

[0047] Figure 3 It is a schematic structural diagram of the computer device provided by the embodiments of this application. Detailed implementation manners

[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the accompanying drawings and the description of the embodiments or the prior art. Obviously, the following description of the accompanying drawing structures is only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings. It should be noted here that the description of these embodiment modes is used to help understand the present invention, but does not constitute a limitation to the present invention.

[0049] It should be understood that although terms such as first and second etc. may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object can be called the second object, and similarly, the second object can be called the first object, without departing from the scope of the exemplary embodiments of the present invention.

[0050] It should be understood that for the term "and / or" that may appear in this article, it is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, B exists alone, or A and B exist simultaneously, etc.; another example, A, B and / or C can represent any one of A, B and C or any combination of them; for the term " / and" that may appear in this article, it is a description of another association object relationship, indicating that two relationships can exist. For example, A / and B can represent: A exists alone or A and B exist simultaneously, etc.; in addition, for the character " / " that may appear in this article, generally, the front and rear associated objects are in an "or" relationship.

[0051] Example:

[0052] As Figure 1 shown, the telecom customer churn prediction method provided in the first aspect of this embodiment and based on the XGBoost algorithm can be, but is not limited to, executed by a computer device with certain computing resources, such as a platform server, a personal computer (Personal Computer, PC, referring to a multi-purpose computer suitable for personal use in terms of size, price, and performance; desktop computers, laptops, small laptops, tablets, and ultrabooks all belong to personal computers), a smart phone, a personal digital assistant (Personal Digital Assistant, PDA), or a wearable device, etc. As Figure 1 shown, the telecom customer churn prediction method can be, but is not limited to, including the following steps S1 to S10.

[0053] S1. Obtain the customer sample data and customer portrait labels of multiple sample customers. Among them, the sample customers refer to the lost telecom customers, and the customer sample data includes, but is not limited to, the static information of the corresponding customers and the dynamic information in the most recent unit period before churn.

[0054] In step S1, the customer sample data of the multiple sample customers is used to participate in the model training work in the form of samples. Generally, it is required that the data samples be as many as possible, the data be diverse, and the data sample quality be relatively high. Specifically, the static information can include, but is not limited to, information such as gender, education level, and / or package content, etc., and the dynamic information can include, but is not limited to, information such as Internet traffic, call duration, and / or the number of text messages sent, etc. These information can be obtained through regular collection. The customer portrait labels of the multiple sample customers are obtained by pre-manual annotation, specifically but not limited to lost elderly customers, unmarried customers, solitary customers, monthly contract customers, high monthly fee customers, customers who prefer value-added services, and / or customers who have subscribed to the Internet service "Fiber optic" and whose payment method is "Electronic check", etc. In addition, the unit period can be, but is not limited to, one week, one month, or one quarter, etc.

[0055] S2. Perform exploratory data analysis on the customer sample data and customer portrait labels of the multiple sample customers to obtain the exploratory data analysis results.

[0056] In the step S2, the Exploratory Data Analysis (EDA, which refers to a data analysis method that explores the structure and rules of existing data through means such as graphing, tabulating, equation fitting, and calculating characteristic quantities with as few prior assumptions as possible) processing is used to effectively help users become familiar with the data and understand the data: initially analyze the mutual relationship between variables and the relationship between variables and predicted values, and can obtain a perceptual understanding of the data set and inspiration for feature engineering. Specifically, perform Exploratory Data Analysis processing on the customer sample data and the customer portrait labels of the multiple sample customers to obtain Exploratory Data Analysis results, including but not limited to the following aspects: (1) Perform an overview data profile viewing process on the customer sample data and the customer portrait labels of the multiple sample customers (which can help users quickly grasp the general range of the data and judge data anomalies, and view the default data situation of each column), and determine whether there is abnormal feature data that needs to be deleted (such as feature data with a variance of 0 or extremely low); (2) Perform a label distribution viewing process on the customer portrait labels of the multiple sample customers to determine whether there is customer sample data and customer portrait labels corresponding to outlier customers that need to be deleted (since this embodiment uses a regression classification model based on the XGBoost algorithm for model training, discrete point data as interference factors need to be excluded to ensure the convergence of model training); (3) Perform a feature distribution viewing process on the customer sample data of the multiple sample customers (such as visual viewing through a pie chart) to determine whether all situations are covered (if not, it is necessary to return to step S1 to re-obtain sample data); (4) Perform a correlation viewing process between features of the customer sample data of the multiple sample customers (such as visual viewing through a heat coefficient graph) to determine whether there is redundant feature data that needs to be deleted; (5) Perform a correlation viewing process between features and labels on the customer sample data and the customer portrait labels of the multiple sample customers (such as visual viewing through the DataFrame.corr() function) to determine whether there is feature data with low correlation that needs to be deleted.

[0057] S3. According to the Exploratory Data Analysis results, perform data preprocessing on the customer sample data of the multiple sample customers to obtain multiple standardized customer sample data that are in one-to-one correspondence with multiple customer portrait labels and have the characteristics of clean and continuous data.

[0058] In the step S3, since the sample data may contain a large number of missing values, noise data, or there are outliers due to manual input errors, which is very unfavorable for the training of the algorithm model. Therefore, it is necessary to clean various dirty data through the data preprocessing to obtain standard, clean, and continuous new sample data for data statistics, data mining, etc. The quality of the data preprocessing determines the accuracy and generalization value of subsequent data analysis, mining, and modeling work. Specifically, the data preprocessing includes, but is not limited to, removing unimportant features irrelevant to data mining, filling in missing data values (specific filling schemes include using the mean or median to fill in missing values for continuous values; using the mode to fill in for discrete values, or treating the missing value as a separate category; for filling in data suitable for time series features that change over time, linear interpolation can be used; in addition, different models can be used to predict missing values, such as using the KNN nearest neighbor method to predict missing values), processing and filtering abnormal data (abnormal data can be found through box plots, more than 3 standard deviations, or custom methods, and then processed through deleting observations, transformation, grouping, estimation, or other statistical methods), eliminating redundant data through data deduplication, uniformly converting data with inconsistent units, reducing the dimension of high-dimensional features (to extract useful features), numericalizing categorical features (for example, methods such as Scikit-learn's LabelEncoder, OrdinalEncoder, OneHotEncoder, Binarizer, KBinsDiscretizer, etc. can be used), and / or dimensionless processing of data (including standardization and discretization processing methods), etc.

[0059] S4. Perform feature engineering analysis and processing on the multiple standardized customer sample data, and select multiple first data feature sets that correspond one-to-one to the multiple customer portrait labels and are suitable for model input.

[0060] In the step S4, after the data preprocessing is completed, it is necessary to select meaningful features and input them into the algorithms and models of machine learning for training. Generally speaking, the selection of features can be considered from the following two aspects: whether the features are divergent, and the correlation between the features and the target. Specifically, the feature engineering analysis and processing includes but is not limited to the following methods for feature selection: the filtering method (Filter, which scores each feature according to divergence or correlation, sets a threshold or the number of thresholds to be selected, and selects features; mainly includes the variance selection method, Pearson correlation coefficient, distance correlation coefficient, F-statistic, mutual information method, chi-square test, F-statistic, and mutual information such as Mutual Information), the wrapper method (Wrapper, which selects several features or excludes several features each time according to the objective function; mainly includes the recursive feature elimination method and the feature interference method, etc.), and / or the embedded method (Embedded, which first uses certain machine learning algorithms and models for training to obtain the weight coefficients of each feature, and selects features according to the coefficients from large to small; mainly includes the feature selection method based on penalty terms and the feature selection method based on tree models).

[0061] S5. Take the multiple first data feature sets as model input items, and take the multiple customer portrait labels as model output items, and import them into a regression classification model based on the XGBoost algorithm for model training to obtain a churn telecom customer portrait model.

[0062] In the step S5, XGBoost is the abbreviation of Exterme Gradient Boosting. It is an ensemble machine learning algorithm based on decision trees, and it takes gradient boosting as the framework. XGBoost is developed from GBDT, and it also uses the additive model and the forward stagewise algorithm to achieve the optimization process of learning. XGBoost has the following advantages: introducing the second derivative, increasing the convergence speed; adding regularization to prevent overfitting; supporting parallelism, increasing the processing speed; Shrinkage weakens the influence of each tree, giving more learning space for the following; supporting column sampling, which can not only reduce overfitting but also reduce calculations; being able to internally handle missing values of feature values; having built-in cross-validation to conveniently obtain the optimal number of Boosting iterations; being able to handle high-dimensional sparse features.

[0063] The objective function of XGBoost consists of two parts: the loss function and the regularization term:

[0064]

[0065] where: ∑ k Ω(f k ) represents the complexity of k trees; It is an expression on a linear space; i is the i-th sample, and k is the k-th tree; is the i-th sample x i predicted value of

[0066] Since then is transformed into the following form:

[0067]

[0068] Perform a second-order Taylor expansion on the objective function, remove the constant term, and optimize the loss function term to obtain:

[0069]

[0070] where the first derivative and the second derivative is

[0071] Then perform a regularization term expansion, remove the constant term, and optimize the regularization term to obtain:

[0072]

[0073] Then combine the first-order term coefficient and the second-order term coefficient to obtain the final objective function:

[0074]

[0075] Define: where G j : The sum of the first-order partial derivatives of the samples contained in leaf node j, which is a constant. H j : The sum of the second-order partial derivatives of the samples contained in leaf node j, which is a constant. γT controls the structure of the tree, where T represents the number of leaf nodes of tree f. w j represents the sample weight of the current leaf node.

[0076] Substitute G j and H j into the objective function to get:

[0077]

[0078] Then the objective function of each leaf node j is:

[0079]

[0080] Since (H j +λ)>0, then f(w j ) reaches the minimum value at , and the minimum value is

[0081] The objective expressions of each leaf node of the XGBoost objective function are independent of each other. That is, when the expressions of each leaf node reach the maximum or minimum points, the entire objective function also reaches the maximum or minimum point. Then the weight of each leaf node and the optimal objective value Obj reached at this time:

[0082]

[0083] Objective value is the smallest, the tree structure is the best, and this is the optimal solution of the objective function at this time.

[0084] In step S5, it is necessary to optimize the machine learning model to ensure the highest accuracy of the model. Specific hyperparameter search algorithms can include but are not limited to the learning curve of manual hyperparameters, grid search, random search, Bayesian optimization, genetic algorithms, etc. In addition, during the model training process, specific regression performance metrics can include but are not limited to mean absolute error, mean squared error, root mean squared error, and R-squared, and specific classification performance metrics can include but are not limited to error rate, accuracy, precision, recall, F1, ROC curve, AUC curve, and R-squared.

[0085] S6. Obtain the static information of the target customer and the dynamic information within the most recent unit period.

[0086] S7. Convert the static information and the dynamic information of the target customer into standardized customer data with the characteristics of clean and continuous data.

[0087] S8. Select a second data feature set from the standardized customer data whose data types are consistent with the first data feature set.

[0088] S9. Import the second data feature set into the churn telecom customer portrait model, and output the confidence levels for classifying the target customer into different customer portrait labels.

[0089] S10. When the confidence level for classifying the target customer into a certain customer portrait label exceeds the preset confidence threshold, push the preset retention plan corresponding to the certain customer portrait label to the current telecom service provider of the target customer.

[0090] In the step S10, for example, when the certain customer portrait label is a lapsed elderly customer, an unmarried customer, a solitary customer, etc., the corresponding preset retention plan may but is not limited to increasing the attention to such customers, or recommending preferential packages for such customers; when the certain customer portrait label is a lapsed customer with a monthly contract, the corresponding preset retention plan may but is not limited to recommending the signing of a 2-year contract; when the certain customer portrait label is a lapsed customer with a preference for value-added services, the corresponding preset retention plan may but is not limited to intensifying the publicity of subsequent network value-added services (such as "online backup service", "technical support service", "device protection service" and "network security service", etc.) to increase the customer's dependence on the service; when the certain customer portrait label is a lapsed customer who has subscribed to the Internet service "Fiber optic" and whose payment method is "Electronic check", the corresponding preset retention plan may but is not limited to conducting a return visit survey to check the customer's feedback. Thus, for some customers who are about to lapse, by implementing the preset retention plan, the customer churn rate can be reduced and the service quality can be improved.

[0091] Based on the foregoing telecommunications customer churn prediction method described in steps S1 to S10, a new solution for realizing telecommunications customer churn prediction by combining portrait modeling and the XGBoost algorithm is provided, that is, first, exploratory data analysis processing, data preprocessing, feature engineering analysis processing, and model training of a regression classification model based on the XGBoost algorithm are sequentially performed on the customer sample data and customer portrait labels of multiple sample customers, and a lapsed telecommunications customer portrait model can be obtained. Then, the test data of the target customer is imported into this portrait model, and the confidence levels of classifying the target customer into different customer portrait labels can be obtained. Finally, when the confidence level of classifying the target customer into a certain customer portrait label exceeds the preset confidence threshold, a preset retention plan corresponding to the certain customer portrait label is pushed to the current telecommunications service provider of the target customer. In this way, some customers who are about to lapse can be predicted in advance, and by pushing the preset retention plan, the customer churn rate can be reduced and the service quality can be improved, which is convenient for practical application and promotion.

[0092] As Figure 2 shown, in the second aspect of this embodiment, a virtual device for implementing the telecommunications customer churn prediction method described in the first aspect is provided, including a customer data acquisition module, an exploratory data analysis module, a data preprocessing module, a feature engineering analysis module, a model training module, a customer data conversion module, a feature data selection module, a model application module, and a retention plan recommendation module;

[0093] The customer data acquisition module is used to acquire customer sample data and customer portrait labels of multiple sample customers. Herein, the sample customers refer to the lost telecom customers, and the customer sample data includes the static information of the corresponding customers and the dynamic information in the most recent unit period before churn.

[0094] The data exploratory analysis module is communicatively connected to the customer data acquisition module and is used to perform data exploratory analysis processing on the customer sample data and the customer portrait labels of the multiple sample customers to obtain a data exploratory analysis result.

[0095] The data preprocessing module is communicatively connected to the data exploratory analysis module and is used to perform data preprocessing on the customer sample data of the multiple sample customers according to the data exploratory analysis result to obtain multiple standardized customer sample data that are in one-to-one correspondence with multiple customer portrait labels and have the characteristics of clean and continuous data.

[0096] The feature engineering analysis module is communicatively connected to the data preprocessing module and is used to perform feature engineering analysis processing on the multiple standardized customer sample data, and select multiple first data feature sets that are in one-to-one correspondence with the multiple customer portrait labels and are suitable for model input.

[0097] The model training module is communicatively connected to the feature engineering analysis module and is used to use the multiple first data feature sets as model input items and the multiple customer portrait labels as model output items, and import them into a regression classification model based on the XGBoost algorithm for model training to obtain a lost telecom customer portrait model.

[0098] The customer data acquisition module is further used to acquire the static information of the target customer and the dynamic information in the most recent unit period.

[0099] The customer data conversion module is communicatively connected to the customer data acquisition module and is used to convert the static information and the dynamic information of the target customer into standardized customer data with the characteristics of clean and continuous data.

[0100] The feature data selection module is communicatively connected to the customer data conversion module and is used to select a second data feature set from the standardized customer data whose data type is the same as that of the first data feature set.

[0101] The model application module is communicatively connected to the model training module and the feature data selection module respectively, and is used to import the second data feature set into the lost telecom customer portrait model and output the confidence levels of classifying the target customer into different customer portrait labels.

[0102] The retention plan recommendation module is communicatively connected to the model application module and is configured to push a preset retention plan corresponding to the certain customer portrait label to the current telecommunications service provider of the target customer when the confidence level of classifying the target customer into a certain customer portrait label exceeds a preset confidence threshold.

[0103] For the working process, working details and technical effects of the foregoing device provided in the second aspect of this embodiment, reference may be made to the telecommunications customer churn prediction method described in the first aspect, which will not be elaborated herein.

[0104] As Figure 3 shown, a computer device for executing the telecommunications customer churn prediction method described in the first aspect is provided in the third aspect of this embodiment, including a memory, a processor, and a transceiver communicatively connected in sequence. Among them, the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the telecommunications customer churn prediction method described in the first aspect. Specifically, for example, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first input first output (FIFO), and / or first input last output (FILO), etc.; the processor may be, but is not limited to, a microprocessor of the STM32F105 series. In addition, the computer device may also include, but is not limited to, a power module, a display screen, and other necessary components.

[0105] For the working process, working details and technical effects of the foregoing computer device provided in the third aspect of this embodiment, reference may be made to the telecommunications customer churn prediction method described in the first aspect, which will not be elaborated herein.

[0106] A computer-readable storage medium storing instructions including the telecommunications customer churn prediction method described in the first aspect is provided in the fourth aspect of this embodiment, that is, instructions are stored on the computer-readable storage medium, and when the instructions run on a computer, the telecommunications customer churn prediction method described in the first aspect is executed. Among them, the computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, computer-readable storage media such as floppy disks, optical discs, hard disks, flash memories, USB flash drives, and / or memory sticks. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices.

[0107] For the working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment, reference may be made to the telecom customer churn prediction method described in the first aspect, which will not be elaborated herein.

[0108] In the fifth aspect of this embodiment, a computer program product containing instructions is provided. When the instructions run on a computer, the computer is caused to execute the telecom customer churn prediction method described in the first aspect. Among them, the computer may be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices.

[0109] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for predicting telecom customer churn based on the XGBoost algorithm, characterized in that, it includes: Obtain the customer sample data and customer portrait labels of multiple sample customers. Among them, the sample customers refer to the telecom customers who have churned, and the customer sample data includes the static information of the corresponding customers and the dynamic information in the most recent unit period before churn; Conduct exploratory data analysis on the customer sample data and the customer portrait labels of the multiple sample customers to obtain the exploratory data analysis results; According to the exploratory data analysis results, preprocess the customer sample data of the multiple sample customers to obtain multiple standardized customer sample data that correspond one-to-one with multiple customer portrait labels and have the characteristics of clean and continuous data; Conduct feature engineering analysis on the multiple standardized customer sample data, and select multiple first data feature sets that correspond one-to-one with the multiple customer portrait labels and are suitable for model input; Use the multiple first data feature sets as model input items, and use the multiple customer portrait labels as model output items, and import them into a regression classification model based on the XGBoost algorithm for model training to obtain a churned telecom customer portrait model; Obtain the static information of the target customer and the dynamic information in the most recent unit period; Convert the static information and the dynamic information of the target customer into standardized customer data with the characteristics of clean and continuous data; Select a second data feature set from the standardized customer data whose data type is the same as that of the first data feature set; Import the second data feature set into the churned telecom customer portrait model, and output the confidence levels for classifying the target customer into different customer portrait labels; When the confidence level for classifying the target customer into a certain customer portrait label exceeds the preset confidence threshold, push a preset retention plan corresponding to the certain customer portrait label to the current telecom service provider of the target customer.

2. The telecom customer churn prediction method according to claim 1, characterized in that, Conduct exploratory data analysis on the customer sample data and the customer portrait labels of the multiple sample customers to obtain the exploratory data analysis results, including: Conduct an overview data profile check on the customer sample data and the customer portrait labels of the multiple sample customers to determine whether there is abnormal feature data that needs to be deleted.

3. The telecom customer churn prediction method according to claim 1, characterized in that, Conduct exploratory data analysis on the customer sample data and the customer portrait labels of the multiple sample customers to obtain the exploratory data analysis results, including: Conduct a label distribution check on the customer portrait labels of the multiple sample customers to determine whether there is customer sample data and customer portrait labels corresponding to outlier customers that need to be deleted.

4. The telecom customer churn prediction method according to claim 1, characterized in that, Conduct a feature distribution check on the customer sample data of the multiple sample customers to determine whether all situations are covered.

5. The telecom customer churn prediction method according to claim 1, characterized in that a correlation check process is performed between features of the customer sample data of the multiple sample customers to determine whether there is redundant feature data that needs to be deleted; and / or, a correlation check process is performed between features and labels of the customer sample data of the multiple sample customers and the customer portrait labels to determine whether there is feature data with low correlation that needs to be deleted.

6. The telecom customer churn prediction method according to claim 1, characterized in that the data preprocessing includes removing features that are not important for data mining, filling missing data values, processing and filtering abnormal data, eliminating redundant data through data deduplication, uniformly converting data with inconsistent units, reducing the dimension of high-dimensional features, numericalizing categorical features, and / or making data dimensionless.

7. The telecom customer churn prediction method according to claim 1, characterized in that the feature engineering analysis process includes feature selection through a filtering method, a wrapper method, and / or an embedding method.

8. A telecom customer churn prediction device based on the XGBoost algorithm, characterized in that it includes a customer data acquisition module, a data exploratory analysis module, a data preprocessing module, a feature engineering analysis module, a model training module, a customer data conversion module, a feature data selection module, a model application module, and a retention plan recommendation module; the customer data acquisition module is used to acquire the customer sample data and customer portrait labels of multiple sample customers, where the sample customers refer to telecom customers who have churned, and the customer sample data includes the static information of the corresponding customers and the dynamic information in the most recent unit period before churn; the data exploratory analysis module is communicatively connected to the customer data acquisition module and is used to perform data exploratory analysis processing on the customer sample data and the customer portrait labels of the multiple sample customers to obtain a data exploratory analysis result; the data preprocessing module is communicatively connected to the data exploratory analysis module and is used to perform data preprocessing on the customer sample data of the multiple sample customers according to the data exploratory analysis result to obtain multiple standardized customer sample data that are in one-to-one correspondence with multiple customer portrait labels and have the characteristics of clean and continuous data; the feature engineering analysis module is communicatively connected to the data preprocessing module and is used to perform feature engineering analysis processing on the multiple standardized customer sample data to select multiple first data feature sets that are in one-to-one correspondence with the multiple customer portrait labels and are suitable for model input; the model training module is communicatively connected to the feature engineering analysis module and is used to use the multiple first data feature sets as model input items and the multiple customer portrait labels as model output items, and import them into a regression classification model based on the XGBoost algorithm for model training to obtain a churned telecom customer portrait model; the customer data acquisition module is further used to acquire the static information of the target customer and the dynamic information in the most recent unit period; The customer data conversion module is communicatively connected to the customer data acquisition module and is used to convert the static information and the dynamic information of the target customer into standardized customer data with the characteristics of clean and continuous data; The feature data selection module is communicatively connected to the customer data conversion module and is used to select a second data feature set from the standardized customer data, where the data type of the second data feature set is consistent with the first data feature set; The model application module is communicatively connected to the model training module and the feature data selection module respectively, and is used to import the second data feature set into the churn telecom customer portrait model and output the confidence levels for classifying the target customer into different customer portrait labels; The retention plan recommendation module is communicatively connected to the model application module and is used to push a preset retention plan corresponding to a certain customer portrait label to the current telecom service provider of the target customer when the confidence level for classifying the target customer into a certain customer portrait label exceeds a preset confidence threshold.

9. A computer device, characterized in that, it includes a memory, a processor and a transceiver that are communicatively connected in sequence, where the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the telecom customer churn prediction method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that , instructions are stored on the computer-readable storage medium, and when the instructions run on a computer, the telecom customer churn prediction method according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Customer portrait-based customer loss prediction and retrieval method and system

    CN112561598A

  • Bank customer loss prediction method and device thereof, storage medium and electronic equipment

    CN114049159A