Prediction method and device for potential network migration churn users

By using the potential mobile network churn user prediction model based on historical user feature data, the potential mobile network churn user is accurately predicted, and the problem of poor prediction accuracy in the existing technology is solved, and effective user retention and operational efficiency improvement is achieved.

CN120020848APending Publication Date: 2025-05-20CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311549801.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The existing technology is difficult to accurately predict potential mobile network churn users, resulting in uncertain investment in operators when retaining users.

Method used

By obtaining the feature data of the target user and inputting it into the potential mobile network churn user prediction model trained based on historical user feature data, a cross-verification stack classifier is used to make predictions to determine whether the user is a potential mobile network churn user.

Benefits of technology

Accurate prediction of potential mobile network churn users is achieved, providing operators with accurate user identification, helping to effectively retain potential churn users, reduce churn rates, improve user satisfaction and operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020848A_ABST
    Figure CN120020848A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for predicting potential mobile network loss users. The method comprises the following steps: acquiring target user feature data in a first time period; inputting the feature data of the target user into a potential mobile network loss user prediction model for prediction to obtain a loss probability prediction value of the target user; the potential network migration loss user prediction model is obtained by performing training and algorithm parameter tuning by adopting a cross validation stacking classifier based on historical user feature data in a second time period; the second time period is longer than the first time period; according to the loss probability prediction value of the target user, judging whether the target user is a potential network migration loss user; and when the loss probability predicted value of the target user is greater than a loss probability preset value, determining that the target user is a potential network migration loss user. The method can accurately predict the potential network migration churn users, provides accurate identification for retention measures, and can effectively retain the potential network migration churn users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and in particular, to a method and apparatus for predicting potential mobile network churn users. Background Art

[0002] At present, there are more and more mobile network churn users among major operators, and the impact of user churn on operators is also increasing.

[0003] For existing mobile network churn users, the technical solutions for data analysis focus on post-event analysis, performing root cause analysis on the churned users, including network analysis and user complaint analysis, in order to find out the reasons for churn and take remedial measures.

[0004] Although the technical solutions for pre-event analysis of churn users cover a wider range of churn users, their accuracy for mobile network churn users is unsatisfactory, and they cannot accurately predict potential mobile network churn users. And the accuracy of prediction directly affects the investment of operators in retaining mobile network churn users in the later stage. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and apparatus for predicting potential mobile network churn users in view of the above deficiencies of the prior art. The method can accurately predict potential mobile network churn users, provide accurate identification for retention measures, and thus can effectively retain potential mobile network churn users.

[0006] In a first aspect, the present invention provides a method for predicting potential mobile network churn users, the method comprising the following steps:

[0007] Step S1: Obtain target user feature data within a first time period;

[0008] Step S2: Input the target user feature data into a potential mobile network churn user prediction model for prediction to obtain a predicted value of the churn probability of the target user;

[0009] Wherein, the potential mobile network churn user prediction model is trained and optimized for algorithm parameters based on historical user feature data within a second time period by using a cross-validation stacking classifier; the second time period is greater than the first time period;

[0010] Step S3: Determine whether the target user is a potential mobile network churn user according to the predicted value of the churn probability of the target user;

[0011] Wherein, when the predicted value of the churn probability of the target user is greater than a preset churn probability value, it is determined that the target user is a potential mobile network churn user; when the predicted value of the churn probability of the target user is less than or equal to the preset churn probability value, it is determined that the target user is marked as a non-potential mobile network churn user.

[0012] Further, step S1 specifically includes the following steps:

[0013] Step S11: Obtain the original communication record information table within the first time period;

[0014] Step S12: Aggregate and filter the data of the original communication record information table to obtain the target user feature data within the first time period.

[0015] Further, step S12 specifically includes the following steps:

[0016] Step A1: Perform category aggregation according to the user dimension of the original communication record information table, and extract the aggregated feature data;

[0017] Among them, the category aggregation according to the user dimension of the original communication record information table includes category aggregation according to the user portrait dimension, category aggregation according to the user package dimension, category aggregation according to the user usage record dimension, and category aggregation according to the user complaint dimension;

[0018] Step A2: Perform feature selection on the aggregated feature data to obtain the retained feature data;

[0019] The feature selection includes:

[0020] Delete the aggregated feature data with a missing value ratio greater than the missing preset value; and / or,

[0021] Delete the aggregated feature data with a category ratio greater than the category ratio preset value; and / or,

[0022] Delete the aggregated feature data with a feature importance value less than the feature importance preset value;

[0023] Step A3: Perform data preprocessing on the retained feature data to obtain the preprocessed feature data;

[0024] The data preprocessing includes:

[0025] After determining that the retained feature data is categorical feature data, perform median filling on the retained feature data and then perform label encoding; and / or,

[0026] After determining that the retained feature data is numerical feature data, perform linear interpolation on the retained feature data and then perform normalization processing; and / or,

[0027] After determining that the retained feature data is neither categorical feature data nor numerical feature data, perform piecewise discretization processing on the retained feature data;

[0028] Step A4: Generate features from the preprocessed feature data to obtain the target user feature data within the first time period;

[0029] Among them, the feature generation includes time feature generation, ratio feature generation, and frequency feature generation;

[0030] The time feature generation includes:

[0031] Generate the user network access year feature and user network access month feature according to the user network access time, and generate the modification year feature and modification month feature according to the modification time;

[0032] The ratio feature generation includes:

[0033] Generate the user network access duration ratio feature according to the user network access duration, and generate the monthly rent cost ratio feature according to the monthly rent cost feature;

[0034] The frequency feature generation includes: generating frequency features according to the label class features.

[0035] Further, in the step A2, deleting the aggregated feature data with the feature importance less than the feature importance preset value specifically includes the following steps:

[0036] Step A21: Calculate the overall Gini coefficient of the aggregated feature data;

[0037]

[0038] Among them, Gini 总体 represents the overall Gini coefficient, and P i represents the proportion of the aggregated feature data of the i-th aggregation class in the overall; n represents the total number of categories;

[0039] Step A22: Calculate the subset Gini coefficient of each aggregated feature;

[0040]

[0041] Among them, Gini 子集 represents the subset Gini coefficient, and P j represents the proportion of the j-th data feature data in the subset; m represents the data composition of the subset;

[0042] Step A23: Calculate the weighted average Gini coefficient of the aggregated feature data;

[0043] The calculation process of the weighted average Gini coefficient is to multiply the subset Gini coefficient of each aggregated feature by the corresponding proportion of each subset, and then sum them up;

[0044] Step A24: Calculate the feature importance value;

[0045] The feature importance value is obtained by subtracting the weighted average Gini coefficient from the overall Gini coefficient;

[0046] V 重要性 = Gini 总体 - Gini 加权平均

[0047] where,

[0048] V 重要性 represents the feature importance value, and Gini 加权平均 represents the weighted average Gini coefficient;

[0049] Step A25: Compare the feature importance value with the feature importance preset value, and delete the aggregated feature data whose feature importance is less than the feature importance preset value.

[0050] Furthermore, after the step S3, there is also a step S4,

[0051] Step S4: Conduct root cause analysis and remedial measures on potential mobile network churn users;

[0052] The step S4 specifically includes the following steps:

[0053] Step S41: After determining that the target user is a potential mobile network churn user, analyze the churn reasons of the potential mobile network churn users to form a list of potential mobile network churn users and a list of potential churn reasons;

[0054] Step S42: Take predicted measures to remedy the churn reasons of the potential mobile network churn users to form a list of remedial measures;

[0055] Step S43: Submit the list of potential mobile network churn users, the list of potential churn reasons, and the list of remedial measures to the marketers so that the marketers can implement the remedial measures according to the list of potential mobile network churn users and the list of potential churn reasons.

[0056] Furthermore, before the step S1, there is also a step S0;

[0057] Step S0: Build a prediction model for potential mobile network churn users;

[0058] The step S0 specifically includes the following steps:

[0059] Step S01: Obtain the data set within the second time period; the data set within the second time period includes historical user features and corresponding historical variables; the historical variables include mobile network churn users and non-mobile network churn users;

[0060] Step S02: Divide the dataset within the second time period into a training dataset and a test dataset;

[0061] Step S03: Use a stacked classifier to combine a first-stage classifier and a second-stage classifier, and perform cross-validation on the training dataset to evaluate the performance of different parameter combinations to obtain a preliminary prediction model; the first-stage classifier is LightGBM, XGBoost, and CatBoost, and the second-stage classifier is a logistic regression classifier;

[0062] Step S04: Optimize the performance of the preliminary prediction model through the test dataset to obtain a potential mobile network churn user prediction model.

[0063] In a second aspect, the present invention provides a device for predicting potential mobile network churn users, and the device includes:

[0064] An acquisition unit, configured to collect target user feature data within the first time period;

[0065] A prediction unit, connected to the acquisition unit, configured to input the target user feature data into a potential mobile network churn user prediction model for prediction to obtain a predicted value of the churn probability of the target user;

[0066] Wherein, the potential mobile network churn user prediction model is based on historical user feature data within the second time period, and is trained and optimized for algorithm parameters using a cross-validation stacked classifier; the second time period is greater than the first time period;

[0067] A determination unit, connected to the prediction unit, configured to determine whether the target user is a potential mobile network churn user according to the predicted value of the churn probability of the target user;

[0068] Wherein, when the predicted value of the churn probability of the target user is greater than the preset churn probability value, it is determined that the target user is a potential mobile network churn user; when the predicted value of the churn probability of the target user is less than or equal to the preset churn probability value, it is determined that the target user is marked as a non-potential mobile network churn user.

[0069] Further, the acquisition unit includes:

[0070] A first acquisition module, configured to acquire an original communication record information table within the first time period;

[0071] An aggregation and screening module, connected to the acquisition module, configured to perform data aggregation and data screening on the original communication record information table to obtain target user feature data within the first time period.

[0072] Further, the aggregation and screening module includes:

[0073] An aggregation sub-module, configured to perform category aggregation according to the user dimension of the original communication record information table, and extract aggregated feature data;

[0074] Among them, the category aggregation according to the user dimension of the original communication record information table includes category aggregation according to the user portrait dimension, category aggregation according to the user package dimension, category aggregation according to the user usage record dimension, and category aggregation according to the user complaint dimension;

[0075] A first screening sub-module, connected to the aggregation sub-module, configured to perform feature selection on the aggregated feature data to obtain retained feature data;

[0076] The feature selection includes:

[0077] Deleting the aggregated feature data with a missing value ratio greater than the missing preset value; and / or,

[0078] Deleting the aggregated feature data with a category ratio greater than the category ratio preset value; and / or,

[0079] Deleting the aggregated feature data with a feature importance value less than the feature importance preset value;

[0080] A second screening sub-module, connected to the first screening sub-module, configured to perform data preprocessing on the retained feature data to obtain preprocessed feature data;

[0081] The data preprocessing includes:

[0082] After determining that the retained feature data is categorical feature data, performing median filling on the retained feature data and then performing label encoding; and / or,

[0083] After determining that the retained feature data is numerical feature data, performing linear interpolation on the retained feature data and then performing normalization processing; and / or,

[0084] After determining that the retained feature data is non-categorical feature data or non-numerical feature data, performing piecewise discretization processing on the retained feature data;

[0085] A generation sub-module, connected to the second screening sub-module, configured to perform feature generation on the preprocessed feature data to obtain target user feature data within a first time period;

[0086] Among them, the feature generation includes time feature generation, ratio feature generation, and frequency feature generation;

[0087] The time feature generation includes:

[0088] Generate the user's network access year feature and user's network access month feature according to the user's network access time, and generate the modification year feature and modification month feature according to the modification time;

[0089] The generation of the proportional feature includes:

[0090] Generate the user's network access duration proportional feature according to the user's network access duration, and generate the monthly rent cost proportional feature according to the monthly rent cost feature;

[0091] The generation of the frequency feature includes: generating the frequency feature according to the tag class feature.

[0092] Further, the device further includes a construction unit, and the construction unit is connected to the prediction unit for constructing a potential mobile network churn user prediction model, so that the prediction unit inputs the target user feature data into the potential mobile network churn user prediction model for prediction;

[0093] The construction unit includes:

[0094] A second acquisition module for acquiring the data set within the second time period; the data set within the second time period includes historical user features and corresponding historical variables; the historical variables include mobile network churn users and non-mobile network churn users;

[0095] A division module, connected to the second acquisition module, for dividing the data set within the second time period into a training data set and a test data set;

[0096] A combined verification module, connected to the division module, for using a stacking classifier to combine a first-stage classifier and a second-stage classifier together, and performing cross-validation on the training data set to evaluate the performance of different parameter combinations to obtain a preliminary prediction model; the first-stage classifier is LightGBM, XGBoost, and CatBoost, and the second-stage classifier is a logistic regression classifier;

[0097] A tuning module, connected to the division module and the combined verification module respectively, for tuning the performance of the preliminary prediction model through the test data set to obtain a potential mobile network churn user prediction model.

[0098] The beneficial effects of the present invention:

[0099] 1. The present invention can accurately predict potential mobile network churn users: through the potential mobile network churn user prediction model for prediction, obtain the predicted value of the churn probability of the target user, and then accurately judge the potential mobile network churn users according to the predicted value of the churn probability, so as to provide accurate identification for the retention measures, and then effectively retain potential mobile network churn users.

[0100] 2. The present invention can reduce the churn rate of mobile networks: By predicting potential mobile network churn users, operators can take targeted measures, such as providing personalized offers, improving service quality, etc., to reduce the churn rate and retain more users.

[0101] 3. The present invention can improve user satisfaction: By predicting potential churn users, operators can timely discover users' problems and needs, improve services targetedly, enhance user satisfaction, and increase user stickiness.

[0102] 4. The present invention can improve operational efficiency: Predicting potential mobile network churn users can help operators optimize resource allocation, improve operational efficiency, and reduce unnecessary marketing costs and customer service costs.

[0103] 5. The present invention can promote refined marketing: Conducting predictive analysis on potential churn users can help operators implement precise marketing strategies, improve marketing effectiveness, and reduce marketing costs.

[0104] 6. The present invention can enhance the competitiveness of operators: By predicting potential mobile network churn users and taking effective retention measures, operators can increase market share, enhance competitiveness, and thus achieve sustainable development. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] Figure 1 It is a schematic diagram of the method for predicting potential mobile network churn users in an embodiment of the present invention;

[0106] Figure 2 It is a schematic diagram of the prediction process of potential mobile network churn users in an embodiment of the present invention;

[0107] Figure 3 It is a schematic diagram of the device for predicting potential mobile network churn users in an embodiment of the present invention;

[0108] Among them, reference numerals: 10, acquisition unit; 20, prediction unit; 30, determination unit. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0109] To enable those skilled in the art to better understand the technical solutions of the present invention, the embodiments of the present invention will be further described in detail below in conjunction with the accompanying drawings.

[0110] It can be understood that the specific embodiments and the accompanying drawings described herein are only used to explain the present invention, rather than limiting the present invention.

[0111] It can be understood that, without conflict, the various embodiments in the present invention and the various features in the embodiments can be combined with each other.

[0112] It can be understood that, for the convenience of description, only the parts related to the present invention are shown in the drawings of the present invention, while the parts unrelated to the present invention are not shown in the drawings.

[0113] It can be understood that each unit and module involved in the embodiments of the present invention may correspond to only one entity structure, or may be composed of multiple entity structures. Alternatively, multiple units and modules may also be integrated into one entity structure.

[0114] It can be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present invention may occur in a different order from that marked in the drawings.

[0115] It can be understood that in the flowcharts and block diagrams of the present invention, the possible system architectures, functions, and operations of the systems, devices, equipment, and methods according to the embodiments of the present invention are shown. Among them, each block in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart may be implemented by a hardware-based system for implementing the specified function, or may be implemented by a combination of hardware and computer instructions.

[0116] It can be understood that the units and modules involved in the embodiments of the present invention may be implemented in software or in hardware. For example, the units and modules may be located in the processor.

[0117] Embodiment 1:

[0118] As Figure 1 and Figure 2 shown, this embodiment provides a method for predicting potential users who may leave the mobile network. The method includes the following steps:

[0119] Step S1: Obtain the target user feature data within the first time period.

[0120] As a specific implementation manner, step S1 specifically includes the following steps:

[0121] Step S11: Obtain the original communication record information table within the first time period; the obtaining may be directly collecting 13 information tables for 12 months such as user basic information, customer real-name derivative information, user integral information, single-user traffic information, single-user voice information, single-user usage information, user bill information, new and lost user information, integrated user information, activity derivative information, stored value record information, bad debt system arrears information, and user complaint record information.

[0122] In the most recent 12 months, going back 12 months from the current training time. For example, for this month's prediction, the range is the data from May 2022 to May 2023. The 13 tables are user information tables in different dimensions. Specifically, it may not be restricted by the number of tables, but by the user information dimension, such as the user's integral information, voice information, etc.

[0123] Step S12: Aggregate and filter the data in the original communication record information table to obtain the target user feature data within the first time period.

[0124] As a more specific implementation manner, step S12 specifically includes the following steps:

[0125] Step A1: Perform category aggregation according to the user dimension of the original communication record information table to extract the aggregated feature data;

[0126] Among them, performing category aggregation according to the user dimension of the original communication record information table includes performing category aggregation according to the user portrait dimension, performing category aggregation according to the user package dimension, performing category aggregation according to the user usage record dimension, and performing category aggregation according to the user complaint dimension.

[0127] Category aggregation is to aggregate the 13 information tables according to the user dimension, and extract 133 user features in categories such as user portrait categories, user package categories, user usage record categories, and user complaint categories.

[0128] For the 13 information tables, first each table is grouped and aggregated (groupby) based on the user code to extract features (select the maximum value, minimum value, median, average value, etc. according to different features), and then using the user basic information table as the basic table, the other 12 tables are associated with the user code to obtain 133 user features.

[0129] Step A2: Perform feature selection on the aggregated feature data to obtain the retained feature data;

[0130] Feature selection includes:

[0131] Deleting the aggregated feature data with the missing value ratio greater than the missing preset value; and / or,

[0132] Deleting the aggregated feature data with the category ratio greater than the category ratio preset value; and / or,

[0133] Deleting the aggregated feature data with the feature importance value less than the feature importance preset value.

[0134] Specifically, this step is to perform feature selection on the above 133 user features, delete features with a missing value ratio greater than 95%, delete categorical features with a category ratio greater than 98%, use a decision tree for feature importance analysis, and delete features with a feature importance value less than 10. After feature selection, 100 user features are finally retained.

[0135] Step A3: Preprocess the retained feature data to obtain the preprocessed feature data.

[0136] Specifically, this step is to perform feature preprocessing on the above 100 user features. If it is a categorical feature, perform median filling and then label encoding; if it is a numerical feature, perform linear interpolation and then normalization; perform piecewise discretization on features such as customer age, customer scale, and cost.

[0137] If it is a categorical feature, perform median filling and then label encoding; if it is a numerical feature, perform linear interpolation and then normalization; perform piecewise discretization on features such as customer age, customer scale, and cost

[0138] Data preprocessing includes:

[0139] After determining that the retained feature data is categorical feature data, perform median filling on the retained feature data and then perform label encoding; and / or,

[0140] After determining that the retained feature data is numerical feature data, perform linear interpolation on the retained feature data and then perform normalization; and / or,

[0141] After determining that the retained feature data is neither categorical feature data nor numerical feature data, perform piecewise discretization on the retained feature data;

[0142] Step A4: Generate features from the preprocessed feature data to obtain the target user feature data within the first time period;

[0143] Among them, feature generation includes time feature generation, ratio feature generation, and frequency feature generation;

[0144] Time feature generation includes:

[0145] Generate the user network access year feature and the user network access month feature based on the user network access time, and generate the modification year feature and the modification month feature based on the modification time;

[0146] Ratio feature generation includes:

[0147] Generate the user network access duration ratio feature based on the user network access duration, and generate the monthly rent cost ratio feature based on the monthly rent cost feature;

[0148] Frequency feature generation includes: generating frequency features based on label class features.

[0149] Specifically, in this step, feature generation is performed on the above 100 user features. Year features and month features are generated for the user network access time and modification time. Ratio features are generated for the user network access duration and monthly rent cost features. Frequency features are generated for the label class features. After the feature generation step, 118 user features are obtained.

[0150] Year features and month features are generated for the user network access time and modification time. Ratio features are generated for the user network access duration and monthly rent cost features. Frequency features are generated for the label class features. After the feature generation step, 118 user features are obtained.

[0151] As a more specific implementation manner, in step A2, the aggregated feature data with feature importance less than the preset feature importance value is deleted, which specifically includes the following steps:

[0152] Step A21: Calculate the overall Gini coefficient of the aggregated feature data;

[0153]

[0154] Among them, Gini 总体 represents the overall Gini coefficient, and P i represents the proportion of the aggregated feature data of the i-th aggregation class in the overall; n represents the total number of categories;

[0155] Step A22: Calculate the subset Gini coefficient of each aggregated feature;

[0156]

[0157] Among them, Gini 子集 represents the subset Gini coefficient, and P j represents the proportion of the j-th data feature data in the subset; m represents the data composition of the subset;

[0158] Step A23: Calculate the weighted average Gini coefficient of the aggregated feature data;

[0159] The calculation process of the weighted average Gini coefficient is to multiply the subset Gini coefficient of each aggregated feature by the corresponding proportion of each subset, and then sum them up;

[0160] Step A24: Calculate the feature importance value;

[0161] The feature importance value is obtained by subtracting the weighted average Gini coefficient from the overall Gini coefficient;

[0162] V 重要性 = Gini 总体-Gini 加权平均

[0163] Among them,

[0164] V 重要性 represents the feature importance value, and Gini 加权平均 represents the weighted average Gini coefficient.

[0165] Specifically, in this step, model training is performed on the above 118 user features to construct a cross-validation stacked classifier. The classifier includes a first-stage classifier: lightgbm classifier, xgboost classifier, catboost classifier, and a second-stage classifier logistic regression classifier. The 118 user features are input into the cross-validation stacked classifier for training and algorithm parameter tuning to obtain a potential mobile network churn user prediction model;

[0166] Model training is to first construct a cross-validation stacked classifier. The classifier includes a first-stage classifier: lightgbm classifier, xgboost classifier, catboost classifier, and a second-stage classifier logistic regression classifier. The parameters of each classifier are set respectively, and the 118 user features are input into the cross-validation stacked classifier for training and algorithm parameter tuning. The sample code example is as follows: #

[0167] lgb_params = {'num_leaves': 30,'max_depth': 6, 'learning_rate': 0.05, 'n_estimators': 600,

[0168] 'n_jobs': 8, 'random_seed': 2022, 'use_label_encoder': False, 'eval_metric': ['logloss', 'auc', 'error']}

[0169] xgb_params = {'max_depth': 6, 'learning_rate': 0.05, 'n_estimators': 600, 'colsample_bytree': 0.7,

[0170] 'min_child_weight': 5, 'n_jobs': 8, 'use_label_encoder': False, 'eval_metric': ['logloss', 'auc', 'error']}

[0171] cab_params = {

[0172] 'iterations': 1000,

[0173] 'learning_rate': 0.05,

[0174] 'depth': 6,

[0175] 'l2_leaf_reg': 6,

[0176] 'silent': True,

[0177] 'thread_count': 8,

[0178] 'random_seed': 2022

[0179] }

[0180] lgb_model = lgb.LGBMClassifier(**lgb_params)

[0181] xgb_model = xgb.XGBClassifier(**xgb_params)

[0182] cab_model = ctb.CatBoostClassifier(**cab_params)

[0183] lr = LogisticRegression()

[0184] sclf = StackingCVClassifier(classifiers=[lgb_model, xgb_model, cab_model], use_probas=True, meta_classifier=lr, cv=5, verbose=1)。

[0185] Step A25: Compare the feature importance values with the preset feature importance values, and delete the aggregated feature data with feature importance less than the preset feature importance values.

[0186] Specifically, in Step A25, the user features for the next 3 months are input into the potential mobile network churn user prediction model for prediction. Those with a prediction probability greater than 0.5 are designated as potential mobile network churn users, and those less than 0.5 are non-potential mobile network churn users, generating a list of potential mobile network churn users.

[0187] The specific prediction process is as follows:

[0188] Load the model generated by training

[0189] Feed the prediction data into the model and output the prediction results

[0190] For those with a predicted probability greater than 0.5, they are designated as potential mobile network churn users, and those less than 0.5 are non-potential mobile network churn users.

[0191] Sample code:

[0192] predict_data = model.predict(test_df, num_iteration = model.best_iteration)

[0193] pred_result = []

[0194] for pred in predict_data:

[0195] if pred >= 0.5:

[0196] pred_result.append(1)

[0197] else:

[0198] pred_result.append(0)

[0199] predict_y = pd.DataFrame(pred_result, columns=['prediction']).

[0200] Step S2: Input the target user feature data into the potential mobile network churn user prediction model for prediction to obtain the predicted churn probability value of the target user.

[0201] The potential mobile network churn user prediction model is based on the historical user feature data within the second time period and is trained and optimized for algorithm parameters using a cross-validation stacking classifier; the second time period is greater than the first time period;

[0202] Step S3: According to the predicted churn probability value of the target user, determine whether the target user is a potential mobile network churn user.

[0203] When the predicted churn probability value of the target user is greater than the preset churn probability value, it is determined that the target user is a potential mobile network churn user; when the predicted churn probability value of the target user is less than or equal to the preset churn probability value, it is determined that the target user is marked as a non-potential mobile network churn user.

[0204] As a specific implementation, after step S3, it further includes step S4,

[0205] Step S4: Conduct root cause analysis and remedial measures for potential mobile network churn users;

[0206] Step S4 specifically includes the following steps:

[0207] Step S41: After determining that the target user is a potential mobile network churn user, analyze the reasons for the churn of potential mobile network churn users to form a list of potential mobile network churn users and a list of potential churn reasons;

[0208] Step S42: Take preventive measures to remedy the reasons for the churn of potential mobile network churn users to form a list of remedial measures;

[0209] Step S43: Submit the list of potential mobile network churn users, the list of potential churn reasons, and the list of remedial measures to the marketers so that the marketers can implement the remedial measures based on the list of potential mobile network churn users and the list of potential churn reasons.

[0210] Specifically, in step S43, for the above-mentioned list of mobile network churn users, analyze the important characteristic information of the users, determine the top 3 characteristics in terms of importance ranking based on the Gini index, generate a list of potential churn reasons and remedial measures for users according to the top 3 characteristic information, and submit the list of potential mobile network churn users and the list of potential churn reasons and remedial measures to the marketers for retention measures.

[0211] As a more specific implementation manner, before step S1, there is also step S0;

[0212] Step S0: Construct a prediction model for potential mobile network churn users;

[0213] Step S0 specifically includes the following steps:

[0214] Step S01: Obtain a data set within a second time period; the data set within the second time period includes historical user characteristics and corresponding historical variables; the historical variables include mobile network churn users and non-mobile network churn users;

[0215] Step S02: Divide the data set within the second time period into a training data set and a test data set;

[0216] Step S03: Use a stacking classifier to combine a first-stage classifier and a second-stage classifier, and perform cross-validation on the training data set to evaluate the performance of different parameter combinations to obtain a preliminary prediction model; the first-stage classifiers are LightGBM, XGBoost, and CatBoost, and the second-stage classifier is a logistic regression classifier;

[0217] Step S04: Optimize the performance of the preliminary prediction model through the test data set to obtain a prediction model for potential mobile network churn users.

[0218] Embodiment 2:

[0219] As shown in Figure 2 and Figure 3 this embodiment provides a prediction device for potential users with mobile network churn. The device includes:

[0220] An acquisition unit 10, configured to collect target user feature data within a first time period.

[0221] As a specific implementation manner, the acquisition unit 10 includes:

[0222] A first acquisition module, configured to acquire an original communication record information table within a first time period;

[0223] An aggregation and screening module, connected to the acquisition module, configured to perform data aggregation and data screening on the original communication record information table to obtain target user feature data within a first time period.

[0224] A prediction unit 20, connected to the acquisition unit 10, configured to input the target user feature data into a potential mobile network churn user prediction model for prediction to obtain a predicted value of the churn probability of the target user;

[0225] Wherein, the potential mobile network churn user prediction model is based on historical user feature data within a second time period, and is trained and algorithm parameter optimized by using a cross-validation stacking classifier; the second time period is greater than the first time period;

[0226] A determination unit 30, connected to the prediction unit 20, configured to determine whether the target user is a potential mobile network churn user according to the predicted value of the churn probability of the target user;

[0227] Wherein, when the predicted value of the churn probability of the target user is greater than the preset churn probability value, it is determined that the target user is a potential mobile network churn user; when the predicted value of the churn probability of the target user is less than or equal to the preset churn probability value, it is determined that the target user is marked as a non-potential mobile network churn user.

[0228] As a specific implementation manner, the aggregation and screening module includes:

[0229] An aggregation sub-module, configured to perform category aggregation according to the user dimension of the original communication record information table to extract aggregated feature data;

[0230] Wherein, the category aggregation according to the user dimension of the original communication record information table includes category aggregation according to the user portrait dimension, category aggregation according to the user package dimension, category aggregation according to the user usage record dimension, and category aggregation according to the user complaint dimension;

[0231] The first screening sub-module, connected to the aggregation sub-module, is used to perform feature selection on the aggregated feature data to obtain the retained feature data;

[0232] The feature selection includes:

[0233] Deleting the aggregated feature data with a missing value ratio greater than the missing preset value; and / or,

[0234] Deleting the aggregated feature data with a category ratio greater than the category ratio preset value; and / or,

[0235] Deleting the aggregated feature data with a feature importance value less than the feature importance preset value;

[0236] The second screening sub-module, connected to the first screening sub-module, is used to perform data preprocessing on the retained feature data to obtain the preprocessed feature data;

[0237] The data preprocessing includes:

[0238] After determining that the retained feature data is categorical feature data, performing median filling on the retained feature data and then performing label encoding; and / or,

[0239] After determining that the retained feature data is numerical feature data, performing linear interpolation on the retained feature data and then performing normalization processing; and / or,

[0240] After determining that the retained feature data is neither categorical feature data nor numerical feature data, performing piecewise discretization processing on the retained feature data;

[0241] The generation sub-module, connected to the second screening sub-module, is used to perform feature generation on the preprocessed feature data to obtain the target user feature data within the first time period;

[0242] Among them, the feature generation includes time feature generation, ratio feature generation, and frequency feature generation;

[0243] The time feature generation includes:

[0244] Generating the user network access year feature and the user network access month feature based on the user network access time, and generating the modification year feature and the modification month feature based on the modification time;

[0245] The ratio feature generation includes:

[0246] Generating the user network access duration ratio feature based on the user network access duration, and generating the monthly rent cost ratio feature based on the monthly rent cost feature;

[0247] The frequency feature generation includes: generating the frequency feature based on the label class feature.

[0248] As a specific implementation manner, the device further includes a construction unit, which is connected to the prediction unit 20 and is used to construct a prediction model for potential mobile network churn users, so that the prediction unit 20 inputs the target user feature data into the prediction model for potential mobile network churn users for prediction;

[0249] The construction unit includes:

[0250] A second acquisition module, which is used to acquire a data set within a second time period; the data set within the second time period includes historical user features and corresponding historical variables; the historical variables include mobile network churn users and non-mobile network churn users;

[0251] A division module, which is connected to the second acquisition module and is used to divide the data set within the second time period into a training data set and a test data set;

[0252] A combined verification module, which is connected to the division module and is used to combine a first-stage classifier and a second-stage classifier using a stacking classifier, and perform cross-validation on the training data set to evaluate the performance of different parameter combinations, so as to obtain a preliminary prediction model; the first-stage classifier is LightGBM, XGBoost, and CatBoost, and the second-stage classifier is a logistic regression classifier;

[0253] A tuning module, which is respectively connected to the division module and the combined verification module and is used to tune the performance of the preliminary prediction model through the test data set to obtain a prediction model for potential mobile network churn users.

[0254] The operation process of the entire device is as follows:

[0255] 1. Data collection: Collect 13 information tables for the previous 12 months, such as user basic information, customer real-name derived information, user integral information, single-user traffic information, single-user voice information, single-user usage information, user bill information, new churn user information, integrated user information, activity-derived information, stored value record information, bad debt system arrears information, and user complaint record information;

[0256] 2. User feature extraction: Aggregate the 13 information tables according to the user dimension, and extract 133 user features in categories such as user portrait, user package, user usage record, and user complaint;

[0257] 3. User feature selection: Perform feature selection on the above 133 user features, delete features with a missing value ratio greater than 95%, delete categorical features with a category ratio greater than 98%, use a decision tree for feature importance analysis, and delete features with a feature importance value less than 10. After feature selection, 100 user features are finally retained;

[0258] 4. User Feature Preprocessing: Preprocess the above 100 user features. For categorical features, fill in the median and then perform label encoding. For numerical features, perform linear interpolation and then normalization. Segment and discretize features such as customer age, customer scale, and fees.

[0259] 5. User Feature Generation: Generate features from the above 100 user features. Generate annual and monthly features from the user network access time and modification time. Generate ratio features from the user network access duration and fee features. Generate frequency features from the label features. After the feature generation step, 118 user features are obtained.

[0260] 6. Model Training: Train a model on the above 118 user features, construct a cross-validation stacked classifier. The classifier includes first-stage classifiers: LightGBM classifier, XGBoost classifier, CatBoost classifier, and a second-stage classifier, the Logistic Regression classifier. Input the 118 user features of the past 12 months into the cross-validation stacked classifier for training and algorithm parameter tuning to obtain a potential mobile network churn user prediction model.

[0261] 7. User Prediction: Input the 118 user features of the next 3 months into the potential mobile network churn user prediction model for prediction. Designate users with a prediction probability greater than 0.5 as potential mobile network churn users, and those less than 0.5 as non-potential mobile network churn users to generate a list of potential mobile network churn users.

[0262] 8. Root Cause Analysis: Analyze the important feature information of the above potential mobile network churn users, determine the top 3 features ranked by importance based on the Gini index, generate a list of potential churn reasons and remedial measures according to the top 3 feature information, and submit the list of potential mobile network churn users and the list of potential churn reasons and remedial measures to the marketing staff for retention measures.

[0263] Note:

[0264] 1. In the user feature extraction phase of Example 1 and Example 2, other quantities of user features in categories such as user portrait, user package, user usage record, and user complaint are extracted. It is acceptable to have 133, more than 133, or less than 133 user features.

[0265] 2. In the user feature selection phase of Example 1 and Example 2, different thresholds are selected for feature selection. For example, select features with a missing value ratio greater than 90% to be deleted, and delete categorical features with a category ratio greater than 90%.

[0266] 3. In Example 1 and Example 2, during the feature preprocessing stage, other normalization methods are selected to process numerical features, such as the standardization method, and different label processing methods are used to process categorical labels, such as one-hot encoding.

[0267] Among them, one-hot encoding is usually used to encode categorical variables in the prediction of potential mobile network churn users so that they can be used in machine learning models. The following are the specific steps for predicting potential mobile network churn users using one-hot encoding:

[0268] Data preparation: First, a dataset of potential mobile network churn users needs to be prepared, including various features (such as gender, age, consumption habits, service usage, etc.) and the target variable (i.e., whether to churn).

[0269] Determine the features that need to be one-hot encoded: In the dataset, find the categorical variables that need to be one-hot encoded, such as gender, user type, etc.

[0270] One-hot encoding: For each categorical variable that needs to be one-hot encoded, perform one-hot encoding on it. The specific steps are as follows: a. For each categorical variable, determine all possible values of the variable. For example, the gender variable may have two values: "male" and "female". b. For each possible value, create a new binary variable to indicate whether the value exists. For example, for the gender variable, two new variables, "gender_male" and "gender_female", can be created, represented by 1 and 0 respectively. c. Replace the original categorical variable with these new binary variables.

[0271] Dataset integration: Integrate the features processed by one-hot encoding with other features in the original dataset to form a new dataset.

[0272] Dataset division: Divide the integrated dataset into a training set and a test set, usually using cross-validation or the hold-out method for division.

[0273] Model training and prediction: Use a machine learning model (such as logistic regression, decision tree, random forest, etc.) to train the training set and make predictions on the test set to obtain the prediction results of potential mobile network churn users.

[0274] 4. In Example 1 and Example 2, for the model training stage, a separate classifier is used for modeling, such as a decision tree classifier, a logistic regression classifier, an xgboost classifier, etc.

[0275] It is understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present invention, but the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also regarded as the protection scope of the present invention.

Claims

1. A method for predicting potential users who have lost due to network migration, characterized in that: The method comprises the following steps: Step S1: Acquire target user characteristic data within a first time period; Step S2: inputting the target user's characteristic data into a potential mobile network churn user prediction model to perform prediction, and obtaining a predicted value of the target user's churn probability; The potential network loss user prediction model is based on historical user feature data in a second time period, and is obtained by training and algorithm parameter tuning using a cross-validation stacked classifier; the second time period is greater than the first time period; Step S3: judging whether the target user is a potential mobile network churn user according to the predicted value of the churn probability of the target user; When the predicted churn probability value of the target user is greater than the preset churn probability value, the target user is determined to be a potential mobile network churn user; when the predicted churn probability value of the target user is less than or equal to the preset churn probability value, the target user is determined to be marked as a non-potential mobile network churn user.

2. The method for predicting potential lost users due to network migration according to claim 1, characterized in that: The step S1 specifically includes the following steps: Step S11: obtaining an original communication record information table within a first time period; Step S12: Aggregate and filter the original communication record information table to obtain target user feature data within the first time period.

3. The method for predicting potential users lost due to network migration according to claim 2, characterized in that: The step S12 specifically includes the following steps: Step A1: performing category aggregation according to the user dimension of the original communication record information table to extract aggregated feature data; The category aggregation according to the user dimension of the original communication record information table includes category aggregation according to the user portrait dimension, category aggregation according to the user package dimension, category aggregation according to the user usage record dimension, and category aggregation according to the user complaint dimension; Step A2: performing feature selection on the aggregated feature data to obtain retained feature data; The feature selection includes: Delete the aggregated feature data whose missing value ratio is greater than the missing value preset value; and / or, Deleting aggregated feature data whose category ratio value is greater than a preset category ratio value; and / or, Delete the aggregated feature data whose feature importance value is less than the preset feature importance value; Step A3: preprocessing the retained feature data to obtain preprocessed feature data; The data preprocessing includes: After determining that the retained feature data is categorical feature data, median padding is performed on the retained feature data, followed by label encoding; and / or, After determining that the retained feature data is numerical feature data, linear interpolation is performed on the retained feature data, followed by normalization processing; and / or, After determining that the retained feature data is non-categorical feature data or non-numerical feature data, performing piecewise discretization processing on the retained feature data; Step A4: performing feature generation on the preprocessed feature data to obtain target user feature data within a first time period; Wherein, the feature generation includes time feature generation, scale feature generation and frequency feature generation; The time feature generation includes: Generate user network access year characteristics and user network access month characteristics according to user network access time, and generate modification year characteristics and modification month characteristics according to modification time; The scale feature generation includes: Generate a user network access duration ratio feature based on the user network access duration, and generate a monthly rental fee ratio feature based on the monthly rental fee feature; The frequency feature generation includes: generating the frequency feature according to the tag-type feature.

4. The method for predicting potential lost users due to network migration according to claim 3 is characterized in that: In the step A2, the aggregated feature data whose feature importance is less than a preset feature importance value is deleted, which specifically includes the following steps: Step A21: Calculate the overall Gini coefficient of the aggregated feature data; Among them, Gini 总体 Represents the overall Gini coefficient, P i represents the proportion of the aggregate feature data of the i-th aggregate class in the total; n represents the total number of categories; Step A22: Calculate the subset Gini coefficient of each aggregated feature; Among them, Gini 子集 represents the subset Gini coefficient, P j represents the proportion of the jth data feature in the subset; m represents the data composition of the subset; Step A23: Calculate the weighted average Gini coefficient of the aggregated feature data; The calculation process of the weighted average Gini coefficient is to multiply the subset Gini coefficient of each aggregated feature by the corresponding proportion of each subset, and then sum them up; Step A24: Calculate the feature importance value; The feature importance value is obtained by subtracting the weighted average Gini coefficient from the overall Gini coefficient; V 重要性 =Guinea 总体 -Guinea 加权平均 in, V 重要性 Represents the feature importance value, Gini 加权平均 represents the weighted average Gini coefficient; Step A25: Compare the feature importance value with the feature importance preset value, and delete the aggregated feature data whose feature importance is less than the feature importance preset value.

5. The method for predicting potential lost users due to network migration according to claim 1, characterized in that: After step S3, step S4 is also included. Step S4: Analyze the root causes and take remedial measures for potential users who have lost due to network migration; The step S4 specifically comprises the following steps: Step S41: after determining that the target user is a potential mobile network lost user, analyzing the loss reasons of the potential mobile network lost user, and forming a list of potential mobile network lost users and a list of potential loss reasons; Step S42: taking expected measures to remedy the causes of loss of potential users due to network migration, and forming a list of remedial measures; Step S43: Submit the list of potential users lost due to network migration, the list of potential reasons for loss and the list of remedial measures to the marketing personnel, so that the marketing personnel can implement remedial measures according to the list of potential users lost due to network migration and the list of potential reasons for loss.

6. The method for predicting potential lost users after network migration according to any one of claims 1 to 5, characterized in that: Before step S1, step S0 is also included; Step S0: construct a prediction model for potential mobile network churn users; The step S0 specifically includes the following steps: Step S01: Acquire a data set within a second time period; the data set within the second time period includes historical user features and corresponding historical variables; the historical variables include users lost through migration and users lost through non-migration; Step S02: dividing the data set in the second time period into a training data set and a test data set; Step S03: using a stacked classifier to combine the one-stage classifier and the two-stage classifier, and performing cross-validation on the training data set to evaluate the performance of different parameter combinations to obtain a preliminary prediction model; the one-stage classifier is LightGBM, XGBoost and CatBoost, and the two-stage classifier is a logistic regression classifier; Step S04: using the test data set, optimizing the performance of the preliminary prediction model to obtain a prediction model for potential network migration churn users.

7. A device for predicting potential users who have lost due to network migration, characterized in that: The device comprises: An acquisition unit, used to collect characteristic data of target users within a first time period; A prediction unit connected to the acquisition unit, used to input the target user feature data into a potential mobile network churn user prediction model to perform prediction and obtain a churn probability prediction value of the target user; The potential network loss user prediction model is based on historical user feature data in a second time period, and is obtained by training and algorithm parameter tuning using a cross-validation stacked classifier; the second time period is greater than the first time period; A determination unit connected to the prediction unit, configured to determine whether the target user is a potential user who has been lost through migration according to the predicted value of the loss probability of the target user; When the predicted churn probability value of the target user is greater than the preset churn probability value, the target user is determined to be a potential mobile network churn user; when the predicted churn probability value of the target user is less than or equal to the preset churn probability value, the target user is determined to be marked as a non-potential mobile network churn user.

8. The device for predicting potential lost users due to network migration according to claim 7, characterized in that: The acquisition unit comprises: A first acquisition module, used to acquire an original communication record information table within a first time period; The aggregation and screening module is connected to the acquisition module and is used to aggregate and screen the original communication record information table to obtain the target user feature data within the first time period.

9. The device for predicting potential users lost due to network migration according to claim 8, characterized in that: The aggregation screening module includes: The aggregation submodule is used to perform category aggregation according to the user dimension of the original communication record information table and extract aggregated feature data; The category aggregation according to the user dimension of the original communication record information table includes category aggregation according to the user portrait dimension, category aggregation according to the user package dimension, category aggregation according to the user usage record dimension, and category aggregation according to the user complaint dimension; A first screening submodule, connected to the aggregation submodule, for performing feature selection on the aggregated feature data to obtain retained feature data; The feature selection includes: Delete the aggregated feature data whose missing value ratio is greater than the missing value preset value; and / or, Deleting aggregated feature data whose category ratio value is greater than a preset category ratio value; and / or, Delete the aggregated feature data whose feature importance value is less than the preset feature importance value; A second screening submodule, connected to the first screening submodule, is used to perform data preprocessing on the retained feature data to obtain preprocessed feature data; The data preprocessing includes: After determining that the retained feature data is categorical feature data, median padding is performed on the retained feature data, followed by label encoding; and / or, After determining that the retained feature data is numerical feature data, linear interpolation is performed on the retained feature data, followed by normalization processing; and / or, After determining that the retained feature data is non-categorical feature data or non-numerical feature data, performing piecewise discretization processing on the retained feature data; A generating submodule, connected to the second screening submodule, for generating features for the preprocessed feature data to obtain feature data of the target user within a first time period; Wherein, the feature generation includes time feature generation, scale feature generation and frequency feature generation; The time feature generation includes: Generate user network access year characteristics and user network access month characteristics according to user network access time, and generate modification year characteristics and modification month characteristics according to modification time; The scale feature generation includes: Generate a user network access duration ratio feature based on the user network access duration, and generate a monthly rental fee ratio feature based on the monthly rental fee feature; The frequency feature generation includes: generating the frequency feature according to the tag-type feature.

10. The device for predicting potential lost users after network migration according to any one of claims 7 to 9, characterized in that: The device further comprises a construction unit, which is connected to the prediction unit and is used to construct a prediction model for potential network loss users, so that the prediction unit inputs the target user feature data into the prediction model for potential network loss users for prediction; The building block comprises: A second acquisition module is used to acquire a data set within a second time period; the data set within the second time period includes historical user characteristics and corresponding historical variables; the historical variables include users lost through migration and users not lost through migration; a division module, connected to the second acquisition module, for dividing the data set in the second time period into a training data set and a test data set; A combined validation module, connected to the partitioning module, for combining the one-stage classifier and the two-stage classifier using a stacked classifier, and performing cross-validation on the training data set to evaluate the performance of different parameter combinations to obtain a preliminary prediction model; the one-stage classifier is LightGBM, XGBoost and CatBoost, and the two-stage classifier is a logistic regression classifier; The tuning module is connected to the partitioning module and the combination verification module respectively, and is used to tune the performance of the preliminary prediction model through the test data set to obtain a potential network migration churn user prediction model.