Household customer group prediction method, device, equipment and medium

CN116992992BActive Publication Date: 2026-09-08CHINA MOBILE GRP GUANGDONG CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211167566.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2026-09-08
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

[0005]本发明提供一种家庭客群预测方法、装置、设备及介质,用以解决现有技术中现有的家庭客群预测不够精准的缺陷,实现提升家庭客群预测的精准度

Benefits of technology

[0051]The present invention provides a method, apparatus, device, and medium for predicting family customer groups. On the one hand, by using basic attribute data as a basis, it enriches the indicator system by adding member performance data, business contribution data, family potential data, complaint behavior data, and credit data, thereby improving the accuracy of existing family customer group predictions. On the other hand, it first divides customers into nine major customer groups through value scoring and stability scoring, and then divides them into two-dimensional value-stability segments. For the four key customer groups, it constructs four major customer group prediction models: high-price high-stability, high-price low-stability, low-price high-stability, and low-price low-stability. These four customer group prediction models are then used for prediction, enriching the customer group prediction system and thus improving the accuracy of existing family customer group predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116992992B_ABST
    Figure CN116992992B_ABST
Patent Text Reader

Abstract

The application provides a household customer group prediction method, device, equipment and medium, comprising: inputting household customer group data into a household value evaluation model and a household stability evaluation model respectively to obtain household value evaluation values and household stability evaluation values; based on the household value evaluation values and the household stability evaluation values, performing value stability two-dimensional slicing on the household customer group data; respectively constructing a random forest classification model, a logistic regression classification model and an extreme gradient boosting tree classification model for high-price high-stability customer group data, high-price low-stability customer group data, low-price high-stability customer group data and low-price low-stability customer group data; after voting based on model prediction results, obtaining a high-price high-stability customer group prediction model, a high-price low-stability customer group prediction model, a low-price high-stability customer group prediction model and a low-price low-stability customer group prediction model. The application realizes improvement of the precision of household customer group prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data mining technology, and in particular to a method, apparatus, equipment and medium for predicting household customer groups. Background Technology

[0002] Limited by policy factors and user saturation, the individual market is highly competitive, while the household market is booming and gradually becoming a new arena for competition. With the decline of the demographic dividend for telecom operators and the gradual decrease in traditional tariff services, household services have become an important revenue driver. In addition to tapping into the value of individual and household customer data, effectively maintaining existing household customers and expanding the household customer base has become a crucial business for telecom operators.

[0003] To rapidly and efficiently support the expansion of the household customer base, a key focus is on accurately assessing and predicting the stability of the household customer group's value. Existing assessment and prediction models often rely on data such as user communication behavior, consumption behavior, and broadband traffic to construct a single assessment or prediction model. The principle behind these existing models is: analyzing the communication behavior of existing or historical users, inputting relevant data from the target user, training a classifier, and then classifying the data to obtain a single assessment or prediction model.

[0004] Existing assessment or prediction models based solely on user or household communication behavior have the following drawbacks: First, the models do not consider actual circumstances. In reality, household customer group prediction is related to various factors in actual circumstances, resulting in inaccurate prediction results. Second, the model classification algorithm is too simplistic. Using only a single classification algorithm does not provide sufficient reliability for prediction, leading to inaccurate prediction results for existing household customer groups. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and medium for predicting household customer groups, in order to address the shortcomings of existing household customer group prediction methods that are not accurate enough, thereby improving the accuracy of household customer group prediction.

[0006] This invention provides a method for predicting household customer groups, comprising:

[0007] The family customer group data is input into the family value assessment model and the family stability assessment model respectively to obtain the family value assessment value output by the family value assessment model and the family stability assessment value output by the family stability assessment model.

[0008] Based on the family value assessment value and the family stability assessment value, the family customer group data is divided into two-dimensional value stability segments to obtain high-price high-stability customer group data, medium-price high-stability customer group data, low-price high-stability customer group data, high-price medium-stability customer group data, medium-price medium-stability customer group data, low-price medium-stability customer group data, high-price low-stability customer group data, medium-price low-stability customer group data, and low-price low-stability customer group data.

[0009] Random forest classification model, logistic regression classification model, and extreme gradient boosting tree classification model are constructed for the high-price and high-stability customer group data, the high-price and low-stability customer group data, the low-price and high-stability customer group data, and the model prediction results of the three models for each customer group data are output respectively.

[0010] After voting on the prediction results of the three models corresponding to the data of each customer group, the following prediction models were obtained: high-price and high-stability customer group prediction model, high-price and low-stability customer group prediction model, low-price and high-stability customer group prediction model, and low-price and low-stability customer group prediction model.

[0011] The high-price, high-stability customer group prediction model is used to predict high-price, high-stability customer groups from the customer group data to be identified; the high-price, low-stability customer group prediction model is used to predict high-price, low-stability customer groups from the customer group data to be identified; the low-price, high-stability customer group prediction model is used to predict low-price, high-stability customer groups from the customer group data to be identified; and the low-price, low-stability customer group prediction model is used to predict low-price, low-stability customer groups from the customer group data to be identified.

[0012] The family value assessment model and the family stability assessment model are constructed based on the family customer data, which includes basic attribute data, member performance data, business contribution data, family potential data, complaint behavior data, and credit data.

[0013] According to a method for predicting household customer groups provided by the present invention, the household value assessment model includes a user influence assessment module, a community value assessment module, a first regional value assessment module, and a first comprehensive assessment module.

[0014] The family customer group data is input into the family value assessment model to obtain the family value assessment value output by the family value assessment model, including:

[0015] The family customer group data is input into the user influence evaluation module to obtain the influence assessment value output by the user influence evaluation.

[0016] The household customer group data is input into the community value assessment module to obtain the community value assessment value output by the community value assessment module;

[0017] The family customer group data is input into the first regional value assessment module to obtain the first regional value assessment value output by the first regional value assessment module;

[0018] The influence assessment value, the community value assessment value, and the first regional value assessment value are input into the first comprehensive assessment module to obtain the family value assessment value output by the first comprehensive assessment module.

[0019] According to a method for predicting household customer groups provided by the present invention, the household customer group data is input into the user influence evaluation module to obtain the influence assessment value output by the user influence evaluation, including:

[0020] Based on family customer group data, a directed graph is determined to represent the calling and called relationships among family users, and a directed graph is determined to represent the binary relation pairs based on the family relationships among family users.

[0021] Based on the directed graph of the calling and called relationships among the family users, user influence is calculated to obtain the first influence evaluation factor for each family user.

[0022] Based on the directed graph of binary relationship pairs among the family users, user influence is calculated to obtain the second influence evaluation factor.

[0023] Based on the first influence evaluation factor and the second influence evaluation factor, a weighted calculation is performed to obtain the influence evaluation index of household users within the household.

[0024] According to a method for predicting household customer groups provided by the present invention, the household customer group data is input into the community value assessment module to obtain the community value assessment value output by the community value assessment module, including:

[0025] Extract community attribute data from the family customer data, wherein the community attribute data includes the average housing price of the community, the average housing price of the area where the community is located, the traffic conditions of the community, the medical resources around the community, the school resources around the community, and the average ARPU value of the community customers.

[0026] Based on the community attribute data, a logistic regression algorithm is used to perform prediction calculations to obtain community value assessment indicators.

[0027] According to a method for predicting household customer groups provided by the present invention, the household stability assessment model includes an external marketing number identification module, a second regional value assessment module, and a second comprehensive assessment module.

[0028] The household customer group data is input into the household stability assessment model to obtain the household stability assessment value output by the model, including:

[0029] The household customer group data is input into the external marketing number identification module to obtain the external marketing number output by the external marketing number identification module. The external marketing number is used to determine the external marketing evaluation value.

[0030] The family customer group data is input into the second regional value assessment module to obtain the second regional value assessment value output by the second regional value assessment module;

[0031] The cross-network marketing assessment value and the second regional value assessment value are input into the second comprehensive assessment module to obtain the family stability assessment value output by the second comprehensive assessment module.

[0032] According to a method for predicting household customer groups provided by the present invention, household customer group data is input into the external marketing number identification module to obtain the external marketing number output by the external marketing number identification module, including:

[0033] Based on the aforementioned family customer data, numbers from other networks that have made calls to numbers from the same network that are related to the anomaly factor are extracted as the first suspected numbers from other networks.

[0034] Based on the aforementioned family customer data, the cross-network numbers that have recently been contacted by ported-out numbers are identified as the second suspected cross-network numbers.

[0035] Based on the percentage of calls made between each of the second suspected cross-network numbers and the ported-out number, the third suspected cross-network number was determined.

[0036] Based on the first suspected out-of-network number and the third suspected out-of-network number, the out-of-network marketing number is determined.

[0037] According to a method for predicting household customer groups provided by the present invention, determining the household customer group data includes:

[0038] Obtain data on the target family customer groups;

[0039] Based on the family customer group data to be screened, calculate the chi-square statistic between each non-negative feature and label of the family customer group data to be screened;

[0040] Based on the chi-square statistics, the features are ranked from high to low, and the top K highest-scoring family customer data are selected as the family customer data.

[0041] The present invention also provides a household customer group prediction device, comprising:

[0042] The evaluation module is used to input family customer group data into the family value evaluation model and the family stability evaluation model respectively, and obtain the family value evaluation value output by the family value evaluation model and the family stability evaluation value output by the family stability evaluation model.

[0043] The slicing and grouping module is used to perform two-dimensional slicing and grouping of the family customer data based on the family value assessment value and the family stability assessment value, to obtain high-price high-stability customer group data, medium-price high-stability customer group data, low-price high-stability customer group data, high-price medium-stability customer group data, medium-price medium-stability customer group data, low-price medium-stability customer group data, high-price low-stability customer group data, medium-price low-stability customer group data, and low-price low-stability customer group data;

[0044] The model building module is used to build random forest classification models, logistic regression classification models, and extreme gradient boosting tree classification models for the high-price and high-stability customer group data, the high-price and low-stability customer group data, the low-price and high-stability customer group data, and the low-price and low-stability customer group data, respectively, and output the model prediction results of the three models for each customer group data.

[0045] The model voting module is used to vote on the prediction results of the three models corresponding to each customer group data to obtain the prediction model for high-price and high-stability customer groups, the prediction model for high-price and low-stability customer groups, the prediction model for low-price and high-stability customer groups, and the prediction model for low-price and low-stability customer groups.

[0046] The high-price, high-stability customer group prediction model is used to predict high-price, high-stability customer groups from the customer group data to be identified; the high-price, low-stability customer group prediction model is used to predict high-price, low-stability customer groups from the customer group data to be identified; the low-price, high-stability customer group prediction model is used to predict low-price, high-stability customer groups from the customer group data to be identified; and the low-price, low-stability customer group prediction model is used to predict low-price, low-stability customer groups from the customer group data to be identified.

[0047] The family value assessment model and the family stability assessment model are constructed based on the family customer data, which includes basic attribute data, member performance data, business contribution data, family potential data, complaint behavior data, and credit data.

[0048] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the household customer prediction method as described above.

[0049] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the household customer group prediction method as described above.

[0050] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the household customer group prediction method as described above.

[0051] The present invention provides a method, apparatus, device, and medium for predicting family customer groups. On the one hand, by using basic attribute data as a basis, it enriches the indicator system by adding member performance data, business contribution data, family potential data, complaint behavior data, and credit data, thereby improving the accuracy of existing family customer group predictions. On the other hand, it first divides customers into nine major customer groups through value scoring and stability scoring, and then divides them into two-dimensional value-stability segments. For the four key customer groups, it constructs four major customer group prediction models: high-price high-stability, high-price low-stability, low-price high-stability, and low-price low-stability. These four customer group prediction models are then used for prediction, enriching the customer group prediction system and thus improving the accuracy of existing family customer group predictions. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0053] Figure 1 This is one of the flowcharts illustrating the household customer group prediction method provided by the present invention;

[0054] Figure 2 This is the second flowchart illustrating the household customer group prediction method provided by the present invention;

[0055] Figure 3 This is the third flowchart illustrating the household customer group prediction method provided by the present invention;

[0056] Figure 4 This is the fourth flowchart illustrating the household customer group prediction method provided by the present invention;

[0057] Figure 5 This is the fifth flowchart illustrating the household customer group prediction method provided by the present invention;

[0058] Figure 6 This is the sixth flowchart illustrating the household customer group prediction method provided by the present invention;

[0059] Figure 7 This is the seventh flowchart of the household customer group prediction method provided by the present invention;

[0060] Figure 8 This is the eighth flowchart of the household customer group prediction method provided by the present invention;

[0061] Figure 9 This is the ninth flowchart of the household customer group prediction method provided by the present invention;

[0062] Figure 10 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0064] The following is combined with Figures 1-9 The present invention describes a method for predicting household customer groups.

[0065] Please refer to Figure 1 The household customer group prediction method provided by this invention includes:

[0066] Step 10: Input the family customer group data into the family value assessment model and the family stability assessment model respectively to obtain the family value assessment value output by the family value assessment model and the family stability assessment value output by the family stability assessment model.

[0067] Step 20: Based on the family value assessment value and the family stability assessment value, the family customer group data is divided into two-dimensional value stability segments to obtain high-price high-stability customer group data, medium-price high-stability customer group data, low-price high-stability customer group data, high-price medium-stability customer group data, medium-price medium-stability customer group data, low-price medium-stability customer group data, high-price low-stability customer group data, medium-price low-stability customer group data, and low-price low-stability customer group data.

[0068] Step 30: Construct random forest classification model, logistic regression classification model, and extreme gradient boosting tree classification model for the high-price and high-stability customer group data, the high-price and low-stability customer group data, the low-price and high-stability customer group data, and the model prediction results of the three models for each customer group data are output respectively.

[0069] Step 40: After voting on the prediction results of the three models corresponding to each customer group data, the following prediction models are obtained: high-price high-stability customer group prediction model, high-price low-stability customer group prediction model, low-price high-stability customer group prediction model, and low-price low-stability customer group prediction model.

[0070] The high-price, high-stability customer group prediction model is used to predict high-price, high-stability customer groups from the customer group data to be identified; the high-price, low-stability customer group prediction model is used to predict high-price, low-stability customer groups from the customer group data to be identified; the low-price, high-stability customer group prediction model is used to predict low-price, high-stability customer groups from the customer group data to be identified; and the low-price, low-stability customer group prediction model is used to predict low-price, low-stability customer groups from the customer group data to be identified.

[0071] The family value assessment model and the family stability assessment model are constructed based on the family customer data, which includes basic attribute data, member performance data, business contribution data, family potential data, complaint behavior data, and credit data.

[0072] This solution constructs a family value assessment model and a family stability assessment model for family customers. The overall process is as follows: First, customer basic attribute data, member performance data, business contribution data, family potential data, complaint behavior data, and credit data are selected to construct the family value assessment model and the family stability assessment model. Then, through two-dimensional slicing of value stability, nine customer groups are obtained: high-price high-stability, medium-price high-stability, low-price high-stability, high-price medium-stability, medium-price medium-stability, low-price medium-stability, high-price low-stability, medium-price low-stability, and low-price low-stability. Finally, based on the data of the four key customer groups (high-price high-stability, high-price low-stability, low-price high-stability, and low-price low-stability), random forest, logistic regression, and XGBoost classification models are constructed for each predicted customer group, and the prediction results of the three models are output respectively. The three results are then voted on to obtain the final prediction model. Thus, prediction models for the four customer groups of high-price high-stability, high-price low-stability, low-price high-stability, and low-price low-stability are constructed. The overall flowchart is shown below. Figure 2 As shown.

[0073] Two-dimensional slices were made on the family value assessment model and the family stability assessment model to generate a value-stability two-dimensional slice grouping wide table, which divided the customer groups into nine categories: high price and high stability, high price and low stability, high price and medium stability, medium price and high stability, medium price and medium stability, medium price and low stability, low price and high stability, low price and medium stability, and low price and low stability.

[0074] The data from four key customer groups is used to extract positive and negative samples for training the prediction models. These samples are used to construct the prediction models for each customer group. Relevant characteristic indicators extracted from the data of these four key customer groups include: whether upgrades were applied for in the past 6 months, average ARPU in the past 3 months, average number of calls to customer service from other networks in the past 3 months, average number of out-of-plan data packages purchased in the past 3 months, average number of tariff-related complaints in the past 6 months, average number of calls to other networks in the past 3 months, month-on-month change in the number of calls to other networks, main plan amount, whether the customer lives in a high-value residential area, and whether the customer is a key member of the household. For example, taking the high-price, high-stability prediction model as an example, samples with a value-stability rating of "high-price, high-stability" are taken as positive samples, while samples with other value-stability ratings are taken as negative samples. The same principle applies to samples from other prediction models.

[0075] The classifier for the family customer group prediction model in this solution is composed of random forest, logistic regression, and extreme gradient boosting tree.

[0076] Random forest is a classifier that uses multiple decision trees. The final output class is determined by the mode of the outputs from these independent decision trees. The advantage of random forest is that it avoids the overfitting problem that can occur with a single decision tree.

[0077] Decision trees are generally generated from top to bottom. Each decision or event (i.e., state of nature) can lead to two or more events, resulting in different outcomes. Visualizing these decision branches resembles the branches of a tree, hence the name "decision tree." The number of child nodes at each node in a decision tree depends on the algorithm used. For example, a decision tree obtained using the CART (Classification and Regression Tree) algorithm has two branches at each node; this type of tree is called a binary tree. Trees that allow nodes to have more than two child nodes are called multi-branch trees. Commonly used binary tree algorithms include CART and ID3, while multi-branch tree algorithms include C5.0 and CHAID.

[0078] Currently popular binary tree algorithms include ID3 and CART, with the branching method determined by the hyperparameter criterion. ID3's information gain metric has two drawbacks: 1) It prioritizes features with more attribute values, which may not be reasonable and can easily lead to overfitting; 2) In ID3, data is split based on attribute values, and these features no longer have an effect, so this rapid splitting method can affect the algorithm's accuracy.

[0079] Therefore, the CART algorithm is chosen instead of ID3 here. Compared to ID3, CART has a wider range of applications, and can be used for both classification and regression. Moreover, CART's use of features is repeatable.

[0080] The CART algorithm consists of the following two steps:

[0081] 1) Decision tree generation: Generate a decision tree based on the training dataset. The generated decision tree should be as large as possible.

[0082] 2) Decision tree pruning: Use the validation dataset to prune the generated tree and select the optimal subtree. In this case, the minimum loss function is used as the pruning criterion.

[0083] In CART classification, the best data splitting feature is selected based on the criterion of minimizing the Gini coefficient. Gini describes purity, similar in meaning to information entropy. Each iteration of CART reduces the Gini coefficient.

[0084] The formula for calculating the Gini coefficient is as follows:

[0085]

[0086]

[0087] The CART generation algorithm stops computation when the number of samples in a node is less than a predetermined threshold, or the Gini coefficient of the sample set is less than a predetermined threshold (the samples are basically of the same class), or there are no more features.

[0088] CART Decision Tree Generation Algorithm Flow: ① Based on the training dataset, starting from the root node, recursively perform the following operations on each node to construct a binary decision tree: ② Let the training dataset of the node be D. Calculate the Gini coefficient of the existing features for that dataset. Then, for each feature A, for each possible value a, divide D into two parts, D1 and D2, based on whether the test for A = a is "yes" or "no" for the sample point, and calculate the Gini coefficient when A = a. ③ Among all possible features A and all their possible split points a, select the feature with the smallest Gini coefficient and its corresponding split point as the optimal feature and optimal split point. Based on the optimal feature and optimal split point, generate two child nodes from the current node, and distribute the training dataset to the two child nodes according to the features. ④ Recursively call steps ② to ③ on the two child nodes until the stopping condition is met. ⑤ Generate the CART decision tree.

[0089] Random forests use sampling with replacement to extract up to m-1 random subsets (m being the number of numbers). Each subset is used to train an independent decision tree. The samples to be identified are then input into these decision trees, and the mode of the identification results is output as the final result.

[0090] Principal component analysis (PCA) only considers decomposing the independent variable matrix to eliminate irrelevant information. However, different classification objectives have different characteristic and confounding information; therefore, the relationship between independent and dependent variables should be considered during the decomposition of the independent variable matrix.

[0091] Partial Least Squares Logistic Regression (PLS-logistic) is a classification algorithm based on the above ideas. This method combines the concepts of logistic regression, principal component analysis (PCA), and canonical correlation analysis (OCC). Before establishing a standard logistic regression model, it decomposes both the independent variable X and the dependent variable Y, simultaneously extracting components (usually called factors) from both variables X and Y to maximize the correlation between the components extracted from X and Y.

[0092] XGBoost is a type of boosting algorithm. The idea behind boosting algorithms is to integrate many weak classifiers together to form a strong classifier. Because XGBoost is a boosting tree model, it integrates many tree models to form a very strong classifier. The tree model used is the CART regression tree model.

[0093] CART regression trees assume the tree is a binary tree and continuously split the features. For example, if the current tree node is split based on the j-th feature value, samples with feature values ​​less than s are assigned to the left subtree, and samples with feature values ​​greater than s are assigned to the right subtree.

[0094] R1(j, s) = {x|x j ≤s}and R2(j,s)={x|x j >s} The CART regression tree essentially partitions the sample space along this feature dimension. Optimizing this spatial partitioning is an NP-hard problem; therefore, heuristic methods are used to solve it in decision tree models. The objective function generated by a typical CART regression tree is:

[0095]

[0096] Therefore, finding the optimal splitting features and the optimal splitting points is transformed into solving for the following objective function:

[0097]

[0098] Therefore, by traversing all the split points of all features, the optimal splitting features and split points can be found. This ultimately yields a regression tree.

[0099] The XGBoost objective function consists of two parts: the first part measures the difference between the predicted score and the true score, and the second part is the regularization term.

[0100]

[0101]

[0102] It can be seen that XGBoost adds L1 and L2 regularization terms to the error function, where the loss function can be either flat loss or logistic loss. The benefit of adding regularization terms is to prevent overfitting, which is reflected in two aspects: first, pre-pruning, because the regularization term limits the number of leaf nodes; second, the coefficient of the L2 modulus score of the leaf score in the regularization term smooths the leaf score. The newly generated tree is to fit the residuals of the previous prediction; that is, after generating trees, the predicted score can be written as...

[0103]

[0104]

[0105]

[0106]

[0107]

[0108] The objective function can be written as:

[0109]

[0110] Find an f t If the objective function can be minimized, it can be approximated as:

[0111]

[0112] If the tree structure is determined, to minimize the objective function, we can set its derivative to 0, and then solve for the optimal prediction score for each leaf node:

[0113]

[0114] Substituting into the objective function, the minimum loss is obtained as:

[0115]

[0116] How much can be reduced at most from the target? This can be called the structure score. This is a function that scores tree structures in a more general way, similar to the Gini coefficient.

[0117] After calculating the three classification results, the final prediction result is determined by voting. If two or more of the three models identify a positive sample, the final output is a positive sample; the same applies to negative samples. The criteria can also be adjusted according to actual needs; for example, only one positive classification result is required to identify a sample as positive.

[0118] The final output includes prediction models for four customer groups: high-price and high-stability, high-price and low-stability, low-price and high-stability, and low-price and low-stability. Taking the high-price and high-stability prediction model as an example, the process is as follows: Figure 3 As shown.

[0119] The present invention provides a method, apparatus, device, and medium for predicting family customer groups. On the one hand, by using basic attribute data as a basis, it enriches the indicator system by adding member performance data, business contribution data, family potential data, complaint behavior data, and credit data, thereby improving the accuracy of existing family customer group predictions. On the other hand, it first divides customers into nine major customer groups through value scoring and stability scoring, and then divides them into two-dimensional value-stability segments. For the four key customer groups, it constructs four major customer group prediction models: high-price high-stability, high-price low-stability, low-price high-stability, and low-price low-stability. These four customer group prediction models are then used for prediction, enriching the customer group prediction system and thus improving the accuracy of existing family customer group predictions.

[0120] In one possible embodiment, please refer to Figure 4 The family value assessment model includes a user influence assessment module, a community value assessment module, a first regional value assessment module, and a first comprehensive assessment module.

[0121] Step 10: Input the family customer group data into the family value assessment model to obtain the family value assessment value output by the family value assessment model, including:

[0122] Step 101: Input the family customer group data into the user influence evaluation module to obtain the influence assessment value output by the user influence evaluation.

[0123] Step 102: Input the family customer group data into the community value assessment module to obtain the community value assessment value output by the community value assessment module;

[0124] Step 103: Input the family customer group data into the first regional value assessment module to obtain the first regional value assessment value output by the first regional value assessment module;

[0125] Step 104: Input the influence assessment value, the community value assessment value, and the first regional value assessment value into the first comprehensive assessment module to obtain the family value assessment value output by the first comprehensive assessment module.

[0126] Family value assessment models can also be built based on user social attribute data, community attribute data, and peer interaction data.

[0127] In this embodiment, a family value assessment model is constructed based on user social attribute data, community attribute data, and peer interaction data, by introducing the identification of key family members and high-value communities. The introduction of the identification of key family members and high-value communities in the family value assessment improves the accuracy of family value assessment.

[0128] In one possible embodiment, please refer to Figure 5 Step 101: Input the family customer group data into the user influence evaluation module to obtain the influence assessment value output by the user influence evaluation, including:

[0129] Step 1011: Based on the family customer group data, determine the directed graph of the calling and called relationships between family users and the directed graph of the binary relation pairs based on the family relationships between family users.

[0130] Step 1012: Based on the directed graph of the calling and called relationships among the family users, calculate the user influence to obtain the first influence evaluation factor for each family user.

[0131] Step 1013: Based on the directed graph of binary relationship pairs of family relationships among the family users, calculate the user influence to obtain the second influence evaluation factor.

[0132] Step 1014: Based on the first influence evaluation factor and the second influence evaluation factor, perform a weighted calculation to obtain the influence evaluation index of the household user within the household.

[0133] Key family member identification uses the PageRank algorithm to identify key family members by inputting family customer attribute data. The analysis of each family member's spending habits, user level, and business transactions is prioritized to identify key family members.

[0134] By using the PageRank algorithm, which combines call behavior among family members with the probability of binary relationships, key individuals within a family can be identified. First, it's important to note that the more frequently a person is contacted within a family, the greater their influence within the family; conversely, the more influential a person is contacted by a family member, the greater their influence within the family.

[0135] The user influence evaluation module is described in detail below:

[0136] 1) From the perspective of calls between family users, the PageRank algorithm is used to calculate user influence on a directed graph based on the caller-caller relationship to obtain an influence evaluation factor.

[0137] 2) From the perspective of the probability of family relationships among users, the PageRank algorithm is used again to calculate user influence on the directed graph based on binary relationship pairs to obtain the influence evaluation factor.

[0138] 3) Taking into account the call behavior and binary relationship probability among members, select appropriate weights and perform weighted calculations on the two influence evaluation factors obtained above to obtain the final user's influence evaluation index rank within the family.

[0139] The expression for this user influence evaluation model is:

[0140] rank = d × rank a +(1-d)×rank b

[0141] This indicates the weight of the influence evaluation factor obtained through inter-member communication relationships in the final comprehensive user influence evaluation index, that is, the degree of influence of inter-family user communication relationships on the final user influence analysis ranking.

[0142] In this embodiment, by identifying key family members, a first influence evaluation factor and a second influence evaluation factor are determined, and then weighted calculations are performed to obtain the influence evaluation index of family users within the family. This adds the influence evaluation factor determined by identifying key family members, further improving the accuracy of family customer group prediction.

[0143] In one possible embodiment, please refer to Figure 6 Step 102: Input the household customer group data into the community value assessment module to obtain the community value assessment value output by the community value assessment module, including:

[0144] Step 1021: Extract the community attribute data from the family customer group data, wherein the community attribute data includes the average housing price of the community, the average housing price of the area where the community is located, the traffic conditions of the community, the medical resources around the community, the school resources around the community, and the average ARPU value of the community customers.

[0145] Step 1022: Based on the community attribute data, use the logistic regression algorithm to perform prediction calculations to obtain the community value assessment index.

[0146] High-value neighborhood identification involves extracting neighborhood attribute data, including the neighborhood's average house price, the average house price in the area where the neighborhood is located, whether the neighborhood has convenient transportation, whether there are supporting hospitals nearby, whether the neighborhood is in a good school district, and the average ARPU (Average Revenue Per Resident) of the neighborhood. Then, a logistic regression algorithm is used for prediction and calculation to ultimately identify high-value neighborhoods.

[0147] Positive and negative samples are extracted for high-value community identification. Data labeled as high-value communities are designated as positive samples, while data with other labels are designated as negative samples. The value contribution of each feature to the high-value community identification prediction is calculated using feature correlation and IV value. This value is then used to filter features, removing those irrelevant to the prediction.

[0148] Logistic regression is used to predict the identification of high-value neighborhoods. Logistic regression is widely used in discriminative models due to its simple structure and easily interpretable coefficients in practical applications. The dependent variable of the extracted positive and negative samples is labeled with 1 and 0 respectively, and all indicators filtered using IV values ​​are entered into the logistic regression model.

[0149] The probability of each neighborhood being classified as a high-value neighborhood can be represented by P. Therefore, the logistic regression model can be expressed as:

[0150]

[0151] Where x i (i = 1, 2, ..., s) is the index. Since P takes values ​​between 0 and 1, and its range can be transformed to any real value after the logit transformation, the solution is to find β = (β0, β1, ..., β...). s ) T The formula for model training and solving is:

[0152]

[0153] In logistic regression model prediction, no particular emphasis is placed on any of the variables entering the model; the model training and solution process is performed in a specific manner. During the process, to ensure that signaling data metrics contribute a high weight to the model, a penalty term is considered.

[0154] After adding the penalty term, in the solution... During the process of solving for the optimal solution of the objective function, it will be done through The relationship constraint between each non-internet content indicator coefficient and the internet content information indicator coefficient makes the internet content data indicators... coefficient It must be greater than the coefficient of other indicators, and This is the penalty coefficient, which is usually a constant.

[0155] In summary, based on the above, we have an adaptive logistic regression model β=(β0,β1,...,β) s ) T The estimator is defined as:

[0156]

[0157] Solve for β = (β0, β1, ..., β) using positive and negative sample data. s ) T Then, an adaptive logistic regression model for assessing whether a community is high-value is obtained. The final model expression is:

[0158]

[0159] In this embodiment, on the one hand, logistic regression algorithm is used to predict the identification of high-value residential communities. Since logistic regression is widely used in discriminative models, its structure is simple, and the role of its coefficients is easily interpreted in business contexts, thus improving the efficiency of constructing predictions for high-value residential communities. On the other hand, community attribute data is extracted, including average housing price, average housing price in the area, whether the community has convenient transportation, whether there are nearby hospitals, whether the community is in a good school district, and the average ARPU (Average Revenue Per User) of the community's residents. Logistic regression algorithm is then used for prediction calculations to ultimately identify high-value residential communities. This adds an assessment of the corresponding high-value communities and improves the accuracy of predictions for family customer groups.

[0160] In one possible embodiment, please refer to Figure 7 The family stability assessment model includes an external marketing number identification module, a second regional value assessment module, and a second comprehensive assessment module.

[0161] Step 10: Input the household customer group data into the household stability assessment model to obtain the household stability assessment value output by the model, including:

[0162] Step 111: Input the family customer group data into the external marketing number identification module to obtain the external marketing number output by the external marketing number identification module. The external marketing number is used to determine the external marketing evaluation value.

[0163] Step 112: Input the family customer group data into the second regional value assessment module to obtain the second regional value assessment value output by the second regional value assessment module;

[0164] Step 113: Input the cross-network marketing assessment value and the second regional value assessment value into the second comprehensive assessment module to obtain the family stability assessment value output by the second comprehensive assessment module.

[0165] Furthermore, due to significant differences in economic levels, user behavior, and household business development across different provinces and cities, the sample distribution varies considerably. Therefore, this study categorizes users within each province based on their local economic development to assess the second regional value, thereby improving the model's predictive accuracy. For example, users can be divided into four main categories: Guangzhou / Shenzhen, Dongguan / Foshan, second-tier cities, and third-tier cities, to evaluate the second regional value.

[0166] In this embodiment, by adding identification of marketing numbers from other networks, regional value assessment, and decision-making; and by improving the accuracy of the family stability assessment value, the accuracy of family customer group prediction is improved.

[0167] In one possible embodiment, please refer to Figure 8 Step 111: Input the household customer group data into the external marketing number identification module to obtain the external marketing number output by the external marketing number identification module, including:

[0168] Step 1111: Based on the family customer group data, extract the numbers from other networks that have made calls to numbers from the same network that are related to the anomaly factor as the first suspected numbers from other networks;

[0169] Step 1112: Based on the family customer group data, obtain the cross-network numbers of the recently contacted ported numbers as the second suspected cross-network numbers;

[0170] Step 1113: Based on the percentage of calls made by each of the second suspected cross-network numbers to the ported-out number, determine the third suspected cross-network number;

[0171] Step 1114: Based on the first suspected out-of-network number and the third suspected out-of-network number, determine the out-of-network marketing number.

[0172] Among them, the numbers related to the abnormal factors include: 1) numbers that have used other apps; 2) numbers that are suspected of using other broadband services; 3) numbers that have recently left the network; 4) numbers that have recently switched networks; and 5) numbers that have recently inquired about or applied for number portability.

[0173] We selected numbers from other networks that had made calls to the above 5 types of numbers from our network as the first suspected numbers from other networks. In other words, we used numbers from other networks that had made calls to the above 5 types of numbers from our network as a basis for judging marketing numbers from other networks.

[0174] First, select the cross-network numbers that have recently contacted the ported-out number as the second suspected cross-network numbers. Then, based on the percentage of calls made by each of the second suspected cross-network numbers to the ported-out number, determine the third suspected cross-network number from among the second suspected cross-network numbers. The percentage of calls made mainly includes: 1) the percentage of calls made to cross-network numbers from T1 to T3 months and from T4 to T6 months (to verify whether the ported-out number has been contacting cross-network numbers more frequently in the past 3 months); 2) the coefficient of variation (to describe whether there is an increasing trend in the number of calls made to cross-network numbers from T-6 to T-1 months); 3) (the number of times the ported-out number called the cross-network number from T1 to T6 months) / (the number of times the ported-out number was called by the cross-network number from T1 to T6 months); 4) the ratio of the number of calls made to the cross-network number during working hours (9-11 am and 13-17 pm) to the total number of calls; 5) the percentage of calls made to cross-network numbers lasting 1-10 minutes.

[0175] Based on the cross-network numbers that have made calls to numbers associated with the local network number related to the anomaly factor, and the cross-network numbers that have recently been contacted and ported out, numbers that appear simultaneously in both methods are identified as cross-network marketing numbers. This improves the accuracy of cross-network marketing numbers and further enhances the accuracy of household customer group prediction.

[0176] In one possible embodiment, step 10, determining the household customer group data, includes:

[0177] Step 121: Obtain the data of the families to be screened;

[0178] Step 122: Based on the family customer group data to be screened, calculate the chi-square statistic between each non-negative feature and label of the family customer group data to be screened;

[0179] Step 123: Rank the features from high to low according to the chi-square statistic, and select the top K highest scores of the family customer group data to be screened as the family customer group data.

[0180] As the scale of the indicator database grows rapidly, the number of indicator data corresponding to each number also increases, including many irrelevant and distracting indicators. To avoid affecting model progress and efficiency, indicator screening is essential. Filtering methods are typically used as a preprocessing step, and feature selection is completely independent of any machine learning algorithm. It selects features based on scores in various statistical tests and various indicators of correlation.

[0181] First, features are selected based on their variance. For example, if a feature has a very small variance, it means that the samples are essentially indistinguishable on that feature; most values ​​of the feature may be the same, or even the entire feature may have the same value. In this case, the feature is of little use in distinguishing the samples. Next, correlation analysis is used to calculate the correlation between the indicators. We have three commonly used methods to assess the correlation between features and labels: chi-square test, F-test, and mutual information.

[0182] Chi-square filtering is specifically designed for relevance filtering in discrete label problems (i.e., classification problems). The chi-square test calculates the chi-square statistic between each non-negative feature and the label, ranking the features from highest to lowest chi-square statistic. By selecting the class of the top K features with the highest scores, we can remove features that are most likely independent of the label and irrelevant to the classification objective. Additionally, if the chi-square test detects that all values ​​in a feature are identical, it will suggest using variance filtering first.

[0183] The F-test, also known as ANOVA (homogeneity of variance test), is a filtering method used to capture the linear relationship between each feature and the label. It can be used for both regression and classification, with two types: F-test for classification and F-test for regression. The F-test for classification is used for data with discrete labels, while the F-test for regression is used for data with continuous labels. It's important to note that the F-test is very stable when the data follows a normal distribution; therefore, if using the F-test for filtering, we first transform the data to conform to a normal distribution. The essence of the F-test is to find a linear relationship between two sets of data, with the null hypothesis being that "there is no significant linear relationship between the data." It returns two statistics: the F-value and the p-value. Similar to chi-square filtering, we want to select features with p-values ​​less than 0.05 or 0.01, as these features are significantly linearly correlated with the label, while features with p-values ​​greater than 0.05 or 0.01 are considered to have no significant linear relationship with the label and should be deleted.

[0184] Mutual information is a filtering method used to capture any relationship (including linear and non-linear relationships) between each feature and the label. Similar to the F-test, it can be used for both regression and classification, and includes two classes: mutual information classification and mutual information regression. The usage and parameters for these two classes are exactly the same as the F-test; however, mutual information is more powerful than the F-test, which can only find linear relationships, while mutual information can find arbitrary relationships.

[0185] In this embodiment, chi-square filtering can be used to specifically filter the relevance of discrete labels (i.e., classification problems) to select family customer data used to build family value assessment models, family stability assessment models, and four major customer group prediction models, thereby improving the accuracy of customer group prediction.

[0186] Furthermore, the screening of family customer data also includes:

[0187] As the scale of the indicator database grows rapidly, the indicator data corresponding to each number also increases, including many irrelevant and interfering indicators. In order to avoid affecting the progress and efficiency of the model, it is necessary to screen the indicators of the household customer group data.

[0188] First, correlation analysis is used to calculate the correlation between indicators. Then, combined with the IV (indicator value) method, indicators with high IV values ​​among highly correlated indicators are retained. This avoids multicollinearity and also filters out irrelevant indicators. The main process for indicator selection for household customer data is as follows: Figure 9 As shown.

[0189] Correlation analysis is a statistical method used to study whether and to what extent features are correlated. To make the algorithm applicable to both categorical and continuous variables, the Spearman rank correlation coefficient method is used in this study. The correlation coefficient ranges from 1 to -1. 1 indicates a perfect positive correlation between the two variables, -1 indicates a perfect negative correlation, and 0 indicates no correlation between the two variables.

[0190] For a sample of size n, the n original data points are transformed into rank data: Two random variables X and Y are sorted (either in ascending or descending order) to obtain two rank sets x and y, where element x... i y i X i Ranking in X and Y i Ranked in Y.

[0191] Subtracting the corresponding elements from sets x and y yields a ranking difference set d, where each element is d. i =x i -y i The Spearman rank correlation coefficient ρ between random variables X and Y is 1≤i≤n.

[0192]

[0193] According to formula (1), the correlation coefficients of each feature are calculated, and then the selection is carried out in combination with subsequent methods.

[0194] IV value method (Information Value) IV (Indicator Value) is a method for evaluating the predictive ability of an indicator. The higher the IV value, the stronger the predictive ability of the indicator.

[0195] ① First, perform chi-square binning on the continuous feature data; for categorical features, IV can be calculated directly.

[0196] ② To calculate the IV, the bin features need to be encoded using WOE (Weight of Evidence). For the data after binning in the i-th group, the WOE calculation formula is as follows:

[0197]

[0198] Where, p y1 This represents the proportion of negative samples in this bin to all negative samples.

[0199] Where, p y0 This represents the proportion of positive samples in this bin to all positive samples.

[0200] Among them, B i B represents the number of negative samples in the bin, where B is the total number of negative samples.

[0201] Among them, G i G represents the number of negative samples in this bin, and G is the total number of negative samples.

[0202] ③ Then calculate the PCT (Percentage) for the i-th group, which is the difference between the proportions of positive and negative samples:

[0203] PCT i =p y1 -p y0 (3)

[0204] ④ Finally, based on (2) and (3), we obtain (4), and calculate the IV value (5) for this group:

[0205] IV i =WOE i ×PCT i (4)

[0206]

[0207] After filtering out highly relevant indicators, the remaining indicators are sorted in descending order based on their IV values. Indicators with IV values ​​greater than 0.5 that have a significant effect on the model are selected for model training, while indicators with weak predictive ability are removed.

[0208] The household customer prediction device provided by the present invention is described below. The household customer prediction device described below can be referred to in correspondence with the household customer prediction method described above.

[0209] The household customer prediction device provided by the present invention includes:

[0210] The evaluation module is used to input family customer group data into the family value evaluation model and the family stability evaluation model respectively, and obtain the family value evaluation value output by the family value evaluation model and the family stability evaluation value output by the family stability evaluation model.

[0211] The slicing and grouping module is used to perform two-dimensional slicing and grouping of the family customer data based on the family value assessment value and the family stability assessment value, to obtain high-price high-stability customer group data, medium-price high-stability customer group data, low-price high-stability customer group data, high-price medium-stability customer group data, medium-price medium-stability customer group data, low-price medium-stability customer group data, high-price low-stability customer group data, medium-price low-stability customer group data, and low-price low-stability customer group data;

[0212] The model building module is used to build random forest classification models, logistic regression classification models, and extreme gradient boosting tree classification models for the high-price and high-stability customer group data, the high-price and low-stability customer group data, the low-price and high-stability customer group data, and the low-price and low-stability customer group data, respectively, and output the model prediction results of the three models for each customer group data.

[0213] The model voting module is used to vote on the prediction results of the three models corresponding to each customer group data to obtain the prediction model for high-price and high-stability customer groups, the prediction model for high-price and low-stability customer groups, the prediction model for low-price and high-stability customer groups, and the prediction model for low-price and low-stability customer groups.

[0214] The high-price, high-stability customer group prediction model is used to predict high-price, high-stability customer groups from the customer group data to be identified; the high-price, low-stability customer group prediction model is used to predict high-price, low-stability customer groups from the customer group data to be identified; the low-price, high-stability customer group prediction model is used to predict low-price, high-stability customer groups from the customer group data to be identified; and the low-price, low-stability customer group prediction model is used to predict low-price, low-stability customer groups from the customer group data to be identified.

[0215] The family value assessment model and the family stability assessment model are constructed based on the family customer data, which includes basic attribute data, member performance data, business contribution data, family potential data, complaint behavior data, and credit data.

[0216] Furthermore, the family value assessment model includes a user influence evaluation module, a community value assessment module, a first regional value assessment module, and a first comprehensive assessment module;

[0217] The evaluation module is also used for:

[0218] The family customer group data is input into the user influence evaluation module to obtain the influence assessment value output by the user influence evaluation.

[0219] The household customer group data is input into the community value assessment module to obtain the community value assessment value output by the community value assessment module;

[0220] The family customer group data is input into the first regional value assessment module to obtain the first regional value assessment value output by the first regional value assessment module;

[0221] The influence assessment value, the community value assessment value, and the first regional value assessment value are input into the first comprehensive assessment module to obtain the family value assessment value output by the first comprehensive assessment module.

[0222] The evaluation module is also used for:

[0223] Based on family customer group data, a directed graph is determined to represent the calling and called relationships among family users, and a directed graph is determined to represent the binary relation pairs based on the family relationships among family users.

[0224] Based on the directed graph of the calling and called relationships among the family users, user influence is calculated to obtain the first influence evaluation factor for each family user.

[0225] Based on the directed graph of binary relationship pairs among the family users, user influence is calculated to obtain the second influence evaluation factor.

[0226] Based on the first influence evaluation factor and the second influence evaluation factor, a weighted calculation is performed to obtain the influence evaluation index of household users within the household.

[0227] The evaluation module is also used for:

[0228] Extract community attribute data from the family customer data, wherein the community attribute data includes the average housing price of the community, the average housing price of the area where the community is located, the traffic conditions of the community, the medical resources around the community, the school resources around the community, and the average ARPU value of the community customers.

[0229] Based on the community attribute data, a logistic regression algorithm is used to perform prediction calculations to obtain community value assessment indicators.

[0230] Furthermore, the family stability assessment model includes an external marketing number identification module, a second regional value assessment module, and a second comprehensive assessment module;

[0231] The evaluation module is also used for:

[0232] The household customer group data is input into the external marketing number identification module to obtain the external marketing number output by the external marketing number identification module. The external marketing number is used to determine the external marketing evaluation value.

[0233] The family customer group data is input into the second regional value assessment module to obtain the second regional value assessment value output by the second regional value assessment module;

[0234] The cross-network marketing assessment value and the second regional value assessment value are input into the second comprehensive assessment module to obtain the family stability assessment value output by the second comprehensive assessment module.

[0235] The evaluation module is also used for:

[0236] Based on the aforementioned family customer data, numbers from other networks that have made calls to numbers from the same network that are related to the anomaly factor are extracted as the first suspected numbers from other networks.

[0237] Based on the aforementioned family customer data, the cross-network numbers that have recently been contacted by ported-out numbers are identified as the second suspected cross-network numbers.

[0238] Based on the percentage of calls made between each of the second suspected cross-network numbers and the ported-out number, the third suspected cross-network number was determined.

[0239] Based on the first suspected out-of-network number and the third suspected out-of-network number, the out-of-network marketing number is determined.

[0240] Furthermore, the household customer group prediction device also includes a household customer group data filtering module, used for:

[0241] Obtain data on the target family customer groups;

[0242] Based on the family customer group data to be screened, calculate the chi-square statistic between each non-negative feature and label of the family customer group data to be screened;

[0243] Based on the chi-square statistics, the features are ranked from high to low, and the top K highest-scoring family customer data are selected as the family customer data.

[0244] Figure 10 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 10As shown, the electronic device may include: a processor 1010, a communications interface 1020, a memory 1030, and a communications bus 1040, wherein the processor 1010, the communications interface 1020, and the memory 1030 communicate with each other through the communications bus 1040. The processor 1010 can call logic instructions in the memory 1030 to execute a household customer group prediction method. This method includes: inputting household customer group data into a household value assessment model and a household stability assessment model respectively, obtaining a household value assessment value output by the household value assessment model and a household stability assessment value output by the household stability assessment model; based on the household value assessment value and the household stability assessment value, performing two-dimensional value stability slicing of the household customer group data to obtain high-price high-stability customer group data, medium-price high-stability customer group data, low-price high-stability customer group data, high-price medium-stability customer group data, medium-price medium-stability customer group data, low-price medium-stability customer group data, high-price low-stability customer group data, medium-price low-stability customer group data, and low-price low-stability customer group data; and constructing a random forest classification model, a logistic regression classification model, and an extreme ladder classification model for the high-price high-stability customer group data, the high-price low-stability customer group data, the low-price high-stability customer group data, and the low-price low-stability customer group data, respectively. The degree of improvement tree classification model is used to output the model prediction results of the three models corresponding to each customer group data. After voting on the model prediction results of the three models corresponding to each customer group data, the following prediction models are obtained: high-price high-stability customer group prediction model, high-price low-stability customer group prediction model, low-price high-stability customer group prediction model, and low-price low-stability customer group prediction model. The high-price high-stability customer group prediction model is used to predict high-price high-stability customer groups based on the customer group data to be identified; the high-price low-stability customer group prediction model is used to predict high-price low-stability customer groups based on the customer group data to be identified; the low-price high-stability customer group prediction model is used to predict low-price high-stability customer groups based on the customer group data to be identified; and the low-price low-stability customer group prediction model is used to predict low-price low-stability customer groups based on the customer group data to be identified. The family value assessment model and the family stability assessment model are constructed based on the family customer group data, which includes basic attribute data, member performance data, business contribution data, family potential data, complaint behavior data, and credit data.

[0245] Furthermore, the logical instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0246] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the household customer group prediction method provided by the above methods. The method includes: inputting household customer group data into a household value assessment model and a household stability assessment model respectively to obtain a household value assessment value output by the household value assessment model and a household stability assessment value output by the household stability assessment model; based on the household value assessment value and the household stability assessment value, performing two-dimensional value stability slicing and grouping of the household customer group data to obtain high-price high-stability customer group data, medium-price high-stability customer group data, low-price high-stability customer group data, high-price medium-stability customer group data, medium-price medium-stability customer group data, low-price medium-stability customer group data, high-price low-stability customer group data, medium-price low-stability customer group data, and low-price low-stability customer group data; and processing the high-price high-stability customer group data, the high-price low-stability customer group data, the low-price high-stability customer group data, and the low-price low-stability customer group data respectively. For low-stability customer group data, random forest classification model, logistic regression classification model, and extreme gradient boosting tree classification model are constructed, and the model prediction results of the three models corresponding to each customer group data are output respectively. After voting on the model prediction results of the three models corresponding to each customer group data, high-price high-stability customer group prediction model, high-price low-stability customer group prediction model, low-price high-stability customer group prediction model, and low-price low-stability customer group prediction model are obtained. Among them, the high-price high-stability customer group prediction model is used to predict high-price high-stability customer group data, the high-price low-stability customer group prediction model is used to predict high-price low-stability customer group data, the low-price high-stability customer group prediction model is used to predict low-price high-stability customer group data, and the low-price low-stability customer group prediction model is used to predict low-price low-stability customer group data. The family value assessment model and the family stability assessment model are constructed based on the family customer group data, which includes basic attribute data, member performance data, business contribution data, family potential data, complaint behavior data, and credit data.

[0247] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the household customer group prediction method provided by the above methods. This method includes: inputting household customer group data into a household value assessment model and a household stability assessment model respectively, to obtain a household value assessment value output by the household value assessment model and a household stability assessment value output by the household stability assessment model; based on the household value assessment value and the household stability assessment value, performing two-dimensional value stability slicing of the household customer group data to obtain high-price high-stability customer group data, medium-price high-stability customer group data, low-price high-stability customer group data, high-price medium-stability customer group data, medium-price medium-stability customer group data, low-price medium-stability customer group data, high-price low-stability customer group data, medium-price low-stability customer group data, and low-price low-stability customer group data; and constructing random forests for the high-price high-stability customer group data, the high-price low-stability customer group data, the low-price high-stability customer group data, and the low-price low-stability customer group data, respectively. The system employs three classification models: classification model, logistic regression classification model, and extreme gradient boosting tree classification model. It outputs the prediction results of each model for each customer group data. After voting on the prediction results of the three models for each customer group data, it obtains four prediction models: high-price high-stability customer group, high-price low-stability customer group, low-price high-stability customer group, and low-price low-stability customer group. Specifically, the high-price high-stability customer group prediction model is used to predict high-price high-stability customer groups based on the customer group data to be identified; the high-price low-stability customer group prediction model is used to predict high-price low-stability customer groups based on the customer group data to be identified; the low-price high-stability customer group prediction model is used to predict low-price high-stability customer groups based on the customer group data to be identified; and the low-price low-stability customer group prediction model is used to predict low-price low-stability customer groups based on the customer group data to be identified. The family value assessment model and the family stability assessment model are constructed based on the family customer group data, which includes basic attribute data, member performance data, business contribution data, family potential data, complaint behavior data, and credit data.

[0248] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0249] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0250] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting household customer groups, characterized in that, include: The family customer group data is input into the family value assessment model and the family stability assessment model respectively to obtain the family value assessment value output by the family value assessment model and the family stability assessment value output by the family stability assessment model. Based on the family value assessment value and the family stability assessment value, the family customer group data is divided into two-dimensional value stability segments to obtain high-price high-stability customer group data, medium-price high-stability customer group data, low-price high-stability customer group data, high-price medium-stability customer group data, medium-price medium-stability customer group data, low-price medium-stability customer group data, high-price low-stability customer group data, medium-price low-stability customer group data, and low-price low-stability customer group data. Random forest classification model, logistic regression classification model, and extreme gradient boosting tree classification model are constructed for the high-price and high-stability customer group data, the high-price and low-stability customer group data, the low-price and high-stability customer group data, and the model prediction results of the three models for each customer group data are output respectively. After voting on the prediction results of the three models corresponding to the data of each customer group, the following prediction models were obtained: high-price and high-stability customer group prediction model, high-price and low-stability customer group prediction model, low-price and high-stability customer group prediction model, and low-price and low-stability customer group prediction model. The high-price, high-stability customer group prediction model is used to predict high-price, high-stability customer groups from the customer group data to be identified; the high-price, low-stability customer group prediction model is used to predict high-price, low-stability customer groups from the customer group data to be identified; the low-price, high-stability customer group prediction model is used to predict low-price, high-stability customer groups from the customer group data to be identified; and the low-price, low-stability customer group prediction model is used to predict low-price, low-stability customer groups from the customer group data to be identified. The family value assessment model and the family stability assessment model are constructed based on the family customer data, which includes basic attribute data, member performance data, business contribution data, family potential data, complaint behavior data, and credit data. The family stability assessment model includes an external marketing number identification module, a second regional value assessment module, and a second comprehensive assessment module. The household customer group data is input into the household stability assessment model to obtain the household stability assessment value output by the model, including: Based on the aforementioned family customer data, numbers from other networks that have made calls to numbers from the same network that are related to the anomaly factor are extracted as the first suspected numbers from other networks. Based on the aforementioned family customer data, the cross-network numbers that have recently been contacted by ported-out numbers are identified as the second suspected cross-network numbers. Based on the percentage of calls made between each of the second suspected cross-network numbers and the ported-out number, the third suspected cross-network number was determined. Based on the first suspected out-of-network number and the third suspected out-of-network number, an out-of-network marketing number is determined, and the out-of-network marketing number is used to determine the out-of-network marketing evaluation value. The family customer group data is input into the second regional value assessment module to obtain the second regional value assessment value output by the second regional value assessment module; The cross-network marketing assessment value and the second regional value assessment value are input into the second comprehensive assessment module to obtain the family stability assessment value output by the second comprehensive assessment module.

2. The method for predicting household customer groups according to claim 1, characterized in that, The family value assessment model includes a user influence assessment module, a community value assessment module, a first regional value assessment module, and a first comprehensive assessment module. The family customer group data is input into the family value assessment model to obtain the family value assessment value output by the family value assessment model, including: The family customer group data is input into the user influence evaluation module to obtain the influence assessment value output by the user influence evaluation. The household customer group data is input into the community value assessment module to obtain the community value assessment value output by the community value assessment module; The family customer group data is input into the first regional value assessment module to obtain the first regional value assessment value output by the first regional value assessment module; The influence assessment value, the community value assessment value, and the first regional value assessment value are input into the first comprehensive assessment module to obtain the family value assessment value output by the first comprehensive assessment module.

3. The method for predicting household customer groups according to claim 2, characterized in that, The family customer group data is input into the user influence evaluation module to obtain the influence assessment value output by the user influence evaluation, including: Based on family customer group data, a directed graph is determined to represent the calling and called relationships among family users, and a directed graph is determined to represent the binary relation pairs based on the family relationships among family users. Based on the directed graph of the calling and called relationships among the family users, user influence is calculated to obtain the first influence evaluation factor for each family user. Based on the directed graph of binary relationship pairs among the family users, user influence is calculated to obtain the second influence evaluation factor. Based on the first influence evaluation factor and the second influence evaluation factor, a weighted calculation is performed to obtain the influence evaluation index of household users within the household.

4. The method for predicting household customer groups according to claim 2, characterized in that, The household customer group data is input into the community value assessment module to obtain the community value assessment value output by the community value assessment module, including: Extract community attribute data from the family customer data, wherein the community attribute data includes the average housing price of the community, the average housing price of the area where the community is located, the traffic conditions of the community, the medical resources around the community, the school resources around the community, and the average ARPU value of the community customers. Based on the community attribute data, a logistic regression algorithm is used to perform prediction calculations to obtain community value assessment indicators.

5. The method for predicting household customer groups according to claim 1, characterized in that, Determining the household customer group data includes: Obtain data on the target family customer groups; Based on the family customer group data to be screened, calculate the chi-square statistic between each non-negative feature and label of the family customer group data to be screened; Based on the chi-square statistics, the features are ranked from high to low, and the top K highest-scoring family customer data are selected as the family customer data.

6. A household customer group prediction device, characterized in that, include: The evaluation module is used to input family customer group data into the family value evaluation model and the family stability evaluation model respectively, and obtain the family value evaluation value output by the family value evaluation model and the family stability evaluation value output by the family stability evaluation model. The slicing and grouping module is used to perform two-dimensional slicing and grouping of the family customer data based on the family value assessment value and the family stability assessment value, to obtain high-price high-stability customer group data, medium-price high-stability customer group data, low-price high-stability customer group data, high-price medium-stability customer group data, medium-price medium-stability customer group data, low-price medium-stability customer group data, high-price low-stability customer group data, medium-price low-stability customer group data, and low-price low-stability customer group data; The model building module is used to build random forest classification models, logistic regression classification models, and extreme gradient boosting tree classification models for the high-price and high-stability customer group data, the high-price and low-stability customer group data, the low-price and high-stability customer group data, and the low-price and low-stability customer group data, respectively, and output the model prediction results of the three models for each customer group data. The model voting module is used to vote on the prediction results of the three models corresponding to each customer group data to obtain the prediction model for high-price and high-stability customer groups, the prediction model for high-price and low-stability customer groups, the prediction model for low-price and high-stability customer groups, and the prediction model for low-price and low-stability customer groups. The high-price, high-stability customer group prediction model is used to predict high-price, high-stability customer groups from the customer group data to be identified; the high-price, low-stability customer group prediction model is used to predict high-price, low-stability customer groups from the customer group data to be identified; the low-price, high-stability customer group prediction model is used to predict low-price, high-stability customer groups from the customer group data to be identified; and the low-price, low-stability customer group prediction model is used to predict low-price, low-stability customer groups from the customer group data to be identified. The family value assessment model and the family stability assessment model are constructed based on the family customer data, which includes basic attribute data, member performance data, business contribution data, family potential data, complaint behavior data, and credit data. The family stability assessment model includes an external marketing number identification module, a second regional value assessment module, and a second comprehensive assessment module; the assessment module is specifically used for: Based on the aforementioned family customer data, numbers from other networks that have made calls to numbers from the same network that are related to the anomaly factor are extracted as the first suspected numbers from other networks. Based on the aforementioned family customer data, the cross-network numbers that have recently been contacted by ported-out numbers are identified as the second suspected cross-network numbers. Based on the percentage of calls made between each of the second suspected cross-network numbers and the ported-out number, the third suspected cross-network number was determined. Based on the first suspected out-of-network number and the third suspected out-of-network number, an out-of-network marketing number is determined, and the out-of-network marketing number is used to determine the out-of-network marketing evaluation value. The family customer group data is input into the second regional value assessment module to obtain the second regional value assessment value output by the second regional value assessment module; The cross-network marketing assessment value and the second regional value assessment value are input into the second comprehensive assessment module to obtain the family stability assessment value output by the second comprehensive assessment module.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the household customer group prediction method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the household customer group prediction method as described in any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the household customer group prediction method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Model training method and device, electronic equipment and storage medium

    CN113554184A