Data analysis method and device

Through data analysis methods, combined with user's historical data and clustering models, the target impact factor of users is determined, which solves the problem of inaccurate prediction of user churn or value decline in the existing technology, and achieves more accurate analysis of the causes of user churn or value decline and timely marketing response.

CN113888226BActive Publication Date: 2025-05-13CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111210635.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-18
Publication Date
2025-05-13
Estimated Expiration
2041-10-18

AI Technical Summary

Technical Problem

The existing technology cannot accurately predict user churn or value decline, resulting in marketing lag, which is not conducive to operators maintaining user ownership and user value.

Method used

Through the data analysis method, after predicting that the user is the target user, the alternative influencing factor is determined based on the user's historical data, the user cluster is determined through the clustering model, the main influencing factor of the user cluster is calculated, and the target influencing factor is determined based on the user's alternative influencing factor.

Benefits of technology

It has achieved more accurate determination of the reasons for user churn or value decline, timely providing matching services to users, and reducing the user churn rate or value decline in the number of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113888226B_ABST
    Figure CN113888226B_ABST
Patent Text Reader

Abstract

The present application discloses a data analysis method and device, which belongs to the field of data processing technology. The data analysis method includes: when a user is predicted to be a target user, determining the user's alternative influencing factors according to the user's historical data; determining the user cluster according to a preset clustering model, the user's historical data and the user's alternative influencing factors; determining the main influencing factor of the user cluster according to the alternative influencing factors of the users belonging to the user cluster; determining the user's target influencing factor according to the main influencing factor of the user cluster and the user's alternative influencing factors. This method can more accurately determine the user's target influencing factor, and then accurately determine the reasons for user loss or value decline, so as to provide users with matching services in a timely manner and reduce the user loss rate or the number of users with declining value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data analysis method and device. Background Art

[0002] When a user who intends to leave the network or reduce the package fee (hereinafter referred to as "downgrading") comes to the business hall and makes a request to leave the network or downgrade, the operator system usually matches the candidate package for the user based on the user's recent consumption status, and the salesperson makes marketing recommendations to the user based on the content of the candidate package to reduce the user churn rate or achieve user value retention. In this technical solution, since it is impossible to accurately predict the reasons for user churn or value reduction, marketing will have a certain lag, which is not conducive to operators maintaining user retention and user value. Summary of the invention

[0003] To this end, the present application provides a data analysis method and device to solve the problem that the reasons for the inability to accurately predict user churn or value decline are not conducive to operators maintaining user retention and user value.

[0004] In order to achieve the above-mentioned purpose, the first aspect of the present application provides a data analysis method, which comprises:

[0005] In the case where the user is predicted to be a target user, determining an alternative influencing factor of the user according to the historical data of the user, wherein the alternative influencing factor is used to characterize an alternative reason affecting the user value;

[0006] Determine a user cluster of the user according to a preset clustering model, historical data of the user, and candidate influencing factors of the user;

[0007] Determining a main influencing factor of the user cluster according to candidate influencing factors of users belonging to the user cluster;

[0008] A target influence factor of the user is determined according to the main influence factor of the user cluster and the candidate influence factors of the user.

[0009] Furthermore, when the user is predicted to be a target user, determining the candidate influencing factors of the user according to the historical data of the user includes:

[0010] When the user is predicted to be a target user, analyzing the user from a preset dimension according to the historical data of the user to obtain an analysis result of the user in the preset dimension;

[0011] According to the analysis result, a candidate influencing factor of the user is determined.

[0012] Further, determining the main influencing factor of the user cluster according to the candidate influencing factors of the users belonging to the user cluster includes:

[0013] Calculating the repetition rate of the candidate impact factors of the users belonging to the user cluster;

[0014] The main influencing factor of the user cluster is determined according to the repetition rate of the candidate influencing factors and a preset repetition rate threshold.

[0015] Further, determining the target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user includes:

[0016] Determine the intersection of the main influencing factor of the user cluster and the candidate influencing factor of the user;

[0017] A target impact factor of the user is determined according to the intersection.

[0018] Furthermore, after determining the intersection of the main influencing factor of the user cluster and the candidate influencing factor of the user, the method further includes:

[0019] When the intersection is an empty set, re-clustering the users to obtain an updated user cluster;

[0020] Determining a main influencing factor of the updating user cluster according to candidate influencing factors of users belonging to the updating user cluster;

[0021] The target influence factor of the user is determined according to the main influence factor of the updated user cluster and the candidate influence factor of the user.

[0022] Furthermore, after determining the target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user, the method further includes:

[0023] Determine a service strategy that matches the target impact factor;

[0024] According to the service strategy, the user's product candidate information is updated.

[0025] Furthermore, in the case where the user is predicted to be a target user, before determining the candidate influencing factors of the user according to the historical data of the user, the method further includes:

[0026] A prediction result of the user is obtained according to the historical data of the user and a preset prediction model, wherein the prediction result is used to characterize whether the user is a target user.

[0027] Furthermore, obtaining the prediction result of the user according to the historical data of the user and a preset prediction model includes:

[0028] Determine a user profile of the user according to the historical data of the user;

[0029] The user portrait is input into the prediction model to obtain the prediction result of the user.

[0030] Furthermore, before obtaining the prediction result of the user according to the historical data of the user and the preset prediction model, the method further includes:

[0031] Constructing an initial prediction model, and training the initial prediction model to obtain an alternative prediction model;

[0032] Evaluating the candidate prediction model to obtain a model evaluation result;

[0033] The prediction model is selected from the candidate prediction models according to the model evaluation result.

[0034] In order to achieve the above-mentioned object, the second aspect of the present application provides a data analysis device, the data analysis device comprising:

[0035] An alternative influencing factor determination module is configured to determine an alternative influencing factor of the user according to the historical data of the user when predicting that the user is a target user, wherein the alternative influencing factor is used to characterize an alternative reason affecting the user value;

[0036] A clustering module, configured to determine a user cluster of the user according to a preset clustering model, historical data of the user and candidate influencing factors of the user;

[0037] A main influencing factor determination module, configured to determine a main influencing factor of the user cluster according to candidate influencing factors of users belonging to the user cluster;

[0038] The target influence factor determination module is configured to determine the target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user.

[0039] This application has the following advantages:

[0040] The data analysis method and device provided by the present application, when predicting that a user is a target user, determine the user's alternative influence factor according to the user's historical data; determine the user's user cluster according to a preset clustering model, the user's historical data and the user's alternative influence factor; determine the main influence factor of the user cluster according to the alternative influence factors of the users belonging to the user cluster; determine the user's target influence factor according to the main influence factor of the user cluster and the user's alternative influence factor. When predicting that there is a possibility of user churn or a downward trend in user value, the method first determines the user's alternative influence factor, and determines the user's target influence factor according to the user's alternative influence factor and the main influence factor of the user cluster to which the user belongs, so that the user's target influence factor can be determined more accurately, and then the reason for the user's churn or value decline can be accurately determined, so as to provide matching services to the user in a timely manner and reduce the user churn rate or the number of users with declining value. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings are used to provide further understanding of the present application and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present application, but do not constitute a limitation to the present application.

[0042] Figure 1 A flowchart of a data analysis method provided in one embodiment of the present application;

[0043] Figure 2 A flowchart of a data analysis method provided in yet another embodiment of the present application;

[0044] Figure 3 A flowchart of a data analysis method provided in yet another embodiment of the present application;

[0045] Figure 4 A flowchart of a data analysis method provided by another embodiment of the present application;

[0046] Figure 5 A block diagram of a data analysis device according to an embodiment of the present application;

[0047] Figure 6 A block diagram of a data analysis device provided in another embodiment of the present application;

[0048] Figure 7 A schematic diagram of a marketing system provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0049] The specific implementation of the present application is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present application, and is not used to limit the present application.

[0050] The disappearance of the demographic dividend has led to increasingly fierce competition among operators. How to promptly discover users who are willing to leave the network or downgrade their packages and maintain the value of these users is particularly important. In related technologies, operators usually use "business pop-ups" to maintain user value. Specifically, the operator system matches the user with packages and products of similar tiers based on the user's recent consumption information. When a user comes to the business hall and makes a request to leave the network or downgrade, the salesperson queries based on the user information, and the query results are displayed in the form of a pop-up window. The pop-up window content includes information on packages and products of similar tiers that match the user. The salesperson makes marketing recommendations to the user based on the pop-up window content to maintain user value. However, this method cannot accurately predict the reasons for user loss or value decline, resulting in a certain lag in marketing, which is not conducive to operators maintaining user value.

[0051] A first aspect of the present application provides a data analysis method. Figure 1 FIG. 1 is a flow chart of a data analysis method provided by an embodiment of the present application, and the data analysis method can be applied to a data analysis device. Figure 1 As shown, the data analysis method includes the following steps:

[0052] Step S101 : when a user is predicted to be a target user, determining candidate influencing factors of the user according to the user's historical data.

[0053] Among them, the target users are users whose user value has a downward trend. For example, users who are willing to leave the network and users who are willing to reduce package charges are target users. Alternative influencing factors are used to characterize alternative reasons that affect user value. In some embodiments, the user's alternative influencing factors include any one or more of user attribute influencing factors, activity influencing factors, consumption influencing factors, service saturation influencing factors, contract term influencing factors, and complaint influencing factors, wherein service saturation refers to the ratio between the amount of service actually used by a user and the amount of service pre-ordered by the user.

[0054] The user's historical data includes basic user data, user business data, and customer service data, etc. Among them, basic user data includes but is not limited to user identification, administrative division information, address information, gender, age, occupation, and work unit information; user business data includes but is not limited to business type, online time, channel type, contract type, billing income, business saturation, call information, user level, and business migration information; customer service data includes but is not limited to complaint reasons, complaint time, number of complaints, and evaluation information.

[0055] In some embodiments, when a user is predicted to be a target user, the user's historical data is analyzed from a preset dimension to obtain the analysis result of the user in the preset dimension; based on the analysis result, the user's candidate influencing factor is determined. The preset dimensions include user attributes, activity, consumption, service saturation, contract period, complaint frequency, complaint type, etc. This application does not limit the specific dimension categories of the preset dimensions.

[0056] For example, the preset dimensions include contract term and service saturation. In the historical data of user A, the contract term is valid from XX / XX / XXXX to YY / YY / YYY, the traffic service saturation is 30%, the voice service saturation is 85%, and it is assumed that YY / YY / YYY is only one month away from the current date.

[0057] Therefore, the analysis result of user A in the contract term dimension is that the user's contract term is about to expire; the analysis result of user A in the service saturation dimension is that the traffic service saturation is relatively low, and the voice service saturation is relatively high. Based on this, it can be determined that the candidate influencing factors of user A include the expiration of the contract term and low traffic service saturation.

[0058] Step S102: determining a user cluster of the user according to a preset clustering model, the user's historical data and the user's candidate influencing factors.

[0059] The clustering model is a model built based on a clustering algorithm, and the clustering algorithm includes but is not limited to a K-means algorithm, a hierarchical clustering algorithm, and an expectation-maximization algorithm (EM algorithm). In this embodiment, the clustering model is used to cluster users to form at least one user cluster, and the users in each user cluster have a high similarity.

[0060] In some embodiments, the user's historical data and the user's candidate influencing factors are input into a clustering model, and the clustering model performs data processing to obtain an output result, and the user cluster of the user can be determined based on the output result.

[0061] Step S103: determining a main influencing factor of the user cluster according to candidate influencing factors of users belonging to the user cluster.

[0062] Among them, the main influencing factor refers to the influencing factor that has the main impact on the user cluster.

[0063] In some embodiments, the repetition rate of the candidate influencing factors of the users belonging to the user cluster is calculated; the main influencing factor of the user cluster is determined according to the repetition rate of the candidate influencing factors and a preset repetition rate threshold. The preset repetition rate threshold can be flexibly set and adjusted according to experience, actual needs and statistical data, and this application does not limit this.

[0064] For example, the user cluster includes 5 users, namely user A, user B, user C, user D and user E; the preset repetition rate threshold is 50%. Among them, the alternative influence factors of user A are x1, x2 and x3; the alternative influence factor of user B is x2; the alternative influence factors of user C are x1 and x3; the alternative influence factors of user D are x1 and x4; the alternative influence factors of user E are x1 and x2.

[0065] It can be seen that in this user cluster, x1 appears 4 times, x2 appears 3 times, x3 appears 2 times, and x4 appears 1 time. Based on this, it can be determined through calculation that the repetition rate of x1 is 4 / 5, the repetition rate of x2 is 3 / 5, the repetition rate of x3 is 2 / 5, and the repetition rate of x4 is 1 / 5. Among them, the alternative influencing factors with a repetition rate greater than 50% include x1 and x2. Therefore, the main influencing factors of this user cluster are x1 and x2.

[0066] Step S104: determining the target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user.

[0067] Among them, the target impact factor is used to characterize the key reasons that affect user value.

[0068] In some embodiments, first, the intersection of the main influence factor of the user cluster and the candidate influence factor of the user is determined; second, the target influence factor of the user is determined according to the intersection.

[0069] For example, the main influencing factors of the user cluster include x1, x2 and x4, and the candidate influencing factors of the user include x1 and x4, then the intersection of the two is {x1, x4}. Therefore, it is determined that the target influencing factors of the user include x1 and x4.

[0070] For another example, the main influencing factors of the user cluster include x1, x2 and x4, and the candidate influencing factors of the user include x1 and x3, then the intersection of the two is {x1}. Therefore, the target influencing factor of the user is determined to be x1.

[0071] In this embodiment, when it is predicted that a user is a target user, the user's alternative influencing factors are determined based on the user's historical data; the user's user cluster is determined based on a preset clustering model, the user's historical data, and the user's alternative influencing factors; the main influencing factor of the user cluster is determined based on the alternative influencing factors of the users belonging to the user cluster; the user's target influencing factor is determined based on the main influencing factor of the user cluster and the user's alternative influencing factors. In the case where it is predicted that there is a possibility of user churn or a downward trend in user value, the method first determines the user's alternative influencing factors, and determines the user's target influencing factors based on the user's alternative influencing factors and the main influencing factors of the user cluster to which the user belongs, so that the user's target influencing factors can be determined more accurately, and then the reasons for the user's churn or value decline can be accurately determined, so as to provide matching services to the user in a timely manner and reduce the user churn rate or the number of users with declining value.

[0072] It should be noted that the data analysis method provided in the embodiment of the present application is applicable to single-user scenarios as well as multi-user scenarios.

[0073] In a single-user scenario, when clustering a certain user, it is also necessary to determine several (for example, 1,000) sample users, input the user's historical data and alternative influencing factors, as well as the historical data and alternative influencing factors of the sample users, into the clustering model to determine the user cluster of the user, and determine the target influencing factor of the user based on the above process.

[0074] In a multi-user scenario, if the number of users is large enough, the historical data and candidate influencing factors of these users are directly input into the clustering model to determine the user cluster of each user, and then the target influencing factor of each user is determined according to the above process. If the number of users is insufficient, it is still necessary to determine a number of sample users (for example, 500), input the historical data and candidate influencing factors of the above users, and the historical data and candidate influencing factors of the sample users into the clustering model, so as to determine the user cluster of each user, and then determine the target influencing factor of each user according to the above process.

[0075] Figure 2 FIG. 1 is a flow chart of a data analysis method provided by another embodiment of the present application, and the data analysis method can be applied to a data analysis device. Figure 2 As shown, the data analysis method includes the following steps:

[0076] Step S201 : when it is predicted that a user is a target user, determining candidate influencing factors of the user according to the user's historical data.

[0077] Step S202: determining a user cluster of the user according to a preset clustering model, the user's historical data and the user's candidate influencing factors.

[0078] Step S203: determining a main influencing factor of the user cluster according to candidate influencing factors of users belonging to the user cluster.

[0079] Step S204: determining the target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user.

[0080] Steps S201 to S204 in this embodiment are the same as steps S101 to S104 in an embodiment of the present application, and are not described again here.

[0081] Step S205: determine a service strategy that matches the target impact factor.

[0082] Among them, service strategies include the ways and methods of providing services to users.

[0083] In some embodiments, a correspondence is pre-established between the service strategy and the target impact factor. When the target impact factor of the user is determined, a service strategy matching the target impact factor can be determined based on the target impact factor and the correspondence.

[0084] In some other embodiments, a service strategy library is pre-established, which includes multiple service strategies. When the target impact factor of the user is determined, the reason for the value reduction of the user is further analyzed, and a service strategy matching the reason for the value reduction is selected from the service strategy library according to the reason for the value reduction.

[0085] Step S206: Update the user's product selection information according to the service policy.

[0086] Among them, alternative products are products that can be ordered and used by users.

[0087] In some embodiments, for each service strategy, a product matching it can be determined in advance, so that after the service strategy is determined, the user's candidate products are updated according to the products matching the service strategy to obtain the user's candidate product information.

[0088] For example, the user's target impact factors include the expiration of the contract period and low traffic service saturation. The service strategy that matches the target impact factor is determined to be to provide a package service with a contract period and reduce the traffic in the package. Therefore, the alternative product provided to the user is determined to be a contract package with less traffic, and the user's alternative product information is updated based on this.

[0089] For sales personnel, they can carry out personalized marketing to users in a timely manner based on the updated alternative product information, thereby maintaining user value and reducing user churn rate.

[0090] In this embodiment, when a user is predicted to be a target user, the user's alternative influence factor is determined based on the user's historical data; the user's user cluster is determined based on the preset clustering model, the user's historical data and the user's alternative influence factor; the main influence factor of the user cluster is determined based on the alternative influence factors of the users belonging to the user cluster; the user's target influence factor is determined based on the main influence factor of the user cluster and the user's alternative influence factor; a service strategy matching the target influence factor is determined; and the user's alternative product information is updated based on the service strategy. This method can more accurately determine the user's target influence factor, and then a more scientific and reasonable service strategy can be determined based on the target influence factor, so as to timely provide users with more suitable alternative products, thereby reducing the user churn rate or the number of users with reduced value.

[0091] Figure 3 FIG. 1 is a flow chart of a data analysis method provided by another embodiment of the present application, and the data analysis method can be applied to a data analysis device. Figure 3 As shown, the data analysis method includes the following steps:

[0092] Step S301 : when it is predicted that a user is a target user, determining candidate influencing factors of the user according to the user's historical data.

[0093] Step S302: determining a user cluster of the user according to a preset clustering model, the user's historical data and the user's candidate influencing factors.

[0094] Step S303: determining a main influencing factor of the user cluster according to candidate influencing factors of users belonging to the user cluster.

[0095] Steps S301 to S303 in this embodiment are the same as steps S101 to S103 in an embodiment of the present application, and are not described again here.

[0096] Step S304: determine the intersection of the main influencing factor of the user cluster and the candidate influencing factor of the user. When the intersection is an empty set, re-cluster the users to obtain an updated user cluster.

[0097] If the intersection is an empty set, it means that the main influencing factor of the user cluster does not have the same influencing factor as the user's candidate influencing factors. In this case, it is considered that the user cluster to which the user currently belongs is inaccurate, so the user is re-clustered to obtain an updated user cluster.

[0098] It should be noted that when re-clustering users, the original clustering model can be directly used for a new round of clustering, or some parameters in the original clustering model can be adjusted and a new round of clustering can be performed using the adjusted clustering model. A new clustering model can also be used or clustering can be performed based on a new clustering algorithm. This application does not limit this.

[0099] Step S305 : determining a main influencing factor of the updated user cluster according to candidate influencing factors of users belonging to the updated user cluster.

[0100] In some embodiments, the repetition rate of candidate influencing factors of users belonging to the updating user cluster is calculated; and the main influencing factor of the updating user cluster is determined according to the repetition rate of the candidate influencing factors and a preset repetition rate threshold.

[0101] Step S306: determining the target influence factor of the user according to the main influence factor of the updated user cluster and the candidate influence factor of the user.

[0102] In some embodiments, first, the intersection of the main influence factor of the updated user cluster and the candidate influence factor of the user is determined; second, the target influence factor of the user is determined according to the intersection.

[0103] It should be noted that if the intersection of the main influencing factor of the updated user cluster and the user's alternative influencing factor is still an empty set, clustering is required again until the intersection of the main influencing factor of the user cluster obtained by clustering and the user's alternative influencing factor is not an empty set. Clustering is stopped and the user's target influencing factor is determined based on the intersection.

[0104] It should also be noted that, in some embodiments, after determining the user's target influence factor through the above steps, it also includes: determining a service strategy that matches the target influence factor; updating the user's alternative product information according to the service strategy, so that sales personnel can carry out personalized marketing to the user in a timely manner based on the updated alternative product information, thereby maintaining user value and reducing user churn rate.

[0105] In this embodiment, when a user is predicted to be a target user, the user's alternative influencing factors are determined according to the user's historical data; the user's user cluster is determined according to a preset clustering model, the user's historical data and the user's alternative influencing factors; the main influencing factor of the user cluster is determined according to the alternative influencing factors of the users belonging to the user cluster; the intersection of the main influencing factor of the user cluster and the alternative influencing factors of the user is determined, and when the intersection is an empty set, the user is re-clustered to obtain an updated user cluster; the main influencing factor of the updated user cluster is determined according to the alternative influencing factors of the users belonging to the updated user cluster; the target influencing factor of the user is determined according to the main influencing factor of the updated user cluster and the alternative influencing factors of the user. In this method, when the intersection of the main influencing factor of the user cluster and the alternative influencing factors of the user is an empty set, the user is re-clustered to determine a more accurate user cluster for the user, thereby obtaining a target influencing factor with higher accuracy.

[0106] Figure 4FIG. 1 is a flow chart of a data analysis method provided by another embodiment of the present application, and the data analysis method can be applied to a data analysis device. Figure 4 As shown, the data analysis method includes the following steps:

[0107] Step S401, obtaining the user's prediction result according to the user's historical data and a preset prediction model.

[0108] The prediction result is used to characterize whether the user is a target user. The target user is a user whose user value is decreasing. For example, users who are willing to leave the network and users who are willing to reduce package charges are target users.

[0109] In some embodiments, a user profile of the user is determined based on the user's historical data; the user profile is input into a prediction model to obtain a prediction result for the user.

[0110] In some specific implementations, data with obvious anomalies in the user's historical data is removed, missing values ​​of the fields are filled, and part of the data is discretized to obtain data with relatively uniform format and relatively complete and accurate content as user portraits. Among them, when filling missing values ​​of fields, it can be filled by referring to the data of similar users or crawling relevant data from the network; the discretization of data includes dividing the value range of the data into intervals, defining the characteristic value of each interval, and taking the interval characteristic value that falls into the interval as the final value. For example, the user's age is divided into intervals of 5 years old, and the characteristic value of the 21-25 years old interval is defined as 1, the characteristic value of the 26-30 years old interval is defined as 2, and so on, to determine the characteristic value of each age interval. When the user is 27 years old, it falls into the 26-30 years old interval, so the value of the user's age field is determined to be "2".

[0111] It should be noted that the user's prediction results include types such as "yes / no" or probability, which are related to the prediction algorithm and configuration parameters of the prediction model, and this application does not limit this.

[0112] For example, if the prediction result of a user is "yes", the user is determined to be a target user.

[0113] For another example, if the prediction result of a user is "0.89", it means that the probability that the user belongs to the target user is 0.89. In this case, whether the user belongs to the target user can be determined by a preset threshold, or by referring to empirical data, statistical data, etc., and this application does not limit this.

[0114] It should be noted that, in some embodiments, before obtaining the user's prediction results based on the user's historical data and a preset prediction model, it also includes: constructing an initial prediction model and training the initial prediction model to obtain an alternative prediction model; evaluating the alternative prediction model to obtain a model evaluation result; and selecting a prediction model from the alternative prediction models based on the model evaluation result.

[0115] Among them, the prediction algorithms of the prediction model include random forest, K-neighbors and support vector machines (SVM) and other algorithms. The evaluation of the candidate prediction model includes evaluation from the aspects of accuracy, precision, recall and F1 score. In some specific implementations, the evaluation values ​​of the candidate prediction model in various aspects can be weighted and summed (the weight coefficient is a preset value), and the comprehensive evaluation result of the candidate prediction model can be determined according to the weighted value.

[0116] Step S402: when it is predicted that the user is a target user, determine candidate influencing factors of the user according to the user's historical data.

[0117] Step S403: determining a user cluster of the user according to a preset clustering model, the user's historical data and the user's candidate influencing factors.

[0118] Step S404: determining a main influencing factor of the user cluster according to candidate influencing factors of users belonging to the user cluster.

[0119] Step S405 , determining the target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user.

[0120] Steps S402 to S405 in this embodiment are the same as steps S101 to S104 in an embodiment of the present application, and are not described again here.

[0121] In this embodiment, the prediction result of the user is obtained according to the historical data of the user and the preset prediction model; when the user is predicted to be the target user, the alternative influencing factor of the user is determined according to the historical data of the user; the user cluster of the user is determined according to the preset clustering model, the historical data of the user and the alternative influencing factor of the user; the main influencing factor of the user cluster is determined according to the alternative influencing factors of the users belonging to the user cluster; the target influencing factor of the user is determined according to the main influencing factor of the user cluster and the alternative influencing factor of the user; the service strategy matching the target influencing factor is determined; and the user's alternative product information is updated according to the service strategy. This method can timely predict whether the user belongs to the target user, and when the user is determined to be the target user, accurately determine the user's target influencing factor, and then accurately determine the reason for user loss or value decline, so as to provide matching services to the user in time and reduce the user loss rate or the number of users with reduced value.

[0122] The step division of the above methods is only for clear description. When implemented, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs without changing the core design of the algorithm and process are all within the scope of protection of this application.

[0123] A second aspect of the present application provides a data analysis device. Figure 5 is a block diagram of a data analysis device provided in one embodiment of the present application. Figure 5 As shown, the data analysis device 500 includes:

[0124] The candidate influence factor determination module 501 is configured to determine the user's candidate influence factors according to the user's historical data when the user is predicted to be a target user.

[0125] Among them, the alternative impact factors are used to characterize the alternative reasons that affect user value.

[0126] In some embodiments, the candidate influencing factor determination module 501 includes an analysis unit and a first determination unit. The analysis unit is configured to analyze the user from a preset dimension based on the user's historical data when predicting that the user is a target user, and obtain the analysis result of the user in the preset dimension; the first determination unit is configured to determine the user's candidate influencing factor based on the analysis result. The preset dimensions include user attributes, activity, consumption, business saturation, contract period, complaint frequency, complaint type, etc.

[0127] The clustering module 502 is configured to determine a user cluster of a user according to a preset clustering model, historical data of the user, and candidate influencing factors of the user.

[0128] The clustering model is a model built based on a clustering algorithm, and the clustering algorithm includes but is not limited to a K-means algorithm, a hierarchical clustering algorithm, and an EM algorithm. In this embodiment, the clustering model is used to cluster users to form at least one user cluster, and the users in each user cluster have a high similarity.

[0129] In some embodiments, the clustering module 502 is specifically used to: input the user's historical data and the user's candidate influencing factors into the clustering model, the clustering model performs data processing, obtains output results, and determines the user cluster of the user according to the output results.

[0130] The main influencing factor determination module 503 is configured to determine the main influencing factor of the user cluster according to the candidate influencing factors of the users belonging to the user cluster.

[0131] Among them, the main influencing factor refers to the influencing factor that has the main impact on the user cluster.

[0132] In some embodiments, the main influencing factor determination module 503 includes a calculation unit and a second determination unit. The calculation unit is configured to calculate the repetition rate of the candidate influencing factors of the users belonging to the user cluster; the second determination unit is configured to determine the main influencing factor of the user cluster according to the repetition rate of the candidate influencing factors and a preset repetition rate threshold.

[0133] The target influence factor determination module 504 is configured to determine the target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user.

[0134] In some embodiments, the target influence factor determination module 504 includes an intersection determination unit and a third determination unit. The intersection determination unit is configured to determine the intersection of the main influence factor of the user cluster and the candidate influence factor of the user; the third determination unit is configured to determine the target influence factor of the user according to the intersection.

[0135] In some other embodiments, the target influence factor determination module 504 is also used for: when the intersection is an empty set, re-clustering the users to obtain an updated user cluster; the intersection determination unit determines the main influence factor of the updated user cluster based on the alternative influence factors of the users belonging to the updated user cluster; the factor determination unit determines the target influence factor of the user based on the main influence factor of the updated user cluster and the alternative influence factors of the users.

[0136] In this embodiment, when predicting that a user is a target user, the alternative influence factor determination module determines the user's alternative influence factor based on the user's historical data; the clustering module determines the user's user cluster based on the preset clustering model, the user's historical data and the user's alternative influence factor; the main influence factor determination module determines the main influence factor of the user cluster based on the alternative influence factors of the users belonging to the user cluster; the target influence factor determination module determines the user's target influence factor based on the main influence factor of the user cluster and the user's alternative influence factor. When predicting that the user is likely to churn or the user's value is in a downward trend, the device first determines the user's alternative influence factor, and determines the user's target influence factor based on the user's alternative influence factor and the main influence factor of the user cluster to which the user belongs, so that the user's target influence factor can be determined more accurately, and then the reason for the user's churn or value decline can be accurately determined, so as to provide matching services to the user in a timely manner and reduce the user churn rate or the number of users with declining value.

[0137] It should be noted that, in some embodiments, the data analysis device 500 further includes: a marketing module, which includes a service strategy determination unit and an update unit. The service strategy determination unit is configured to determine a service strategy that matches the target influencing factor; the update unit is configured to update the user's candidate product information according to the service strategy.

[0138] Figure 6 This is a block diagram of the composition of a data analysis device provided in yet another embodiment of the present application.

[0139] like Figure 6 As shown, the data analysis device 600 includes:

[0140] The prediction module 601 is configured to obtain the prediction result of the user according to the historical data of the user and the preset prediction model.

[0141] In some embodiments, the prediction module 601 includes a portrait unit and a model prediction unit. The portrait unit is configured to determine the user portrait of the user based on the user's historical data; the model prediction unit is configured to input the user portrait into the prediction model to obtain the user's prediction result.

[0142] It should be noted that the user's prediction results include types such as "yes / no" or probability, which are related to the prediction algorithm and configuration parameters of the prediction model.

[0143] The candidate influence factor determination module 602 is configured to determine the user's candidate influence factors according to the user's historical data when the user is predicted to be a target user.

[0144] The clustering module 603 is configured to determine a user cluster of the user according to a preset clustering model, the user's historical data and the user's candidate influencing factors.

[0145] The main influencing factor determination module 604 is configured to determine the main influencing factor of the user cluster according to the candidate influencing factors of the users belonging to the user cluster.

[0146] The target influence factor determination module 605 is configured to determine the target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user.

[0147] In this embodiment, the prediction module obtains the prediction result of the user according to the historical data of the user and the preset prediction model; the alternative influence factor determination module determines the alternative influence factor of the user according to the historical data of the user when predicting that the user is the target user; the clustering module determines the user cluster according to the preset clustering model, the historical data of the user and the alternative influence factor of the user; the main influence factor determination module determines the main influence factor of the user cluster according to the alternative influence factors of the users belonging to the user cluster; the target influence factor determination module determines the target influence factor of the user according to the main influence factor of the user cluster and the alternative influence factor of the user. The device can timely predict whether the user belongs to the target user, and accurately determine the target influence factor of the user when determining that the user is the target user, and then accurately determine the reason for user loss or value decline, so as to provide matching services to the user in time and reduce the user loss rate or the number of users with value decline.

[0148] Figure 7 This is a schematic diagram of a marketing system provided in an embodiment of the present application. Figure 7 As shown, the marketing system 700 includes a storage module 701 , a prediction model 702 , a clustering model 703 , a service strategy library 704 and a marketing module 705 .

[0149] Among them, the storage module 701 is used to provide storage functions, which can store user historical data, alternative influencing factors and other information; the prediction model 702 is used to predict whether the user belongs to the target user; the clustering model 703 is used to determine the user cluster to which the user belongs; the service strategy library 704 includes several service strategies; the marketing model 705 is used to carry out marketing work based on the clustering results and the service strategy library.

[0150] In some embodiments, the prediction model 702 is an untrained model, and therefore, the prediction model 702 needs to be trained first to obtain a prediction model with more accurate prediction results.

[0151] Specifically, historical data of all users within the current six months are obtained from storage module 701, and users who meet the conditions of "package change occurs", "package resources after change are lower than the original package charges", and "the average billing income after change is lower than the average billing income in the three months before change" are selected.

[0152] For the historical data of the selected users, obviously unreasonable outliers are eliminated, missing values ​​are filled by referring to similar user data or crawling data from the Internet, and characteristic values ​​of fields such as age are discretized to obtain the processed historical data of each user. This data is the user portrait of each user.

[0153] The user portrait of each user is used as the training data of the prediction model 702, and the prediction model 702 is trained to obtain the trained prediction model 702, and the trained prediction model 702 is cross-validated and tested to obtain the accuracy, precision, recall rate and F1-score of the prediction model. The prediction performance of the prediction model 702 is evaluated by the accuracy, precision, recall rate and F1-score, and the prediction model with good evaluation results is selected as the final model for user prediction.

[0154] After determining the prediction model 702, when there is a prediction demand, the historical data of the user to be predicted is input into the prediction model 702, and the prediction model 702 processes the data and obtains an output result. According to the output result of the prediction model 702, it can be determined whether the user belongs to the target user.

[0155] When the user is determined as a target user by the prediction model 702, the candidate influencing factors of the user are determined according to the user's historical data, and the user's historical data and the candidate influencing factors of the user are input into the clustering model 703. The clustering model 703 performs a clustering operation according to the input data to obtain an output result, and the user cluster to which the user belongs can be determined according to the output result.

[0156] After determining the user cluster to which the user belongs, the repetition rate of the alternative influencing factors of the user belonging to the user cluster is calculated, and the main influencing factor of the user cluster is determined based on the repetition rate of the alternative influencing factors and a preset repetition rate threshold, and the intersection of the main influencing factor of the user cluster and the alternative influencing factor of the user is determined, and the target influencing factor of the user is determined based on the intersection.

[0157] When the intersection is an empty set, the users are clustered again to obtain the updated user cluster. The repetition rate of the candidate influencing factors of the users belonging to the updated user cluster is calculated, and the main influencing factor of the updated user cluster is determined based on the repetition rate of the candidate influencing factors and the preset repetition rate threshold. The intersection of the main influencing factor of the updated user cluster and the candidate influencing factor of the user is determined, and the target influencing factor of the user is determined based on the intersection.

[0158] If the intersection of the main influencing factor of the updated user cluster and the user's alternative influencing factor is still an empty set, clustering is required again until the intersection of the main influencing factor of the user cluster obtained by clustering and the user's alternative influencing factor is not an empty set, and then the user's target influencing factor is determined based on the intersection.

[0159] According to the user's target influencing factors, the reasons for the decline in user value are analyzed, and a service strategy is matched for the user according to the service strategy library 704. Finally, the marketing module 705 performs personalized marketing for the user according to the matched service strategy, such as recommending a more suitable package or product to the user.

[0160] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application, but this does not mean that there are no other units in this embodiment.

[0161] It is to be understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of the present application, but the present application is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and substance of the present application, and these modifications and improvements are also considered to be within the scope of protection of the present application.

Claims

1. A data analysis method, characterized in that: include: In the case where the user is predicted to be a target user, determining an alternative influencing factor of the user according to the historical data of the user, wherein the alternative influencing factor is used to characterize an alternative reason affecting the user value; Determine a user cluster of the user according to a preset clustering model, historical data of the user, and candidate influencing factors of the user; Determining a main influencing factor of the user cluster according to candidate influencing factors of users belonging to the user cluster; Determining a target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user; Wherein, when predicting that a user is a target user, determining the candidate influencing factors of the user according to the historical data of the user includes: When the user is predicted to be a target user, analyzing the user from a preset dimension according to the historical data of the user to obtain an analysis result of the user in the preset dimension; Determining, according to the analysis result, an alternative influencing factor of the user; The determining of the main influencing factor of the user cluster according to the candidate influencing factors of the users belonging to the user cluster includes: Calculating the repetition rate of the candidate impact factors of the users belonging to the user cluster; The main influencing factor of the user cluster is determined according to the repetition rate of the candidate influencing factors and a preset repetition rate threshold.

2. The data analysis method according to claim 1, characterized in that: The determining the target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user includes: Determine the intersection of the main influencing factor of the user cluster and the candidate influencing factor of the user; A target impact factor of the user is determined according to the intersection.

3. The data analysis method according to claim 2, characterized in that: After determining the intersection of the main influencing factor of the user cluster and the candidate influencing factor of the user, the method further includes: When the intersection is an empty set, re-clustering the users to obtain an updated user cluster; Determining a main influencing factor of the updating user cluster according to candidate influencing factors of users belonging to the updating user cluster; The target influence factor of the user is determined according to the main influence factor of the updated user cluster and the candidate influence factor of the user.

4. The data analysis method according to claim 1, characterized in that: After determining the target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user, the method further includes: Determine a service strategy that matches the target impact factor; According to the service strategy, the user's candidate product information is updated.

5. The data analysis method according to claim 1, characterized in that: In the case where the user is predicted to be a target user, before determining the candidate influencing factors of the user according to the historical data of the user, the method further includes: A prediction result of the user is obtained according to the historical data of the user and a preset prediction model, wherein the prediction result is used to characterize whether the user is a target user.

6. The data analysis method according to claim 5, characterized in that: The obtaining the prediction result of the user according to the historical data of the user and the preset prediction model includes: Determine a user profile of the user according to the historical data of the user; The user portrait is input into the prediction model to obtain the prediction result of the user.

7. The data analysis method according to claim 5 or 6, characterized in that: Before obtaining the prediction result of the user according to the historical data of the user and the preset prediction model, the method further includes: Constructing an initial prediction model, and training the initial prediction model to obtain an alternative prediction model; Evaluating the candidate prediction model to obtain a model evaluation result; The prediction model is selected from the candidate prediction models according to the model evaluation result.

8. A data analysis device, characterized in that: include: An alternative influencing factor determination module is configured to determine an alternative influencing factor of the user according to the historical data of the user when predicting that the user is a target user, wherein the alternative influencing factor is used to characterize an alternative reason affecting the user value; A clustering module, configured to determine a user cluster of the user according to a preset clustering model, historical data of the user and candidate influencing factors of the user; A main influencing factor determination module, configured to determine a main influencing factor of the user cluster according to candidate influencing factors of users belonging to the user cluster; a target influence factor determination module, configured to determine a target influence factor of the user according to the main influence factor of the user cluster and the candidate influence factor of the user; Wherein, when predicting that a user is a target user, determining the candidate influencing factors of the user according to the historical data of the user includes: When the user is predicted to be a target user, analyzing the user from a preset dimension according to the historical data of the user to obtain an analysis result of the user in the preset dimension; Determining, according to the analysis result, an alternative influencing factor of the user; The determining of the main influencing factor of the user cluster according to the candidate influencing factors of the users belonging to the user cluster includes: Calculating the repetition rate of the candidate impact factors of the users belonging to the user cluster; The main influencing factor of the user cluster is determined according to the repetition rate of the candidate influencing factors and a preset repetition rate threshold.

Citation Information

Patent Citations

  • Data analysis method, device and equipment and computer storage medium

    CN110852780A

  • Target scheme acquisition method, system, electronic equipment and medium

    CN113377967A