A customer complaint risk prediction method based on the random forest algorithm
By adjusting the depth of the decision tree in the stochastic forest algorithm, combining the relevant indicators of influencing factors and complaint types, the problem of unreasonable depth construction of decision tree in the existing technology is solved, and the accuracy and reliability of customer complaint risk prediction is improved.
Patent Information
- Application Number
- CN202411318220.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-09-20
AI Technical Summary
When using the random forest algorithm to predict customer complaint frequency, the degree of correlation between influencing factors for different complaint types and the necessity of parent nodes in the data segmentation process are ignored, resulting in unreasonable construction of decision tree depth and reducing the accuracy of complaint prediction.
By obtaining complaint timing data and influencing factors in different local areas, we calculate the relevant indicators of influencing factors and complaint types, and adjust its initial depth according to the performance of the decision tree to build a decision tree more reasonably and improve the accuracy of complaint prediction.
By adjusting the depth of the decision tree, the impact of influencing factors on different complaint types is more reasonably reflected, and the accuracy and reliability of customer complaint risk prediction is improved.
Smart Images

Figure CN119295090B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of complaint risk prediction, and particularly to a method for predicting customer complaint risk based on a random forest algorithm. Background Art
[0002] Customers can file complaints about problems encountered during power supply and use through the customer service hotline. Customer complaints help to identify deficiencies in their own services. Customer complaints can be classified into five types, namely frequent power outages, repair beyond the time limit, long-term abnormal voltage quality, business expansion and installation beyond the time limit, and other types. Then, targeted improvement measures can be taken for each type of complaint. By predicting customer complaints, it helps to quickly repair the fault points, ensure the power quality of users, and reduce the power service risk.
[0003] Since the complaint frequency of complaint types is affected by various influencing factors. For example, factors such as the severity of the weather, the electricity consumption, and the dissatisfaction with the manual service will all affect the customer's satisfaction with the power, and thus affect the complaint frequency of the complaint types. The random forest algorithm is a common means to predict the complaint frequency of customers by combining influencing factors. However, when using the random forest algorithm to predict the complaint frequency, the differences in the degree of correlation between influencing factors for different complaint types and the necessity of the parent node in the data segmentation process are ignored, resulting in an unreasonable construction of the depth of the decision tree and reducing the accuracy of complaint prediction. Summary of the Invention
[0004] In order to solve the technical problem that the depth of the decision tree is unreasonably constructed when predicting the complaint frequency, reducing the accuracy of complaint prediction, the purpose of the present invention is to provide a method for predicting customer complaint risk based on a random forest algorithm. The specific technical solution adopted is as follows:
[0005] A method for predicting customer complaint risk based on a random forest algorithm, the method comprising the following steps:
[0006] Obtain the complaint time-series data of all types of complaints in different local areas, and the influence time-series data of all types of influencing factors; the complaint time-series data includes the complaint data at all historical moments; the influence time-series data includes the influence data at all historical moments;
[0007] Take any local area as the target area; in the target area, according to the correlation between the complaint time-series data and the influence time-series data, and the difference between the target area and the complaint data corresponding to all local areas, obtain the correlation indexes of the influencing factors and complaint types in the target area;
[0008] In the same complaint type in the target area, according to the complaint data at all historical moments and the impact data of all influencing factors at all historical moments, obtain the data sets of all parent nodes of each decision tree and the initial depth of each decision tree; in the decision tree, according to the numerical distribution of all the complaint data in the data set of the parent node, the relevant indicators of each influencing factor and complaint type, and the performance of the decision tree, adjust the initial depth of the decision tree to obtain the adjusted depth of the decision tree.
[0009] According to the adjusted depth of each decision tree corresponding to the same complaint type in the target area, obtain the predicted complaint data of each complaint type in each local area.
[0010] Further, the method for obtaining the relevant indicators specifically includes:
[0011] According to the correlation between the complaint time series data and the impact time series data, obtain the correlation factors of the influencing factors and complaint types in the target area.
[0012] In the same complaint type, according to the central tendency of the complaint data in each local area, obtain the complaint concentration factor of each local area.
[0013] According to the difference between the target area and all the local areas in the complaint concentration factors and the complaint concentration factor of the target area, obtain the adjustment parameter of the target area in the complaint type.
[0014] According to the adjustment parameter, adjust the correlation factors of the influencing factors and complaint types in the target area to obtain the relevant indicators of the influencing factors and complaint types in the target area; the adjustment parameter, the relevant factor and the relevant indicator are all positively correlated.
[0015] Further, the method for obtaining the relevant factor includes:
[0016] Calculate the absolute value of the Pearson correlation coefficient between the complaint time series data and the impact time series data to obtain the correlation factors of the influencing factors and complaint types in the target area.
[0017] Further, the method for obtaining the adjustment parameter includes:
[0018] In the same complaint type, calculate the absolute value of the difference between the complaint concentration factor of the target area and the complaint concentration factors of each local area to obtain the local deviation factor of each local area.
[0019] Calculate the mean value of the local deviation factors of all the local areas to obtain the complaint deviation parameter of the target area.
[0020] Obtain the adjustment parameter of the target area for the complaint type according to the complaint deviation parameter of the target area and the complaint concentration factor of the target area; the complaint deviation parameter is positively correlated with the adjustment parameter; the complaint concentration factor of the target area is negatively correlated with the adjustment parameter.
[0021] Further, the method for obtaining the adjusted depth includes:
[0022] Obtain the simplifiable index of the parent node according to the numerical distribution of all the complaint data in the data set of the parent node and the relevant indicators of each influencing factor and complaint type.
[0023] Adjust the initial depth of the decision tree according to the simplifiable index of the parent node and the performance of the decision tree to obtain the adjusted depth of the decision tree.
[0024] Further, the method for obtaining the simplifiable index includes:
[0025] Obtain the distribution concentration factor of the parent node according to the distribution concentration situation of all the complaint data in the data set of the parent node.
[0026] Obtain the distribution dispersion factor of the parent node according to the distribution difference situation of all the complaint data in the data set of the parent node.
[0027] Obtain the simplifiable index of the parent node according to the relevant indicators, distribution concentration factor and distribution dispersion factor of the influencing factor corresponding to the parent node and the complaint type; the relevant indicators and the distribution dispersion factor are both negatively correlated with the simplifiable index; the distribution concentration factor is positively correlated with the simplifiable index.
[0028] Further, the method for obtaining the distribution concentration factor includes:
[0029] In the data set of the parent node, sort all the complaint data in ascending order of the complaint data to obtain the complaint data sequence of the parent node; calculate the mean value of the difference values of all the complaint data points in the complaint data sequence to obtain the overall dispersion parameter; perform a negative correlation mapping on the overall dispersion parameter to obtain the distribution concentration factor of the parent node.
[0030] Further, the method for obtaining the distribution dispersion factor includes:
[0031] In the data set of the parent node, obtain the absolute value of the difference between every two complaint data, and take the largest absolute value of the difference as the distribution dispersion factor of the parent node.
[0032] Further, the method for obtaining the adjusted depth includes:
[0033] Obtain the evaluation function of the decision tree and the levels of the decision tree; according to the evaluation function of the decision tree and the levels of the decision tree, obtain the performance factor of the decision tree; the evaluation function and the performance factor are negatively correlated, and the levels and the performance factor are negatively correlated;
[0034] Integrate the reducibility indicators of all parent nodes in the decision tree to obtain the overall reducibility parameter of the decision tree;
[0035] According to the overall reducibility parameter and the performance factor of the decision tree, adjust the initial depth of the decision tree to obtain the adjusted depth of the decision tree; the initial depth, the performance factor are both positively correlated with the adjusted depth, and the overall reducibility parameter is negatively correlated with the adjusted depth.
[0036] Furthermore, the method for obtaining the overall reducibility parameter includes:
[0037] Calculate the cumulative value of the reducibility indicators of all parent nodes in the decision tree and perform normalization to obtain the overall reducibility parameter of the decision tree.
[0038] The present invention has the following beneficial effects:
[0039] Since in the local area, the complaint types are affected by the influencing factors in the local area; for example, when the power consumption is at the peak period, it will cause excessive load and abnormal voltage quality, and the complaint types corresponding to the abnormal voltage quality for a long time often have a significant increase in the number of complaints; when the bad weather causes the failure of external transmission equipment and results in power outages, the complaint types corresponding to frequent power outages often have a significant increase in the number of complaints; when it is a problem of manual service resulting in overtime repair or packaging, the complaint types corresponding to the over-limit of business expansion and installation often have a significant increase in the number of complaints; for the above different complaint types, there are certain differences in the degree of influence of the influencing factors. Obtain the relevant indicators of the influencing factors and complaint types in the target area. The larger the relevant indicator, the greater the degree of influence of the influencing factor on the complaint type in the target area.
[0040] In order to predict the complaint data of all types of complaints in the local area, first use the random forest algorithm to construct a decision tree and obtain the initial depth of the decision tree. Considering that the initial depth of the decision tree ignores the correlation degree between the influencing factors and the complaint types and ignores the necessity of some parent nodes in the data segmentation process, since the depth of the decision tree affects the prediction result, adjust the initial depth of the decision tree to make the adjusted depth more reasonable and improve the reliability of the prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 Flowchart of a customer complaint risk prediction method based on the random forest algorithm provided by an embodiment of the present invention;
[0043] Figure 2 Flowchart of a method for obtaining relevant indicators provided by an embodiment of the present invention;
[0044] Figure 3 Flowchart of a method for obtaining simplified indicators provided by an embodiment of the present invention;
[0045] Figure 4 Flowchart of a method for obtaining the adjusted depth provided by an embodiment of the present invention. Detailed implementation manners
[0046] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features and effects of a customer complaint risk prediction method based on the random forest algorithm proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0048] The following specifically describes the specific solution of a customer complaint risk prediction method based on the random forest algorithm provided by the present invention in combination with the accompanying drawings.
[0049] Please refer to Figure 1 , which shows the flowchart of a customer complaint risk prediction method based on the random forest algorithm provided by an embodiment of the present invention. The method includes the following steps:
[0050] Step S1, obtain the complaint time series data of all types of complaints in different local areas, as well as the influence time series data of all types of influencing factors; the complaint time series data includes the complaint data at all historical moments; the influence time series data includes the influence data at all historical moments.
[0051] Considering that the complaint behaviors in different local areas are different, in order to predict the complaint behaviors of each complaint type in a local area, it is necessary to deeply analyze the complaint behaviors of each complaint type in the history of different local areas and analyze various influencing factors.
[0052] Specifically, considering that customer complaints can be classified by types of frequent power outage complaints, repair over-time limit complaints, long-term abnormal voltage quality complaints, business expansion and installation over-time limit complaints, and other complaint types, these five complaint types are regarded as all complaint types in the present invention. There are power equipment points distributed in different regions, and the differences between the complaint behaviors in different regions can reflect the sensitivity of regional power equipment, which affects the reliability of subsequent analysis of the correlation between influencing factors and complaint types. Therefore, it is necessary to separately study the equipment with different spatial distributions, that is, obtain the local area to which the data belongs from the complaint monitoring system.
[0053] In order to analyze the number of user complaints of different complaint types in different local areas, for the same complaint type in the same local area, use the monitoring system to accumulate all the complaint times corresponding to the complaint type in the local area within a preset time interval as the complaint data for the preset time interval; then obtain the complaint data for each preset time interval in the historical reference interval; regard the preset time interval as a historical moment, and sequentially count the complaint data for all historical moments in the historical reference interval in the order of historical moments as the complaint time series data of this complaint type in the local area.
[0054] Considering that the number of complaints is affected by various influencing factors, the situations of influencing factors in different local areas are different and may also be different at different times. For example, influencing factors such as the severity of bad weather, electricity consumption, and dissatisfaction with manual service. It is necessary to obtain the complaint time series data of all complaint types in different local areas. In an embodiment of the present invention, the severity of bad weather, electricity consumption, and dissatisfaction with manual service are regarded as all influencing factors; the method for obtaining the influence time series data of all influencing factors in the local area is as follows:
[0055] For the influencing factor of the severity of bad weather, use the monitoring system to obtain the score value of the severity of bad weather, and regard the score value of the severity of bad weather in the local area within the preset time interval as the influence data for the preset time interval; for the influencing factor of electricity consumption, use the monitoring system to obtain the electricity consumption, and accumulate the electricity consumption in the local area within the preset time interval as the influence data for the preset time interval; for the influencing factor of dissatisfaction with manual service, use the monitoring system to obtain the score value of dissatisfaction with manual service; regard the average value of the score values of dissatisfaction with manual service in the local area within the preset time interval as the influence data for the preset time interval;
[0056] For any influencing factor in a local area, obtain the influencing data for each preset time interval in the historical reference interval; take the preset time interval as a historical moment, and sequentially count the influencing data for all historical moments in the historical reference interval in the order of historical moments, as the influencing time-series data of this influencing factor in the local area; finally, obtain the influencing time-series data of all types of influencing factors in the local area.
[0057] In an embodiment of the present invention, the preset time interval is from the start to the end of a day, the historical reference interval is 100 days, and the last day of the historical reference interval is the day before the current moment, which can be set by the implementer according to the implementation scenario. It should be noted that the present invention analyzes any local area and obtains the predicted complaint data for each complaint type corresponding to the local area.
[0058] It should be noted that for the convenience of calculation, all index data involved in the operations in the embodiments of the present invention have undergone data preprocessing, thereby eliminating the influence of dimensions. The specific means for eliminating the dimension influence are well-known technical means to those skilled in the art and will not be limited here.
[0059] Step S2, take any local area as the target area; in the target area, according to the correlation between the complaint time-series data and the influencing time-series data, and the difference between the target area and the complaint data corresponding to all local areas, obtain the relevant indicators of the influencing factors and complaint types in the target area.
[0060] Since in a local area, the complaint type is affected by the influencing factors in the local area; for example, when the power consumption is at a peak, it will cause excessive load and abnormal voltage quality, and the complaint type corresponding to the long-term abnormal voltage quality often has a significant increase in the number of complaints; when bad weather causes a failure of external power transmission equipment and results in a power outage, the complaint type corresponding to frequent power outages often has a significant increase in the number of complaints; when the manual service problem causes an over-time repair or packaging over-time, the complaint type corresponding to the over-time limit of business expansion and installation often has a significant increase in the number of complaints; for the above different complaint types, the degree of influence by the influencing factors is different, and obtain the relevant indicators of the influencing factors and complaint types in the target area. The larger the relevant indicator, the greater the degree of influence of the influencing factor on the complaint type in the target area.
[0061] Please refer to Figure 2 , which shows a flowchart of a method for obtaining relevant indicators provided by an embodiment of the present invention. Preferably, in an embodiment of the present invention, since the stability of power equipment and the quality of power services in different local areas are different, the degree of influence of the influencing factors on the complaint type will be different. To obtain the relevant indicators of the influencing factors and complaint types in the target area, the method for obtaining the relevant indicators specifically includes:
[0062] Step S201: Obtain the correlation factors of the influencing factors and complaint types in the target area according to the correlation between the complaint time series data and the influencing time series data;
[0063] Obtain the correlation factors through the correlation between the complaint time series data and the influencing time series data. The larger the correlation factor, the more relevant the complaint time series data and the influencing time series data are.
[0064] In an embodiment of the present invention, calculate the absolute value of the Pearson correlation coefficient between the complaint time series data and the influencing time series data to obtain the correlation factors of the influencing factors and complaint types in the target area, and reflect the correlation through the absolute value of the Pearson correlation coefficient. It should be noted that the Pearson correlation coefficient is a well-known prior art to those skilled in the art and will not be elaborated here. In other embodiments of the present invention, the absolute value of the Chebyshev correlation coefficient between the complaint time series data and the influencing time series data can also be calculated to obtain the correlation factors of the influencing factors and complaint types in the target area.
[0065] Step S202: In the same type of complaint, obtain the complaint concentration factors of each local area according to the central tendency of the complaint data in each local area;
[0066] The complaint concentration factor can reflect the overall concentration characteristics of all complaint data in the local area in the complaint type. In an embodiment of the present invention, calculate the mean value of all complaint data of the same complaint type in the local area to obtain the complaint concentration factor of the local area, and reflect the central tendency through the mean value. In other embodiments of the present invention, the mode or median of all complaint data in the local area is used as the complaint concentration factor of the local area, and the central tendency is reflected through the mode or median.
[0067] Step S203: Obtain the adjustment parameter of the target area in the complaint type according to the difference between the target area and all local areas in the complaint concentration factors and the complaint concentration factor of the target area;
[0068] In an embodiment of the present invention, in the same type of complaint, calculate the absolute value of the difference between the complaint concentration factor of the target area and the complaint concentration factors of each local area to obtain the local deviation factor of the local area; reflect the difference through the absolute value of the difference. The larger the local deviation factor, the greater the difference in complaint data between the local area and the target area in the complaint type.
[0069] Calculate the mean value of the local deviation factors of all local areas to obtain the complaint deviation parameter of the target area; the larger the complaint deviation parameter, the greater the difference in complaint data between all local areas and the target area;
[0070] Obtain the adjustment parameter of the target area according to the complaint deviation parameter of the target area and the complaint concentration factor of the target area; the complaint deviation parameter and the adjustment parameter are positively correlated; the complaint concentration factor of the target area and the adjustment parameter are negatively correlated.
[0071] In an embodiment of the present invention, calculate the ratio of the complaint deviation parameter and the complaint concentration factor of the target area, normalize the ratio, and calculate the sum value of 1 and the normalized result to obtain the adjustment parameter of the target area. The specific formula includes:
[0072] Where, TZ is the adjustment parameter of the target area for the complaint type; PC is the complaint deviation parameter of the target area for the complaint type; SJ is the complaint concentration factor of the target area for the complaint type; norm() is the normalization function.
[0073] In the formula, the complaint data reflects the dissatisfaction degree of customers with power equipment. The power equipment in different local areas has different sensitivities to influencing factors, resulting in different influencing degrees of the influencing factors in different areas on the same complaint type. For example, the score value of the bad weather degree in some local areas is very large, that is, the weather is bad, and the power equipment is not very sensitive to the weather, that is, it is difficult for bad weather to cause power equipment failures, so that the complaint type of frequent power outages corresponds to a lower number of complaints or no complaints. Its frequency change performance may be somewhat different from that of other local areas, resulting in the inaccuracy of the above calculated correlation factors in representing the influence degree of influencing factors on the complaint type. The larger the complaint deviation parameter, the greater the difference in complaint data between the target area and all local areas, and the less sensitive the target area is to external factors. At this time, it is necessary to increase the relevant factors, and the larger the adjustment parameter; the smaller the complaint concentration factor, the lower the complaint frequency of the area, indicating that the power equipment in the local area operates stably and has a low sensitivity to influencing factors. At this time, it is necessary to increase the relevant factors, and the larger the adjustment parameter; It reflects the adjustment according to the sensitivity of the local area on the basis of retaining the original relevant factors; the larger the adjustment parameter, the lower the sensitivity of the power equipment to the influencing factors.
[0074] Step S204: Adjust the relevant factors of the influencing factors and complaint types in the target area according to the adjustment parameter to obtain the relevant indicators of the influencing factors and complaint types in the target area; the adjustment parameter, the relevant factors and the relevant indicators are positively correlated.
[0075] Through the adjustment of the adjustment parameter, the relevant indicators can more accurately reflect the influence degree of the influencing factors on the complaint type in the local area. The larger the value, the greater the influence degree.
[0076] In one embodiment of the present invention, the product of the calculation adjustment parameter and the relevant factor is calculated and the product is normalized to obtain the relevant index of the influencing factor and the complaint type in the target area. In one embodiment of the present invention, norm() can be used for normalization, so that the value is normalized between 0 and 1, and norm() is a normalization function.
[0077] Step S3, in the same complaint type in the target area, according to the complaint data at all historical moments and the influence data of all influencing factors at all historical moments, obtain the data sets of all parent nodes of each decision tree and the initial depth of each decision tree; in the decision tree, according to the numerical distribution of all complaint data in the data set of the parent node, the relevant index of each influencing factor and complaint type, and the performance of the decision tree, adjust the initial depth of the decision tree to obtain the adjusted depth of the decision tree.
[0078] In order to predict the complaint data of all complaint types in the local area, first use the random forest algorithm to construct a decision tree and obtain the initial depth of the decision tree. Considering that the initial depth of the decision tree ignores the correlation degree between the influencing factors and the complaint types and ignores the necessity of some parent nodes in the data splitting process, since the depth of the decision tree affects the prediction result, adjust the initial depth of the decision tree to obtain the adjusted depth of the decision tree. Make the adjusted depth more reasonable and improve the reliability of the prediction model.
[0079] Specifically, when predicting the complaint data of the complaint type in the target area through the random forest algorithm, first construct a decision tree, and according to the complaint data at all historical moments and the influence data of all influencing factors at all historical moments, obtain each decision tree of the complaint type in the target area, and then obtain the data sets of all parent nodes of each decision tree and the initial depth of each decision tree; it should be noted that constructing a decision tree by the random forest algorithm is a well-known prior art to those skilled in the art, and only a brief description is given here: the decision tree in the random forest is a tree structure composed of a root node, a parent node, a child node and a leaf node. The data set corresponding to the parent node containing the complaint data is split into subsets according to the value of the influencing factor and passed to the child node; the child node continues to split according to the value of other influencing factors until the leaf node is reached. In the random forest, each decision tree is independently constructed, so their parent node and child node structures may be different. Among them, the root node: the top node of the decision tree, usually contains the samples of the entire data set; the leaf node: the bottom node of the decision tree, the node without child nodes, representing the final classification result; the parent node: also called the internal node, used to split the data set into subsets according to the value of a certain feature attribute; the child node: split from the parent node according to a certain rule, and each child node represents a more specific classification or a smaller data set.
[0080] To analyze the necessity of the parent node in the data segmentation process, according to the numerical distribution of all complaint data in the data set of the parent node and the relevant indicators of each influencing factor and complaint type, obtain the reducible indicator of the parent node; the larger the value of the reducible indicator, the greater the degree to which the parent node can be simplified, and the smaller the necessity of the parent node for segmentation.
[0081] Please refer to Figure 3 , which shows a flowchart of a method for obtaining a reducible indicator provided by an embodiment of the present invention. Preferably, in an embodiment of the present invention, the method for obtaining a reducible indicator includes:
[0082] Step S301: Obtain the concentration factor of the distribution of the parent node according to the concentration situation of the distribution of all complaint data in the data set of the parent node;
[0083] In an embodiment of the present invention, in the data set of the parent node, sort all complaint data in ascending order of complaint data to obtain the complaint data sequence of the parent node; calculate the mean value of the difference values of all complaint data points in the complaint data sequence to obtain the overall dispersion parameter; perform a negative correlation mapping on the overall dispersion parameter to obtain the concentration factor of the distribution of the parent node. By calculating the mean value of the difference values of all complaint data and performing a negative correlation mapping, the concentration situation of the distribution is reflected. It should be noted that in an embodiment of the present invention, exp(-x) is used to perform a negative correlation mapping on x, and exp() is an exponential function with e as the base. During the sorting process, there may be cases where the sizes of the complaint data are the same. In this case, sort them in the order of the historical moments corresponding to the complaint data from the earliest to the latest.
[0084] In other embodiments of the present invention, it is also possible to calculate the variance of all complaint data in the data set of the parent node and perform a negative correlation mapping on the variance to obtain the concentration factor of the distribution of the parent node. By performing a negative correlation mapping on the variance, the concentration situation of the distribution is reflected.
[0085] Step S302: Obtain the dispersion factor of the distribution of the parent node according to the distribution difference situation of all complaint data in the data set of the parent node;
[0086] In an embodiment of the present invention, in the data set of the parent node, obtain the absolute value of the difference between every two complaint data, and take the largest absolute value of the difference as the dispersion factor of the distribution of the parent node; the distribution difference situation is reflected by the largest absolute value of the difference. In other embodiments of the present invention, it is also possible to calculate the variance of all complaint data in the data set of the parent node to obtain the dispersion factor of the distribution of the parent node.
[0087] Step S303: Obtain the reducible index of the parent node according to the influencing factor corresponding to the parent node, the relevant index of the complaint type, the distribution concentration factor, and the distribution dispersion factor; both the relevant index and the distribution dispersion factor are negatively correlated with the reducible index; the distribution concentration factor is positively correlated with the reducible index.
[0088] In an embodiment of the present invention, negative correlation mapping is respectively performed on the relevant index and the distribution dispersion factor, the product of the two negative correlation mapping results is calculated, and the distribution concentration factor is multiplied by the product to obtain the reducible index of the parent node; the specific formula includes:
[0089] Wherein, YS is the reducible index of the parent node; XG is the relevant index of the influencing factor corresponding to the parent node and the complaint type; JZ is the distribution concentration factor of the parent node; FS is the distribution dispersion factor of the parent node; ε1 is the first denominator adjustment factor; ε2 is the second denominator adjustment factor. In an embodiment of the present invention, the value of the first denominator adjustment factor is 0.001, and the value of the second denominator adjustment factor is 0.0001, which is used to prevent the denominator from being 0, and this is not limited herein.
[0090] In the formula, since in the process of dividing the decision tree, the parent node divides the data according to the influencing factor, and since different influencing factors have different degrees of influence on the complaint type and different necessities in the division process, considering that the larger the distribution concentration factor of the parent node, the more concentrated the distribution of the internal set of the parent node, representing that the data is relatively pure, so further division may not produce much effect and can be more simplified, and the larger the reducible index; considering that the larger the distribution dispersion factor of the parent node, the more dispersed the distribution of the internal set of the parent node, the less it can be simplified, and the smaller the reducible index; considering that the larger the relevant index of the influencing factor corresponding to the node and the complaint type, the more important the influencing factor is in the division process, the less it can be simplified, and the smaller the reducible index; obtaining the reducible index, the smaller the reducible index, the more important the parent node is in the process of data segmentation and the less it can be simplified.
[0091] In order to adjust the initial depth of the decision tree, it is necessary to analyze the performance of the decision tree corresponding to the initial depth and the reducibility degree of the parent node, adjust the initial depth of the decision tree, and obtain the adjusted depth of the decision tree. Make the adjusted depth more reasonable and improve the reliability of the prediction model.
[0092] Please refer to Figure 4 , which shows a flowchart of a method for obtaining an adjusted depth provided by an embodiment of the present invention. Preferably, in an embodiment of the present invention, the method for obtaining the adjusted depth includes:
[0093] Step S311: Obtain the evaluation function of the decision tree and the level of the decision tree; according to the evaluation function of the decision tree and the level of the decision tree, obtain the performance factor of the decision tree; the evaluation function and the performance factor are negatively correlated, and the level and the performance factor are negatively correlated;
[0094] It should be noted that obtaining the evaluation function of the decision tree and the level of the decision tree is well-known prior art to those skilled in the art and will not be elaborated here. The smaller the evaluation function of the decision tree and the fewer the levels of the decision tree, the better the performance of the decision tree, and the larger the performance factor. At this time, it means that the initial depth of the decision tree is more appropriate, and the degree of adjustment required for the initial depth should be smaller.
[0095] In an embodiment of the present invention, calculate the product of the evaluation function of the decision tree and the level of the decision tree, and perform a negative correlation mapping on the product to obtain the performance factor of the decision tree. For example, in an embodiment of the present invention, use norm(-x) to perform a negative correlation mapping on x.
[0096] Step S312: Synthesize the simplifiable indicators of all parent nodes in the decision tree to obtain the overall simplifiable parameter of the decision tree;
[0097] The overall simplifiable parameter reflects the overall simplifiable degree of the decision tree. The larger the value, the more the decision tree can be simplified, and the initial depth can be adjusted smaller.
[0098] In an embodiment of the present invention, calculate the cumulative value of the simplifiable indicators of all parent nodes in the decision tree and perform normalization to obtain the overall simplifiable parameter of the decision tree, and reflect the overall trend through the cumulative value. In an embodiment of the present invention, calculate the mean value of the simplifiable indicators of all parent nodes in the decision tree and perform normalization to obtain the overall simplifiable parameter of the decision tree. For example, in an embodiment of the present invention, use the linear normalization method for normalization, and normalize it between 0 and 1.
[0099] Step S313: Adjust the initial depth of the decision tree according to the overall simplifiable parameter and the performance factor of the decision tree to obtain the adjusted depth of the decision tree; the initial depth, the performance factor, and the adjusted depth are positively correlated, and the overall simplifiable parameter and the adjusted depth are negatively correlated.
[0100] Calculate the ratio of the performance factor to the overall simplifiable parameter, and perform normalization on the ratio to obtain the normalization result; calculate the product of the normalization result and the initial depth to obtain the adjusted depth of the decision tree; in an embodiment of the present invention, the formula for the adjusted depth includes:
[0101] Eout is the adjusted depth of the decision tree; Ein is the initial depth of the decision tree; XG is the performance factor of the decision tree; ZYS is the overall reducible parameter of the decision tree; norm() is the normalization function.
[0102] In the formula, the overall reducible parameter reflects the overall reducibility of the decision tree. The larger the value, the more the decision tree can be simplified, and the initial depth can be adjusted smaller. The larger the performance factor, the more appropriate the initial depth of the decision tree is at this time, indicating that the degree of adjustment of the initial depth should be smaller. This makes the adjusted depth more reasonable and improves the reliability of the prediction model.
[0103] Step S4: According to the adjusted depth of each decision tree corresponding to the same complaint type in the target area, obtain the predicted complaint data of each complaint type in each local area.
[0104] Using the random forest method, according to the adjusted depth of each decision tree corresponding to the same complaint type in the target area, predict the complaint data to obtain the predicted complaint data of each complaint type in each local area. It should be noted that the random forest method is a well-known technical means for those skilled in the art and will not be elaborated here.
[0105] Furthermore, in the local area, calculate the cumulative value of the predicted complaint data of all complaint types in the local area, calculate the ratio of the predicted complaint data of the complaint type in the local area to the cumulative value, and obtain the complaint risk probability of the complaint type in the local area; analyze the possible reasons for the complaint type, conduct early investigation, repair, and maintenance in advance, achieve "foresight" of complaints, enhance the company's early warning ability for complaint work orders, and based on this, carry out more targeted service improvement, reduce service risks, and achieve a reduction in the company's complaint work order volume.
[0106] In summary, the embodiment of the present invention provides a method for predicting customer complaint risks based on the random forest algorithm. First, obtain the relevant indicators of influencing factors and complaint types in the target area; adjust the initial depth of the decision tree according to the numerical distribution of all complaint data in the data set of the parent node, the relevant indicators of each influencing factor and complaint type, and the performance of the decision tree, and obtain the adjusted depth of the decision tree; according to the adjusted depth of each decision tree corresponding to the same complaint type in the target area, obtain the predicted complaint data of each complaint type in each local area. In the embodiment of the present invention, by combining the differences in the correlation degree between complaint types due to influencing factors and the necessity of the parent node in the data segmentation process, the adjusted depth is made more reasonable and the prediction reliability is improved.
[0107] It should be noted that the above order of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0108] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.
Claims
1. A customer complaint risk prediction method based on random forest algorithm, characterized in that: The method comprises the following steps: Obtain complaint time series data of all complaint types in different local areas, and influence time series data of all influencing factors; the complaint time series data includes complaint data at all historical moments; the influence time series data includes influence data at all historical moments; Any local area is taken as a target area; in the target area, according to the correlation between the complaint time series data and the impact time series data, and the difference between the complaint data corresponding to the target area and all local areas, the relevant indicators of the influencing factors and complaint types in the target area are obtained; In the same complaint type in the target area, based on the complaint data at all historical moments and the impact data of all influencing factors at all historical moments, the data set of all parent nodes of each decision tree and the initial depth of each decision tree are obtained; in the decision tree, based on the numerical distribution of all the complaint data in the data set of the parent node, the relevant indicators of each influencing factor and complaint type, and the performance of the decision tree, the initial depth of the decision tree is adjusted to obtain the adjusted depth of the decision tree; According to the adjusted depth of each decision tree corresponding to the same complaint type in the target area, the predicted complaint data of each complaint type in each local area is obtained; the method for obtaining the relevant indicators specifically includes: Calculate the absolute value of the Pearson correlation coefficient between the complaint time series data and the impact time series data to obtain the correlation factor between the impact factor and the complaint type in the target area; Calculate the mean of all complaint data of the same complaint type in the local area to obtain the complaint concentration factor of the local area; In the same complaint type, the absolute value of the difference between the complaint concentration factor of the target area and the complaint concentration factor of each local area is calculated to obtain the local deviation factor of each local area; Calculate the mean of the local deviation factors of all the local areas to obtain the complaint deviation parameter of the target area; According to the complaint deviation parameter of the target area and the complaint concentration factor of the target area, an adjustment parameter of the target area in the complaint type is obtained; the complaint deviation parameter and the adjustment parameter are positively correlated; the complaint concentration factor of the target area and the adjustment parameter are negatively correlated; Calculate the product of the adjustment parameter and the relevant factor and normalize the product to obtain the relevant indicators of the influencing factors and complaint types in the target area; The method for obtaining the adjusted depth includes: In the data set of the parent node, all the complaint data are sorted in order from small to large to obtain the complaint data sequence of the parent node; the mean of the difference values of all complaint data points in the complaint data sequence is calculated to obtain the overall dispersion parameter; the overall dispersion parameter is negatively correlated to obtain the distribution concentration factor of the parent node; In the data set of the parent node, the absolute value of the difference between every two complaint data is obtained, and the maximum absolute value of the difference is used as the distribution dispersion factor of the parent node; Negative correlation mapping is performed on the relevant index and the distribution dispersion factor respectively, the product of the two negative correlation mapping results is calculated, and the distribution concentration factor and the product are multiplied to obtain the simplifiable index of the parent node; Obtain the evaluation function of the decision tree and the level of the decision tree; calculate the product of the evaluation function of the decision tree and the level of the decision tree, and perform negative correlation mapping on the product to obtain the performance factor of the decision tree; Calculate the cumulative value of the reducible index of all parent nodes in the decision tree and normalize it to obtain the overall reducible parameter of the decision tree; The ratio of the performance factor to the overall simplifiable parameter is calculated, and the ratio is normalized to obtain a normalized result; the product of the normalized result and the initial depth is calculated to obtain the adjusted depth of the decision tree.
Citation Information
Patent Citations
Random forest-based power failure sensitive user prediction method and system, storage medium and computer equipment
CN112766550A
Cooling system intelligent control method and system based on fluorinated liquid performance degradation evaluation
CN117479510A