Customer service fault prevention method and device based on data mining, electronic equipment, storage medium and program product
By constructing a business failure prediction model that integrates strategy variables, result variables and residual fit models, the problem of inaccurate predictions in the existing technology is solved, and more accurate and targeted fault prediction and prevention strategy formulation is achieved.
Patent Information
- Application Number
- CN202411852314.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-06
AI Technical Summary
When using a single machine learning model to predict customer service business failures, the existing technology ignores the complex causal relationship between the policy variables and the result variables, resulting in the prediction results being inaccurate enough and it is difficult to provide support for formulating effective failure prevention strategies.
By obtaining historical service record data from the customer service system, cleaning and preprocessing, building a strategy variable model, a result variable model and a residual fitting model, integrating these models to build a business failure prediction model, and optimizing the model hyperparameters through K-fold cross-validation, and finally predicting real-time data based on the trained model, identifying service failure risks and formulating prevention strategies.
It improves the accuracy and pertinence of predicting failures in customer service business, and can more effectively support the formulation and implementation of fault prevention strategies, thereby improving customer satisfaction and corporate reputation.
Smart Images

Figure CN119940908A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of customer service business failure processing, and in particular to a customer service business failure prevention method, device, electronic device, storage medium and program product based on data mining. Background Art
[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims. No description herein is admitted to be prior art by inclusion in this section.
[0003] Customer service is an important bridge between enterprises and customers. Its efficiency and quality are directly related to customer satisfaction and corporate reputation. Various failures encountered in the process of customer service, such as slow response and improper handling, will damage customer experience and may lead to customer loss, thus affecting the long-term development of the enterprise.
[0004] In recent years, with the rapid development of data mining and machine learning technologies, attempts have been made to apply these technologies to the prevention of customer service failures. Most of the related technologies use a single machine learning model, such as decision trees, random forests, or neural networks, to predict customer service failures.
[0005] However, related technologies ignore the complex causal relationship between strategy variables and outcome variables, resulting in inaccurate prediction results and difficulty in providing strong support for the formulation of effective fault prevention strategies. Summary of the invention
[0006] In view of this, the purpose of the present disclosure is to propose a customer service business failure prevention method, device, electronic device, storage medium and program product based on data mining, which at least solves one of the technical problems in the related technology to a certain extent.
[0007] Based on the above purpose, the first aspect of the exemplary embodiment of the present disclosure provides a customer service business failure prevention method based on data mining, including:
[0008] Acquire historical service record data from the customer service system, clean and preprocess the historical service record data, and obtain a basic data set;
[0009] Performing service request and customer complaint correlation analysis and customer group cluster analysis on the basic data set to obtain analysis results, and merging the basic data set and the analysis results to obtain a training data set;
[0010] Constructing a strategy variable model and a result variable model, using the residuals of the strategy variable model and the result variable model to construct a residual fitting model, integrating the strategy variable model, the result variable model and the residual fitting model to obtain a business failure prediction model;
[0011] The service failure prediction model is trained based on the training data set, and the hyperparameters of the service failure prediction model are optimized based on the K-fold cross validation method to obtain a trained service failure prediction model;
[0012] Acquire real-time customer service business data from the customer service system, and predict the real-time customer service business data based on the trained business failure prediction model to obtain a service failure risk identification result;
[0013] Based on the service failure risk identification result, a failure prevention strategy is constructed, and the failure prevention strategy is applied to the customer service system.
[0014] In some exemplary embodiments, cleaning the historical service record data includes:
[0015] Determining missing values and noise data in the historical service record data;
[0016] Filling the missing values based on a linear interpolation algorithm;
[0017] For the noise data, data denoising is performed based on a sliding average filtering algorithm;
[0018] Preprocessing the historical service record data includes:
[0019] Dividing the historical service record data into continuous data and categorized data;
[0020] Performing standardization on the continuous data;
[0021] The categorized data is subjected to one-hot encoding conversion processing.
[0022] In some exemplary embodiments, performing a correlation analysis between service requests and customer complaints on the basic data set includes:
[0023] Based on the association rule mining algorithm, the basic data set is analyzed for association between service requests and customer complaints to obtain the service request type corresponding to the customer complaint;
[0024] Performing customer group clustering analysis on the basic data set, including:
[0025] Based on the K-Means algorithm, customer group clustering analysis is performed on the basic data set to obtain a high-satisfaction customer group and a customer group prone to complaints.
[0026] In some exemplary embodiments, constructing a strategy variable model and an outcome variable model includes:
[0027] The strategy variable model takes the strategy variable as an input feature, performs fitting based on the random forest algorithm, and outputs the predicted value of the strategy variable;
[0028] The outcome variable model uses the strategy variables and other covariates as input features, performs fitting based on the random forest algorithm, and outputs the predicted value of the outcome variable;
[0029] The method of constructing a residual fitting model by using the residuals of the strategy variable model and the outcome variable model includes:
[0030] Based on the residuals of the strategy variable model and the outcome variable model, the causal effect of the strategy variable on the outcome variable is estimated based on a linear regression algorithm.
[0031] In some exemplary embodiments, the training of the service failure prediction model based on the training data set, and optimizing the hyperparameters of the service failure prediction model based on the K-fold cross validation method to obtain the trained service failure prediction model include:
[0032] Step a: Divide the training data set into K subsets;
[0033] Step b: For each cross-validation, select K minus one subset as the training set and the remaining subset as the validation set;
[0034] Step c: training the service failure prediction model based on the training set to obtain a trained service failure prediction model;
[0035] Step d: verifying the trained service failure prediction model based on the verification set to obtain a verification error of the current fold;
[0036] Step e: Repeat steps b to d for a total of K times, and select a different validation set for each cross-validation;
[0037] Step f: Determine the average validation error of K-fold cross validation;
[0038] Step g: Select a model or a hyperparameter combination based on the average verification error as the trained business failure prediction model.
[0039] In some exemplary embodiments, constructing a fault prevention strategy based on the service fault risk identification result includes:
[0040] Based on the service failure risk identification results, adjust customer service staff allocation, optimize processing priorities, improve automated response strategies, and adjust service processes;
[0041] After applying the fault prevention strategy to the customer service system, the method further includes:
[0042] Monitoring the effectiveness of the implementation of the fault prevention strategy;
[0043] The fault prevention strategy is adjusted based on the implementation effect.
[0044] Based on the same inventive concept, the second aspect of the exemplary embodiment of the present disclosure provides a customer service business failure prevention device based on data mining, including:
[0045] A basic data set acquisition module is configured to acquire historical service record data from a customer service system, clean and pre-process the historical service record data, and obtain a basic data set;
[0046] A training data set acquisition module is configured to perform service request and customer complaint association analysis and customer group cluster analysis on the basic data set to obtain analysis results, and merge the basic data set with the analysis results to obtain a training data set;
[0047] A service failure prediction model building module is configured to build a policy variable model and a result variable model, use the residuals of the policy variable model and the result variable model to build a residual fitting model, and integrate the policy variable model, the result variable model and the residual fitting model to obtain a service failure prediction model;
[0048] A service failure prediction model training module is configured to train the service failure prediction model based on the training data set, optimize the hyperparameters of the service failure prediction model based on the K-fold cross-validation method, and obtain a trained service failure prediction model;
[0049] A service failure prediction module is configured to obtain real-time customer service business data from the customer service system, predict the real-time customer service business data based on the trained service failure prediction model, and obtain a service failure risk identification result;
[0050] The fault prevention strategy construction and application module is configured to construct a fault prevention strategy based on the service fault risk identification result and apply the fault prevention strategy to the customer service system.
[0051] Based on the same inventive concept, a third aspect of the exemplary embodiment of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in the first aspect is implemented.
[0052] Based on the same inventive concept, a fourth aspect of the exemplary embodiments of the present disclosure provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method described in the first aspect.
[0053] Based on the same inventive concept, a fifth aspect of the exemplary embodiments of the present disclosure provides a computer program product, including computer program instructions. When the computer program instructions are executed on a computer, the computer executes the method described in the first aspect.
[0054] As can be seen from the above, the customer service business failure prevention method, device, electronic device, storage medium and program product based on data mining provided by the embodiments of the present disclosure include: obtaining historical service record data from the customer service system, cleaning and preprocessing the historical service record data to obtain a basic data set; performing service request and customer complaint association analysis and customer group clustering analysis on the basic data set to obtain analysis results, and merging the basic data set and the analysis results to obtain a training data set; constructing a policy variable model and a result variable model, using the residuals of the policy variable model and the result variable model to construct a residual fitting model, integrating the policy variable model, the result variable model and the residual fitting model to obtain a business failure prediction model; training the business failure prediction model based on the training data set, optimizing the hyperparameters of the business failure prediction model based on the K-fold cross validation method, and obtaining a trained business failure prediction model; obtaining real-time customer service business data from the customer service system, predicting the real-time customer service business data based on the trained business failure prediction model, and obtaining a service failure risk identification result; constructing a failure prevention strategy based on the service failure risk identification result, and applying the failure prevention strategy to the customer service system. The present disclosure analyzes the association between the policy variable, namely the service request, and the result variable, namely the customer complaint, thereby improving the prediction accuracy and pertinence of customer service business failures. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0056] Figure 1 A schematic diagram of an application scenario of a customer service business failure prevention method based on data mining provided by an exemplary embodiment of the present disclosure;
[0057] Figure 2A flow chart of a customer service business failure prevention method based on data mining provided by an exemplary embodiment of the present disclosure;
[0058] Figure 3 Another flowchart of a customer service business failure prevention method based on data mining provided by an exemplary embodiment of the present disclosure;
[0059] Figure 4 Another flowchart of a customer service business failure prevention method based on data mining provided by an exemplary embodiment of the present disclosure;
[0060] Figure 5 A flow chart of a method for constructing a business failure prediction model provided by an exemplary embodiment of the present disclosure;
[0061] Figure 6 A flow chart of a training method of a service failure prediction model provided by an exemplary embodiment of the present disclosure;
[0062] Figure 7 A schematic diagram of a structure of a customer service business failure prevention device based on data mining provided by an exemplary embodiment of the present disclosure;
[0063] Figure 8 A schematic structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0064] It is understandable that before using the technical solutions disclosed in the embodiments of this application, the type, scope of use, usage scenarios, etc. of the personal information involved in this application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0065] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present application according to the prompt message.
[0066] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0067] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation method of the present application. Other methods that meet the relevant laws and regulations may also be applied to the implementation method of the present application.
[0068] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0069] In order to make the purpose, technical solutions and advantages of the present disclosure more clear, the principles and spirit of the present disclosure will be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0070] It should be understood herein that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction rather than having any limiting meaning.
[0071] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not represent any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connecting" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly. The article "one" or "a" before an element does not exclude the existence of multiple such elements.
[0072] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure.
[0073] As described in the background technology, in today's highly competitive business environment, customer service is an important bridge for communication between enterprises and customers, and its efficiency and quality are directly related to customer satisfaction and corporate reputation. At present, various failures are often encountered in the process of customer service, such as slow response speed, customer complaints caused by improper handling, and problems that cannot be solved in one go. These problems not only damage the customer experience, but may also lead to customer loss, thereby affecting the long-term development of the enterprise.
[0074] In recent years, with the rapid development of data mining and machine learning technologies, some companies have begun to try to apply these technologies to the prevention of customer service business failures.
[0075] However, the inventors of the present disclosure have found that most of the related technologies use a single machine learning model, such as a decision tree, random forest or neural network. Although these models can improve the accuracy of fault prediction to a certain extent, they often ignore the complex causal relationship between policy variables (such as customer service staff allocation, processing priority, etc.) and outcome variables (such as customer satisfaction, complaint rate, etc.), resulting in inaccurate prediction results and difficulty in providing strong support for the formulation of effective fault prevention strategies. Taking the customer service staff allocation strategy as an example, different customer service staff may produce completely different results when handling the same service request due to differences in experience, skills and personality. A single machine learning model often cannot capture such subtle differences, resulting in inaccurate prediction results. Similarly, processing priority strategies, automated response strategies and service process strategies also have an important impact on the results of customer service business, but these factors are often ignored or simplified in related technologies.
[0076] In order to solve the above problems, the present disclosure provides a customer service business failure prevention solution based on data mining, which specifically includes: obtaining historical service record data from a customer service system, cleaning and preprocessing the historical service record data to obtain a basic data set; performing service request and customer complaint correlation analysis and customer group clustering analysis on the basic data set to obtain analysis results, and merging the basic data set and the analysis results to obtain a training data set; constructing a policy variable model and a result variable model, using the residuals of the policy variable model and the result variable model to construct a residual fitting model, integrating the policy variable model, the result variable model and the residual fitting model to obtain a business failure prediction model; training the business failure prediction model based on the training data set, optimizing the hyperparameters of the business failure prediction model based on the K-fold cross-validation method to obtain a trained business failure prediction model; obtaining real-time customer service business data from the customer service system, predicting the real-time customer service business data based on the trained business failure prediction model to obtain a service failure risk identification result; constructing a failure prevention strategy based on the service failure risk identification result, and applying the failure prevention strategy to the customer service system.
[0077] The present disclosure analyzes the association between the policy variable, namely the service request, and the result variable, namely the customer complaint, thereby improving the prediction accuracy and pertinence of customer service business failures.
[0078] In some exemplary embodiments, customer group clustering analysis is also performed, and the analysis results are merged with the basic data set to form a more comprehensive training data set, which helps the model to more accurately understand the needs and expectations of different customer groups for services, thereby improving the accuracy and pertinence of the predictions.
[0079] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure are described in detail below.
[0080] refer to Figure 1 , which is a schematic diagram of an application scenario of a customer service business failure prevention method based on data mining provided by an exemplary embodiment of the present disclosure.
[0081] This application scenario includes a terminal device 101, a server 102, and a data storage system 103. The terminal device 101, the server 102, and the data storage system 103 may be connected via a wired or wireless communication network to achieve data interaction.
[0082] The terminal device 101 may be an electronic device with data transmission and multimedia input / output functions close to the user side, including but not limited to a desktop computer, a mobile phone, a mobile computer, a tablet computer, a media player, a smart wearable device, a personal digital assistant (PDA) or other electronic devices capable of realizing the above functions. The electronic device may include a processor and a display screen with a touch input function, the display screen is used to present a graphical user interface, the graphical user interface can display an application interface, and the processor is used to process application data, generate a graphical user interface, and control the display of the graphical user interface on the display screen.
[0083] Server 102 and data storage system 103 can both be independent physical servers, or a server cluster or distributed system composed of multiple physical servers. They can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0084] In some exemplary embodiments, the customer service business failure prevention method based on data mining may be run on the server 102 .
[0085] When the customer service business failure prevention method based on data mining runs on the server 102, the server 102 obtains historical service record data from the customer service system, cleans and preprocesses the historical service record data to obtain a basic data set; performs service request and customer complaint association analysis and customer group clustering analysis on the basic data set to obtain analysis results, and merges the basic data set and the analysis results to obtain a training data set; constructs a policy variable model and a result variable model, uses the residuals of the policy variable model and the result variable model to construct a residual fitting model, integrates the policy variable model, the result variable model and the residual fitting model to obtain a business failure prediction model; trains the business failure prediction model based on the training data set, optimizes the hyperparameters of the business failure prediction model based on the K-fold cross-validation method, and obtains a trained business failure prediction model; obtains real-time customer service business data from the customer service system, predicts the real-time customer service business data based on the trained business failure prediction model, and obtains a service failure risk identification result; constructs a failure prevention strategy based on the service failure risk identification result, and applies the failure prevention strategy to the customer service system.
[0086] Server 102 is used to provide customer service to users of terminal device 101. A client that communicates with server 102 is installed in terminal device 101. Users can input questions through the client, and the client sends the question to server 102. Server 102 generates an answer to the question, and server 102 sends the answer to the client, which displays the answer to the user.
[0087] The data storage system 103 stores historical service record data or a training data set obtained by processing the historical service record data.
[0088] Combine the following Figure 1 The present invention describes the customer service business fault prevention method based on data mining according to the exemplary embodiment of the present invention by using the application scenario. It should be noted that the above application scenario is only shown to facilitate understanding of the spirit and principle of the present invention, and the embodiments of the present invention are not limited in this respect. On the contrary, the embodiments of the present invention can be applied to any applicable scenario.
[0089] refer to Figure 2 and Figure 3 , which is a flow chart of a customer service business failure prevention method based on data mining provided by an exemplary embodiment of the present disclosure.
[0090] The customer service business failure prevention method based on data mining includes the following steps:
[0091] Step S210: Acquire historical service record data from the customer service system, clean and preprocess the historical service record data, and obtain a basic data set.
[0092] During specific implementation: historical service record data is collected from the customer service system, including but not limited to basic customer information, service request content, customer service personnel information, processing records, customer satisfaction feedback, complaint records, etc. The collected data may have problems such as missing, duplicated or outliers, so the data needs to be cleaned and preprocessed to ensure the accuracy and consistency of the data, and finally form a basic data set.
[0093] In this exemplary embodiment, cleaning the historical service record data includes:
[0094] Determining missing values and noise data in the historical service record data;
[0095] Filling the missing values based on a linear interpolation algorithm;
[0096] For the noise data, data denoising is performed based on a sliding average filtering algorithm.
[0097] In this exemplary embodiment, preprocessing the historical service record data includes:
[0098] Dividing the historical service record data into continuous data and categorized data;
[0099] Performing standardization on the continuous data;
[0100] The categorized data is subjected to one-hot encoding conversion processing.
[0101] In specific implementation: standardize the continuous data to eliminate the dimension difference, use the Z-score method to standardize the data, subtract the mean of each numerical data and divide it by its standard deviation, so that the processed data conforms to the standard normal distribution. The formula for standardization is: Among them, X is the original data, μ is the mean of the original data, σ is the standard deviation of the original data, and Z is the standardized data.
[0102] Step S220: performing service request and customer complaint correlation analysis and customer group clustering analysis on the basic data set to obtain analysis results, and merging the basic data set and the analysis results to obtain a training data set.
[0103] In this exemplary embodiment, performing a correlation analysis between service requests and customer complaints on the basic data set includes:
[0104] Based on the association rule mining algorithm, the basic data set is analyzed for association between service requests and customer complaints to obtain the service request type corresponding to the customer complaint.
[0105] In specific implementation: use association rule mining algorithm to conduct association analysis on service request records and customer complaint records in the customer service system; set minimum support threshold and minimum confidence threshold to filter out strong association rules between frequently occurring service request types and customer complaints;
[0106] Calculate the support and confidence of the association rules, where the support represents the frequency of service requests and customer complaints occurring simultaneously, and the confidence represents the probability of customer complaints occurring when service requests occur;
[0107] Identify the types of service requests that are frequently accompanied by customer complaints, providing key features for subsequent model building.
[0108] In the specific implementation, the association rule mining algorithm Apriori is used to realize the association analysis between service requests and customer complaints. The specific steps include:
[0109] Service request records and customer complaint records are extracted from the customer service system to form two independent data sets. Service request records should include information such as the type, time, and processing results of the service request; customer complaint records should include information such as the content, time, type of complaint, and identifiers related to the service request (such as the service request ID).
[0110] The Apriori algorithm is used to mine association rules on the preprocessed data. First, the minimum support threshold and the minimum confidence threshold are set. The minimum support threshold is used to filter out the frequently occurring combinations of service request types and customer complaints, while the minimum confidence threshold is used to measure the strength of association between these combinations.
[0111] The Apriori algorithm is used to perform correlation analysis on service request records and customer complaint records. The algorithm traverses all possible combinations of service request types and customer complaints and calculates their support and confidence. Support indicates the frequency of service requests and customer complaints appearing at the same time, while confidence indicates the probability of customer complaints appearing when service requests appear.
[0112] According to the set minimum support threshold and minimum confidence threshold, strong association rules between frequently occurring service request types and customer complaints are screened out. The strong association rules reveal which service request types are more likely to cause customer complaints, providing key features for subsequent model construction.
[0113] In order to illustrate the implementation of the above method in more detail, an example is given below.
[0114] Assume that the customer service system contains the following service request types: A (balance inquiry), B (transfer), C (password change) and D (consulting service). At the same time, the customer complaint record contains the following complaint types: E (slow service speed), F (operation error) and G (inaccurate information).
[0115] Through association rule mining, the following strong association rules can be obtained:
[0116] Rule 1: A&B→E (support = 0.2, confidence = 0.8)
[0117] This means that when the service request types of balance inquiry (A) and transfer (B) appear at the same time, the probability of customers complaining about slow service speed (E) is higher.
[0118] Rule 2: C→F (support = 0.15, confidence = 0.75)
[0119] This means that when the service request type is password change (C), the probability of customer complaints about operational errors (F) is higher.
[0120] Rule 3: D→G (support = 0.1, confidence = 0.6)
[0121] This means that when the service request type is consulting business (D), the probability of customers complaining about inaccurate information (G) is higher.
[0122] In this exemplary embodiment, performing customer group cluster analysis on the basic data set includes:
[0123] Based on the K-Means algorithm, customer group clustering analysis is performed on the basic data set to obtain a high-satisfaction customer group and a customer group prone to complaints.
[0124] In specific implementation: construct customer feature vectors based on multi-dimensional data of customers’ historical service records, satisfaction ratings, and number of complaints; apply cluster analysis algorithm K-Means to perform cluster analysis on customer feature vectors;
[0125] Divide customers into different groups by calculating the similarity or distance between them;
[0126] Based on the clustering results, customer groups with high satisfaction and customer groups prone to complaints are identified to provide more representative features for subsequent model construction.
[0127] In the specific implementation, the K-Means clustering analysis algorithm was used to realize the customer group clustering analysis. The specific steps include:
[0128] Extract multi-dimensional data of customers' historical service records, satisfaction ratings, and number of complaints from the customer service system. Historical service records should include information such as the type, time, and processing results of service requests; satisfaction ratings should include customer satisfaction ratings for each service; and number of complaints should record the total number of complaints made by customers over a period of time. Construct a customer feature vector based on the extracted data. The feature vector should contain key indicators that can reflect customer behavior and service experience, such as the frequency of service request types, average satisfaction ratings, number of complaints, etc. For each customer, convert the corresponding data into a feature vector.
[0129] The feature vectors are standardized to eliminate the dimensional differences between different indicators and make the data comparable.
[0130] The K-Means clustering analysis algorithm is used to perform cluster analysis on customer feature vectors. First, set the number of clusters K, that is, the number of groups into which customers are expected to be divided. Then, randomly select K customers as the initial cluster centers, calculate the similarity or distance between each customer and the cluster center, and divide the customers into the clusters with the closest distance. Next, update the cluster center, calculate the mean of the feature vectors of the customers in each cluster, and use it as the new cluster center. Repeat the above steps until the cluster center no longer changes or the preset number of iterations is reached.
[0131] According to the clustering results, customers are divided into different groups. Each group represents a class of customers with similar characteristics and behavior patterns. By calculating the satisfaction scores and number of complaints of customers in each group, high-satisfaction customer groups and customer groups prone to complaints can be identified.
[0132] Assume that the customer service system contains 1,000 customers, extract their historical service records, satisfaction ratings, and number of complaints, and construct a customer feature vector. The feature vector includes the following indicators: frequency of service request type A, frequency of service request type B, average satisfaction score, and number of complaints.
[0133] Set the number of clusters K to 3, that is, we want to divide customers into 3 groups. Through K-Means cluster analysis, we get the following clustering results:
[0134] Group 1: The frequency of service request type A is high, the frequency of service request type B is low, the average satisfaction score is high, and the number of complaints is low. This type of customers can be identified as a high satisfaction customer group.
[0135] Group 2: The frequency of service request types A and B is moderate, the average satisfaction score is medium, and the number of complaints is high. This type of customer can be regarded as a normal customer group.
[0136] Group 3: The frequency of service request type B is high, the frequency of service request type A is low, the average satisfaction score is low, and the number of complaints is high. This type of customer can be identified as a customer group prone to complaints.
[0137] refer to Figure 4 , step S230, constructing a strategy variable model and a result variable model, using the residuals of the strategy variable model and the result variable model to construct a residual fitting model, integrating the strategy variable model, the result variable model and the residual fitting model to obtain a business failure prediction model.
[0138] In this exemplary embodiment, the construction of the strategy variable model and the outcome variable model includes:
[0139] The strategy variable model takes the strategy variable as an input feature, performs fitting based on the random forest algorithm, and outputs the predicted value of the strategy variable;
[0140] The outcome variable model takes the strategy variables and other covariates as input features, performs fitting based on the random forest algorithm, and outputs the predicted value of the outcome variable.
[0141] During specific implementation: strategy variables include customer service staff allocation strategy, processing priority strategy, automated response strategy and service process strategy; outcome variables include customer satisfaction, complaint rate, resolution time and first resolution rate.
[0142] In specific implementation: a dual machine learning model design is adopted to construct a strategy variable model and an outcome variable model respectively;
[0143] The strategy variable model takes customer service staff allocation strategy, processing priority strategy, automated response strategy and service process strategy as input features, uses the random forest algorithm for fitting, and outputs the predicted values of the strategy variables; the result variable model takes customer satisfaction, complaint rate, resolution time and first resolution rate as output targets, and also uses the random forest algorithm for fitting, combining the influence of strategy variables and other covariates to output the predicted values of the result variables.
[0144] The random forest algorithm is implemented by building multiple decision trees. For each tree, a subset of features and a subset of samples are randomly selected from the training data set for training. The decision tree construction process involves selecting the best features for splitting until the stopping condition is met (such as the tree reaches the maximum depth, the number of samples in the node is less than the minimum number of samples, etc.). For new input data, each tree will give a prediction result.
[0145] The specific design of the random forest algorithm is:
[0146] For a given training dataset where x i is the input feature vector, y i is the output target value;
[0147] Construct multiple decision trees, and each tree is trained by randomly selecting a feature subset and a sample subset to obtain a different tree model T 1 ,T 2 ,…,T k ;
[0148] For new input data x, the prediction result of random forest is the average or majority vote result of all decision tree prediction results;
[0149] In the policy variable model, the input feature vector x iIt includes customer service staff allocation strategy, processing priority strategy, automated response strategy and service process strategy, and the output is the predicted value of the strategy variable;
[0150] In the outcome variable model, the input feature vector includes not only the strategic variables but also other covariates, and the output is the predicted values of customer satisfaction, complaint rate, resolution time, and first resolution rate.
[0151] In this exemplary embodiment, the residual fitting model is constructed by using the residuals of the strategy variable model and the outcome variable model, including:
[0152] Based on the residuals of the strategy variable model and the outcome variable model, the causal effect of the strategy variable on the outcome variable is estimated based on a linear regression algorithm.
[0153] In specific implementation, the specific design of the linear regression algorithm is:
[0154] y=β 0 +β 1 x 1 +β 2 x 2 +…+β p x p +∈;
[0155] Among them, y is the residual target value, x 1 ,x 2 ,…,x p is the residual eigenvalue, β 0 ,β 1 ,…,β p is the regression coefficient, and ∈ is the error term.
[0156] The purpose of model training is to estimate the regression coefficient β 0 ,β 1 ,…,β p , so that the error between the predicted value and the true value is minimized; when initializing the model parameters, for each feature x in the linear regression model i The corresponding parameter β i Assign an initial value, which is set to a random number close to 0, and the intercept term β 0 An initial value is also given; these initial parameter values will serve as the starting point for model training and will be continuously optimized and adjusted through subsequent training processes.
[0157] Step S240: training the service failure prediction model based on the training data set, optimizing the hyperparameters of the service failure prediction model based on a K-fold cross-validation method, and obtaining a trained service failure prediction model.
[0158] In this exemplary embodiment, the service failure prediction model is trained based on the training data set, and the hyperparameters of the service failure prediction model are optimized based on the K-fold cross validation method to obtain the trained service failure prediction model, including:
[0159] Step a: Divide the training data set into K subsets;
[0160] Step b: For each cross-validation, select K minus one subset as the training set and the remaining subset as the validation set;
[0161] Step c: training the service failure prediction model based on the training set to obtain a trained service failure prediction model;
[0162] Step d: verifying the trained service failure prediction model based on the verification set to obtain a verification error of the current fold;
[0163] Step e: Repeat steps b to d for a total of K times, and select a different validation set for each cross-validation;
[0164] Step f: Determine the average validation error of K-fold cross validation;
[0165] Step g: Select a model or a hyperparameter combination based on the average verification error as the trained business failure prediction model.
[0166] refer to Figure 5 , when implemented: Step a: randomly divide the training data set D into K equal or approximately equal subsets D 1 ,D 2 ,…,D k , ensuring that each subset contains roughly the same number of data points and that the distribution characteristics of the data points remain consistent;
[0167] Step b: For each cross-validation, select K-1 subsets as training sets, i.e. D train =DD k , where k is the index of the current validation set, k∈{1,2,...,K}, and the remaining subset D k As a validation set;
[0168] Step c: Use training set D train The business fault prediction model is trained to obtain model parameters;
[0169] Step d: Apply the trained model to the validation set D k , the mean square error algorithm is used to calculate the error between the predicted value and the true value as the verification error of the current fold;
[0170] Step e: Repeat steps b to d for a total of K times, selecting a different validation set each time, ensuring that each subset has been used as a validation set once;
[0171] Step f: Calculate the average validation error of K-fold cross validation, the formula is Where E k is the validation error of the kth fold; the average validation error E avg As an evaluation indicator of model performance, it is used to compare the performance of different models or hyperparameter combinations;
[0172] Step g: Select the model or hyperparameter combination with the best performance according to the average validation error as the final business failure prediction model.
[0173] Assume that there is a training data set D containing 1000 records, and a 5-fold cross validation (K=5) is planned. The following describes the implementation steps of the above training method in detail with a specific example:
[0174] Step a: First, the dataset D needs to be randomly divided into 5 equal or approximately equal subsets, that is, each subset contains about 200 records. To ensure the consistency of the distribution characteristics of the data points, a stratified random sampling method can be used to stratify according to a key feature in the dataset (such as customer type, fault type, etc.), and then perform random sampling in each layer to ensure that the proportion of data points in each category in each subset is close to that of the original dataset.
[0175] Step b: For each cross-validation, select 4 subsets as training sets and the remaining subset as validation set. For example, in the first cross-validation, select D 2 ,D 3 ,D 4 ,D 5 As the training set, D 1 As a validation set. Then, the business failure prediction model is trained on the training set using the selected random forest algorithm. During the training process, the algorithm learns the model parameters based on the data in the training set to minimize the prediction error.
[0176] Step c: During the training process, the parameters of the random forest algorithm, such as the number of trees, tree depth, and learning rate, must be set first. Then, the data in the training set is input into the algorithm, which builds multiple decision trees and obtains the final prediction model through ensemble learning. Each tree predicts the output target (such as customer satisfaction, complaint rate, etc.) based on the input features (such as customer service staff allocation strategy, processing priority strategy, etc.).
[0177] Step d: After training is completed, the trained model is applied to the validation set to calculate the error between the predicted value and the true value. Using the mean square error (MSE) as the error metric, the square of the difference between the predicted value and the true value of each data point in the validation set needs to be calculated, and then the average is calculated to obtain the validation error of the current fold.
[0178] Step e: Repeat the above steps 5 times, choosing a different validation set each time. For example, in the second cross-validation, choose D 1 ,D 3 ,D 4 ,D 5 As the training set, D 2 As the validation set; in the third cross-validation, select D 1 ,D 2 ,D 4 ,D 5 As the training set, D 3 As a validation set, and so on.
[0179] Step f: Calculate the average validation error of the 5-fold cross validation. The formula is Where E k is the validation error of the kth fold. Assuming that the validation errors of the 5 folds are 0.1, 0.12, 0.11, 0.09, and 0.13 respectively, then the average validation error is The average validation error is used as the evaluation indicator of model performance to compare the performance of different models or hyperparameter combinations.
[0180] Step g: Select the model or hyperparameter combination with the best performance based on the average validation error, and select the hyperparameter combination with the smallest average validation error as the final business failure prediction model.
[0181] Step S250: Acquire real-time customer service business data from the customer service system, and predict the real-time customer service business data based on the trained business failure prediction model to obtain a service failure risk identification result.
[0182] During specific implementation: Based on the output of the business failure prediction model, identify service requests or customer groups with potential service failure risks; for the identified risk points, analyze how adjustments to different strategy variables affect the outcome variables.
[0183] The following describes the implementation in detail with reference to specific examples:
[0184] First, use the trained dual machine learning model to predict real-time customer service business data. Assume that the model outputs the following prediction results:
[0185] Customer satisfaction forecast: 85% (lower than the preset threshold of 90%).
[0186] Complaint rate forecast: 3% (higher than the preset threshold of 2%).
[0187] Estimated resolution time: 48 hours (longer than the preset target of 40 hours).
[0188] Predicted first-time resolution rate: 70% (lower than the preset target of 80%).
[0189] At the same time, the model also provides estimates of the causal effects of the strategy variables on the outcome variables, such as:
[0190] The causal effect of customer service staff allocation strategy on customer satisfaction: -0.05 (indicating that the current strategy has a negative impact on customer satisfaction).
[0191] The causal effect of the processing priority strategy on the complaint rate: 0.02 (indicating that the current strategy has a positive impact on the complaint rate, but the contribution is small).
[0192] The causal effect of the automated response strategy on the resolution time is -0.10 (indicating that the current strategy has a significant negative impact on the resolution time).
[0193] The causal effect of service process strategy on first-time resolution rate is 0.08 (indicating that the current strategy has a positive impact on the first-time resolution rate, but there is limited room for improvement).
[0194] Step S260: construct a fault prevention strategy based on the service fault risk identification result, and apply the fault prevention strategy to the customer service system.
[0195] In this exemplary embodiment, the constructing of a fault prevention strategy based on the service fault risk identification result includes:
[0196] Based on the service failure risk identification results, adjust customer service staff allocation, optimize processing priorities, improve automated response strategies, and adjust service processes;
[0197] After applying the fault prevention strategy to the customer service system, the method further includes:
[0198] Monitoring the effectiveness of the implementation of the fault prevention strategy;
[0199] The fault prevention strategy is adjusted based on the implementation effect.
[0200] During implementation: Based on the analysis results, formulate targeted fault prevention strategies, including adjusting customer service staff allocation to match customer needs, optimizing processing priorities to reduce waiting time, improving automated response strategies to improve processing efficiency, and adjusting service processes to eliminate inefficient bottlenecks.
[0201] During specific implementation: continuously monitor the implementation effect of the strategy, and adjust the formulated strategy in a timely manner according to the monitoring results to continuously optimize the fault prevention effect.
[0202] refer to Figure 6 The following describes the implementation in detail with reference to specific examples:
[0203] Based on the prediction results, the following potential service failure risk points can be identified:
[0204] Low customer satisfaction: Current strategies result in customer satisfaction levels falling below a preset threshold and require priority improvement.
[0205] High complaint rate: Although the processing priority strategy has a positive impact on the complaint rate, the overall complaint rate is still higher than the preset threshold and needs further optimization.
[0206] Long resolution time: Automated response strategies have a significant negative impact on resolution time and are the main cause of long resolution time.
[0207] Low first-time resolution rate: Although the service process strategy has a positive impact on the first-time resolution rate, the first-time resolution rate is still lower than the preset target and there is room for improvement.
[0208] Develop the following targeted fault prevention strategies based on the identified risk points:
[0209] Adjust customer service staff allocation strategy: Analyze the workload and professional skills of current customer service staff and reallocate staff to ensure efficient processing of customer requests. For example, assign experienced customer service staff to handle high-priority or complex requests to improve customer satisfaction.
[0210] Optimize the processing priority strategy: Dynamically adjust the processing priority according to the urgency and importance of the request. For example, introduce an intelligent sorting algorithm to automatically adjust the priority according to factors such as request type, customer level and waiting time to reduce the complaint rate.
[0211] Improve automated response strategies: Upgrade and optimize the automated response system to increase the proportion and accuracy of automated processing. For example, introduce natural language processing technology to enhance the intelligence level of automated response, reduce manual intervention and shorten resolution time.
[0212] Optimize service process strategy: sort out and optimize service processes to eliminate unnecessary links and bottlenecks. For example, simplify customer feedback processes, improve internal collaboration efficiency, etc., to increase the first-time resolution rate.
[0213] Encode the developed prevention strategies into executable rules or algorithms and integrate them into the decision support module of the customer service system. By real-time monitoring of business data and strategy execution effects, timely adjust strategies to adapt to business changes. For example, regularly evaluate the impact of strategies on indicators such as customer satisfaction, complaint rate, resolution time, and first resolution rate, and fine-tune the strategies based on the evaluation results.
[0214] From the above, it can be seen that the customer service business failure prevention method based on data mining provided by the embodiment of the present disclosure first collects historical service record data from the customer service system, and forms a basic data set after cleaning and preprocessing. Next, the basic data set is subjected to association analysis and cluster analysis to form a training data set. Then, a dual machine learning model is constructed to fit the strategy variables and the result variables respectively, and the residual is used to construct a fitting model to accurately predict the failure risk. By training and optimizing the model, real-time data is used for prediction, potential service failure risks are identified, and targeted prevention strategies are formulated. Finally, the strategy is applied to the actual customer service system to monitor its effective execution.
[0215] This disclosure optimizes customer service resource allocation and improves customer satisfaction and loyalty through accurate prediction and strategy formulation.
[0216] The present disclosure analyzes the association between the policy variable, namely the service request, and the result variable, namely the customer complaint, thereby improving the prediction accuracy and pertinence of customer service business failures.
[0217] The present disclosure also conducts customer group cluster analysis and merges these analysis results with the basic data set to form a more comprehensive training data set, which helps the model to more accurately understand the needs and expectations of different customer groups for services, thereby improving the accuracy and pertinence of predictions.
[0218] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.
[0219] It should be noted that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0220] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a customer service business failure prevention device based on data mining.
[0221] refer to Figure 7 , which is a structural diagram of a customer service business failure prevention device based on data mining provided by an exemplary embodiment of the present disclosure.
[0222] The customer service business failure prevention device based on data mining includes the following modules:
[0223] A basic data set acquisition module 710 is configured to acquire historical service record data from a customer service system, clean and pre-process the historical service record data, and obtain a basic data set;
[0224] The training data set acquisition module 720 is configured to perform service request and customer complaint association analysis and customer group cluster analysis on the basic data set to obtain analysis results, and merge the basic data set and the analysis results to obtain a training data set;
[0225] The service failure prediction model construction module 730 is configured to construct a policy variable model and a result variable model, use the residuals of the policy variable model and the result variable model to construct a residual fitting model, and integrate the policy variable model, the result variable model and the residual fitting model to obtain a service failure prediction model;
[0226] The service failure prediction model training module 740 is configured to train the service failure prediction model based on the training data set, optimize the hyperparameters of the service failure prediction model based on the K-fold cross validation method, and obtain a trained service failure prediction model;
[0227] The service failure prediction module 750 is configured to obtain real-time customer service business data from the customer service system, predict the real-time customer service business data based on the trained service failure prediction model, and obtain a service failure risk identification result;
[0228] The fault prevention strategy construction and application module 760 is configured to construct a fault prevention strategy based on the service fault risk identification result, and apply the fault prevention strategy to the customer service system.
[0229] In some exemplary embodiments, the basic data set acquisition module 710 is specifically configured to:
[0230] Determining missing values and noise data in the historical service record data;
[0231] Filling the missing values based on a linear interpolation algorithm;
[0232] For the noise data, data denoising is performed based on a sliding average filtering algorithm;
[0233] Preprocessing the historical service record data includes:
[0234] Dividing the historical service record data into continuous data and categorized data;
[0235] Performing standardization on the continuous data;
[0236] The categorized data is subjected to one-hot encoding conversion processing.
[0237] In some exemplary embodiments, the training data set acquisition module 720 is specifically configured to:
[0238] Based on the association rule mining algorithm, the basic data set is analyzed for association between service requests and customer complaints to obtain the service request type corresponding to the customer complaint;
[0239] Performing customer group clustering analysis on the basic data set, including:
[0240] Based on the K-Means algorithm, customer group clustering analysis is performed on the basic data set to obtain a high-satisfaction customer group and a customer group prone to complaints.
[0241] In some exemplary embodiments, the service failure prediction model building module 730 is specifically configured to:
[0242] The strategy variable model takes the strategy variable as an input feature, performs fitting based on the random forest algorithm, and outputs the predicted value of the strategy variable;
[0243] The outcome variable model uses the strategy variables and other covariates as input features, performs fitting based on the random forest algorithm, and outputs the predicted value of the outcome variable;
[0244] Based on the residuals of the strategy variable model and the outcome variable model, the causal effect of the strategy variable on the outcome variable is estimated based on a linear regression algorithm.
[0245] In some exemplary embodiments, the service failure prediction model training module 740 is specifically configured to:
[0246] Step a: Divide the training data set into K subsets;
[0247] Step b: For each cross-validation, select K minus one subset as the training set and the remaining subset as the validation set;
[0248] Step c: training the service failure prediction model based on the training set to obtain a trained service failure prediction model;
[0249] Step d: verifying the trained service failure prediction model based on the verification set to obtain a verification error of the current fold;
[0250] Step e: Repeat steps b to d for a total of K times, and select a different validation set for each cross-validation;
[0251] Step f: Determine the average validation error of K-fold cross validation;
[0252] Step g: Select a model or a hyperparameter combination based on the average verification error as the trained business failure prediction model.
[0253] In some exemplary embodiments, the fault prevention strategy construction and application module 760 is specifically configured to:
[0254] Based on the service failure risk identification results, adjust customer service staff allocation, optimize processing priorities, improve automated response strategies, and adjust service processes;
[0255] The fault prevention strategy construction and application module 760 is further configured to:
[0256] Monitoring the effectiveness of the implementation of the fault prevention strategy;
[0257] The fault prevention strategy is adjusted based on the implementation effect.
[0258] For the convenience of description, the above device is described by dividing it into various modules according to its functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0259] The device of the above embodiment is used to implement the corresponding customer service business failure prevention method based on data mining in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0260] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the customer service business fault prevention method based on data mining described in any of the above embodiments is implemented.
[0261] Figure 8A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.
[0262] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0263] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0264] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0265] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0266] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0267] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0268] The electronic device of the above embodiment is used to implement the corresponding customer service business failure prevention method based on data mining in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0269] The memory 1020 stores machine-readable instructions executable by the processor 1010. When the electronic device is running, the processor 1010 communicates with the memory 1020 via the bus 1030, so that the processor 1010 executes the following instructions when running:
[0270] Acquire historical service record data from the customer service system, clean and preprocess the historical service record data, and obtain a basic data set;
[0271] Performing service request and customer complaint correlation analysis and customer group cluster analysis on the basic data set to obtain analysis results, and merging the basic data set and the analysis results to obtain a training data set;
[0272] Constructing a strategy variable model and a result variable model, using the residuals of the strategy variable model and the result variable model to construct a residual fitting model, integrating the strategy variable model, the result variable model and the residual fitting model to obtain a business failure prediction model;
[0273] The service failure prediction model is trained based on the training data set, and the hyperparameters of the service failure prediction model are optimized based on the K-fold cross validation method to obtain a trained service failure prediction model;
[0274] Acquire real-time customer service business data from the customer service system, and predict the real-time customer service business data based on the trained business failure prediction model to obtain a service failure risk identification result;
[0275] Based on the service failure risk identification result, a failure prevention strategy is constructed, and the failure prevention strategy is applied to the customer service system.
[0276] In a possible implementation manner, the instructions executed by the processor 1010 to clean the historical service record data include:
[0277] Determining missing values and noise data in the historical service record data;
[0278] Filling the missing values based on a linear interpolation algorithm;
[0279] For the noise data, data denoising is performed based on a sliding average filtering algorithm;
[0280] Preprocessing the historical service record data includes:
[0281] Dividing the historical service record data into continuous data and categorized data;
[0282] Performing standardization on the continuous data;
[0283] The categorized data is subjected to one-hot encoding conversion processing.
[0284] In a possible implementation manner, the instructions executed by the processor 1010 include performing a correlation analysis between service requests and customer complaints on the basic data set, including:
[0285] Based on the association rule mining algorithm, the basic data set is analyzed for association between service requests and customer complaints to obtain the service request type corresponding to the customer complaint;
[0286] Performing customer group clustering analysis on the basic data set, including:
[0287] Based on the K-Means algorithm, customer group clustering analysis is performed on the basic data set to obtain a high-satisfaction customer group and a customer group prone to complaints.
[0288] In a possible implementation manner, in the instructions executed by the processor 1010, the step of constructing the strategy variable model and the result variable model includes:
[0289] The strategy variable model takes the strategy variable as an input feature, performs fitting based on the random forest algorithm, and outputs the predicted value of the strategy variable;
[0290] The outcome variable model uses the strategy variables and other covariates as input features, performs fitting based on the random forest algorithm, and outputs the predicted value of the outcome variable;
[0291] The method of constructing a residual fitting model by using the residuals of the strategy variable model and the outcome variable model includes:
[0292] Based on the residuals of the strategy variable model and the outcome variable model, the causal effect of the strategy variable on the outcome variable is estimated based on a linear regression algorithm.
[0293] In a possible implementation manner, in the instructions executed by the processor 1010, the training of the service fault prediction model based on the training data set, optimizing the hyperparameters of the service fault prediction model based on the K-fold cross validation method, and obtaining the trained service fault prediction model include:
[0294] Step a: Divide the training data set into K subsets;
[0295] Step b: For each cross-validation, select K minus one subset as the training set and the remaining subset as the validation set;
[0296] Step c: training the service failure prediction model based on the training set to obtain a trained service failure prediction model;
[0297] Step d: verifying the trained service failure prediction model based on the verification set to obtain a verification error of the current fold;
[0298] Step e: Repeat steps b to d for a total of K times, and select a different validation set for each cross-validation;
[0299] Step f: Determine the average validation error of K-fold cross validation;
[0300] Step g: Select a model or a hyperparameter combination based on the average verification error as the trained business failure prediction model.
[0301] In a possible implementation manner, in the instructions executed by the processor 1010, the constructing a fault prevention strategy based on the service fault risk identification result includes:
[0302] Based on the service failure risk identification results, adjust customer service staff allocation, optimize processing priorities, improve automated response strategies, and adjust service processes;
[0303] After applying the fault prevention strategy to the customer service system, the method further includes:
[0304] Monitoring the effectiveness of the implementation of the fault prevention strategy;
[0305] The fault prevention strategy is adjusted based on the implementation effect.
[0306] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the customer service business fault prevention method based on data mining as described in any of the above embodiments.
[0307] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0308] The above-mentioned non-transitory computer-readable storage medium can be any available medium or data storage device that can be accessed by a computer, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)), etc.
[0309] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the customer service business failure prevention method based on data mining as described in any embodiment in the above exemplary method part, and have the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0310] Based on the same inventive concept, corresponding to the customer service business fault prevention method based on data mining described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer so that the computer and / or the processor executes the customer service business fault prevention method based on data mining. Corresponding to the execution subject corresponding to each step in each embodiment of the customer service business fault prevention method based on data mining, the processor that executes the corresponding step may belong to the corresponding execution subject.
[0311] The computer program product of the above embodiment is used to enable the computer and / or the processor to execute the customer service business fault prevention method based on data mining as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0312] Those skilled in the art will appreciate that the embodiments of the present disclosure may be implemented as a system, method, or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." In addition, in some embodiments, the present disclosure may also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.
[0313] Any combination of one or more computer-readable media may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive examples) of computer-readable storage media may include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
[0314] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, which carry computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0315] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0316] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0317] It should be understood that each box in the flowchart and / or block diagram and the combination of boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine, and these computer program instructions are executed by a computer or other programmable data processing device to produce a device that implements the functions / operations specified in the boxes in the flowchart and / or block diagram.
[0318] These computer program instructions may also be stored in a computer-readable medium that enables a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable medium produce a product that includes an instruction device that implements the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0319] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby enabling the instructions executed on the computer or other programmable device to provide a process for implementing the functions / operations specified in the blocks in the flowchart and / or block diagram.
[0320] In addition, although the operations of the disclosed method are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all the operations shown must be performed to achieve the desired results. On the contrary, the steps depicted in the flow chart can be performed in a different order. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step, and / or one step can be decomposed into multiple steps.
[0321] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0322] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.
[0323] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0324] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present application difficult to understand, the known power supply / ground connection with the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform to be implemented in the embodiments of the present application (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present application, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0325] Although the present application has been described in conjunction with specific embodiments of the present application, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0326] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present application.
[0327] Although the spirit and principle of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims. The scope of the attached claims conforms to the broadest interpretation, thereby including all such modifications and equivalent structures and functions.
Claims
1. A customer service business failure prevention method based on data mining, characterized in that: include: Acquire historical service record data from the customer service system, clean and preprocess the historical service record data, and obtain a basic data set; Performing service request and customer complaint correlation analysis and customer group cluster analysis on the basic data set to obtain analysis results, and merging the basic data set and the analysis results to obtain a training data set; Constructing a strategy variable model and a result variable model, using the residuals of the strategy variable model and the result variable model to construct a residual fitting model, integrating the strategy variable model, the result variable model and the residual fitting model to obtain a business failure prediction model; The service failure prediction model is trained based on the training data set, and the hyperparameters of the service failure prediction model are optimized based on the K-fold cross validation method to obtain a trained service failure prediction model; Acquire real-time customer service business data from the customer service system, and predict the real-time customer service business data based on the trained business failure prediction model to obtain a service failure risk identification result; Based on the service failure risk identification result, a failure prevention strategy is constructed, and the failure prevention strategy is applied to the customer service system.
2. The method according to claim 1, characterized in that Cleaning the historical service record data includes: Determining missing values and noise data in the historical service record data; Filling the missing values based on a linear interpolation algorithm; For the noise data, data denoising is performed based on a sliding average filtering algorithm; Preprocessing the historical service record data includes: Dividing the historical service record data into continuous data and categorized data; Performing standardization on the continuous data; The categorized data is subjected to one-hot encoding conversion processing.
3. The method according to claim 1, characterized in that Performing correlation analysis between service requests and customer complaints on the basic data set, including: Based on the association rule mining algorithm, the basic data set is analyzed for association between service requests and customer complaints to obtain the service request type corresponding to the customer complaint; Performing customer group clustering analysis on the basic data set, including: Based on the K-Means algorithm, customer group clustering analysis is performed on the basic data set to obtain a high-satisfaction customer group and a customer group prone to complaints.
4. The method according to claim 1, characterized in that: The constructing of the strategy variable model and the result variable model includes: The strategy variable model takes the strategy variable as an input feature, performs fitting based on the random forest algorithm, and outputs the predicted value of the strategy variable; The outcome variable model uses the strategy variables and other covariates as input features, performs fitting based on the random forest algorithm, and outputs the predicted value of the outcome variable; The method of constructing a residual fitting model by using the residuals of the strategy variable model and the outcome variable model includes: Based on the residuals of the strategy variable model and the outcome variable model, the causal effect of the strategy variable on the outcome variable is estimated based on a linear regression algorithm.
5. The method according to claim 1, characterized in that The training of the service failure prediction model based on the training data set and the optimization of the hyperparameters of the service failure prediction model based on the K-fold cross validation method to obtain the trained service failure prediction model include: Step a: Divide the training data set into K subsets; Step b: For each cross-validation, select K minus one subset as the training set and the remaining subset as the validation set; Step c: training the service failure prediction model based on the training set to obtain a trained service failure prediction model; Step d: verifying the trained service failure prediction model based on the verification set to obtain a verification error of the current fold; Step e: Repeat steps b to d for a total of K times, and select a different validation set for each cross-validation; Step f: Determine the average validation error of K-fold cross validation; Step g: Select a model or a hyperparameter combination based on the average verification error as the trained business failure prediction model.
6. The method according to claim 1, characterized in that The constructing a fault prevention strategy based on the service fault risk identification result includes: Based on the service failure risk identification results, adjust customer service staff allocation, optimize processing priorities, improve automated response strategies, and adjust service processes; After applying the fault prevention strategy to the customer service system, the method further includes: Monitoring the effectiveness of the implementation of the fault prevention strategy; The fault prevention strategy is adjusted based on the implementation effect.
7. A customer service business failure prevention device based on data mining, characterized in that: include: A basic data set acquisition module is configured to acquire historical service record data from a customer service system, clean and pre-process the historical service record data, and obtain a basic data set; A training data set acquisition module is configured to perform service request and customer complaint association analysis and customer group cluster analysis on the basic data set to obtain analysis results, and merge the basic data set with the analysis results to obtain a training data set; A service failure prediction model building module is configured to build a policy variable model and a result variable model, use the residuals of the policy variable model and the result variable model to build a residual fitting model, and integrate the policy variable model, the result variable model and the residual fitting model to obtain a service failure prediction model; A service failure prediction model training module is configured to train the service failure prediction model based on the training data set, optimize the hyperparameters of the service failure prediction model based on the K-fold cross-validation method, and obtain a trained service failure prediction model; A service failure prediction module is configured to obtain real-time customer service business data from the customer service system, predict the real-time customer service business data based on the trained service failure prediction model, and obtain a service failure risk identification result; The fault prevention strategy construction and application module is configured to construct a fault prevention strategy based on the service fault risk identification result and apply the fault prevention strategy to the customer service system.
8. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 6.