Customer Service Risk Assessment and Early Warning Method and System Based on Historical Data

By using a two-layer decision-making framework and HMDP model based on historical data in customer service risk assessment, combined with the DQN method, the problem of lack of comprehensiveness, dynamicity and flexibility in risk assessment in the existing technology is solved, and personalized and multi-level risk management services are realized, improving customer service quality and enterprise data security.

CN119089235BActive Publication Date: 2025-05-27BEIJING HUADANG ZHIYUAN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410500222.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-03-21
Filing Date
2024-04-24
Publication Date
2025-05-27
Estimated Expiration
2044-04-24

AI Technical Summary

Technical Problem

The existing technology lacks comprehensiveness, dynamicity and flexibility in customer service risk assessment, cannot provide personalized and multi-level services, and there is a problem of incomplete risk coverage and lag in predefined rules in enterprise data and archive management.

Method used

By establishing a risk assessment model based on historical data, combining a two-layer decision-making framework, a two-level Markov decision-making process (HMDP) model is constructed, a risk assessment algorithm is constructed, and risk assessment and early warning is used to use the deep Q network (DQN) method.

Benefits of technology

It realizes dynamic risk assessment and real-time early warning feedback, provides personalized risk management services, improves the intelligence and automation level of risk management, enhances the level of enterprise data security protection and file management, and improves customer service quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119089235B_ABST
    Figure CN119089235B_ABST
Patent Text Reader

Abstract

The present application discloses a method and system for customer service risk assessment and early warning based on historical data, including: collecting historical data of customer service; constructing a two-level Markov decision process HMDP model; constructing a risk assessment algorithm based on the HMDP model; and using the risk assessment algorithm to analyze the historical data so as to assess and early warn the risk of customer service. The method and system for customer service risk assessment and early warning based on historical data provided by the present application realize dynamic risk assessment and real-time early warning feedback by establishing a risk assessment model based on historical data and combining a two-layer decision-making framework. It solves the problems in the prior art that the risk assessment lacks dynamics and flexibility, and cannot provide personalized and multi-level services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information security technology, and particularly to a method and system for customer service risk assessment and early warning based on historical data. Background Art

[0002] In the modern business environment, the operation of the customer service system is accompanied by various uncertain factors, which may affect the operation of the enterprise and customer satisfaction. Therefore, an effective risk assessment and early warning system is crucial for improving service quality, preventing potential problems, and maintaining customer relationships. Most current risk assessment methods are based on static data analysis and determine risk levels according to historical experience. However, these methods are often insufficient to cope with rapidly changing market conditions and customer behaviors and lack dynamism.

[0003] In enterprise data management, risk assessment and early warning also play an important role. Enterprises need to promptly detect and address various data security risks, such as data leakage, illegal access, malicious tampering, etc., to ensure the security of data assets and business continuity. Traditional data risk control methods mainly rely on rule engines and expert experience, and have limitations such as lagging predefined rules and incomplete risk coverage. In addition, for the file management aspect in the enterprise operation process, most existing file management systems adopt shallow analysis methods such as keyword retrieval and word frequency statistics, and cannot fully explore the value of files. Summary of the Invention

[0004] In view of this, this application provides a method and system for customer service risk assessment and early warning based on historical data. By establishing a risk assessment model based on historical data and combining a two-layer decision-making framework, dynamic risk assessment and real-time early warning feedback in the process of enterprise data and file management are realized. It solves the problems in the prior art that risk assessment lacks comprehensiveness, dynamism, and flexibility, and cannot provide personalized and multi-level services.

[0005] This application provides a method for customer service risk assessment and early warning based on historical data, including:

[0006] Collect historical data of customer service;

[0007] Construct a two-level Markov decision process HMDP model;

[0008] Based on the HMDP model, construct a risk assessment algorithm;

[0009] Use the risk assessment algorithm to analyze the historical data in order to assess and early warn of the risks of customer service.

[0010] Optionally, constructing the HMDP model includes:

[0011] Build the top-level decision-making model of HMDP, which is used to identify and classify different customer service scenarios;

[0012] Build the bottom-level decision-making model of HMDP, which is used to build specific risk assessment and response plans.

[0013] Optionally, building the top-level decision-making model of HMDP includes:

[0014] Classify historical service scenarios;

[0015] According to different classified historical service scenarios, integrate corresponding risk factors, and build the environmental state, state transition probability matrix and reward function;

[0016] Build a decision-making strategy using the greedy algorithm.

[0017] Optionally, building the bottom-level decision-making model of HMDP includes:

[0018] Define the fine-grained actions of the response plan;

[0019] Design reward and punishment criteria for each of the fine-grained actions.

[0020] Optionally, if the risk assessment algorithm is the deep Q-network DQN method, then based on the HMDP model, build a risk assessment algorithm, including:

[0021] Obtain the environmental state S, fine-grained action set A of the response plan, and reward function R defined by the HMDP model;

[0022] Initialize the Q-value table.

[0023] Optionally, using the risk assessment algorithm to analyze the historical data includes:

[0024] Classify the historical data by service scenario, and obtain the environmental state St of the first scenario at time t;

[0025] Based on the Q-value at time t, select the action in the fine-grained action set A of the response plan that can maximize the expected reward;

[0026] Execute the action that can maximize the expected reward, obtain the reward function Rt, and the state St+1 at time t+1;

[0027] According to the reward function Rt and the state St+1 at time t+1, use the Bellman equation to update the Q-value.

[0028] Optionally, before using the risk assessment algorithm to analyze the historical data, the method further includes:

[0029] Obtain a historical data training set, and use the historical data training set to train the DQN model;

[0030] Perform parameter adjustment and DQN model verification, and use grid search to find the best parameter combination;

[0031] Optimize the DQN model.

[0032] Optionally, after analyzing the historical data using the risk assessment algorithm, the method further includes:

[0033] Improve the accuracy of risk prediction through the gradual annealing and random sampling algorithms.

[0034] Optionally, improving the accuracy of risk prediction through the gradual annealing and random sampling algorithms includes:

[0035] Use the simulated annealing technique to optimize the decision-making process of the risk assessment algorithm;

[0036] Use the Monte Carlo simulation method to test the robustness of the risk assessment algorithm in a diverse environment;

[0037] Adjust the parameters of the risk assessment algorithm according to the test results.

[0038] An embodiment of the present invention further provides a customer service risk assessment and early warning system based on historical data, including:

[0039] A collection module for collecting historical data of customer service;

[0040] A construction module for constructing a two-level Markov decision process HMDP model; based on the HMDP model, construct a risk assessment algorithm;

[0041] An analysis module for using the risk assessment algorithm to analyze the historical data in order to assess and early warn of the risks of customer service.

[0042] This application provides a method and system for customer service risk assessment and early warning based on historical data. By establishing a risk assessment model based on historical data and combining a two-layer decision-making framework, dynamic risk assessment and real-time early warning feedback are achieved. Among them, the top-level model is responsible for scenario analysis and macro decision-making of risks, while the lower-level model formulates detailed emergency response plans for specific situations. It can not only provide personalized risk management services, but also improve the intelligence and automation levels of risk management. In terms of enterprise data and file management, it can strengthen the enterprise's data security protection and file management level internally, and improve the customer service quality and user experience externally, providing strong support for the enterprise's intelligent transformation and sustainable development, and ultimately providing a more multi-level, reliable and efficient customer service experience for the enterprise and customers. Brief Description of the Drawings

[0043] To more clearly illustrate the disclosed embodiments in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0044] Figure 1 It is a schematic flowchart of the method for customer service risk assessment and early warning based on historical data provided by the embodiments of this application;

[0045] Figure 2 It is a schematic structural diagram of the customer service risk assessment and early warning system based on historical data provided in an embodiment. Detailed Embodiments

[0046] The following will describe the embodiments of the present disclosure in detail with reference to the drawings.

[0047] It should be clear that the following specific examples illustrate the implementation manners of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0048] Note that the following description relates to various aspects of embodiments within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of the aspects set forth herein can be used to implement an apparatus and / or practice a method. Additionally, this apparatus and / or method can be implemented using other structures and / or functionality in addition to one or more of the aspects set forth herein.

[0049] It should also be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present disclosure. Only the components related to the present disclosure are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0050] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the aspects described herein can be practiced without these specific details.

[0051] Embodiment 1

[0052] An embodiment of the present application provides a customer service risk assessment and early warning system based on historical data, which is deployed in a cloud service cluster. The system improves the awareness and response ability to potential service risks by implementing a hierarchical Markov decision process (HMDP). Using a large amount of historical interaction data, machine learning and statistical modeling techniques are employed to predict and respond to possible service risks, and at the same time, effective emergency response plans are formulated.

[0053] The embodiments of the present application are applicable to customer services such as e-commerce and finance, and can also be applied to fields such as healthcare, transportation, and smart cities to achieve dynamic risk assessment and real-time early warning. These fields also involve a large amount of customer services and decision-making processes.

[0054] Figure 1 FIG. is a schematic flowchart of a customer service risk assessment and early warning method based on historical data provided for an embodiment of the present disclosure. The customer service risk assessment and early warning method based on historical data includes steps S1 to S4.

[0055] S1. Collect historical data of customer service;

[0056] In order to build a reliable risk assessment and early warning system, it is first necessary to collect and preprocess the historical data of customer service to establish a good data foundation.

[0057] Specifically, first, extract customer interaction history data from the service platform database. This service platform database can be an e-commerce data platform, a financial data platform, or other data platforms related to customer service. Second, access the database of the service platform and select the corresponding data export mode to ensure data integrity and consistency. Finally, extract the customer's interaction records, query logs, and service results, including timestamps, frequencies, durations, customer feedback, etc. In addition, for various types of data, the correlation of various types of data can be pre-analyzed to identify the data types and features that are most valuable for risk assessment. Thus, other irrelevant types can be selected or discarded.

[0058] After collecting historical data, data cleaning techniques can also be used to remove noise and handle missing values. For example, apply a data cleaning program to remove duplicate records, incorrect data entries, and irrelevant data to increase the availability of the data; handle missing data, such as using an interpolation algorithm to fill in the blanks or removing records with a large amount of missing data; perform data normalization, such as normalizing numerical features and encoding categorical features, to reduce the impact of data dimensions; finally, optimize the features through data transformation to adapt to the requirements of the Markov decision process. For example, use dimensionality reduction techniques, such as principal component analysis (PCA), to extract the most informative features.

[0059] S1 is the basic work for the establishment of the subsequent risk assessment and early warning system. High-quality data collection and preprocessing directly affect the accuracy and efficiency of the model. Through thorough data preprocessing, ensure that the Markov decision process can run based on a clean and rich dataset, which will provide reliable input for the final risk assessment and enable the early warning system to respond quickly and accurately to potential service risks.

[0060] S2. Construct a two-level Hierarchical Markov Decision Process (HMDP) model;

[0061] The core of this step is to establish a hierarchical decision-making model, namely the Hierarchical Markov Decision Process (HMDP), for complex risk assessment and emergency planning.

[0062] The two-level Hierarchical Markov Decision Process (HMDP) model is a hierarchical decision-making model, including the top layer and the bottom layer. The top layer is responsible for macro risk assessment and strategy guidance; the bottom layer focuses on specific operations and response measures.

[0063] The HMDP model decomposes the problem into two levels: the top layer and the bottom layer:

[0064] Top-level model: Responsible for macroscopically evaluating different customer service scenarios and situations, and deciding on the direction for further detailed analysis or the macro strategies to be implemented. For example, the top level may decide whether there is a high service risk for a certain customer group.

[0065] Bottom-level model: Under the guidance of the direction or strategy specified by the top level, the bottom-level model enters specific operations, formulates specific service decisions and response measures. This part usually requires more detailed risk assessment and more specific response actions. For example, after identifying a customer group with high risk, the bottom-level model works on how to reduce this risk, perhaps through actions such as sending coupons, providing additional services, etc.

[0066] The construction of the HMDP model described in S2 includes the following steps S21 - S22:

[0067] S21. Construct the top-level decision-making model of HMDP, which is used to identify and classify different customer service scenarios;

[0068] Specifically, the construction of the top-level decision-making model of HMDP in S21 includes steps S211 - S213:

[0069] S211. Classify historical service scenarios;

[0070] Specifically, use clustering algorithms such as K-Means, hierarchical clustering, or DBSCAN to process historical service data, automatically divide the data into several categories, and each category represents a service scenario.

[0071] Evaluate and adjust the clustering results based on business logic to ensure that each classification conforms to the actual service scenario.

[0072] Taking K-Means as an example, the process of classifying historical service scenarios using the K-Means clustering algorithm includes:

[0073] A1. Data preprocessing

[0074] A11. Clean the data, remove outliers and incorrect fields.

[0075] A12. Standardize numerical data to avoid excessive influence of certain features on distance calculation due to dimension.

[0076] A13. Encode non-numerical data, such as using one-hot encoding.

[0077] A2. Determine the value of K

[0078] A21. Determine the optimal number of clusters K using the Elbow Method. This method calculates the sum of squared errors (SSE) corresponding to different values of K and selects the point where the rapid decline of SSE slows down as the value of K.

[0079] A22. The Silhouette Method can be used as an auxiliary to verify the rationality of the value of K.

[0080] A3. Perform K-Means clustering

[0081] A31. Randomly select K cluster centers.

[0082] A32. Assign each data point to the nearest cluster center to form K clusters.

[0083] A33. Calculate the average value of each cluster and use it as the new cluster center.

[0084] A34. Repeat steps A32 - A33 until the cluster centers are stable or the maximum number of iterations is reached.

[0085] A4. Evaluation and optimization of clustering results

[0086] A41. Evaluate the clustering effect: Use indicators such as the Silhouette Coefficient or Cohesion / Separation within and between classes for evaluation.

[0087] A42. Business logic verification: Collaborate with the business team to confirm whether the clustering results conform to business understanding.

[0088] A43. Adjust the model parameters or re - define features according to the feedback from the business team to achieve the best business application effect.

[0089] S212. Integrate the corresponding risk factors according to different historical service scenarios after classification, and construct the environmental state, state transition probability matrix, and reward function;

[0090] Specifically, S212 includes:

[0091] Determine the state space: Convert the clustering results into the state space in the Markov decision process, where each service scenario corresponds to one or more states. Assume that the customer service scenarios are divided into four environmental states: initial consultation, further consultation, targeted answer, and service termination. Each state represents a different stage of the customer in the service process.

[0092] Construct the transition probability matrix: Statistically analyze the frequency of transitions between different scenarios, and use this as the transition probability. Specifically, count the number of times a customer transitions from one service state to another in historical data, and estimate the probability based on the transition frequency. For example, the transition probability from the initial consultation to the further consultation is P(C2|C1). In this way, the transition probabilities of different states form the transition probability matrix.

[0093] Definition of the reward function: Determine the reward value based on the business metrics affected by each service scenario, such as customer satisfaction, processing duration, etc. For example, determine the reward value based on the business key performance indicators (KPIs), such as customer satisfaction, quick response time, etc. Set a positive reward for each terminal state (such as C4), and set a smaller negative reward or zero reward for intermediate states (such as C1, C2) to encourage the system to move towards the terminal state.

[0094] The following is a specific example for description:

[0095] The state space includes:

[0096] State "C1": The customer submits a service request and has not received a response yet.

[0097] State "C2": The customer receives a preliminary reply but needs more information.

[0098] State "C3": The customer has received a targeted answer and the problem is partially solved.

[0099] State "C4": The customer's problem is completely solved and the service ends.

[0100] The state transition probability matrix is shown in Table 1, and the sum of each row is 1:

[0101] Table 1

[0102]

[0103] For example, P(C2|C1) = 0.7 means that the probability of transitioning from C1 to C2 is 70%.

[0104] The reward values include:

[0105] For the C4 state (problem solved), a positive reward of +100 is provided.

[0106] For the C3 state, a small positive reward of +10 is obtained (because the problem is developing towards being solved).

[0107] For the C2 state, there is neither reward nor punishment, 0 (remaining neutral and encouraging information collection).

[0108] For the C1 state, due to the delayed service, a small punishment of -10 is given (to encourage quick response).

[0109] S213. Construct a decision-making strategy using the greedy algorithm.

[0110] Apply the greedy strategy in reinforcement learning, and each time select the action that maximizes the reward (or long-term return) in the current state as part of the strategy.

[0111] This strategy may perform well in the initial stage of modeling, but it needs to be optimized over time to obtain the global optimal strategy.

[0112] Using the greedy strategy (Greedy Policy) of reinforcement learning for decision-making in S213 mainly includes the following steps:

[0113] B1. Initialize the state and action value functions

[0114] Suppose there is a Q-table for the action value function, which records the expected returns for taking different actions in each state.

[0115] B2. Select the action strategy

[0116] At each decision point (current state s), select the action a that maximizes Q(s,a) as the current optimal action strategy.

[0117] That is, in state s, look up the Q-values corresponding to all possible actions and select the action with the highest Q-value to execute.

[0118] B3. Update the Q-value table

[0119] Execute the action and observe the result to see if the expected reward is achieved.

[0120] Update the Q-value corresponding to the current state and action according to the obtained immediate reward and the maximum Q-value of the next state to improve the strategy.

[0121] Suppose in a certain customer service state, we need to select the best response action to improve customer satisfaction:

[0122] Current state s: The customer has submitted a request to return a product.

[0123] Possible actions a: A1 Provide a return label, A2 Provide instant online customer service chat, A3 Provide the option to return in-store.

[0124] The Q-table of the action value function is initialized as shown in Table 2:

[0125] Table 2

[0126]

[0127] In this state s, action A2 has the maximum Q value, that is, Q(s, A2) = 8. Therefore, according to the greedy strategy, action A2 (providing instant online customer service chat) is selected as the optimal strategy.

[0128] After performing the action, assume that a new immediate reward of +10 is obtained, and the customer satisfaction score also improves. Then the Q value of action A2 in state s can be updated. If the Bellman equation is used:

[0129] Q(s, A2) = original Q value + learning rate × (immediate reward + discount rate × maximum next Q value - original Q value)

[0130] Assume that the learning rate is 0.5, the discount rate is 0.8, and the maximum Q value of the next state is 12. Then the updated calculation process is:

[0131] Q(s, A2) = 8 + 0.5 × (10 + 0.8 × 12 - 8) = 15

[0132] Finally, the updated Q table according to the greedy strategy is shown in Table 3:

[0133] Table 3

[0134]

[0135] In this way, the greedy strategy will continuously drive the strategy to evolve towards generating higher immediate rewards and long-term returns, and finally form an efficient decision-making strategy.

[0136] S22. Construct the underlying decision-making model of HMDP, which is used to construct specific risk assessment and response plans.

[0137] In step S22, the underlying decision-making model will be implemented, involving the definition of actions related to specific risk assessment and the design of reward and punishment criteria for these actions. This is the underlying part of the HMDP model, aiming to make more detailed operation decisions according to the guidance of the top-level strategy.

[0138] Among them, constructing the underlying decision-making model of HMDP in S22 includes steps S211 - S212:

[0139] S221. Define the fine-grained actions of the response plan;

[0140] Assume that in the scenario of customer service, the following fine-grained actions are defined:

[0141] D1: Provide an instant online answer.

[0142] D2: Send a detailed instruction document.

[0143] D3: Arrange a callback service.

[0144] D4: Submit to the technical support team.

[0145] First, analyze the historical service records to identify common micro - steps in the problem - solving process. Second, communicate with the customer service team to understand which micro - steps can have a positive impact on customer satisfaction. Finally, determine the action set, list all possible service response plan actions, and ensure that each action can point to different service goals or problem - solving methods.

[0146] S222. Design reward and punishment criteria for each of the described fine - grained actions.

[0147] Exemplarily, formulate reward and punishment criteria accordingly:

[0148] D1: If the problem is solved immediately, the reward is +10; if the reply times out, the punishment is - 5.

[0149] D2: If customer satisfaction improves, the reward is +8; if the document is not helpful in solving the problem, the punishment is - 3.

[0150] D3: Timely service can get a +15 reward; failure to contact the customer within the agreed time is punished - 10.

[0151] D4: The technical team gets +20 for solving the problem; if there is no feedback after exceeding the predetermined duration after submission, the punishment is - 15.

[0152] Specifically, S222 includes:

[0153] Determine the reward - punishment mechanism and set corresponding rewards according to the contribution degree of each micro - action to the overall goal.

[0154] Monitor the implementation effect in real - time and monitor and evaluate the results after the actions are executed.

[0155] Customer feedback analysis, customer satisfaction surveys or feedback analysis provide a measure of the effectiveness of the plan actions.

[0156] Dynamically adjust the reward - punishment criteria and adjust the reward - punishment criteria in a timely manner according to real - time data and customer feedback.

[0157] In summary, the constructed HMDP model can effectively grasp and guide the risk decision - making in the customer service process, thus realizing the dynamic assessment and early warning of customer service risks. The top - level model provides a macro - control of the risks in different service scenarios, while the low - level model can formulate refined response strategies for each specific situation.

[0158] S3. Based on the HMDP model, construct a risk assessment algorithm;

[0159] The risk assessment calculation algorithm works under the guidance of this HMDP model. It selects the best action by calculating the expected reward based on the current state and the action options provided by the HMDP, and then updates the model (learns) according to the result of the action, with the expectation of making better risk assessments and early warnings when encountering similar situations in the future.

[0160] Among them, the risk assessment calculation algorithm is a specific algorithm used to perform risk assessment and early warning based on the established hierarchical Markov decision process (HMDP) model. That is, the HMDP model provides the framework for risk assessment, and the risk assessment calculation algorithm is the method of operating within this framework.

[0161] In an embodiment of the present invention, exemplarily, the risk assessment algorithm is the deep Q-network DQN method. Then, in S3, based on the HMDP model, a risk assessment algorithm is constructed, which specifically includes:

[0162] S31. Obtain the environmental state S defined by the HMDP model, the fine-grained action set A of the response plan, and the reward function R;

[0163] Illustrative example:

[0164] Environmental state S: We assume there are four states: new customer consultation (S1), order processing (S2), customer service question (S3), and after-sales service evaluation (S4).

[0165] Fine-grained action set A of the response plan: The action set includes: providing product information (A1), confirming order details (A2), answering service questions (A3), and collecting customer feedback (A4).

[0166] Reward function R: The reward corresponding to each action is determined according to user feedback and business goals. For example, if the state changes from S2 to user satisfaction through action A2, there is a positive reward R_pos; if the service speed is delayed, it is a penalty R_neg.

[0167] S31 specifically includes:

[0168] Define the environmental state S, covering all key customer service nodes.

[0169] List all possible actions A, which represent all possible service response measures.

[0170] Construct the reward function R, and combine business metrics and customer feedback to determine the reward value or penalty value of each action.

[0171] S32. Initialize the Q-value table.

[0172] Create a Q-table to store and update the values of state-action pairs. Each record in this table represents the expected reward for a certain action in a given state.

[0173] Initialize all Q-values to 0 or a small random number to provide a starting point for the learning process.

[0174] Ensure that the update mechanism is ready so that after each interaction, the Q-table can update the values corresponding to the current state and action.

[0175] The Q-values are shown in Table 4:

[0176] Table 4

[0177]

[0178] Through this initialization process, the DQN algorithm will have a starting point to explore the effects of various state-action combinations, learn, and optimize the target policy.

[0179] S4. Use the risk assessment algorithm to analyze the historical data in order to evaluate and warn of the risks of customer service.

[0180] S4 specifically includes steps S41 - S44:

[0181] S41. Classify the historical data by service scenario and obtain the environmental state S of the first scenario at time t t ;

[0182] Automatically classify the service scenarios using a classification algorithm (such as K-Means clustering, see S211); assign a state label to each class of service scenarios so that each scenario corresponds to a state in the HMDP model; at time t, based on the current customer interaction or service request, determine the state S of the current service scenario. t . For example, divide the service interactions in the historical data into multiple scenarios, such as product consultation, order query, complaint handling, etc. At time t, the environmental state can be that the customer is querying an order.

[0183] S42. Based on the Q-values at time t, select the action in the fine-grained action set A of the response plan that can maximize the expected reward;

[0184] Access the Q-value table at time t and check the Q-values of each action in the current state S t . Select the action A t with the highest Q-value as the best response strategy. In state S t , select the action A t with the highest expected reward according to the Q-value table.

[0185] S43. Execute the action that can maximize the expected reward, and obtain the reward function R t , and the state S at time t + 1 t+1 ;

[0186] Actually execute action A t , and monitor the customer's reaction or interaction result. Determine the immediate reward R obtained after executing the action t . Record the change in the customer service state after the action is executed, and identify it as the new state S t+1 . Specifically,

[0187] Execute action A t , such as responding to the customer for order inquiry and providing the required information.

[0188] According to the execution result, obtain the reward R t , for example, a positive reward is obtained if the customer satisfaction improves.

[0189] The observed new state is S t+1 , such as transferring from order inquiry to confirmed receipt status.

[0190] S44. According to the reward function Rt and the state St+1 at time t + 1, use the Bellman equation to update the Q value.

[0191] According to the Bellman equation, the learning rate (α), and the discount rate (γ), update the value of Q(St, At). The new Q value is based on the old Q value, combined with the actually obtained reward Rt and the Q value expectation of the best action for the next state S t+1 .

[0192] In addition, optionally, for different application scenarios, the embodiments of the present application can also use other reinforcement learning methods such as Proximal Policy Optimization (PPO) and Asynchronous Advantage Actor-Critic (A3C) for risk assessment. Exemplarily:

[0193] In the inventory and order flow management during the promotion period, in order to cope with the highly volatile demand, PPO can be used to replace the original DQN algorithm:

[0194] Use PPO for dynamic inventory management. It reduces the change in the learning rate (step size) within an update phase and runs the optimization of multiple data through resampling (importance sampling) of the previous behavior, improving the sample efficiency and stability of the algorithm.

[0195] When the order volume surges, the PPO algorithm not only predicts the risk of inventory shortage but also evaluates the long-term benefits under various promotional strategies, so as to select the optimal replenishment plan and pricing strategy to balance immediate sales and inventory costs.

[0196] Compared with DQN, the PPO algorithm can learn more effectively in the multi-action space of inventory management, quickly adapt to the dynamic changes during the promotion period, and perform policy iteration and optimization before the next promotion cycle.

[0197] In the monitoring and early warning of abnormal purchase behavior, A3C can process and respond faster:

[0198] The A3C algorithm uses multiple worker copies to explore and learn in parallel, quickly collecting diverse enough experiences to train the model, which helps to speed up the training process of the abnormal behavior detection model.

[0199] When an abnormal transaction is detected, the "Actor" component of A3C can give decision suggestions, such as temporarily freezing the order or performing further identity verification; while the "Critic" component evaluates the quality of the decision and quickly updates the policy.

[0200] Compared with DQN, A3C can adapt to newly emerging fraud techniques faster and continuously improve the efficiency of identifying and preventing fraud behavior during continuous transactions.

[0201] In addition, before analyzing the historical data using the risk assessment algorithm, the method further includes:

[0202] Obtain a historical data training set and use the historical data training set to train the DQN model;

[0203] For example, use the customer service interaction data of the past year, including the initial state of each service, the actions taken, the rewards obtained, and the final state of the service. Clean and preprocess the data to ensure data quality. Use this data set to train the DQN model. Overcome the correlation problem between consecutive samples by replaying historical data (Experience Replay), and improve the training efficiency through mini-batch updates.

[0204] Perform parameter adjustment and DQN model verification, and use grid search to find the best parameter combination;

[0205] Use the standard training, validation, and test data set splitting method. Optimize the hyperparameters during the DQN training process using cross-validation and grid search techniques to determine the best parameter combination. Performance evaluation on the validation set helps to select the most suitable model and use it for the final evaluation on the test set.

[0206] Optimize the DQN model.

[0207] Use technical means such as Early Stopping and Regularization to prevent overfitting. Adjust the Exploration-Exploitation strategy, such as using the epsilon-greedy strategy, to ensure a balance between the exploration and exploitation of the model. Monitor metrics such as average reward and moving average, and adjust and optimize them as the training iterations progress. Continuously iterate the training according to the model's performance to achieve the required performance metrics.

[0208] In addition, after analyzing the historical data using the risk assessment algorithm, the method further includes:

[0209] Improve the accuracy of risk prediction through stepwise annealing and random sampling algorithms, specifically including:

[0210] Use simulated annealing technology to optimize the decision-making process of the risk assessment algorithm;

[0211] Simulated annealing technology is a heuristic search algorithm that solves optimization problems by mimicking the annealing process in physical processes. In the optimization of the DQN model, it can help jump out of local optima and search for better solutions in global searches.

[0212] Detailed steps:

[0213] Initialize the temperature and decay rate: Set a starting high temperature and determine a decay rate to gradually reduce the temperature as the iterations progress.

[0214] Select the initial solution: Randomly or strategically select an initial parameter set of the DQN as the current solution.

[0215] Iterative optimization process: In each iteration, make small random changes to the current solution and calculate the cost function of the new solution (such as the increment of the expected reward).

[0216] Accept the new solution according to the probability: If the new solution is better, accept the new solution; if the cost function of the new solution is higher, still accept the new solution with a certain probability, and the probability is related to the temperature and the cost difference.

[0217] Reduce the temperature: As the iterations progress, gradually reduce the temperature to reduce the probability of accepting a poor solution, thus converging to the global optimal solution or an approximate solution.

[0218] Use the Monte Carlo simulation method to test the robustness of the risk assessment algorithm in diverse environments;

[0219] Monte Carlo simulation is a numerical calculation method based on probability statistics. In risk assessment, this method can be used to evaluate the performance of algorithms under different scenarios.

[0220] The detailed steps are as follows:

[0221] Simulation environment construction: Create a variety of simulated customer service environments to reflect the diverse scenarios that may be encountered in reality.

[0222] Run the simulation: Conduct multiple model experiments according to the probability distribution, and estimate the performance of the algorithm through these independent simulation experiments.

[0223] Collect data: Summarize the experimental results, especially focusing on the performance of the algorithm under extreme conditions.

[0224] Robustness evaluation: Evaluate the average performance and robustness of the DQN strategy based on the data, that is, the performance consistency under various environments.

[0225] According to the test results, adjust the parameters of the risk assessment algorithm.

[0226] Analyze the simulation data: Identify the environmental factors where the algorithm performance varies significantly.

[0227] Parameter fine-tuning: Adjust the parameters of DQN according to the simulation results, such as the exploration rate, discount factor, learning rate, etc.

[0228] Repeat the simulation: The adjusted model is subjected to Monte Carlo simulation again to verify the effect after parameter adjustment.

[0229] Continue to optimize: If the performance of the adjusted model is significantly improved, fix these parameters. If the effect is not obvious or the model performance deteriorates, consider other parameter combinations or optimization strategies.

[0230] In summary, the embodiments of the present invention can be applied to various industries. Taking an e-commerce platform as an example:

[0231] 1. Monitoring and warning of abnormal purchase behaviors

[0232] Customer A continuously makes a large number of purchases of high-value goods within a short period of time, and this behavior is very different from his past purchase history. The system immediately identifies the abnormal pattern and triggers the risk assessment algorithm.

[0233] In response to this situation, the system obtains the optimal strategy through DQN, which may be to temporarily freeze this order and send a verification request to Customer A to confirm the authenticity of the order, or to notify the risk management team for manual review.

[0234] 2. Inventory and order flow management during promotions

[0235] Before a large-scale promotion event on e-commerce platform B, the system evaluates the upcoming traffic and potential order volume.

[0236] The underlying decision-making model calculates the possible risk of insufficient inventory and automatically adjusts the predicted inventory threshold. If the risk assessment algorithm predicts that the inventory depletion is too fast, the system automatically places an order for replenishment or adjusts the promotion strategy to prevent out-of-stock situations.

[0237] 3. Monitoring and handling negative feedback in customer service interactions

[0238] Customer C expressed dissatisfaction with the product quality during a chat with the customer service. The underlying decision-making model captured this risk signal and conducted a detailed sentiment analysis on it.

[0239] Based on the evaluation by the DQN algorithm, the system determines the best response strategy: it may be to offer an immediate price discount to Customer C to compensate for their dissatisfaction, or to report this situation to the senior customer service team to provide a more personalized service solution.

[0240] 4. Personalized recommendation and risk prevention

[0241] Through the historical purchase data of customer D, the system learns about the customer's preference for a certain type of product and predicts future purchase trends based on the HMDP model.

[0242] If the early warning system identifies that a certain trend may lead to inventory backlog, it will notify the supply chain in advance to adjust the procurement plan, or push promotion activities to these customers in advance through the personalized recommendation algorithm to balance supply and demand and reduce risks.

[0243] This application provides a method for customer service risk assessment and early warning based on historical data. By establishing a risk assessment model based on historical data and combining a two-layer decision-making framework, it realizes dynamic risk assessment and real-time early warning feedback. Among them, the top-level model is responsible for scenario analysis and macro decision-making of risks, while the lower-level model formulates detailed emergency response plans for specific situations. It can not only provide personalized risk management services, but also improve the intelligence and automation levels of risk management. In terms of the enterprise's data and file management, it can strengthen the enterprise's data security protection and file management level internally, and improve the customer service quality and user experience externally, providing strong support for the enterprise's intelligent transformation and sustainable development, and ultimately providing a more multi-level, reliable and efficient customer service experience for the enterprise and customers.

[0244] Example 2

[0245] Figure 2 This is a schematic structural diagram of a customer service risk assessment and early warning system based on historical data provided by an embodiment of the present disclosure. The customer service risk assessment and early warning system 200 based on historical data includes:

[0246] A collection module 201 for collecting historical data of customer service;

[0247] A construction module 202 for constructing a two-level Markov decision process HMDP model; based on the HMDP model, constructing a risk assessment algorithm;

[0248] An analysis module 203 for using the risk assessment algorithm to analyze the historical data so as to evaluate and give early warnings of the risks of customer service.

[0249] The construction module 201 is used for constructing the HMDP model, specifically including:

[0250] Constructing the top-level decision model of HMDP, where the top-level decision model is used to identify and classify different customer service scenarios;

[0251] Constructing the bottom-level decision model of HMDP, where the bottom-level decision model is used to construct specific risk assessment and response plans.

[0252] The construction module 201 is used for constructing the top-level decision model of HMDP, including:

[0253] Classifying historical service scenarios;

[0254] According to the classified different historical service scenarios, integrating corresponding risk factors, and constructing an environmental state, a state transition probability matrix, and a reward function;

[0255] Constructing a decision-making strategy using the greedy algorithm.

[0256] The construction module 201 is used for constructing the bottom-level decision model of HMDP, including:

[0257] Defining fine-grained actions of the response plan;

[0258] Designing reward and punishment criteria for each of the fine-grained actions.

[0259] Optionally, if the risk assessment algorithm is the deep Q-network DQN method, then based on the HMDP model, constructing the risk assessment algorithm includes:

[0260] Obtaining the environmental state S defined by the HMDP model, the set A of fine-grained actions of the response plan, and the reward function R;

[0261] Initializing the Q-value table.

[0262] Among them, the analysis module 203 is used for using the risk assessment algorithm to analyze the historical data, including:

[0263] Classifying the historical data by service scenarios, and obtaining the environmental state St of the first scenario at time t;

[0264] Select an action in the fine-grained action set A of the response plan that can maximize the expected reward based on the Q value at time t;

[0265] Execute the action that can maximize the expected reward, obtain the reward function Rt, and the state St+1 at time t+1;

[0266] Update the Q value using the Bellman equation based on the reward function Rt and the state St+1 at time t+1.

[0267] In addition, the system further includes:

[0268] A training module for obtaining a historical data training set and training the DQN model using the historical data training set;

[0269] A verification module for performing parameter adjustment and DQN model verification, and using grid search to find the best parameter combination;

[0270] An optimization module for optimizing the DQN model.

[0271] Optionally, the system further includes:

[0272] A prediction module for improving the accuracy of risk prediction through a gradual annealing and random sampling algorithm.

[0273] Optionally, improving the accuracy of risk prediction through a gradual annealing and random sampling algorithm includes:

[0274] Using simulated annealing technology to optimize the decision-making process of the risk assessment algorithm;

[0275] Using the Monte Carlo simulation method to test the robustness of the risk assessment algorithm in a diverse environment;

[0276] Adjust the parameters of the risk assessment algorithm according to the test results.

[0277] This application provides a customer service risk assessment and early warning system based on historical data. By establishing a risk assessment model based on historical data and combining a two-layer decision-making framework, it realizes dynamic risk assessment and real-time early warning feedback. Among them, the top-level model is responsible for scenario analysis and macro decision-making of risks, while the lower-level model formulates detailed emergency response plans for specific situations, which can not only provide personalized risk management services, but also improve the intelligence and automation levels of risk management, and ultimately provide a more multi-level, reliable and efficient customer service experience for enterprises and customers.

[0278] The system of the embodiments of the present disclosure can execute the methods provided by the embodiments of the present disclosure, and their implementation principles are similar. The actions performed by each module in the systems of the embodiments of the present disclosure correspond to the steps in the methods of the embodiments of the present disclosure. For the detailed function descriptions of each module of the system, reference can specifically be made to the descriptions in the corresponding methods shown above, and details will not be repeated here.

[0279] The above are only optional implementation manners of some implementation scenarios of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present disclosure, other similar implementation means based on the technical idea of the present disclosure also fall within the protection scope of the embodiments of the present disclosure.

Claims

1. A customer service risk assessment and early warning method based on historical data, characterized in that: The method comprises: Collect historical data on customer service; Construct a two-level Markov decision process HMDP model; Based on the HMDP model, construct a risk assessment algorithm; Analyzing the historical data using the risk assessment algorithm to assess and warn of risks in customer service; The construction of a two-level Markov decision process HMDP model includes: Constructing a top-level decision model of HMDP, wherein the top-level decision model is used to identify and classify different customer service scenarios; Constructing an underlying decision model of HMDP, wherein the underlying decision model is used to construct a specific risk assessment and response plan; The construction of the underlying decision model of HMDP includes: defining fine-grained actions of the response plan; designing reward and penalty criteria for each of the fine-grained actions; The fine-grained actions for defining response plans include: analyzing historical service records to identify common micro-steps in the problem-solving process; communicating with the customer service team to understand which micro-steps can have a positive impact on customer satisfaction; determining an action set and listing all possible service response plan actions to ensure that each action can point to a different service goal or problem-solving method; The reward and punishment standards are designed for each fine-grained action, including: determining the reward and punishment mechanism, setting corresponding rewards according to the contribution of each micro action to the overall goal; real-time monitoring of the implementation effect, monitoring and evaluating the results after the action is executed; customer feedback analysis, customer satisfaction surveys or feedback analysis to measure the effectiveness of the plan action; dynamic adjustment of reward and punishment standards, timely adjustment of reward and punishment standards according to real-time data and customer feedback.

2. The method according to claim 1, characterized in that Construct the top-level decision model of HMDP, including: Categorize historical service scenarios; According to the classified historical service scenarios, the corresponding risk factors are integrated to construct the environment state, state transition probability matrix and reward function; A greedy approach is used to construct the decision strategy.

3. The method according to claim 1, characterized in that The risk assessment algorithm is a deep Q network DQN method, and based on the HMDP model, a risk assessment algorithm is constructed, including: Obtaining the environment state S, the fine-grained action set A of the response plan, and the reward function R defined by the HMDP model; Initialize the Q value table.

4. The method according to claim 3, characterized in that Analyzing the historical data using the risk assessment algorithm includes: The historical data is classified into service scenarios, and the environment state S of the first scenario at time t is obtained. t; Based on the Q value at time t, select an action that can maximize the expected reward from the fine-grained action set A of the response plan; Execute the action that can maximize the expected reward and obtain the reward function R t , and the state S at time t+1 t+1 ; According to the reward function R t and the state S at time t+1 t+1 , use the Bellman equation to update the Q value.

5. The method according to claim 3, characterized in that: Before analyzing the historical data using the risk assessment algorithm, the method further includes: Obtain a historical data training set, and use the historical data training set to perform model training on the DQN; Perform parameter tuning and DQN model validation, using grid search to find the best parameter combination; The DQN model is optimized.

6. The method according to claim 1, characterized in that After analyzing the historical data using the risk assessment algorithm, the method further includes: The accuracy of risk prediction is improved through stepwise annealing and random sampling algorithms.

7. The method according to claim 6, characterized in that The accuracy of risk prediction is improved through stepwise annealing and random sampling algorithms, including: Using simulated withdrawal technology to optimize the decision-making process of the risk assessment algorithm; Monte Carlo simulation method is used to test the robustness of the risk assessment algorithm in diverse environments; Based on the test results, the parameters of the risk assessment algorithm are adjusted.

8. A customer service risk assessment and early warning system based on historical data, characterized in that: include: Collection module, used to collect historical data of customer service; Building module, used to build a two-level Markov decision process HMDP model; Based on the HMDP model, construct a risk assessment algorithm; An analysis module, used to analyze the historical data using the risk assessment algorithm, so as to assess and warn of risks in customer service; The construction of the double-level Markov decision process HMDP model includes: constructing a top-level decision model of the HMDP, the top-level decision model is used to identify and classify different customer service scenarios; constructing a bottom-level decision model of the HMDP, the bottom-level decision model is used to construct a specific risk assessment and response plan; The construction of the underlying decision model of HMDP includes: defining fine-grained actions of the response plan; designing reward and penalty criteria for each of the fine-grained actions; The fine-grained actions for defining response plans include: analyzing historical service records to identify common micro-steps in the problem-solving process; communicating with the customer service team to understand which micro-steps can have a positive impact on customer satisfaction; determining an action set and listing all possible service response plan actions to ensure that each action can point to a different service goal or problem-solving method; The reward and punishment standards are designed for each fine-grained action, including: determining the reward and punishment mechanism, setting corresponding rewards according to the contribution of each micro action to the overall goal; real-time monitoring of the implementation effect, monitoring and evaluating the results after the action is executed; customer feedback analysis, customer satisfaction surveys or feedback analysis to measure the effectiveness of the plan action; dynamic adjustment of reward and punishment standards, timely adjustment of reward and punishment standards according to real-time data and customer feedback.

Citation Information

Patent Citations

  • Product recommendation method and device, computer equipment and storage medium

    CN113256390A

  • Electric vehicle charging guide strategy method based on hierarchical deep reinforcement learning

    CN114117910A

  • Power distribution network risk dynamic early warning method, system and device and storage medium

    CN114219045A