Intelligent customer service work order distribution method based on machine learning and related equipment

By modeling the work order allocation process in the customer service system as a Markov decision-making process and using the deep Q network model for iterative training, the problem of low work order allocation efficiency in traditional customer service systems is solved, and the intelligent allocation of work orders and the overall performance improvement of the customer service system is achieved.

CN119940773APending Publication Date: 2025-05-06FIBRLINK NETWORKS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411852306.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The work order allocation method in traditional customer service systems relies on a simple rule engine, resulting in insufficient processing efficiency in complex and changeable customer service scenarios.

Method used

Using the intelligent customer service work order allocation method based on machine learning, the work order allocation process is modeled as a Markov decision-making process, the state space and action space are defined, the reward function is designed, and the deep Q network model is built for iterative training to optimize the work order allocation strategy.

Benefits of technology

It realizes the automation and intelligent distribution of work orders, improves work order processing efficiency and customer satisfaction, significantly improves the overall performance and user experience of the customer service system, and balances the workload of customer service personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940773A_ABST
    Figure CN119940773A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent customer service work order allocation method based on machine learning and related equipment. Comprising the following steps: modeling a work order allocation process as a Markov decision process; defining a state space and an action space; designing a reward function; the state space comprises a work single vector and a customer service staff vector; the work order vector comprises a work order type vector, a work order emergency degree vector and a work order creation time and deadline vector; the customer service personnel vectors comprise professional skill vectors of customer service personnel, types and quantity vectors of historical processing work orders, current workload vectors of the customer service personnel and online state vectors of the customer service personnel; the action space comprises a decision for distributing work orders to different customer service personnel, setting work order priorities and deciding whether transfer is needed or not; constructing a deep Q network model; acquiring historical data of customer service work order distribution; performing iterative training on the deep Q network model; and deploying the deep Q network model after iteration training to a customer service system, and intelligently distributing the real-time work orders.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent customer service technology, and in particular to an intelligent customer service work order allocation method and related equipment based on machine learning. Background Art

[0002] In modern customer service systems, efficient and accurate allocation of work orders is crucial to improving customer satisfaction and service efficiency. However, traditional work order allocation methods mostly rely on simple rule engines. Although this allocation method has achieved automation to a certain extent, it has the problem of insufficient processing efficiency when facing complex and changing customer service scenarios. Summary of the invention

[0003] In view of this, the purpose of this application is to propose an intelligent customer service work order allocation method and related equipment based on machine learning.

[0004] Based on the above objectives, this application provides an intelligent customer service work order allocation method based on machine learning, including:

[0005] Model the work order allocation process as a Markov decision process; define the state space and action space; design a reward function; wherein the state space includes a work order vector and a customer service staff vector; the work order vector includes a work order type vector, a work order urgency vector, and a work order creation time and deadline vector; the customer service staff vector includes a customer service staff professional skill vector, a historically processed work order type and quantity vector, a customer service staff current workload vector, and a customer service staff online status vector; the action space includes the decision to allocate work orders to different customer service staff, set work order priorities, and decide whether to transfer; the reward function is designed based on the work order processing speed, work order processing quality, customer satisfaction, and the balance of customer service staff workload;

[0006] Build a deep Q network model;

[0007] Acquire historical data of customer service work order allocation, and preprocess the historical data to obtain a training data set; the historical data includes a historical state space vector and a corresponding historical action vector; the historical space state vector includes a work order vector to be allocated and a customer service personnel vector to be allocated;

[0008] Iteratively train the deep Q network model using the training data set;

[0009] The iteratively trained deep Q network model is deployed to the customer service system to intelligently allocate real-time work orders.

[0010] In some embodiments, the reward function of the deep Q network model is R(s,a)=α×processing speed reward+β×processing quality reward+γ×customer satisfaction reward+δ×workload balancing reward; wherein s is the current state, a represents the action taken in state s; α is the weight coefficient of the work order processing speed, β is the weight coefficient of the work order processing quality, γ is the weight coefficient of the customer satisfaction and δ is the weight coefficient of the workload balancing; α+β+γ+δ=1;

[0011]

[0012] In some embodiments, the iterative training of the deep Q network model using the training dataset includes:

[0013] Initialize a deep Q network model and an experience replay pool; the experience replay pool is used to store state transition samples;

[0014] Determine action a;

[0015] Execute action a, observe and record the corresponding reward r and next state s', obtain the state transition sample (s, a, r, s') and store it in the experience replay pool;

[0016] Randomly extract state transition samples from the experience replay pool, train the deep Q network model, and update the parameters of the deep Q network model;

[0017] Update the parameters of the target Q network model so that the parameters of the target Q network model are the same as the parameters of the current deep Q network model;

[0018] Return to initialize the experience replay pool until the number of training times reaches the preset value.

[0019] In some embodiments, the loss function of the deep Q network model is: Where N is the number of batch samples, γ is the discount factor; Q(s′ i , a′; θ - ) is the Q value output by the target Q network model, which is a copy of the deep Q network model; Q(S i ,a i ; θ) is the Q value output by the current Q network model;

[0020] The determining action a includes, in each training round, for a given state s, selecting action a using an ε-greedy strategy.

[0021] In some embodiments, the preprocessing includes: at least one of missing value processing, denoising processing, standardization processing and encoding processing; wherein the standardization processing is performed by formula Among them, X is the original data, u is the mean of the original data, σ is the standard deviation of the original data, and Z is the standardized data;

[0022] The historical data also includes work order processing result data; the work order processing result data includes work order processing status data, work order processing quality score data and customer satisfaction score data.

[0023] In some embodiments, the iterative training of the deep Q network model using the training data set further comprises:

[0024] Selecting experimental subjects and time periods, and randomly dividing the work orders in the customer service system into an experimental group and a control group;

[0025] The experimental group is assigned work orders based on the work order allocation scheme recommended by the deep Q network model obtained through iterative training; the control group is assigned work orders based on the work order allocation scheme of the customer service system;

[0026] Record and obtain a first evaluation data set of the experimental group and a second evaluation data set of the control group; the first evaluation data set and the second evaluation data set both include work order processing efficiency, customer satisfaction, and customer service staff workload;

[0027] Comparing the differences between the first evaluation data set and the second evaluation data set; in response to determining that there is no significant difference between the first evaluation data set and the second evaluation data set, determining that the deep Q network model obtained by the iterative training can be deployed in the customer service system.

[0028] In some of the embodiments, the method further includes: during the model training process, tuning the learning rate and discount factor of the model.

[0029] An embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the methods described above when executing the program.

[0030] An embodiment of the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any of the methods described above.

[0031] An embodiment of the present application further provides a computer program product, comprising computer program instructions, which, when executed on a computer, enable the computer to execute any of the methods described above.

[0032] From the above, it can be seen that the intelligent customer service work order allocation method and related equipment based on machine learning provided by the present application model the work order allocation process as a Markov decision process; define the state space and the action space; design the reward function; wherein the state space includes a work order vector and a customer service staff vector; the work order vector includes a work order type vector, a work order urgency vector, and a work order creation time and deadline vector; the customer service staff vector includes the customer service staff's professional skills vector, the type and number vector of historically processed work orders, the customer service staff's current workload vector, and the customer service staff's online status vector; the action space includes the decision to allocate work orders to different customer service staff, set work order priorities, and decide whether to transfer; the reward function is based on the work order processing speed, work order processing quality, customer satisfaction, and customer service staff. Workload balance design; build a deep Q network model; obtain historical data on customer service work order allocation, and preprocess the historical data to obtain a training data set; the historical data includes historical state space vectors and corresponding historical action vectors; the historical space state vector includes work order vectors to be allocated and customer service personnel vectors to be allocated; use the training data set to iteratively train the deep Q network model; deploy the iteratively trained deep Q network model to the customer service system, and intelligently allocate real-time work orders. It can analyze the professional skills, current workload, online status, and type and urgency of work orders of customer service personnel in real time, realize automatic and intelligent allocation of work orders, thereby improving work order processing efficiency and customer satisfaction, significantly improving the overall performance and user experience of the customer service system, and balancing the workload of customer service personnel. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the present application or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0034] Figure 1 A flowchart of a method for allocating intelligent customer service work orders based on machine learning according to an embodiment of the present application;

[0035] Figure 2 This is another flow chart of the intelligent customer service work order allocation method based on machine learning according to an embodiment of the present application;

[0036] Figure 3 A schematic diagram of the design of the reward function of an embodiment of the present application;

[0037] Figure 4 A schematic diagram of a process for iterative training of a deep Q network model according to an embodiment of the present application;

[0038] Figure 5 Another schematic diagram of a process for iterative training of the deep Q network model according to an embodiment of the present application;

[0039] Figure 6 A flowchart of the effect evaluation of the deep Q network model according to an embodiment of the present application;

[0040] Figure 7 A schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0041] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0042] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be understood by people with ordinary skills in the field to which the present application belongs. The "first", "second" and similar words used in the embodiments of the present application do not represent any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0043] In the current customer service system, most customer service systems use a work order allocation solution based on a rule engine. This solution allocates work orders based on preset conditions and rules, such as the professional field of the customer service staff, the urgency of the work order, etc.

[0044] Although this traditional allocation method can achieve the distribution of work orders to a certain extent and is simple and easy to implement, it has many defects, including that it can usually only be allocated according to preset conditions and rules, lacks the ability of real-time adjustment and intelligent optimization, and cannot be dynamically adjusted according to the real-time situation and the actual status of customer service personnel. As a result, in some cases, work orders may be assigned to customer service personnel who are not suitable for handling, thereby affecting processing efficiency and quality, seriously affecting the efficiency of the customer service system and customer satisfaction.

[0045] Therefore, the current customer service system lacks flexibility: the rule engine cannot be dynamically adjusted according to the real-time situation and the actual status of the customer service staff, resulting in improper allocation; limited intelligence: it is impossible to optimize the allocation strategy based on historical data and real-time feedback, and it is difficult to adapt to the ever-changing customer service needs; and inefficiency: inappropriate allocation may lead to delays in work order processing, affecting customer satisfaction and the work efficiency of customer service staff, etc.

[0046] Based on this, in view of the problems of inflexible work order allocation, limited intelligence and low efficiency in the prior art, this application proposes an intelligent customer service work order allocation algorithm based on machine learning. The algorithm first models the work order allocation process as a Markov decision process, and uses a deep Q network algorithm to iteratively train on actual customer service work order data by designing a comprehensive reward mechanism to optimize the work order allocation strategy. During the training process, the ε-greedy strategy is used to balance exploration and utilization, and the learning rate and discount factor are tuned. Through A / B test comparison experiments, the actual application effect of the model in improving work order processing efficiency, increasing customer satisfaction and reducing the workload of customer service personnel is objectively evaluated and verified. Finally, the trained model is deployed to the actual customer service system, and the professional skills, current workload, online status, and type and urgency of the customer service personnel are analyzed in real time to realize the automation and intelligent allocation of work orders, thereby improving the work order processing efficiency and customer satisfaction, significantly improving the overall performance and user experience of the customer service system, and balancing the workload of customer service personnel.

[0047] like Figure 1 and Figure 2 As shown, the embodiment of the present application provides an intelligent customer service work order allocation method based on machine learning, which may include:

[0048] Step S100, modeling the work order allocation process as a Markov decision process; defining a state space and an action space; designing a reward function; wherein the state space includes a work order vector and a customer service staff vector; the work order vector includes a work order type vector, a work order urgency vector, and a work order creation time and deadline vector; the customer service staff vector includes a customer service staff professional skill vector, a historically processed work order type and quantity vector, a customer service staff current workload vector, and a customer service staff online status vector; the action space includes the decision of allocating work orders to different customer service staff, setting work order priorities, and deciding whether to transfer; the reward function is designed based on the balance of work order processing speed, work order processing quality, customer satisfaction, and customer service staff workload;

[0049] Step S200, constructing a deep Q network model;

[0050] Step S300, obtaining historical data of customer service work order allocation, and preprocessing the historical data to obtain a training data set; the historical data includes a historical state space vector and a corresponding historical action vector; the historical space state vector includes a work order vector to be allocated and a customer service staff vector to be allocated;

[0051] Step S400, iteratively training the deep Q network model using the training data set;

[0052] Step S500, deploy the iteratively trained deep Q network model to the customer service system to intelligently allocate real-time work orders.

[0053] In some embodiments, in step S100, the work order type may specifically include account login problems and order query problems. The urgency of the work order may specifically include three levels of identification, such as high, medium, and low. The creation time of the work order may include a specific time in the form of year, month, day, and time when the work order is created, and the deadline of the work order may include a specific time in the form of year, month, day, and time when the work order is closed (i.e., the processing is completed). For example, if a work order is created at 10:00 on March 1, 2023, the deadline is 10:00 on March 2, 2023.

[0054] In some of the embodiments, the professional skills of the customer service personnel may include areas of expertise, such as technical support, financial management, product consulting, etc. The type and number of historically processed work orders may be a first number of work orders of the first type processed within a preset time and a second number of work orders of the second type processed. For example, customer service personnel A processed 20 technical support work orders and 10 product consulting work orders in the past month. The current workload of the customer service personnel may include the number of work orders being processed, the number of completed work orders, and the average processing time of the work orders. For example, customer service personnel A currently has 3 work orders being processed. Customer service personnel A completed 50 work orders in the past week. The average time for customer service personnel A to process work orders is 15 minutes. The online status of the customer service personnel may include online, offline, busy, or resting. For example, the current status of customer service personnel A is online.

[0055] In this way, by modeling the work order allocation process as a Markov decision process, the state space is defined to include the customer service staff's professional skills, current workload, online status, and vector representation of the type and urgency of the work order, so that the spatial state information can fully reflect the real-time situation of the customer service system and work orders, thereby providing accurate decision-making basis for the deep Q network model.

[0056] In this way, by defining the action space including allocating work orders to different customer service personnel, setting work order priorities, and deciding whether to transfer work orders, the main operations in the work order allocation process can be covered, allowing the model to select appropriate actions according to different states.

[0057] In some embodiments, the work order processing speed may include the work order processing time. The work order processing quality may include the work order processing quality score. The customer satisfaction may include the customer satisfaction score. The balance of the customer service staff workload may include the standard deviation of the workload.

[0058] The reward mechanism is a key part of the model training process, which directly affects the learning effect and decision-making strategy of the model. In this way, by designing a reward function for evaluating action results based on work order processing efficiency, customer satisfaction, and customer service staff workload balance indicators, multiple factors can be comprehensively considered. These indicators are interrelated and mutually constrained, and together constitute the optimization goal of model training. The rules for selecting actions under a given state can be obtained, and the cumulative rewards can be maximized through subsequent learning optimization, so that the deep Q network model can learn the optimal allocation strategy during the training process, and the deep Q network model can achieve the best effect in the long-term decision-making process.

[0059] In some embodiments, such as Figure 3 As shown, the reward function of the deep Q network model can be R(s,a)=α×processing speed reward+β×processing quality reward+γ×customer satisfaction reward+δ×workload balance reward. Where s is the current state, and a represents the action taken in state s. α is the weight coefficient of the work order processing speed, β is the weight coefficient of the work order processing quality, γ is the weight coefficient of customer satisfaction, and δ is the weight coefficient of workload balance; α+β+γ+δ=1.

[0060] In some embodiments, the calculation formula for the processing speed reward may be: For example, assuming that the fastest processing time for a certain type of work order is 5 minutes, and the current work order processing time is 10 minutes, the processing speed reward is 0.5. In this way, the deep Q network model should tend to choose actions that can shorten the work order processing time to obtain a higher processing speed reward. Therefore, the processing speed reward can motivate the deep Q network model to choose actions that can increase the speed of work order processing.

[0061] In some embodiments, the processing quality reward may be determined based on the quality evaluation score of the work order processing result, and the calculation formula may be: For example, assuming that the quality score of ticket processing ranges from 0 to 100, the quality score of a certain ticket processing is 80 points, and the highest quality score is 100 points, then the processing quality reward is 0.8. In this way, the deep Q network model should tend to choose actions that can improve the quality of ticket processing in order to obtain higher processing quality rewards. Therefore, the processing quality reward can motivate the deep Q network model to choose actions that can improve the quality of ticket processing.

[0062] In some embodiments, the customer satisfaction reward may be determined based on the customer satisfaction survey results, and the calculation formula may be: For example, assuming that the customer satisfaction rating ranges from 1 to 5 stars, the result of a customer satisfaction survey is 4 stars, and the highest satisfaction rating is 5 stars, then the customer satisfaction reward is 0.8. In this way, the deep Q network model should tend to choose actions that can improve customer satisfaction in order to obtain higher customer satisfaction rewards. Therefore, the customer satisfaction reward can motivate the deep Q network model to choose actions that can improve customer satisfaction.

[0063] In some embodiments, the workload balance reward can be calculated based on the degree of balance of the customer service staff's workload, measured by the standard deviation index, and the calculation formula can be: For example, suppose there are 5 customer service staff, and their workloads (in terms of the number of work orders) are 10, 12, 8, 15, and 5 respectively. The standard deviation of the workload is 4.08 (using the sample standard deviation calculation formula). Assume that in an extreme case, all work orders are assigned to the same customer service staff. In this case, the standard deviation of the workload is the largest, which is 7.91 (i.e. The absolute value of Therefore, the workload balancing reward is In this way, the deep Q network model should tend to choose actions that can balance the workload of customer service personnel in order to obtain higher workload balancing rewards. Therefore, the workload balancing reward can motivate the deep Q network model to choose actions that can balance the workload of customer service personnel.

[0064] In this way, through the above reward function design, the effect of action a in state s can be comprehensively evaluated, thereby guiding the model to learn and optimize to maximize the cumulative reward.

[0065] In some of the embodiments, in step S200, the deep Q network model may include an input layer, a hidden layer, and an output layer. The input layer is mainly used to receive the vector representation of the state space as described above, which includes the work order type vector, the work order urgency vector, the work order creation time and deadline vector, etc., the customer service staff's professional skills vector, the type and quantity vector of the historically processed work orders, the customer service staff's current workload vector, and the customer service staff's online state vector. The hidden layer mainly processes the input information through a multi-layer neural network, extracts features, and learns the relationship between state and action. The output layer is used to output the Q value corresponding to each action, that is, the expected benefit of taking an action in this state.

[0066] By using the deep Q network (DQN) algorithm for iterative training, the Markov decision process can be solved. The DQN algorithm combines the advantages of deep learning and reinforcement learning, and has the advantages of being able to handle high-dimensional state space and action space, and learn complex decision-making strategies.

[0067] In some of the embodiments, in step S300, historical data of customer service work order allocation can be regularly collected through the customer service system background database. The work order allocation model is optimized by preprocessing the historical customer service work order allocation data and using it as a training data set. Generally, the collected historical data of customer service work order allocation can cover multiple dimensions, including customer service personnel-related data such as professional skill data of customer service personnel, types and quantity data of historically processed work orders, current workload data of customer service personnel, and online status data of customer service personnel. The collected historical data of customer service work order allocation can also include type data of work orders, urgency data of work orders, creation time and deadline data of work orders, etc.

[0068] In some embodiments, the historical data also includes work order processing result data. The work order processing result data may include work order processing status data, work order processing quality score data and customer satisfaction score data. It should be understood that the work order processing result data is the processing result data corresponding to the action corresponding to the work order in the state space.

[0069] In some embodiments, the work order processing status may include resolved, pending, or in progress. The work order processing quality score data may include a specific score for the work order, for example, the processing quality score of work order B is 90 points (out of 100 points). The customer satisfaction score may include a satisfaction score for the customer of work order B, for example, the customer satisfaction score of work order B is 4.5 stars (out of 5 stars).

[0070] In some embodiments, preprocessing the historical data may include: performing at least one of missing value processing, denoising, standardization and encoding processing on the historical data; wherein the standardization processing is performed by formula Wherein, X is the original data, u is the mean of the original data, σ is the standard deviation of the original data, and Z is the standardized data. That is, the preprocessing may include at least one of missing value processing, denoising processing, standardization processing and encoding processing.

[0071] In some embodiments, the missing value processing may include, for the missing value, using a linear interpolation algorithm to fill. The denoising processing may include, for the noisy data, using a sliding average filter algorithm to perform data denoising. In this way, the historical data can be cleaned and the uniformity of the data can be improved.

[0072] In some possible embodiments, the standardization process may include standardizing continuous data. Specifically, the Z-score method may be used to standardize the data, mainly including: subtracting the mean of each numerical data and dividing it by its standard deviation, so that the processed data conforms to the standard normal distribution. This can eliminate dimensional differences. Continuous data may include work order processing quality score data and customer satisfaction score data, etc.

[0073] In some possible embodiments, the standardization process may include standardizing continuous data. Continuous data may include work order processing quality score data, customer satisfaction score data, and work order creation time and deadline, etc. Specifically, the Z-score method may be used to standardize the data, mainly including: subtracting the mean of each numerical data and dividing it by its standard deviation, so that the processed data conforms to the standard normal distribution. This can eliminate dimensional differences.

[0074] In some possible embodiments, the encoding process may include performing one-hot encoding conversion on the classified data. The classified data may include work order type data, work order urgency data, work order processing status data, customer service personnel's professional skills data and customer service personnel's online status data. It should be noted that the one-hot encoding conversion in the embodiment of the present application may be an existing one-hot encoding conversion technology, and the embodiment of the present application does not involve improvements to the existing one-hot encoding conversion.

[0075] In this way, by iteratively training the model on actual customer service work order data, the model can gradually learn the strategy of taking the optimal action under different states.

[0076] In some embodiments, such as Figure 4 and Figure 5As shown, in step S400, the iterative training of the deep Q network model using the training data set may include:

[0077] Step S410, initializing a deep Q network model and an experience replay pool; the experience replay pool is used to store state transition samples;

[0078] Step S420, determining action a;

[0079] Step S430, execute action a, observe and record the corresponding reward r and next state s', obtain a state transition sample (s, a, r, s') and store it in the experience replay pool;

[0080] Step S440, randomly extracting state transition samples from the experience replay pool, training the deep Q network model, and updating the parameters of the deep Q network model;

[0081] Step S450, updating the parameters of the target Q network model so that the parameters of the target Q network model are the same as the parameters of the current deep Q network model;

[0082] Step S460, returning to the initialized experience replay pool until the number of training times reaches a preset value.

[0083] In some of the embodiments, in step S410, the experience replay pool may be used to store state transition samples (s, a, r, s'), where s is the current state, a is the action taken, r is the reward, and s' is the next state.

[0084] In some embodiments, in step S420, determining action a may include: in each training round, for a given state s, selecting action a using an ε-greedy strategy. Specifically, executing action a may include randomly selecting an action with probability ε; and selecting the action with the largest current Q value with probability 1-ε.

[0085] In this way, by randomly selecting an action with probability ε, we can explore new action spaces and prevent the deep Q network model from falling into the local optimum; by randomly selecting the action with the largest current Q value with probability 1-ε, we can use known information to improve the performance of the model. Therefore, we can balance the model's exploration behavior and the ability to use known information, and prevent the model from falling into the local optimal solution during the training process.

[0086] In some embodiments, in step S430, when a new state transition sample is added to the experience replay pool, the earliest sample is replaced to ensure the diversity and timeliness of the state transition samples. At the same time, the problem of limited size of the experience replay pool can also be avoided.

[0087] In some embodiments, in step S440, when training the deep Q network model, a mean square error loss function may be used for back propagation training. The loss function may be: Where N is the number of batch samples, γ is the discount factor; Q(s′ i ,a′;θ - ) is the Q value output by the target Q network model, which is a copy of the deep Q network; Q(S i ,A i ; θ) is the Q value output by the current Q network model.

[0088] In this way, by minimizing the loss function and updating the parameters of the deep Q network model, the deep Q network model can better approximate the optimal Q value function.

[0089] In some embodiments, in step S550, the target Q network model is a copy of the deep Q network model, and its parameters are regularly copied from the deep Q network model, mainly for stabilizing the training process. Usually, after each update of the parameters of the deep Q network model, the parameters of the target Q network model need to be updated to maintain its consistency with the current deep Q network model.

[0090] In some embodiments, during the training of the deep Q network model, the learning rate and discount factor of the deep Q network model are continuously tuned. The learning rate determines the step size of the model parameter update, and the discount factor affects the degree of influence of future rewards on the current decision. By tuning these two parameters, the model can achieve better convergence effect and performance during the training process.

[0091] In some embodiments, such as Figure 6 As shown, the iterative training of the deep Q network model using the training data set may also include:

[0092] Step S471, selecting experimental subjects and time periods, and randomly dividing the work orders in the customer service system into an experimental group and a control group;

[0093] Step S472, assigning work orders to the experimental group according to the work order assignment scheme recommended by the deep Q network model obtained through iterative training; assigning work orders to the control group according to the work order assignment scheme of the customer service system;

[0094] Step S473, recording and obtaining a first evaluation data set of the experimental group and a second evaluation data set of the control group; the first evaluation data set and the second evaluation data set both include work order processing efficiency, customer satisfaction, and customer service staff workload;

[0095] Step S474, comparing the differences between the first evaluation data set and the second evaluation data set; in response to determining that there are significant differences between the first evaluation data set and the second evaluation data set, updating the iteratively trained deep Q network model according to the evaluation data of the second evaluation data set to obtain an updated deep Q network model.

[0096] In some of the embodiments, in step S471, the customer service system of a large e-commerce platform can be specifically selected as the experimental object. The system receives a large number of consultation and complaint work orders from users every day. The experimental time period can be set to one month. A random number generator can be used to randomly divide the work orders to be processed in the system into two groups, namely an experimental group and a control group. The experimental group contains 50% of the work orders, and the control group contains the other 50% of the work orders. In this way, the two groups of work orders maintain similar distributions in terms of type, urgency, etc., which can make the experiment have good fairness.

[0097] In some embodiments, in step S472, for the experimental group, the deep Q network model obtained by iterative training of the embodiment of the present application is used to allocate work orders. The deep Q network model automatically recommends the optimal allocation plan based on factors such as the professional skills, current workload, online status, and type and urgency of the customer service personnel. For the control group, the allocation plan currently implemented by the customer service system continues to be used, that is, work orders are allocated based on the current simple rules.

[0098] In some embodiments, in step S473, the work order processing efficiency can be specifically measured by the average processing time of the work order as the main indicator to measure the effect of the algorithm on the improvement of the processing speed. Customer satisfaction can be specifically evaluated through user feedback and satisfaction surveys to evaluate the impact of the algorithm on customer satisfaction. The workload of customer service personnel can be specifically measured by the number of work orders processed by customer service personnel every day and the average working hours as indicators to measure the balancing effect of the algorithm on the workload.

[0099] In some embodiments, in step S474, the overall effect of the algorithm can be evaluated by calculating the mean and standard deviation of each index of the experimental group and the control group, and then performing statistical analysis such as t-test to test whether there is a significant difference between the two groups.

[0100] In some embodiments, in response to determining that there is a significant difference between the first evaluation data set and the second evaluation data set, the source of the difference can be further analyzed, and the model can be retrained or the model parameters can be optimized. For example, the model performance can be further improved by adjusting hyperparameters such as learning rate and discount factor, or improving the design of the reward function. At the same time, new customer service work order data will continue to be collected for subsequent iterative training and updating of the model to obtain an updated deep Q network model. The updated deep Q network model is then deployed to the customer service system.

[0101] In some embodiments, in response to determining that there is no significant difference between the first evaluation data set and the second evaluation data set, there is no need to update the deep Q network model after iterative training, and it can be determined that the deep Q network model obtained by the iterative training can be deployed in the customer service system.

[0102] In some of the embodiments, in order to verify the effect of the deep Q network model obtained by iterative training, specific examples can also be designed. For example, suppose there are three customer service personnel A, B, and C, who are good at handling different types of problems. At the same time, there are three work orders D, E, and F, and the types and urgency of the three work orders are different. In the initial state, a vector representation of the state space is constructed based on the professional skills, current workload, online status, and type and urgency of the customer service personnel, and input into the deep Q network model. The deep Q network model will output the Q value corresponding to each action, and select the optimal action based on the Q value, that is, which customer service personnel the work order is assigned to.

[0103] In this way, by adopting the comparative experimental method of A / B testing, the work order allocation scheme recommended by the deep Q network model is compared in parallel with the currently implemented allocation scheme, and the data changes of the two groups in terms of improved work order processing efficiency, increased customer satisfaction, and reduced workload of customer service personnel are recorded and analyzed respectively. Through objective data comparison and analysis, the advantages and disadvantages and effects of the model in practical applications can be evaluated and verified.

[0104] In some possible embodiments, the iterative training of the deep Q network model using the training data set may also include:

[0105] Obtaining an action vector output by the deep Q network model obtained after iterative training, a corresponding state space vector, and corresponding evaluation data to obtain a first evaluation data set;

[0106] Acquire an action vector obtained according to a preset rule, a corresponding state space vector, and corresponding evaluation data to obtain a second evaluation data set;

[0107] In response to determining that the evaluation data of the first evaluation data set is significantly smaller than the evaluation data of the second evaluation data set, the iteratively trained deep Q network model is updated according to the evaluation data of the second evaluation data set to obtain an updated deep Q network model.

[0108] In some of the embodiments, in step S500, after the iteratively trained deep Q network model is deployed to the customer service system, when a new work order arrives, it is only necessary to input the current state into the model of the deep Q network model to obtain the optimal work order allocation plan, thereby realizing the automated and intelligent allocation of work orders.

[0109] In some embodiments, the iteratively trained deep Q network model is deployed to the customer service system, and the intelligent allocation of real-time work orders may specifically include: obtaining in real time the vector of the state space corresponding to the work order to be allocated; the vector of the state space includes the work order vector to be allocated and the customer service staff vector to be allocated; wherein the work order vector includes the work order type vector and the work order urgency vector; the customer service staff vector includes the professional skills vector of the customer service staff, the type and quantity vector of the historically processed work orders, the current workload vector of the customer service staff, and the online status vector of the customer service staff;

[0110] The vector of the state space obtained in real time is input into the iteratively trained deep Q network model to obtain an output action; the output action includes allocating to different customer service personnel, setting the work order priority and determining whether it needs to be transferred.

[0111] The embodiment of the present application can be applied to the customer service system of a large e-commerce platform by constructing and training a customer service work order allocation model (i.e., a deep Q network model), and can dynamically allocate work orders in real time according to the professional skills, current workload, online status, and type and urgency of the customer service personnel. This intelligent allocation method can enable work orders to be promptly and accurately allocated to the most suitable customer service personnel, reduce processing delays, and avoid improper allocation caused by insufficient flexibility of traditional rule engines, thereby significantly improving the processing efficiency of work orders. Compared with the traditional rule engine-based allocation method, the embodiment of the present application adopts advanced machine learning technologies such as deep Q network models, making the system more flexible and intelligent. It can continuously optimize the allocation strategy according to real-time conditions and historical data to adapt to the ever-changing customer service needs. It can improve customer satisfaction, improve processing quality by optimizing the work order allocation strategy, and thus improve customer satisfaction. It can also balance the workload of customer service personnel, comprehensively consider the workload and online status of customer service personnel, and reasonably allocate work orders to avoid uneven workload.

[0112] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.

[0113] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can independently choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0114] As an optional but non-limiting implementation, in response to receiving the user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0115] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0116] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the described method.

[0117] It should be noted that the above describes some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0118] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the intelligent customer service work order allocation method based on machine learning as described in any of the above embodiments is implemented.

[0119] Figure 7 A more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment is shown, and the device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 in the device.

[0120] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0121] The memory 1020 may be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0122] The input / output interface 1030 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0123] The communication interface 1040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).

[0124] The bus 1050 includes a path that transmits information between the various components of the device (eg, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0125] It should be noted that, although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.

[0126] The electronic device of the above-mentioned embodiment is used to implement the corresponding intelligent customer service work order allocation method based on machine learning in any of the aforementioned embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0127] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the machine learning-based intelligent customer service work order allocation method as described in any of the above embodiments.

[0128] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0129] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the intelligent customer service work order allocation method based on machine learning as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0130] Based on the same inventive concept, corresponding to the intelligent customer service work order allocation method based on machine learning described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer so that the computer and / or the processor execute the intelligent customer service work order allocation method based on machine learning. Corresponding to the execution subject corresponding to each step in each embodiment of the intelligent customer service work order allocation method based on machine learning, the processor that executes the corresponding step may belong to the corresponding execution subject.

[0131] The computer program product of the above embodiment is used to enable the computer and / or the processor to execute the intelligent customer service work order allocation method based on machine learning as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0132] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the concept of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0133] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present application difficult to understand, the known power supply / ground connection with the integrated circuit (IC) chip and other components may or may not be shown in the provided drawings. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform to be implemented in the embodiments of the present application (that is, these details should be fully within the scope of understanding of those skilled in the art). In the case of elaborating specific details (e.g., circuits) to describe exemplary embodiments of the present application, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.

[0134] Although the present application has been described in conjunction with specific embodiments of the present application, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0135] The embodiments of the present application are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the present application.

Claims

1. An intelligent customer service work order allocation method based on machine learning, characterized in that: include: Model the work order assignment process as a Markov decision process; Define state space and action space; design reward function; wherein, the state space includes work order vector and customer service staff vector; the work order vector includes work order type vector, work order urgency vector and work order creation time and deadline vector; the customer service staff vector includes customer service staff professional skills vector, historical work order type and quantity vector, customer service staff current workload vector and customer service staff online status vector; the action space includes the decision of allocating work orders to different customer service staff, setting work order priority and deciding whether to transfer; the reward function is designed according to the balance of work order processing speed, work order processing quality, customer satisfaction and customer service staff workload; Build a deep Q network model; Acquire historical data of customer service work order allocation, and preprocess the historical data to obtain a training data set; the historical data includes a historical state space vector and a corresponding historical action vector; the historical space state vector includes a work order vector to be allocated and a customer service personnel vector to be allocated; Iteratively train the deep Q network model using the training data set; The iteratively trained deep Q network model is deployed to the customer service system to intelligently allocate real-time work orders.

2. The intelligent customer service work order allocation method based on machine learning according to claim 1 is characterized in that: The reward function of the deep Q network model is R(s,a)=α×processing speed reward+β×processing quality reward+γ×customer satisfaction reward+δ×workload balance reward; where s is the current state, a represents the action taken in state s; α is the weight coefficient of work order processing speed, β is the weight coefficient of work order processing quality, γ is the weight coefficient of customer satisfaction and δ is the weight coefficient of workload balance; α+β+γ+δ=1; 3. The intelligent customer service work order allocation method based on machine learning according to claim 1 is characterized in that: The iterative training of the deep Q network model using the training data set includes: Initialize a deep Q network model and an experience replay pool; the experience replay pool is used to store state transition samples; Determine action a; Execute action a, observe and record the corresponding reward r and next state s', obtain the state transition sample (s, a, r, s') and store it in the experience replay pool; Randomly extract state transition samples from the experience replay pool, train the deep Q network model, and update the parameters of the deep Q network model; Update the parameters of the target Q network model so that the parameters of the target Q network model are the same as the parameters of the current deep Q network model; Return to initialize the experience replay pool until the number of training times reaches the preset value.

4. The intelligent customer service work order allocation method based on machine learning according to claim 3 is characterized in that: The loss function of the deep Q network model is: Where N is the number of batch samples, γ is the discount factor; Q(s′ i ,a′;θ - ) is the Q value output by the target Q network model, which is a copy of the deep Q network model; Q(S i ,a i ; θ) is the Q value output by the current Q network model; The determining action a includes, in each training round, for a given state s, selecting action a using an ε-greedy strategy.

5. The intelligent customer service work order allocation method based on machine learning according to claim 4 is characterized in that: The preprocessing includes: at least one of missing value processing, denoising processing, standardization processing and encoding processing; wherein the standardization processing is performed by the formula Among them, X is the original data, u is the mean of the original data, σ is the standard deviation of the original data, and Z is the standardized data; The historical data also includes work order processing result data; the work order processing result data includes work order processing status data, work order processing quality score data and customer satisfaction score data.

6. The intelligent customer service work order allocation method based on machine learning according to claim 1 is characterized in that: The iterative training of the deep Q network model using the training data set also includes: Selecting experimental subjects and time periods, and randomly dividing the work orders in the customer service system into an experimental group and a control group; The experimental group is assigned work orders based on the work order allocation scheme recommended by the deep Q network model obtained through iterative training; the control group is assigned work orders based on the work order allocation scheme of the customer service system; Record and obtain a first evaluation data set of the experimental group and a second evaluation data set of the control group; the first evaluation data set and the second evaluation data set both include work order processing efficiency, customer satisfaction, and customer service staff workload; Comparing the differences between the first evaluation data set and the second evaluation data set; in response to determining that there is no significant difference between the first evaluation data set and the second evaluation data set, determining that the deep Q network model obtained by the iterative training can be deployed in the customer service system.

7. The intelligent customer service work order allocation method based on machine learning according to claim 1 is characterized in that: The method also includes: during the model training process, tuning the learning rate and discount factor of the model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.

10. A computer program product, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Online shopping mall platform intelligent management method and system

    CN120894093A