Online customer service user incoming line queuing and distribution method and system

By employing a hybrid scheduling strategy and a multi-objective incentive mechanism, the problems of singular state awareness and singular incentive orientation in user line entry queuing allocation methods are solved, achieving efficient and reliable user allocation and optimizing resource utilization and user experience.

CN120911887BActive Publication Date: 2026-04-24POINT CONTROL CLOUD (BEIJING) INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
POINT CONTROL CLOUD (BEIJING) INTELLIGENT TECH CO LTD
Filing Date
2025-08-01
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

The existing user queue allocation method suffers from a lack of state awareness, ignores the dynamic differences between users and customer service representatives, and lacks flexible adjustment capabilities, resulting in poor allocation reliability. Furthermore, the incentive orientation is singular, the optimization goals are unbalanced, the user churn rate is high, and the allocation effect is poor.

Method used

By capturing service gaps through deviation quantification, a hybrid scheduling strategy is adopted to balance the rigidity of rules and the flexibility of intelligence. A dynamic matching-load balancing mechanism is introduced, a multi-objective incentive mechanism is constructed, a segmented waiting time incentive and a load imbalance penalty are designed, and a tiered anomaly response mechanism is adopted to optimize resource utilization and user experience.

Benefits of technology

It improves the reliability and effectiveness of allocation, reduces user churn by dynamically adjusting resource utilization, balances user experience and resource efficiency, avoids hard queue jumping and extreme waiting, and achieves load balancing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911887B_ABST
    Figure CN120911887B_ABST
Patent Text Reader

Abstract

The application discloses an online customer service user incoming line queuing distribution method and system, which comprises user state capturing, mixed scheduling strategy, incentive mechanism design, scheduling model training and distribution exception response. The application belongs to the field of intelligent scheduling, and specifically relates to an online customer service user incoming line queuing distribution method and system. The scheme is based on user cutting-in weight, customer service matching recommendation degree and service time length buffer adjustment value, fine-tuned outside the rule framework, avoids the abruptness of hard cutting-in, realizes fine distribution with high skill matching degree and load balance based on a dynamic matching-load balance mechanism, avoids taking an aggressive strategy for extremely shortening overtime waiting through segmented design of waiting time incentive, guarantees efficient use of resources through load imbalance degree to build load balance punishment, and reduces user loss caused by distribution failure based on a ladder type exception response, and further improves distribution effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent scheduling technology, specifically to a method and system for allocating online customer service users in an online queue. Background Technology

[0002] User call queuing and allocation methods refer to the rules and mechanisms by which a customer service system queues users waiting for service after they initiate an inquiry and assigns them to suitable customer service representatives. The core principle is to balance service efficiency, user experience, and resource load by rationally arranging the queuing order and matching customers, ensuring users receive timely and high-quality service. However, general user call queuing and allocation methods suffer from a lack of state awareness, ignoring the dynamic differences between users and customer service representatives, and lacking flexible adjustment capabilities, leading to poor allocation reliability. Furthermore, general user call queuing and allocation methods have a single incentive orientation, resulting in an imbalanced optimization goal, making it difficult to balance user experience, resource efficiency, and service quality, leading to high user churn rates and ultimately poor allocation effectiveness. Summary of the Invention

[0003] To address the above issues and overcome the shortcomings of existing technologies, this invention provides an online customer service user queue allocation method and system. Addressing the problems of general user queue allocation methods, such as limited state awareness, neglect of dynamic differences between users and customer service representatives, and lack of flexible adjustment capabilities, leading to poor allocation reliability, this solution uses deviation quantification to capture service gaps. Based on a hybrid scheduling strategy, it balances rule rigidity with intelligent flexibility, fine-tuning outside the rule framework based on user queue-jumping weights, customer service representative matching recommendation degrees, and service duration buffer adjustment values ​​to avoid abrupt hard queue-jumping. Based on a dynamic matching-load balancing mechanism, it achieves refined allocation with high skill matching and load balancing. A dynamic matching coefficient is introduced, dynamically adjusted according to the real-time load of customer service representatives. This solution aims to optimize resource utilization and improve allocation reliability. Addressing the issue of simplistic incentive-driven allocation methods for general user queues, which lead to unbalanced optimization goals and difficulty in balancing user experience, resource efficiency, and service quality, resulting in high user churn and poor allocation performance, this solution addresses these problems by constructing a multi-objective incentive mechanism to balance core metrics and avoid biases from singular optimization. It employs segmented waiting time incentives to avoid aggressive strategies aimed at drastically shortening waiting times; load balancing penalties based on load imbalance to mitigate situations where a few customer service representatives are overloaded while the majority are idle, ensuring efficient resource utilization; emotion-based rewards to reduce user dissatisfaction; and a tiered anomaly response system to reduce user churn due to allocation failures, thereby improving allocation effectiveness.

[0004] The technical solution adopted by this invention is as follows: This invention provides an online customer service user call-in queuing allocation method, which includes the following steps:

[0005] Step S1: User state capture;

[0006] Step S2: Hybrid scheduling strategy;

[0007] Step S3: Incentive mechanism design;

[0008] Step S4: Training the scheduling model;

[0009] Step S5: Assign exception handling.

[0010] Further, in step S1, the user state capture constructs a user state vector, including: the current user's actual waiting time. User's target waiting threshold c) Problem complexity estimation; g) Sentiment prediction; g) Current load of candidate customer service representatives. Skill matching between users and customer service representatives ;

[0011] Construct the deviation vector, represented as: ; ; Combined into a state vector s, it can be represented as: Only the top K customer service candidates will be accepted; among them, It is a waiting time deviation; It is an emotional deviation; It is a load deviation; This is the average load of candidate customer service representatives; This is standardized processing; j is the candidate customer service index.

[0012] Furthermore, in step S2, the hybrid scheduling strategy specifically includes the following:

[0013] Step S21: Build basic rule units; calculate waiting ratio , is represented as: Design rule decisions are expressed as: ;in, and It is the threshold for rule-based decision-making, used to classify user queuing strategies;

[0014] Step S22: Intelligent scheduling unit; Let action a be a continuous vector, including: p, the priority adjustment factor for the user; The customer service matching recommendation rate for candidate customer service j; Service duration buffer adjustment value;

[0015] ; It is the Sigmoid function; It is a shearing function; , and It involves adjusting the weights;

[0016] The customer service matching and recommendation system introduces a dynamic matching and load balancing mechanism, which first calculates a base score. , is represented as: ; , and It is the allocation coefficient; It is the historical allocation success rate; calculating the fusion matching degree. , is represented as: Ultimately, customer service matching and recommendation scores are obtained. , is represented as: ;in, It is the base scaling factor; and Customer service The degree of fusion matching and the basic score;

[0017] ; It is the basic adjustment coefficient; This is the baseline load for customer service;

[0018] Step S23: The output of the scheduling model is represented as follows: ; Estimating the action value function using Critic Target Q value The update is represented as: The loss function is expressed as: ;in, It is the policy function of the scheduling model. These are policy network parameters; It is the action vector at step t; and It is the customer service matching recommendation rate of the candidate customer service; This is the current reward, and t is the step index; It is a discount factor; It is a target value network with parameters as follows: ,and The structure is the same, and the parameters of Q are copied every Y steps; and These are the state vectors at step t+1 and step t, respectively; N is the total number of samples, and i is the sample index; by maximizing Q... Gradient updates.

[0019] Furthermore, in step S3, the incentive mechanism design involves constructing a multi-objective reward mechanism and defining the waiting bias. , is represented as: Let the tolerance threshold T be set, and construct a piecewise incentive mechanism. , is represented as: Where C is the positive reward benchmark; It is the penalty coefficient when the target is far away; it also applies to load imbalance penalties. , is represented as: ; If user sentiment prediction improves, an additional reward will be given. , is represented as: Where Mg represents the load imbalance. It is the emotional reward coefficient;

[0020] The final reward r is represented as: .

[0021] Further, in step S4, the scheduling model is trained using historical customer service interaction data, and the scheduling model is pre-trained offline. The pre-trained scheduling model is deployed to a real customer service system, where new samples are generated through interaction with the environment, and the strategy is continuously optimized online. Exploration is introduced during the online phase, and Gaussian noise N(0,0.1) is superimposed on the output action. As the number of training steps increases, the noise is attenuated, as shown below: ; is the noise standard deviation at step t; T is the total number of decay steps.

[0022] Furthermore, in step S5, the allocation anomaly response is a tiered recovery mechanism to handle allocation failures, including rapid re-searching and activating a long-term solution in patrol mode. If a user is not successfully allocated, a re-search is first performed, followed by a short-term, rapid retry while maintaining the current user state, gradually relaxing matching constraints according to a predefined search pattern. If this fails, the search is expanded to the next highest skill level. If this also fails, a short-term overload allocation is considered, with temporary weighted load balancing. If allocation still fails after a time threshold, the system switches to patrol mode, dynamically adjusts user priorities, and initiates resource expansion notifications.

[0023] This invention provides an online customer service user inbound queuing allocation system, including a user status capture module, a hybrid scheduling strategy design module, an incentive mechanism design module, a scheduling model training module, and an allocation anomaly response module;

[0024] The user state capture module constructs a user state vector and obtains an deviation vector;

[0025] The hybrid scheduling strategy design module, based on user state vectors and deviation vectors, achieves priority and customer service matching scheduling through threshold rule units and intelligent scheduling units;

[0026] The incentive mechanism design module designs a multi-objective reward function consisting of three parts: waiting deviation segmented incentive, load imbalance penalty, and emotion prediction reward, and constructs a comprehensive reward signal.

[0027] The scheduling model training module first pre-trains the scheduling model based on historical interaction data, and then optimizes the scheduling model online by superimposing Gaussian noise on the actions and following a noise attenuation strategy.

[0028] The allocation anomaly handling module employs a tiered recovery mechanism to address allocation anomalies.

[0029] The beneficial effects achieved by the present invention using the above solution are as follows:

[0030] (1) In view of the problem that the general user queuing allocation method has a single state perception, ignores the dynamic differences between users and customer service, lacks flexible adjustment capability, and thus leads to poor allocation reliability, this solution captures service gaps through deviation quantification; based on the hybrid scheduling strategy, it balances the rigidity of rules and the flexibility of intelligence, and fine-tunes outside the rule framework based on the user queue-jumping weight, customer service matching recommendation degree and service duration buffer adjustment value to avoid the abruptness of hard queue-jumping; based on the dynamic matching-load balancing mechanism, it realizes fine allocation with high skill matching degree and load balance; introduces dynamic matching coefficient, which is dynamically adjusted according to the real-time load of customer service to optimize resource utilization; thereby improving allocation reliability.

[0031] (2) In view of the problem that the general user queuing allocation method has a single incentive orientation and an unbalanced optimization goal, it is difficult to take into account user experience, resource efficiency and service quality, resulting in high user churn rate and poor allocation effect. This solution constructs a multi-objective incentive mechanism to balance core indicators and avoid single optimization bias; it designs a segmented waiting time incentive to avoid taking aggressive strategies to shorten the waiting time to the extreme; it constructs a load balancing penalty through load imbalance to suppress the situation of a few customer service representatives being overloaded and most being idle, thus ensuring efficient use of resources; it reduces user dissatisfaction based on emotional rewards; and it reduces user churn caused by allocation failure based on tiered anomaly response; thus improving allocation effect. Attached Figure Description

[0032] Figure 1 This is a flowchart illustrating an online customer service user call-in queuing allocation method provided by the present invention.

[0033] Figure 2 This is a schematic diagram of an online customer service user queuing allocation system provided by the present invention.

[0034] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation

[0035] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0036] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the system or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0037] Example 1, see Figure 1 The present invention provides a method for allocating online customer service users in a queue, the method comprising the following steps:

[0038] Step S1: User state capture; construct the user state vector and obtain the deviation vector;

[0039] Step S2: Hybrid scheduling strategy; Based on user state vector and deviation vector, priority and customer service matching scheduling is achieved through threshold rule unit and intelligent scheduling unit;

[0040] Step S3: Incentive mechanism design; Design a multi-objective reward function consisting of three parts: waiting bias segmented incentive, load imbalance penalty, and emotion prediction reward, and construct a comprehensive reward signal;

[0041] Step S4: Scheduling model training; First, the scheduling model is pre-trained based on historical interaction data, and then the scheduling model is optimized online by superimposing Gaussian noise on the actions and following the noise attenuation strategy.

[0042] Step S5: Handling allocation anomalies; a tiered recovery mechanism is used to handle allocation anomalies.

[0043] Example 2, see Figure 1 This embodiment is based on the above embodiment. In step S1, the user state is captured to construct a user state vector, including: the current user's actual waiting time. User's target waiting threshold We selected the average user wait time and the problem complexity estimate c, used the BERT model to classify the initial problem description and predict the sentiment g, and used Alibaba Cloud's speech sentiment recognition to output 0-1 values ​​and the current load of candidate customer service representatives. Skill matching between users and customer service representatives Cosine similarity between customer service skill tags and user question categories;

[0044] Quantify the various dimensions in which users deviate from the ideal service state into deviations, and focus on adjusting the factors that lead to service quality gaps;

[0045] Construct the deviation vector, represented as: ; ; Combined into a state vector s, it can be represented as: Only accept the top K customer service candidates to avoid a surge in search volume; among them, It is a waiting time deviation; It is an emotional deviation; It is a load deviation; This is the average load of candidate customer service representatives; It is a standardized process, using max-min standardization; j is the candidate customer service index.

[0046] Example 3, see Figure 1 This embodiment is based on the above embodiment. In step S2, the hybrid scheduling strategy specifically includes the following:

[0047] Step S21: Build basic rule units; calculate waiting ratio , is represented as: Design rule decisions are expressed as: ;in, and It is the threshold for rule-based decision-making, used to classify user queuing strategies;

[0048] Step S22: Intelligent scheduling unit; Let action a be a continuous vector, including: p, the priority adjustment factor for users, which affects the final queue insertion weight; , where represents the customer service matching recommendation degree for candidate customer service j, and represents the recommendation assignment probability; Service duration buffer adjustment value, which adjusts the expected service duration and is used to dynamically allocate the buffer;

[0049] ; It is the Sigmoid function; It is a shearing function; , and It involves adjusting the weights;

[0050] The customer service matching and recommendation system introduces a dynamic matching and load balancing mechanism, which first calculates a base score. , is represented as: ; , and It is the allocation coefficient; It is the historical allocation success rate; calculating the fusion matching degree. , is represented as: Ultimately, customer service matching and recommendation scores are obtained. , is represented as: ;in, It is the base scaling factor; and Customer service The degree of fusion matching and the basic score;

[0051] The system dynamically adjusts its allocation based on the real-time load of each customer service representative, enabling fine-grained allocation that suppresses high loads and allows for fault tolerance based on high matching scores. This avoids load imbalance caused by simply allocating based on matching scores, as well as service quality degradation caused by simply allocating based on load.

[0052] ; It is the basic adjustment coefficient; This is the baseline load for customer service;

[0053] Step S23: The output of the scheduling model is represented as follows: ; Estimating the action value function using Critic Target Q value The update is represented as: The loss function is expressed as: ;in, It is the policy function of the scheduling model. These are policy network parameters; It is the action vector at step t; and It is the customer service matching recommendation rate of the candidate customer service; This is the current reward, and t is the step index; It is a discount factor; It is a target value network with parameters as follows: ,and The structure is the same, and the parameters of Q are copied every Y steps; and These are the state vectors at step t+1 and step t, respectively; N is the total number of samples, and i is the sample index; by maximizing Q... Gradient updates allow the scheduling model to fine-tune the continuous weights of priority and customer service matching outside of the rules, achieving flexible scheduling instead of hard allocation; smooth weight changes avoid jitter caused by frequent switching.

[0054] By performing the above operations, this solution addresses the problems of poor allocation reliability caused by the general user queue allocation method's singular state perception, neglect of the dynamic differences between users and customer service representatives, and lack of flexible adjustment capabilities. It captures service gaps through deviation quantification; balances rule rigidity and intelligent flexibility based on a hybrid scheduling strategy; fine-tunes outside the rule framework based on user queue-jumping weights, customer service representative matching recommendation degrees, and service duration buffer adjustment values ​​to avoid the abruptness of hard queue-jumping; achieves refined allocation with high skill matching and load balancing based on a dynamic matching-load balancing mechanism; and introduces a dynamic matching coefficient that dynamically adjusts with real-time customer service load to optimize resource utilization, thereby improving allocation reliability.

[0055] Example 4, see Figure 1 This embodiment is based on the above embodiment. In step S3, the incentive mechanism design involves constructing a multi-objective reward mechanism to balance waiting time, load balancing, and user satisfaction, guiding the scheduling model to generate an optimization strategy; and defining the waiting bias. , is represented as: Let the tolerance threshold T be set, and construct a piecewise incentive mechanism. , is represented as: Where C is the positive reward benchmark; It is the penalty coefficient when the target is far away; it also applies to load imbalance penalties. , is represented as: ; If user sentiment prediction improves, an additional reward will be given. , is represented as: Where Mg represents the load imbalance. It is the emotional reward coefficient;

[0056] The final reward r is represented as: ;

[0057] When waiting within an acceptable range, a strong positive incentive is used to bring the time closer, while a mild linear negative penalty is used to prevent overreaction when the time exceeds the acceptable range. At the same time, load and satisfaction terms are added to form a multi-objective optimization signal, which prompts the scheduling model to achieve a trade-off between timeliness and balance, and avoids unstable strategies caused by extreme waiting due to overreaction to extreme delays.

[0058] Example 5, see Figure 1 This embodiment is based on the above embodiment. In step S4, the scheduling model is trained using historical customer service interaction data, and the scheduling model is pre-trained offline. The pre-trained scheduling model is deployed to a real customer service system, and new samples are generated through interaction with the environment to continuously optimize the strategy online. To avoid the scheduling model getting stuck in local optima, exploration is introduced in the online stage, and Gaussian noise N(0,0.1) is superimposed on the output action. As the number of training steps increases, the noise is attenuated, as shown below: ; is the noise standard deviation at step t; T is the total number of decay steps; the noise standard deviation gradually decreases from 0.1 to 0.01, balancing the exploration of new strategies with the utilization of known optimal strategies.

[0059] Example 6, see Figure 1 This embodiment is based on the above embodiment. In step S5, the handling of allocation anomalies is a tiered recovery mechanism to deal with allocation failures. It involves rapid source tracing via re-search, and the activation of a long-term solution in patrol mode to reduce user churn. If a user is not successfully allocated, including when all candidate customer service representatives are instantly busy, matching fails, or the allocation request times out, a re-search is first performed for a short, quick retry while maintaining the current user state. Matching constraints are gradually relaxed according to a predefined search pattern, including: first, searching for available customer service representatives in the skill-high matching pool; if that fails, expanding to the next highest skill level; if that also fails, short-term overload allocation is considered, with temporary weighted load balancing; if allocation still fails after a time threshold, the system switches to patrol mode, including: pushing self-service solutions to the user; dynamically adjusting user priorities; and initiating resource expansion notifications.

[0060] By implementing the above operations, we address the issues of a single-minded incentive approach and unbalanced optimization goals in the general user call queue allocation method. This results in a failure to balance user experience, resource efficiency, and service quality, leading to high user churn and poor allocation performance. This solution addresses these problems by constructing a multi-objective incentive mechanism to balance core metrics and avoid biases from singular optimization. It also addresses the issues of segmented waiting time incentives to avoid aggressive strategies aimed at drastically shortening waiting times; load balancing penalties based on load imbalance to prevent situations where a few customer service representatives are overloaded while the majority are idle, ensuring efficient resource utilization; reducing user dissatisfaction through emotional rewards; and minimizing user churn due to allocation failures through tiered anomaly responses. Ultimately, this improves allocation effectiveness.

[0061] Example 7, see Figure 2 Based on the above embodiments, this embodiment provides an online customer service user inbound queuing allocation system, including a user status capture module, a hybrid scheduling strategy design module, an incentive mechanism design module, a scheduling model training module, and an allocation anomaly response module.

[0062] The user state capture module constructs a user state vector and obtains an deviation vector;

[0063] The hybrid scheduling strategy design module, based on user state vectors and deviation vectors, achieves priority and customer service matching scheduling through threshold rule units and intelligent scheduling units;

[0064] The incentive mechanism design module designs a multi-objective reward function consisting of three parts: waiting deviation segmented incentive, load imbalance penalty, and emotion prediction reward, and constructs a comprehensive reward signal.

[0065] The scheduling model training module first pre-trains the scheduling model based on historical interaction data, and then optimizes the scheduling model online by superimposing Gaussian noise on the actions and following a noise attenuation strategy.

[0066] The allocation anomaly handling module employs a tiered recovery mechanism to address allocation anomalies.

[0067] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0068] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.

[0069] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A method for allocating online customer service users in a queue, characterized in that: The method includes the following steps: Step S1: User state capture; construct the user state vector and obtain the deviation vector; Step S2: Hybrid scheduling strategy; Based on user state vector and deviation vector, priority and customer service matching scheduling is achieved through threshold rule unit and intelligent scheduling unit; Step S3: Incentive mechanism design; Design a multi-objective reward function consisting of three parts: waiting bias segmented incentive, load imbalance penalty, and emotion prediction reward, and construct a comprehensive reward signal; Step S4: Scheduling model training; First, the scheduling model is pre-trained based on historical interaction data, and then the scheduling model is optimized online by superimposing Gaussian noise on the actions and following the noise attenuation strategy. Step S5: Handling Allocation Anomalies; A tiered recovery mechanism is used to handle allocation anomalies; In step S1, the user state capture constructs a user state vector, including: the current user's actual waiting time. User's target waiting threshold c) Problem complexity estimation; g) Sentiment prediction; g) Current load of candidate customer service representatives. Skill matching between users and customer service representatives ; Construct the deviation vector, represented as: ; ; Combined into a state vector s, it can be represented as: Only the top K customer service candidates will be accepted; among them, It is a waiting time deviation; It is an emotional deviation; It is a load deviation; This is the average load of candidate customer service representatives; This is standardized processing; j is the candidate customer service index; In step S2, the hybrid scheduling strategy specifically includes the following: Step S21: Build basic rule units; calculate waiting ratio , is represented as: Design rule decisions are expressed as: ;in, and It is the threshold for rule-based decision-making, used to classify user queuing strategies; Step S22: Intelligent scheduling unit; Let action a be a continuous vector, including: p, the priority adjustment factor for the user; The customer service matching recommendation rate for candidate customer service j; Service duration buffer adjustment value; ; It is the Sigmoid function; It is a shearing function; , and It involves adjusting weights; a dynamic matching-load balancing mechanism is introduced for customer service matching and recommendation, first calculating the base score. , is represented as: ; , and It is the allocation coefficient; It is the historical allocation success rate; calculating the fusion matching degree. , is represented as: Ultimately, customer service matching and recommendation scores are obtained. , is represented as: ;in, It is the base scaling factor; and Customer service The degree of fusion matching and the basic score; ; It is the basic adjustment coefficient; This is the baseline load for customer service; Step S23: The output of the scheduling model is represented as follows: ; Estimating the action value function using Critic Target Q value The update is represented as: The loss function is expressed as: ;in, It is the policy function of the scheduling model. These are policy network parameters; It is the action vector at step t; and It is the customer service matching recommendation rate of the candidate customer service; This is the current reward, and t is the step index; It is a discount factor; It is a target value network with parameters as follows: ,and The structure is the same, and the parameters of Q are copied every Y steps; and These are the state vectors at step t+1 and step t, respectively; N is the total number of samples, and i is the sample index; by maximizing Q... Gradient update; In step S3, the incentive mechanism design involves constructing a multi-objective reward mechanism and defining the waiting bias. , is represented as: Let the tolerance threshold T be set, and construct a piecewise incentive mechanism. , is represented as: Where C is the positive reward benchmark; It is the penalty coefficient when the target is far away; it also applies to load imbalance penalties. , is represented as: ; If user sentiment prediction improves, an additional reward will be given. , is represented as: Where Mg represents the load imbalance. It is the emotional reward coefficient; The final reward r is represented as: ; In step S4, the scheduling model is trained using historical customer service interaction data, and the scheduling model is pre-trained offline. The pre-trained scheduling model is deployed to a real customer service system, where new samples are generated through interaction with the environment, and the strategy is continuously optimized online. During the online phase, exploration is introduced by superimposing Gaussian noise N(0,0.1) onto the output action. As the number of training steps increases, the noise is attenuated, as shown below: ; is the noise standard deviation at step t; T is the total number of decay steps.

2. The online customer service user queuing allocation method according to claim 1, characterized in that: In step S5, the allocation anomaly response is a tiered recovery mechanism to handle allocation failures, including rapid re-searching and activating a long-term solution in patrol mode. If a user is not successfully allocated, a re-search is first performed, followed by a short-term, rapid retry while maintaining the current user status and gradually relaxing matching constraints according to a predefined search pattern. If this fails, the search is expanded to the next highest skill level. If this also fails, a short-term overload allocation is considered, with temporary weighted load balancing. If allocation still fails after a time threshold, the system switches to patrol mode, dynamically adjusts user priorities, and initiates resource expansion notifications.

3. An online customer service user call-in queuing and allocation system, used to implement the online customer service user call-in queuing and allocation method as described in any one of claims 1-2, characterized in that: It includes a user state capture module, a hybrid scheduling strategy design module, an incentive mechanism design module, a scheduling model training module, and an allocation anomaly handling module; The user state capture module constructs a user state vector and obtains an deviation vector; The hybrid scheduling strategy design module, based on user state vectors and deviation vectors, achieves priority and customer service matching scheduling through threshold rule units and intelligent scheduling units; The incentive mechanism design module designs a multi-objective reward function consisting of three parts: waiting deviation segmented incentive, load imbalance penalty, and emotion prediction reward, and constructs a comprehensive reward signal. The scheduling model training module first pre-trains the scheduling model based on historical interaction data, and then optimizes the scheduling model online by superimposing Gaussian noise on the actions and following a noise attenuation strategy. The allocation anomaly handling module employs a tiered recovery mechanism to address allocation anomalies.

Citation Information

Patent Citations

  • Service reservation queuing system and method

    CN118446342A

  • Online customer service user incoming line queuing and distributing device

    CN118612344A

  • Multimodal transport intelligent scheduling optimization method, apparatus and device, and storage medium

    CN119167789A