Call center call assisting method and device based on large model, medium and product

By acquiring multi-dimensional profiles of call center agents and analyzing large language models, precise matching of call center agents was achieved, solving the problem of unreasonable agent allocation in existing technologies and improving problem-solving efficiency and customer service quality.

CN121728191APending Publication Date: 2026-03-24上海井星信息科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

The existing call center's agent allocation mechanism is simplistic and cannot accurately match agents based on their business capabilities and processing strengths, resulting in low problem-solving efficiency and a poor customer service experience.

Method used

By acquiring multi-dimensional agent profiles of human agents, analyzing user voice call content in real time, recognizing user intent and emotional state using a large language model, selecting the most suitable agent for transfer based on preset rules, and optimizing allocation through a comprehensive scoring strategy.

Benefits of technology

This enabled precise matching of agents, improved problem-solving efficiency and customer satisfaction, ensured service consistency and high quality, and optimized resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121728191A_ABST
    Figure CN121728191A_ABST
Patent Text Reader

Abstract

The invention discloses a call center coordinated call method and device based on a large model, a medium and a product. The method comprises the following steps: acquiring a multi-dimensional seat portrait of each seat account in a manual seat group; user voice is collected; converting the user voice into a user utterance text, converting the robot voice into a robot answer text, and generating a structured text dialogue; inputting the structured text dialogue into a large language model to obtain a user intention, an emotional state and a multi-dimensional question abstract; judging whether a manual intervention process is triggered or not based on the user intention and the emotional state; when it is judged that the manual intervention process is triggered, candidate seat accounts are screened out according to the target entity information; calculating a matching score of each seat in the candidate seat accounts, and determining a target idle seat account; and pushing the structured text dialogue and the multi-dimensional question abstract to a workbench of the target idle seat account. By implementing the technical scheme, the technical defect that the existing call center seat distribution mechanism is single and mechanical is effectively overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent communication, specifically to a call center collaborative calling method, device, medium, and product based on a large model. Background Technology

[0002] In existing human-machine collaboration solutions for call centers, intelligent robots typically handle incoming calls first. When the system identifies preset transfer conditions, such as specific keywords or multiple failed rounds of dialogue, it transfers the call to a pre-defined business skill group. Subsequently, within the skill group, the system generally uses a polling or longest idle time allocation strategy to route the call to an idle human agent.

[0003] However, the aforementioned existing technologies have significant technical flaws in the agent allocation process. Their allocation mechanism relies solely on the availability of agents, ignoring the differences in actual business capabilities and processing strengths among different agents. This "one-size-fits-all" allocation method is mechanical and indiscriminate, failing to guarantee that a specific call is matched with the most suitable agent, often resulting in low problem-solving efficiency, increased call duration, and difficulty in providing a consistent, high-quality customer service experience. Summary of the Invention

[0004] To address the aforementioned technical issues, this application provides a call center collaborative calling method, device, medium, and product based on a large model.

[0005] The first aspect of this application provides a call center collaborative calling method based on a large model, employing the following technical solution: Obtain a multi-dimensional agent profile for each agent account in the human agent group. The multi-dimensional agent profile includes skill tags, ability level, and historical problem resolution success rate. Establish a voice call between the user and the robot, and collect the user's voice during the voice call in real time; The user's voice is converted into user speech text, and the preset robot voice is converted into robot response text. The user speech text and the robot response text are integrated to generate a structured text dialogue. The structured text dialogue is input into a preset large language model to obtain user intent, emotional state, and multi-dimensional question summaries; Based on the user's intent and emotional state, and according to preset human intervention rules, determine whether to trigger the human intervention process; When the manual intervention process is triggered, candidate agent accounts with matching skill tags are selected from the human agent group based on the user intent and the target entity information in the multi-dimensional question summary. Among the candidate agent accounts, based on the emotional state, the ability level, and the historical problem-solving success rate, a preset agent allocation strategy is used to calculate the agent matching score for each candidate agent account, and the target available agent account is determined based on the agent matching score. The voice call is transferred to the target available agent account, and the structured text dialogue and the multi-dimensional question summary are pushed to the workbench of the target available agent account.

[0006] By adopting the above technical solution, the shortcomings of the existing call center agent allocation mechanism—its singular and mechanical nature—are effectively overcome. By introducing multi-dimensional agent profiling and large language model analysis, intelligent and precise agent matching is achieved. Specifically, this solution analyzes call content in real time to accurately identify user intent, emotional state, and core issues. Based on this, it dynamically filters candidate agents with matching skills and further combines agent ability levels and historical success rates for optimal allocation. This significantly improves the efficiency and accuracy of problem-solving, shortens call duration, and ensures a consistent and high-quality service experience, achieving a technological leap from blind allocation to precise matching.

[0007] Optionally, the step of inputting the structured text dialogue into a preset large language model to obtain user intent, emotional state, and multi-dimensional question summaries includes: Construct a prompt template, which divides the prompt template into a task instruction area and a data input area; The structured text dialogue is filled into the data input area, and a preset task execution example is integrated into the task instruction area to generate prompt text; The prompt text is input into the large language model to obtain the raw response data; By matching the preset user intent field, emotional state field, and question summary field in the original response data, the user intent, the emotional state, and the multi-dimensional question summary are obtained.

[0008] By adopting the above technical solution, structured prompt templates and task examples are used to guide the large language model in analysis, effectively standardizing the model's output format and content focus, and ensuring the accuracy and consistency of information extraction such as user intent, sentiment, and question summaries. This method not only reduces the randomness and instability of directly applying the large language model in industrial scenarios, but also achieves a reliable conversion from unstructured text to standardized business data through a pre-processing step of pre-defined field matching. This provides high-quality and reliable decision-making basis for subsequent accurate agent matching, thereby improving the automation level and decision reliability of the entire system.

[0009] Optionally, the step of determining whether to trigger the manual intervention process based on the user's intent and emotional state, and according to preset manual intervention rules, includes: The user intent is analyzed, and when the content of the user intent belongs to a preset type of manual service request, it is determined that the first intervention condition is met. The emotional state is analyzed, and when the type of the emotional state belongs to a preset negative emotion category and the confidence score of the emotional state is higher than a preset emotion threshold, it is determined that the second intervention condition is met. Create and associate a repeat intent count value for the voice call, and initialize the repeat intent count value to a preset initial value; The user intent is compared with historical user intents in a preset historical database in a preset time sequence, and the duplicate intent count value is adjusted according to the comparison result. When the repeat intent count value exceeds the preset repeat count threshold, it is determined that the third intervention condition is met; When any one of the first intervention condition, the second intervention condition, or the third intervention condition is met, the manual intervention process is determined to be triggered.

[0010] By adopting the above technical solution, a multi-dimensional and dynamic human intervention judgment mechanism has been constructed, effectively overcoming the drawbacks of traditional single and rigid rules. This solution achieves precise control over the timing of human intervention by comprehensively considering three key dimensions: the user's direct service request, emotional state fluctuations, and the frequency of repeated intents. It can not only respond promptly to users' explicit needs and negative emotions, but also intelligently identify repeated inquiries caused by unresolved issues. Thus, while ensuring the necessary service experience, it effectively avoids the waste of resources caused by excessive intervention, significantly improving the accuracy of human agent dispatch and overall service efficiency.

[0011] Optionally, the step of filtering candidate agent accounts with matching skill tags from the human agent group based on the user intent and the target entity information in the multi-dimensional question summary includes: Multiple target entity information is parsed and extracted from the multi-dimensional problem summary; Based on multiple target entity information, a target skill tag set is searched and determined in a preset entity skill mapping table, and the demand weight corresponding to each skill tag in the target skill tag set is obtained, with one target entity information corresponding to one skill tag; Iterate through each agent account in the human agent group, compare the skill tags of the agent account with the target skill tag set, accumulate the demand weights corresponding to the successfully matched skill tags, and generate the comprehensive skill matching score of the agent account. The agent accounts whose comprehensive skill matching score is greater than the preset skill score threshold are identified as candidate agent accounts.

[0012] By employing the aforementioned technical solution, a refined matching process from problem to agent skills is achieved. This method analyzes key entities in the problem summary and transforms them into weighted skill tags using a mapping table. This shifts the assessment of agent capabilities from a single, vague judgment to a multi-dimensional, quantifiable comprehensive score. This weighted matching mechanism accurately identifies the candidate agent group whose skill combinations best match the current user's problem, providing a high-quality pool for the final selection and laying the foundation for accurate and efficient service from the outset.

[0013] Optionally, the step of calculating the agent matching score for each candidate agent account based on the emotional state, ability level, and historical problem-solving success rate using a preset agent allocation strategy, and determining the target available agent account based on the agent matching score, includes: Based on the emotional state, the affinity weight is obtained by querying the preset relationship between emotion and affinity. The candidate agent accounts are traversed, and the ability level and historical problem-solving success rate of each candidate agent account are input into a preset agent ability calculation formula to generate a basic agent score. The basic agent score is then combined with the affinity weight through a preset weighted calculation to generate the agent matching score. Within a preset statistical period, the total communication duration and total number of communications served by each candidate agent account are used to generate an agent load index through a preset load calculation function. The agent matching score is then adjusted based on the agent load index to obtain the final allocation score. From all candidate agent accounts whose communication status is idle, select the candidate agent account with the highest final allocation score as the target idle agent account.

[0014] By adopting the above technical solution, multi-dimensional comprehensive evaluation and optimal matching of candidate agents are achieved. This method not only considers the basic business ability score of agents, but also innovatively introduces an affinity weight based on the user's emotional state, ensuring that emotionally sensitive customers are assigned to agents with better communication skills. At the same time, by introducing an agent load index to dynamically adjust the matching score, the problem of overburdening highly capable agents is effectively avoided. While ensuring service quality and efficiency, it also takes into account a reasonable balance of agent workload, thereby achieving the optimal allocation and utilization of service resources.

[0015] Optionally, the method further includes: After the call is transferred to the target available agent account, the human service record generated by the target available agent account is obtained. The human service record includes the problem handling result and service summary tags. Based on the service summary tags, and under the preset tag update rules, update the skill tags of the target idle agent account; Based on the problem handling results, the historical problem resolution success rate of the target idle agent account is updated using a preset success rate calculation function; The structured text dialogue, the multi-dimensional question summary, and the human service record are associated to generate training samples, which are then stored in a preset training database. The large language model is then updated based on the training database.

[0016] By adopting the above technical solution, a complete closed-loop learning and optimization system was constructed. This method collects processing results after each human service interaction and dynamically updates the agent's skill profile and success rate data accordingly. This allows agent capability assessment to continuously evolve with practical experience, ensuring the timeliness and accuracy of agent profiles. Simultaneously, real interaction data is transformed into training samples for model iteration, enabling the large language model to self-optimize and continuously enhance its processing capabilities and intent recognition. This, in turn, gives the entire call center system a data-driven capability for continuous self-improvement.

[0017] Optionally, updating the historical problem resolution success rate of the target idle agent account based on the problem handling result using a preset success rate calculation function includes: The problem processing result is compared with a preset set of successful processing types, and a success status identifier is generated based on the comparison result. Based on the seat identifier of the target idle seat account, obtain the total number of existing successful services and the total number of existing processed services within the preset evaluation period; Based on the success status identifier, update the existing total number of successful services and the existing total number of processed services to obtain the updated total number of successful services and the updated total number of processed services; Substitute the total number of successful services after the update and the total number of services processed after the update into the success rate calculation function to generate the updated historical problem resolution success rate.

[0018] By adopting the above technical solution, dynamic, quantitative, and automated updates to the historical problem-solving success rate of agents are achieved. This method generates a clear success status identifier by comparing the result of each service with preset standards, ensuring the objectivity of the data source. Then, based on this identifier, the number of successful transactions and the total number of processes within a period are accurately accumulated, and the success rate is recalculated using a mathematical formula. This mechanism ensures that agent capability assessment data reflects their recent service performance in real time, maintaining high timeliness and accuracy in agent profiles. This provides more reliable and realistic data for subsequent agent matching decisions, driving continuous optimization of the entire system's matching accuracy.

[0019] A second aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the foregoing.

[0020] A third aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described in any of the preceding descriptions.

[0021] A fourth aspect of this application provides a computer program product that, when run on an electronic device, causes the electronic device to perform the method as described in any of the preceding claims.

[0022] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: By constructing a precise matching mechanism that integrates intelligent analysis of large language models with multi-dimensional agent profiles, the system achieves intelligent decision-making throughout the entire process, from user intent recognition and emotion perception to agent skill matching. It innovatively designs multi-dimensional human intervention judgment rules and comprehensive weight allocation strategies, effectively improving problem-solving efficiency and customer satisfaction. Simultaneously, through a closed-loop learning system, it continuously optimizes agent capability assessment and model performance, enabling the entire call center system to possess self-evolving data-driven capabilities, fundamentally solving the technical problems of mechanical and simplistic agent allocation and low service efficiency in existing technologies. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the system architecture of an embodiment of a call center collaborative calling method or a call center collaborative calling system based on a large model, which applies this application. Figure 2 This is a flowchart illustrating a call center collaborative calling method based on a large model disclosed in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application.

[0024] Explanation of reference numerals in the attached figures: 100, System architecture; 101, First terminal device; 102, Second terminal device; 103, Third terminal device; 104, Network; 105, Server; 301, Processor; 302, Communication bus; 303, User interface; 304, Network interface; 305, Memory. Detailed Implementation

[0025] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0026] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0027] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0028] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0029] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as model training applications, video recognition applications, web browser applications, social platform software, etc.

[0030] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP3 (Moving Picture Experts Group Audio Layer IV) players, laptops, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules (e.g., multiple software programs or software modules used to provide distributed services) or as a single software program or software module. No specific limitations are imposed here.

[0031] This embodiment discloses a call center collaborative calling method based on a large model. Figure 2 This is a flowchart illustrating a call center collaborative calling method based on a large model disclosed in an embodiment of this application, as shown below. Figure 2 As shown, the method includes the following steps: S201. Obtain a multi-dimensional agent profile of each agent account in the human agent group. The multi-dimensional agent profile includes skill tags, ability level and historical problem resolution success rate. This multi-dimensional agent profile is essentially a dynamically maintained structured dataset for each agent account, stored in a server-side database. Its purpose is to transform agent capabilities from vague, subjective impressions into quantifiable and comparable objective indicators. Specifically, there are several alternative ways to acquire the skill tags, ability levels, and historical problem-solving success rates included in the profile. Regarding skill tags, which aim to identify an agent's business expertise, one approach is for the system administrator to manually configure them initially based on the agent's onboarding training or business line, such as configuring tags like "product consultation" or "complaint handling." Another option is to allow agents to self-report their skills, which are then approved by their supervisors. A more preferred approach is for the system to automatically extract keywords and topics based on past service orders or call recordings using natural language processing technology. When an agent is detected frequently and successfully handling interactions related to "installment payments," the system can automatically associate or increase the weight of the "installment processing" skill tag, thus achieving dynamic tag updates. Regarding competency levels, the aim is to quantify the overall business skills of agents. Qualitative levels, such as "Junior," "Intermediate," and "Advanced," can be determined through performance evaluations by supervisors or quality control teams. More preferably, a numerical comprehensive scoring system can be adopted. For example, a pre-set weighted formula can be used to calculate a specific score, ranging from 1 to 100, combining multiple indicators such as agent's business knowledge test score, average call processing time, first-time problem resolution rate, and customer satisfaction. This allows for more refined quantitative assessment. Regarding historical problem resolution success rate, as objective data, its calculation logic is the number of times problems were successfully resolved within a pre-set statistical period (e.g., the past 30 days) divided by the total number of problems handled. The system can define "successfully resolved" in several ways: for example, based on the agent's active selection of a "problem resolved" checkbox on the system interface after the interaction; or by the system automatically sending a satisfaction survey to the user after the call ends and counting "resolved" or "satisfied" responses as successful; or by the system monitoring within a pre-set time window (e.g., 48 hours) that the user has not initiated repeated calls on the same topic, thus assuming the problem has been successfully resolved. The system's backend statistics service will periodically update the success rate data for each agent account based on these judgment results, ensuring the timeliness and accuracy of agent profiles and laying a solid data foundation for subsequent precise matching.

[0032] S202. Establish a voice call between the user and the robot, and collect the user's voice during the voice call in real time; In this embodiment, the system establishes a voice call connection between the user and the robot through an intelligent voice gateway component. When the user dials the call center's service number through a terminal device (e.g., a smartphone or landline), the voice request is routed to the intelligent voice gateway. The intelligent voice gateway converts the user's voice signal into a digital voice stream in real time and connects to the call center's server via a communication network (e.g., SIP or WebRTC protocol). On the server side, the Automatic Speech Recognition (ASR) module acquires the user's voice signal stream in real time and segments it into frame format data adapted to the sampling rate (typically 16kHz, 16-bit PCM format). Simultaneously, the system calls a preset voice playback module to play the robot's response voice. The response voice can be selected from a preset response library or directly synthesized by the TTS engine based on dynamically generated text for the user. The ASR module processes the acquired user voice in real time, generating intermediate or final text results of the user's voice, and integrates them synchronously with the robot's response text, laying the foundation for subsequent structured dialogue processing. Furthermore, the user's original voice stream is encrypted and stored on the server (e.g., using AES encryption algorithm) to ensure the security and confidentiality of the voice data. This embodiment ensures the real-time performance and accuracy of voice calls through the above combination of methods, while providing high-quality data support for subsequent large language model interaction and service optimization.

[0033] S203. Convert the user's voice into user speech text and convert the preset robot voice into robot response text, integrate the user speech text and the robot response text to generate a structured text dialogue; In this embodiment, after acquiring user speech, the system uses a built-in speech recognition engine to convert the user's speech in real time, transforming continuous speech signals into textual expressions of the user's words. Specifically, this process is implemented through speech recognition algorithms, such as a combination of acoustic modeling and language modeling based on deep neural networks, ensuring high accuracy and low latency in speech recognition. Simultaneously, the preset speech content generated by the robot is synchronously converted into corresponding robot response text. The conversion process directly generates the text based on the text input during speech synthesis, ensuring consistency with the user's interactive content. Subsequently, the system sequentially sorts and integrates the user's speech text and the robot's response text, generating a structured text dialogue based on the contextual relationships between the speech. The text content includes timestamp information, the identifier of the speech sender (such as user or robot), and the specific content of the user's speech and the robot's response. The structured processing uses rule matching to identify and store each round of dialogue between the user and the robot as a complete dialogue unit, while marking abnormal speech (such as interruptions or silence) for subsequent analysis and optimization. Furthermore, the system supports generating structured text output in various formats (such as JSON or XML) according to business needs, facilitating reading and processing by other application modules. Through the above implementation method, this embodiment can efficiently and accurately convert the voice interaction between the user and the robot into structured text records, providing complete information support for interaction analysis and business processing, and has high operability and scalability.

[0034] S204. Input the structured text dialogue into a preset large language model to obtain user intent, emotional state, and multi-dimensional question summary; Specifically, the system takes the structured text dialogue as input and submits it to a pre-defined large language model through a constructed prompt. This prompt not only includes the complete dialogue content but also includes explicit instructions, such as: "Please analyze the following dialogue, extract the user's core intent and current emotional state, and generate a multi-dimensional summary containing the core of the problem, key information, and attempted solutions." Regarding the implementation of the large language model, several alternative technical solutions exist: a common approach is to call a general-purpose large language model service provided through an API interface, such as OpenAI's GPT series or similar commercial models; a more preferred approach is to use a dedicated model fine-tuned in the customer service domain. This model, by learning from massive amounts of historical service tickets and labeled data, can more accurately identify industry terminology and user intent in specific business scenarios. After receiving and processing the input, the large language model generates a response containing the requested information. The system then parses this response into a structured data object, which explicitly contains the user's intent (e.g., "check repayment status"), emotional state (e.g., "anxious" or "calm"), and a multi-dimensional question summary, providing high-quality, multi-dimensional analysis results for subsequent intelligent decision-making and accurate allocation.

[0035] Optionally, the step of inputting the structured text dialogue into a preset large language model to obtain user intent, emotional state, and multi-dimensional question summary includes: constructing a prompt template, wherein the prompt template is divided into a task instruction area and a data input area; filling the structured text dialogue into the data input area and integrating a preset task execution example into the task instruction area to generate prompt text; inputting the prompt text into the large language model to obtain raw response data; and separating the user intent, the emotional state, and the multi-dimensional question summary by matching preset user intent fields, emotional state fields, and question summary fields in the raw response data.

[0036] Specifically, the system first constructs a structured prompt template. This template is logically divided into two core areas: a task instruction area and a data input area. The task instruction area defines the model's role, task objectives, and expected output format. For example, the instruction could be set as: "Please act as a professional customer service quality inspection expert, analyze the following dialogue, and return the user's core intent, emotional state, and a multi-dimensional summary information in JSON (JavaScript Object Notation) format." The data input area serves as a placeholder (e.g., using...).<dialogue_content> Tags are reserved for later filling in the actual dialogue text. This templated design aims to separate unchanging instructions from changing input data, greatly enhancing the stability and reproducibility of interaction with large language models.

[0037] Furthermore, the system will dynamically generate the final complete prompt text to be input. This process includes two key actions: First, the structured text dialogue generated in step S203 (which already includes speakers, timing, and text content) is formatted and completely filled into the placeholders in the data input area of ​​the prompt template. Second, to further guide the large language model to understand the task and follow the specified output format—that is, to achieve so-called contextual learning—the system will also integrate one or more preset task execution examples in the task instruction area. Each example is itself an input-output pair; for example, the example input is a short simulated dialogue, and the example output is a JSON structure that perfectly conforms to the expected format. By displaying these "success stories" in front of real tasks, the probability of the model outputting incorrect format or deviating from the expected content can be significantly reduced, thereby generating a final prompt text that is complete in information and highly guiding.

[0038] Furthermore, after generating the prompt text containing complete instructions, task examples, and real-time dialogue data, the system uses it as core input to call a pre-defined large language model to obtain the analysis results. This call process is typically completed through a standard API interface (such as an HTTP-based RESTful API). During the call, in addition to transmitting the prompt text, the system can configure a series of parameters that affect the model's generation behavior. For example, a lower "temperature" parameter value (such as 0.1 or 0.2) can be set to make the model's output more deterministic and consistent, avoiding unnecessary creative oversights; simultaneously, a maximum token count can be set to ensure that the response content is not excessively long or truncated. Upon receiving the request, the large language model performs deep semantic analysis and reasoning based on the prompt text, ultimately generating a raw response data. This raw response data is typically a text string, and ideally, its content and format will strictly adhere to the requirements of the task instructions and examples in the prompt text.

[0039] Furthermore, to transform the raw response data returned by the large language model into structured information that the application can directly utilize, the system needs to parse and extract its fields. Since the prompt template explicitly requires JSON output, the primary and preferred approach is for the system to first attempt to parse the raw response data string using a standard JSON parser. If parsing is successful, a program-operable data object is obtained. Subsequently, the system accurately extracts and separates the corresponding values ​​by matching preset user intent fields (e.g., user_intent), emotion state fields (e.g., emotion_state), and problem summary fields (e.g., problem_summary) as keys within this data object. As a necessary fault tolerance mechanism, if JSON parsing fails (e.g., the model outputs a non-strict JSON format), the system can activate an alternative regular expression-based matching scheme to capture the content of the required fields from the raw string using preset patterns. This ensures that even with flawed model output, effective user intent, emotion state, and multi-dimensional problem summaries can be extracted to the greatest extent possible.

[0040] S205. Based on the user's intent and emotional state, and according to preset human intervention rules, determine whether to trigger the human intervention process; After obtaining accurate user intent and emotional state, the system enters the intelligent decision-making stage. Its core is to determine, based on preset human intervention rules, whether the current conversation needs immediate transfer to a human agent, thereby enabling real-time intervention in high-risk or complex conversations. These rules are a set of logical conditions that can be flexibly configured by the administrator in the background. The system makes a final judgment by matching the user intent and emotional state output by the large language model with these rules. The specific judgment mechanism can be implemented in several ways, such as, but not limited to, the following three: The first embodiment uses a direct trigger rule based on keyword matching. This rule base defines a series of high-priority emotional states (such as anger, crying, roaring) and high-risk user intents (such as complaining, threatening to sue, demanding compensation). The system only needs to detect that the current conversation matches any item in the list to immediately trigger the human intervention process. This method is simple to implement, responds quickly, and can effectively intercept the most urgent negative situations. The second embodiment uses a more refined combination judgment logic. This rule base consists of a series of intent-emotion binary pairs. Intervention is only triggered when the user's current intent and emotional state simultaneously meet a preset combination. For example, the rule can be configured to: (trigger when the intent is to query a failed order and the emotion is anxiety or disappointment), or (trigger when the intent is to request a refund and the emotion is anger). This approach avoids unnecessary transfers caused by minor negative emotions in general consultations, significantly improving the processing efficiency of the automated system. The third embodiment implements a dynamic threshold judgment mechanism based on weighted scoring. In this scheme, the system presets different risk scores for different user intents and emotional states (e.g., 50 points for complaint intent, 50 points for anger; 5 points for query intent, 0 points for calm emotion). After each round of dialogue analysis, the system calculates the sum of the current intent score and emotion score. Once the total risk score exceeds a preset threshold (e.g., 80 points), the system triggers manual intervention. This approach provides high flexibility, allowing for the quantification and prioritization of the urgency of different situations based on business needs. Regardless of the method used, once the judgment result is triggered, the system marks the conversation as requiring manual handling and initiates subsequent transfer or assignment processes.

[0041] Optionally, the step of determining whether to trigger the manual intervention process based on the user intent and the emotional state, and according to preset manual intervention rules, includes: parsing the user intent; when the content of the user intent belongs to a preset type of manual service request, determining that a first intervention condition is met; parsing the emotional state; when the type of the emotional state belongs to a preset negative emotion category and the confidence score of the emotional state is higher than a preset emotion threshold, determining that a second intervention condition is met; creating and associating a repeated intent count value for the voice call, and initializing the repeated intent count value to a preset initial value; comparing the user intent with historical user intents in a preset historical database according to a preset time order, and adjusting the repeated intent count value according to the comparison result; when the repeated intent count value exceeds a preset repetition count threshold, determining that a third intervention condition is met; and determining that the manual intervention process is triggered when any one of the first intervention condition, the second intervention condition, or the third intervention condition is met.

[0042] Specifically, the system pre-configures and maintains a list of human service request types. This list is essentially a set of keywords or intent tags that can be updated by the administrator, containing service scenarios that inherently require, or that the user explicitly requests, human intervention. For example, this list may include, but is not limited to, intent tags such as: "I want to complain," "Transfer to human assistance," "Find a manager," "I want to sue," and "Handle complex business A." After parsing the user's intent in step S204, the system performs a precise or fuzzy match between the intent and this list. Once the determined user intent falls within any item in this pre-set list, the system determines that the first intervention condition has been met. The technical purpose of this solution is to directly intercept requests that automated processes cannot effectively handle or that users have lost patience with at the business logic level, thereby providing the highest priority intervention.

[0043] Furthermore, to accurately detect negative emotions arising from poor user experience, the system independently determines whether the second intervention condition is met. This process involves two verification steps: First, the system determines whether the emotional state type parsed from the large language model belongs to a preset set of negative emotion categories. This set can include explicit negative emotion labels such as anger, anxiety, disappointment, sadness, and aggression. Only when the emotion type matches this set will the system perform the second verification step, obtaining the confidence score corresponding to the emotional state (usually provided by the large language model in the output, ranging from 0 to 1) and comparing it with a preset emotion threshold (e.g., 0.85). Only when the emotion is a negative category and its confidence score is higher than the threshold will the system ultimately determine that the second intervention condition is met. This dual verification mechanism is designed to filter out negative emotions that the model may misjudge or that are mild, ensuring that human intervention is targeted at those scenarios where the model is highly confident in the significant outbreak of negative emotions, thereby improving the accuracy of the intervention.

[0044] Furthermore, to address scenarios where users repeatedly encounter the same unresolved problem, the system intervenes by determining if a third intervention condition is met. To this end, when a new voice call begins, the system creates a separate repeat intent count in memory or cache, initializing it to a preset value, typically 0. During each round of the call, after parsing the current user intent, the system compares it to historical user intents in a preset "historical database." This database can simply contain intent records from previous rounds of the call. The comparison logic is as follows: if the current intent is the same as the immediately preceding intent, the repeat intent count associated with that voice call is incremented by 1; if they are different, the count is reset to its initial value or decremented by 1. The system then compares the adjusted repeat intent count to a preset repetition threshold (e.g., a threshold of 3). Once the count exceeds this threshold, it means the user may have repeatedly expressed the same intent without receiving a satisfactory response, at which point the system determines that the third intervention condition is met. This approach can effectively identify and escalate conversations that are stuck in a deadlock, improving the user experience.

[0045] Furthermore, the system logically integrates the three parallel and independent judgment conditions mentioned above to make a final decision on whether to trigger the human intervention process. This integration logic employs the OR gate logic, meaning that if any one of the first, second, or third intervention conditions is satisfied, the system will set the overall judgment result to trigger. In other words, whether the user directly requests human intervention, displays strong negative emotions, or repeatedly gets stuck on the same problem, any of these situations is sufficient to initiate the human intervention process. This multi-dimensional, multi-path triggering mechanism ensures that the system can comprehensively monitor dialogue quality from three levels: business needs, emotional perception, and process status. This greatly improves the capture rate and response timeliness of potentially risky conversations, fundamentally guaranteeing the service baseline in complex scenarios.

[0046] S206. When it is determined that the manual intervention process is triggered, according to the user intent and the target entity information in the multi-dimensional question summary, candidate agent accounts with matching skill tags are selected from the manual agent group. The key target entity information contained in this summary, such as product model, order number, or business type, together with the user intent, constitutes the matching input conditions. Specific filtering and matching strategies can include, but are not limited to, the following three implementation methods: The first embodiment employs a direct matching strategy based on rigid labels. In this scheme, the system pre-configures a set of explicit skill labels for each agent account, such as refund processing, product A expert, or complaint handling. The system directly uses the parsed user intent (e.g., requesting a refund) and target entity information (e.g., product A) as search terms to find agents with the exact same skill labels in their accounts. This method is simple to implement, fast in matching, and suitable for scenarios with clear business divisions. The second embodiment employs a hierarchical matching strategy based on a structured skill tree. The system pre-establishes a tree-like skill classification system, for example, using after-sales service as the parent node and refunds and exchanges as its child nodes. The user's intent and target entity are mapped to specific leaf nodes of this skill tree, while the agent's skill label can be a leaf node or a higher-level parent node. During the screening process, the system not only searches for agents whose skill tags perfectly match the needs, but also for agents who possess the skills of the parent node of the need. For example, when a user's need is a refund, agents with after-sales service skills are also considered candidates. This approach greatly increases the flexibility and success rate of matching. The third implementation is a dynamic optimal matching strategy based on weighted scoring. The system decomposes the user's intent and target entity information into multiple need points and assigns weights to each need point. Simultaneously, it sets a proficiency level for each agent's skill tag, such as beginner, intermediate, or expert. During matching, the system iterates through all agents and calculates a comprehensive score based on the match between their skill tags and need points, as well as their own proficiency level. For example, a scoring model could be designed where intent matching accounts for 60% and entity information matching accounts for 40%. Ultimately, the system filters out agents with a comprehensive score higher than a preset threshold and can sort them by score, ensuring that customers are always transferred to the most suitable candidate agent.

[0047] Optionally, the step of selecting candidate agent accounts with matching skill tags from the human agent group based on the user intent and the target entity information in the multi-dimensional question summary includes: parsing and extracting multiple target entity information from the multi-dimensional question summary; searching and determining a set of target skill tags in a preset entity skill mapping table based on the multiple target entity information, and obtaining the demand weight corresponding to each skill tag in the target skill tag set, where one target entity information corresponds to one skill tag; traversing each agent account in the human agent group, comparing the skill tags of the agent account with the set of target skill tags, accumulating the demand weights corresponding to the successfully matched skill tags, and generating a comprehensive skill matching score for the agent account; and identifying agent accounts with a comprehensive skill matching score greater than a preset skill score threshold as candidate agent accounts.

[0048] As a more refined implementation of the aforementioned agent screening process, the system first needs to parse and extract multiple target entity information from the generated multi-dimensional question summary. The purpose of this step is to convert the user's unstructured spoken description into a set of precise, structured data labels that can be used for subsequent machine matching. This target entity information consists of core elements of the question, such as product model, order number, service package name, or specific error code. Specific extraction techniques can include: The first approach is rule-based extraction, such as matching strings with specific formats using pre-defined regular expressions, like recognizing 12-digit numbers starting with SN as order numbers. The second approach is extraction based on natural language understanding models, such as using a pre-trained named entity recognition model to automatically identify and label predefined entity categories from the summary text; this approach has stronger generalization capabilities. The extracted target entity information, such as {Product Model: Model A, Service Type: Package Upgrade}, will serve as input for the next step of precise matching.

[0049] Furthermore, after successfully extracting the target entity information, the system will search for and construct a set of target skill tags based on these entities, and quantify the importance of each skill in the set. To achieve this function, the system will pre-configure and maintain an entity-skill mapping table. This table is a key data structure that defines the specific skill tag corresponding to each target entity and the weight of that skill in solving the problem. For example, the mapping table may contain entries {Entity: Model A, Skill: Model A Hardware Support, Weight: 8} and {Entity: Package Upgrade, Skill: Package Service Processing, Weight: 10}. When the system extracts the entities Model A and Package Upgrade, it will determine the target skill tag set as {Model A Hardware Support, Package Service Processing} and obtain their respective demand weights of 8 and 10. The setting of demand weights can be static, pre-configured by business experts based on experience; or it can be dynamic, where the system can learn from historical data, for example, analyzing the skills that played a key role in past successful cases and dynamically increasing their corresponding weights, thereby making the weight allocation more scientific and closer to reality.

[0050] Furthermore, after determining the target skill tag set and their respective demand weights, the system begins to traverse each available agent account in the human agent group to calculate its comprehensive skill matching score. This process is a cyclical scoring process. For each agent account, the system first obtains its assigned skill tag list. Then, the system compares the agent's skill list with the target skill tag set determined in the previous step, one by one. If the agent's skill list contains a skill tag from the target skill set, the demand weight corresponding to that skill is added to the agent's temporary score. For example, if an agent possesses both the skills of "Model A Hardware Support" and "Package Service Processing," their comprehensive skill matching score is 8 + 10 = 18 points; if another agent only possesses the skill of "Package Service Processing," their score is 10 points. The comparison method here can be a completely exact match or a match that includes hierarchical relationships. For example, if an agent holds the parent skill of "Advanced Service Processing," it can also be considered as matching the child skill of "Package Service Processing," and can obtain all or part of its weight score.

[0051] Furthermore, based on the calculated comprehensive skill matching score, the system will perform a screening process to ultimately determine the range of candidate agent accounts. The system uses a preset skill score threshold as the screening criterion, which can be configured by the administrator according to service quality requirements. The system will compare the comprehensive skill matching score of each agent with this threshold. All agent accounts with a score greater than or equal to the threshold will be identified as candidate agents capable of resolving the current user's problem. For example, if the skill score threshold is set to 9, then the aforementioned agents with a score of 18 and 10 will both be selected as candidate agents. The final list of candidate agent accounts can be further sorted from highest to lowest matching score so that the subsequent allocation process prioritizes agents with the highest matching degree. This mechanism based on weighted scoring and threshold filtering ensures that service requests transferred to human agents are handled by agents with sufficient and relevant skills, effectively improving the first-time problem resolution rate and customer satisfaction.

[0052] S207. Among the candidate agent accounts, based on the emotional state, the ability level and the historical problem-solving success rate, calculate the agent matching score of each candidate agent account through a preset agent allocation strategy, and determine the target available agent account based on the agent matching score. Specifically, after identifying candidate agent accounts with the relevant skills, the system enters a more refined secondary matching stage, aiming to select the target agent most suitable for the current conversation scenario from the candidates. The core of this stage is to calculate the comprehensive matching score for each candidate agent based on three dimensions: the user's emotional state, the robot's own ability level, and the candidate agent's historical problem-solving success rate, using a preset agent allocation strategy. Specific allocation strategies can employ various technical solutions: The first embodiment is a scoring scheme based on a dynamic weighted model. This scheme sets baseline weights for multiple factors such as emotional state, ability level, and success rate, and dynamically adjusts them according to real-time conditions. For example, when the system detects that the user's emotion is anger, it automatically increases the weight of the agent's historical problem-solving success rate and empathy ability tags, thus giving experienced agents skilled at calming users a higher matching score. The second embodiment is an assignment scheme based on multi-level screening rules. The system filters candidate agents layer by layer according to priority, like a funnel. The first layer of rules might be: if a user is emotionally agitated, prioritize agents with a historical negative review rate below a preset value (e.g., 1%). The second layer might be: if a high-level bot still cannot resolve the issue, indicating complexity, further filter the remaining agents to those with expert or senior titles. Through these layers of filtering, one or a few optimal agents are ultimately identified. A third implementation is an optimization scheme based on a comprehensive utility function. This scheme views agent allocation as a process of maximizing overall service utility. Its function inputs include user waiting time, agent idle time, average processing time for this type of problem, and historical success rate. For example, the system might calculate utility scores for an expert agent who can quickly resolve the issue but requires an extra 30 seconds of waiting, and a regular agent who can connect immediately but may take longer. The system ultimately selects the idle agent that achieves the optimal overall utility (e.g., a balance between customer satisfaction and operating costs) as the target agent account and assigns a session to it.

[0053] Optionally, the step of calculating the agent matching score for each candidate agent account based on the emotional state, the ability level, and the historical problem-solving success rate using a preset agent allocation strategy, and determining the target available agent account based on the agent matching score, includes: based on the emotional state, querying and obtaining the affinity weight from a preset emotion-affinity correspondence; traversing the candidate agent accounts, inputting the ability level and historical problem-solving success rate of each candidate agent account into a preset agent ability calculation formula, and generating... The agent base score is combined with the affinity weight through a preset weighted calculation to generate the agent matching score. Within a preset statistical period, the total communication time and total number of communications served by each candidate agent account are queried. An agent load index is generated using a preset load calculation function, and the agent matching score is adjusted based on the agent load index to obtain the final allocation score. From all candidate agent accounts with an idle communication status, the candidate agent account with the highest final allocation score is selected as the target idle agent account.

[0054] Specifically, as a preferred embodiment of the agent allocation strategy, the system first quantifies the user's need for agent friendliness based on their emotional state. To achieve this, the system pre-configures an emotion-friendliness mapping table. This table is essentially a key-value database, where the key is the category of emotional state, such as anger, anxiety, or neutral, and the value is the corresponding friendliness weight. When the system identifies the user's current emotion as anger in the preceding steps, it will look up the higher friendliness weight corresponding to the anger state in this table, for example, 0.8. Conversely, if the user's emotion is neutral or positive, the retrieved friendliness weight will be lower, for example, 0.2. The friendliness weight can be pre-determined by experts based on psychological research and customer service operation experience, or the system can dynamically optimize and adjust it by analyzing user evaluations of agent service under different emotions in historical interactions through machine learning, thereby making the weight assignment more scientific. The purpose is to ensure that emotionally agitated users are preferentially assigned to agents with greater soothing ability and empathy.

[0055] Further, the system iterates through the candidate agent account list, calculating the objective business capabilities of each agent, i.e., the agent's base score. This score is calculated using a preset agent capability calculation formula, whose inputs are the agent's capability level and their historical problem-solving success rate. For example, a specific calculation formula could be: Agent Base Score = α × Quantitative Value of Agent Capability Level + β × Historical Problem-Solving Success Rate. Here, α and β are configurable coefficients used to adjust the proportion of capability level and success rate in the total score. Agent capability levels can be internally assessed, such as beginner, intermediate, and expert levels, which the system converts into numerical values ​​(e.g., 1, 3, 5) for calculation. After obtaining the agent base score, the system combines this base score with the affinity weight obtained in the previous step through a preset weighted calculation to generate an agent matching score. For example, the weighted formula could be designed as: Agent Matching Score = (1 - Affinity Weight) × Agent Base Score. This design ensures that when a user's emotions require high affinity (affinity is heavily weighted), the influence of the agent's base score on the final score will decrease accordingly, and vice versa. This achieves a dynamic balance between the agent's hard and soft skills.

[0056] Furthermore, to achieve more balanced order dispatch and avoid overloading of top agents while other agents remain idle, the system introduces an agent load index to adjust the matching score. The system queries the total communication duration and total number of communications for each candidate agent account within a preset statistical period (e.g., the past hour or the current workday). This data serves as input, fed into a preset load calculation function to generate the agent load index. A feasible load calculation function is: Agent Load Index = (Total Communication Duration / Total Duration of Statistical Period) × λ + (Total Number of Communications / Preset Maximum Number of Communications) × μ, where λ and μ are adjustment coefficients. A higher index indicates a busier agent. The system then dynamically adjusts the agent matching score based on this load index to obtain the final allocation score. The adjustment can employ a penalty mechanism, for example: Final Allocation Score = Agent Matching Score × (1 - Agent Load Index). In this way, agents with higher loads will have their final scores reduced, thus lowering their priority in the allocation queue.

[0057] Furthermore, based on the calculated final allocation score, the system makes a final decision among all candidate agent accounts with an idle communication status. The system filters out all agents currently in an accessible state and compares their final allocation scores. The candidate agent account with the highest score is identified as the target idle agent account for this session. The system then routes the user's session request to that target agent. This series of refined calculations, corrections, and comparisons not only considers the agent's skill matching but also integrates multiple factors such as user sentiment and real-time agent load. This achieves optimal allocation of service resources on a macro level and matches each user with the most suitable agent at the moment on a micro level, significantly improving service quality and operational efficiency.

[0058] S208. Transfer the voice call line to the target idle agent account, and push the structured text dialogue and the multi-dimensional question summary to the workbench of the target idle agent account.

[0059] Specifically, once a target available agent account is identified, the system executes a closed-loop operation of call transfer and information synchronization to ensure seamless human service. This process can be implemented in several ways: The first embodiment is a collaborative solution based on Computer Telephone Integration (CTI) and APIs. The system first sends a transfer command to the communication switch via CTI middleware. This command includes the session identifier of the current call and the extension number of the target agent, thereby redirecting the user's voice line. Immediately afterward, or almost simultaneously, the system's business logic layer encapsulates the generated structured text dialogue and multi-dimensional question summaries into a JSON data packet. This packet is then pushed via an HTTP POST request to the RESTful API of the target agent's workbench backend service, enabling a pop-up display of information. The second embodiment uses a real-time push solution based on WebSocket. A long-lived connection is maintained between the agent's workbench frontend and the backend server. After the transfer command is issued, the server immediately pushes a composite message containing context information and call events directly to the target agent's client via the WebSocket channel. Upon receiving the message, the client responds to the call via integrated WebRTC functionality or by controlling the local softphone client. Simultaneously, it immediately renders the complete conversation history and summary information on the interface, ensuring an exceptional audio-visual synchronization experience. The third implementation is a decoupled architecture based on an enterprise-grade message bus (e.g., Kafka or RabbitMQ). The transfer decision module publishes a transfer event to a designated topic. This event message body contains call line information, the target agent's identifier, and all contextual data. An independent CTI service and an agent workbench service both subscribe to this topic. Upon hearing the event, the CTI service performs a physical call transfer, while the agent workbench service retrieves the data from the event and pushes it to the corresponding agent. This approach achieves loose coupling between system modules, improving system robustness and scalability. Regardless of the chosen solution, the core objective is to enable agents to fully grasp the user's historical requests the moment the call is answered, eliminating the need for the user to repeat their statements and thus providing efficient and professional service.

[0060] Optionally, the method further includes: after the call is transferred to the target available agent account, obtaining the human service record generated by the target available agent account, the human service record including problem handling results and service summary tags; updating the skill tags of the target available agent account according to the service summary tags and under a preset tag update rule; updating the historical problem resolution success rate of the target available agent account based on the problem handling results and through a preset success rate calculation function; associating the structured text dialogue, the multi-dimensional problem summary, and the human service record to generate training samples, storing the training samples in a preset training database, and updating the large language model according to the training database.

[0061] Specifically, when a human agent's service session ends, the system automatically triggers a service summary module on the agent's workbench interface. This module requires the agent to record the interaction, which typically includes two core parts. The first part is the issue resolution result, where the agent selects a status from a predefined dropdown menu or option group. These statuses indicate the final handling status of the service, such as whether the issue is resolved, requires follow-up, or has been escalated to the second-line support team. The second part is the service summary tags, where the agent selects one or more keywords highly relevant to the service content from a dynamic tag pool. These keywords cover common business scenarios, such as account permission issues, payment failure handling, and product feature inquiries. The system's backend service captures this structured record submitted by the agent and uses it for subsequent automated processing.

[0062] Furthermore, the system analyzes this human service record to dynamically and accurately update the agent's profile. On one hand, the system adjusts the agent's skill tag spectrum based on the service summary tags selected by the agent and a set of preset tag update rules. For example, a specific rule could be that when an agent handles a specific type of problem a certain number of times within a statistical period reaches a preset threshold, and the problem handling results are mostly resolved, the system automatically increases the agent's weight in that skill area or adds a new skill expertise tag. On the other hand, the system uses information from the problem handling results to refresh the agent's historical problem-solving success rate using a preset success rate calculation function. This function can use a moving average algorithm, calculating only data from the most recent period to reflect the agent's real-time performance and ability fluctuations, making the success rate indicator more timely.

[0063] Furthermore, to enable continuous learning and iteration of the large language model, the system encapsulates each complete user interaction process, from robot reception to human termination, into a high-quality training sample. Specifically, the system establishes a strong data-level correlation between the structured text dialogue and multi-dimensional question summaries generated by the robot in the front end of the interaction and the human service records generated by the human agent in the back end. This correlation is achieved through a unique session ID that runs throughout the entire process. The resulting data entity is a complete record encompassing the question context, the machine's initial understanding, the human intervention process, the final processing result, and authoritative content tags. This record constitutes a complete portrayal of a real-world problem and its standard solution.

[0064] Furthermore, these newly generated training samples will be systematically stored in a pre-defined dedicated training database. The system can be configured with two triggering mechanisms to update the underlying large language model. The first is batch triggering, where the training task is automatically started when the number of new samples accumulated in the database reaches a preset batch size, such as one thousand. The second is periodic triggering, for example, during the business off-peak period each weekend, the system automatically pulls all new samples for model training within that period. The training process itself can be full training or more efficient incremental fine-tuning. By continuously learning from real-world service cases that have been manually verified, the large language model can continuously calibrate and improve its accuracy in semantic understanding, summary generation, and preliminary solution recommendation in specific business scenarios, thereby providing users with more accurate and efficient automated services in future interactions, forming a self-reinforcing intelligent closed loop.

[0065] Optionally, updating the historical problem resolution success rate of the target idle agent account based on the problem handling result and using a preset success rate calculation function includes: comparing the problem handling result with a preset set of successful handling types, and generating a success status identifier based on the comparison result; obtaining the total number of existing successful services and the total number of existing processed services within a preset evaluation period based on the agent identifier of the target idle agent account; updating the total number of existing successful services and the total number of existing processed services based on the success status identifier to obtain the updated total number of successful services and the updated total number of processed services; and substituting the updated total number of successful services and the updated total number of processed services into the success rate calculation function to generate the updated historical problem resolution success rate.

[0066] In a more specific embodiment, the process of updating the historical problem resolution success rate of the target idle agent account is broken down into a series of precise automated steps. First, the system needs to perform a standardized evaluation of the problem handling results submitted by the agents. The system pre-configures a set of successful handling types, which defines all states considered successful services, such as resolved, user-confirmed satisfaction, and one-time resolution. When the system obtains the problem handling result for this service, for example, if the agent selects "resolved," the system compares this result with the members in the successful handling type set. If the comparison is successful, the system generates a success status identifier representing success, which can be a Boolean value of true or a numerical value of 1; conversely, if there is no match, for example, if the agent selects "unresolved and awaiting follow-up," an identifier representing unsuccessful service is generated. This step transforms the agent's subjective input into a binary logical judgment that can be directly processed by the machine.

[0067] Furthermore, to perform incremental calculations, the system needs to obtain the agent's historical service data within the current evaluation period. Based on the agent's identifier (e.g., employee number A007), the system initiates a query request to the backend agent profile database or performance statistics module. The purpose of this request is to obtain two key cumulative values: the total number of successful services and the total number of services processed. The preset evaluation period is a configurable parameter, such as the current calendar month, the last 30 days (rolling cycle), or the current quarter, to adapt to different performance management needs. After executing the query, the system obtains these two basic count values, which serve as the basis for calculating the success rate before the update.

[0068] Furthermore, after obtaining the success status flag and historical count value, the system performs a count update operation. This is a simple conditional accumulation process. The system always increments the current total number of processed services by one, because regardless of whether the service is successful or not, it represents the completion of one service processing step. Then, the system checks the success status flag generated in the first step. If the flag indicates success, the system also increments the current total number of successful services by one. If the flag indicates failure, the current total number of successful services remains unchanged. Through this step, the system obtains two updated values: the updated total number of successful services and the updated total number of processed services. These two values ​​accurately reflect the latest cumulative situation, including the current service.

[0069] Furthermore, the system substitutes the updated values ​​into a preset success rate calculation function to generate the final updated historical problem resolution success rate. In a simple implementation, this success rate calculation function can be a basic division formula: the total number of successful services after the update divided by the total number of services processed after the update, then multiplied by 100%. In a more complex implementation, this function can also introduce a time decay factor or weights for different problem types to obtain a weighted success rate. The calculated new success rate value, for example, an increase from 85.00% to 85.14%, is immediately written back to the agent profile database, overwriting the old value. In this way, the system completes a real-time, accurate, and automated update of the agent's historical problem resolution success rate, ensuring the timeliness and accuracy of the agent capability assessment data.

[0070] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0071] This embodiment also discloses an electronic device, as shown in the reference. Figure 3 The electronic device may include: at least one processor 301, at least one communication bus 302, a user interface 303, a network interface 304, and at least one memory 305. The communication bus 302 is used to enable communication between these components. The user interface 303 may include a display screen or a camera; optionally, the user interface 303 may also include a standard wired interface or a wireless interface. The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0072] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.

[0073] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory 305 may include a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. Figure 3 As shown, the memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a call center collaborative calling method based on a large model.

[0074] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the disclosure in this specification. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A call center collaborative calling method based on a large model, characterized in that, Applied to a server, the method includes: Obtain a multi-dimensional agent profile for each agent account in the human agent group. The multi-dimensional agent profile includes skill tags, ability level, and historical problem resolution success rate. Establish a voice call between the user and the robot, and collect the user's voice during the voice call in real time; The user's voice is converted into user speech text, and the preset robot voice is converted into robot response text. The user speech text and the robot response text are integrated to generate a structured text dialogue. The structured text dialogue is input into a preset large language model to obtain user intent, emotional state, and multi-dimensional question summaries; Based on the user's intent and emotional state, and according to preset human intervention rules, determine whether to trigger the human intervention process; When the manual intervention process is triggered, candidate agent accounts with matching skill tags are selected from the human agent group based on the user intent and the target entity information in the multi-dimensional question summary. Among the candidate agent accounts, based on the emotional state, the ability level, and the historical problem-solving success rate, a preset agent allocation strategy is used to calculate the agent matching score for each candidate agent account, and the target available agent account is determined based on the agent matching score. The voice call is transferred to the target available agent account, and the structured text dialogue and the multi-dimensional question summary are pushed to the workbench of the target available agent account.

2. The method according to claim 1, characterized in that, The process involves inputting the structured text dialogue into a pre-defined large language model to obtain user intent, emotional state, and multi-dimensional question summaries, including: Construct a prompt template, which divides the prompt template into a task instruction area and a data input area; The structured text dialogue is filled into the data input area, and a preset task execution example is integrated into the task instruction area to generate prompt text; The prompt text is input into the large language model to obtain the raw response data; By matching the preset user intent field, emotional state field, and question summary field in the original response data, the user intent, the emotional state, and the multi-dimensional question summary are obtained.

3. The method according to claim 1, characterized in that, The step of determining whether to trigger a human intervention process based on the user's intent and emotional state, and according to preset human intervention rules, includes: The user intent is analyzed, and when the content of the user intent belongs to a preset type of manual service request, it is determined that the first intervention condition is met. The emotional state is analyzed, and when the type of the emotional state belongs to a preset negative emotion category and the confidence score of the emotional state is higher than a preset emotion threshold, it is determined that the second intervention condition is met. Create and associate a repeat intent count value for the voice call, and initialize the repeat intent count value to a preset initial value; The user intent is compared with historical user intents in a preset historical database in a preset time sequence, and the duplicate intent count value is adjusted according to the comparison result. When the repeat intent count value exceeds the preset repeat count threshold, it is determined that the third intervention condition is met; When any one of the first intervention condition, the second intervention condition, or the third intervention condition is met, the manual intervention process is determined to be triggered.

4. The method according to claim 1, characterized in that, The step of selecting candidate agent accounts with matching skill tags from the human agent group based on the user intent and the target entity information in the multi-dimensional question summary includes: Multiple target entity information is parsed and extracted from the multi-dimensional problem summary; Based on multiple target entity information, a target skill tag set is searched and determined in a preset entity skill mapping table, and the demand weight corresponding to each skill tag in the target skill tag set is obtained, with one target entity information corresponding to one skill tag; Iterate through each agent account in the human agent group, compare the skill tags of the agent account with the target skill tag set, accumulate the demand weights corresponding to the successfully matched skill tags, and generate the comprehensive skill matching score of the agent account. The agent accounts whose comprehensive skill matching score is greater than the preset skill score threshold are identified as candidate agent accounts.

5. The method according to claim 1, characterized in that, The process of calculating a seat matching score for each candidate agent account based on its emotional state, ability level, and historical problem-solving success rate using a preset agent allocation strategy, and determining a target available agent account based on the seat matching score, includes: Based on the emotional state, the affinity weight is obtained by querying the preset relationship between emotion and affinity. The candidate agent accounts are traversed, and the ability level and historical problem-solving success rate of each candidate agent account are input into a preset agent ability calculation formula to generate a basic agent score. The basic agent score is then combined with the affinity weight through a preset weighted calculation to generate the agent matching score. Within a preset statistical period, the total communication duration and total number of communications served by each candidate agent account are used to generate an agent load index through a preset load calculation function. The agent matching score is then adjusted based on the agent load index to obtain the final allocation score. From all candidate agent accounts whose communication status is idle, select the candidate agent account with the highest final allocation score as the target idle agent account.

6. The method according to claim 1, characterized in that, The method further includes: After the call is transferred to the target available agent account, the human service record generated by the target available agent account is obtained. The human service record includes the problem handling result and service summary tags. Based on the service summary tags, and under the preset tag update rules, update the skill tags of the target idle agent account; Based on the problem handling results, the historical problem resolution success rate of the target idle agent account is updated using a preset success rate calculation function; The structured text dialogue, the multi-dimensional question summary, and the human service record are associated to generate training samples, which are then stored in a preset training database. The large language model is then updated based on the training database.

7. The method according to claim 6, characterized in that, The step of updating the historical problem resolution success rate of the target idle agent account based on the problem handling result, using a preset success rate calculation function, includes: The problem processing result is compared with a preset set of successful processing types, and a success status identifier is generated based on the comparison result. Based on the seat identifier of the target idle seat account, obtain the total number of existing successful services and the total number of existing processed services within the preset evaluation period; Based on the success status identifier, update the existing total number of successful services and the existing total number of processed services to obtain the updated total number of successful services and the updated total number of processed services; Substitute the total number of successful services after the update and the total number of services processed after the update into the success rate calculation function to generate the updated historical problem resolution success rate.

8. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.

10. A computer program product, characterized in that, When the computer program product is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1-7.