Intelligent seat assisting method and system

By combining real-time monitoring and multi-dimensional dialogue deadlock entropy assessment with a dual-objective dialogue value network, the intelligent agent assistance system solves the problems of single dialogue strategy objectives and inaccurate decision-making in existing technologies. It enables earlier and more accurate intervention and strategy optimization, thereby improving customer experience and trust in human-machine collaboration.

CN121985071APending Publication Date: 2026-05-05SHANGHAI HAOYI INFORMATION SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI HAOYI INFORMATION SCI & TECH CO LTD
Filing Date
2026-02-04
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing intelligent agent assistance systems suffer from problems such as a single objective in dialogue strategies, inaccurate judgment of intervention timing in complex dialogue scenarios, and inability to intelligently and securely arbitrate when there are conflicts in strategy suggestions from multiple modules. These issues lead to a decline in customer experience and low agent trust.

Method used

By monitoring the dialogue process in real time, calculating the multidimensional dialogue deadlock entropy, generating rule-based first auxiliary information and deep inference-based second auxiliary information, and introducing a dual-objective dialogue value network to evaluate task-related value and user experience-related value, combining risk thresholds for arbitration, generating a final auxiliary strategy, and providing strategy explanation.

Benefits of technology

It achieves the goal of ensuring task completion efficiency while taking into account customer emotional experience, improving the timeliness and effectiveness of dialogue intervention, enhancing trust in human-machine collaboration, and solving the problems of inaccurate decision-making and improper conflict handling in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985071A_ABST
    Figure CN121985071A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent seat assisting method and system, and relates to the technical field of artificial intelligence. The method comprises the following steps: monitoring a dialogue in real time and calculating a dialogue state measurement index; generating first auxiliary information based on rules; when the index exceeds a trigger threshold value, generating an evaluation result containing a task related value and a user experience related value for the candidate action, and selecting one candidate action as second auxiliary information; when the seat action type suggested by the first auxiliary information is different from the seat action type suggested by the second auxiliary information, comparing the user experience related value of the second auxiliary information with a risk threshold to determine a final auxiliary strategy; and presenting the final strategy to the seat. According to the method, by introducing the dual-target value evaluation and the risk-based conflict arbitration mechanism, the emotional experience of the customer can be considered while the task efficiency is guaranteed, the man-machine cooperation credibility is enhanced through strategy interpretation, and the timeliness of intervention and the security of decision making are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an intelligent agent assistance method and system. Background Technology

[0002] In the field of intelligent customer service and call centers, intelligent agent assistance systems aim to improve service efficiency and quality. Existing technologies mainly provide assistance in two ways: one is based on rules or simple models to recommend immediate dialogue based on the user's current utterance; the other is to introduce models such as reinforcement learning to generate dialogue strategies by optimizing a single long-term, business-related objective (such as task completion rate).

[0003] However, existing technologies have significant shortcomings in practical applications. First, most reinforcement learning-based systems are task-completion oriented, ignoring the customer's emotional experience during the conversation. This may lead the system to recommend "cold" or "brutal" strategies to achieve business goals, potentially resulting in a decline in customer experience even if the task is ultimately completed. Second, while some systems attempt to introduce dynamic triggering mechanisms, their triggering conditions are usually based on isolated events or simple thresholds, making it difficult to accurately and early identify complex "dialogue deadlocks" caused by multiple factors, leading to delayed system intervention. Finally, when suggestions from different modules within the system contradict each other, existing technologies generally lack an intelligent arbitration mechanism to resolve such conflicts and cannot explain to agents why a seemingly "counterintuitive" strategy was chosen. This "black box" decision-making process makes it difficult for agents to build trust in the system, thereby reducing the adoption rate of effective suggestions. Summary of the Invention

[0004] This application aims to address the technical problems existing in the prior art, such as the single objective of dialogue strategies, inaccurate judgment of the timing of intervention in complex dialogue scenarios, and the inability to conduct intelligent and secure arbitration when there are conflicts in the strategy suggestions of multiple modules.

[0005] To address the aforementioned technical problems, this application provides an intelligent agent assistance method, comprising: real-time monitoring of the dialogue process between the agent and the customer, and calculating a dialogue state metric representing the dialogue state; generating rule-based first assistance information; when the dialogue state metric exceeds a preset trigger threshold, generating an evaluation result for at least one candidate action, the evaluation result including task-related value and user experience-related value, and selecting one of the candidate actions as second assistance information based on the evaluation result; when the type of agent action suggested by the first assistance information is different from the type of agent action suggested by the second assistance information, comparing the user experience-related value corresponding to the second assistance information with a preset risk threshold to determine a final assistance strategy; when the type of agent action suggested by the first assistance information is the same as the type of agent action suggested by the second assistance information, selecting the first assistance information as the final assistance strategy; and presenting the final assistance strategy to the agent.

[0006] Optionally, the dialogue state metric is a multidimensional dialogue deadlock entropy, the calculation of which integrates at least two of the following features: intent information entropy calculated based on the dialogue intent distribution, the proportion of abnormal silence duration on the client side, and the variance of the client's speech rate variation. By integrating features from multiple dimensions such as semantics, acoustics, and behavior, signs of a dialogue deadlock can be identified more accurately and earlier, thereby enabling timely intervention.

[0007] Optionally, the task-related value and the user experience-related value are generated by a bi-objective dialogue value network; the task-related value is the long-term reward Q-value for task completion, and the user experience-related value is the expected future emotional value E-Value. This approach enables parallel quantitative evaluation of task completion and user experience, providing a more comprehensive value reference for subsequent decision-making.

[0008] Optionally, comparing the user experience-related value corresponding to the second auxiliary information with a preset risk threshold to determine the final auxiliary strategy further includes: if the user experience-related value corresponding to the second auxiliary information is lower than the risk threshold, then the first auxiliary information is selected as the final auxiliary strategy; if the user experience-related value corresponding to the second auxiliary information is equal to or higher than the risk threshold, then the second auxiliary information is selected as the final auxiliary strategy. This arbitration mechanism, by assessing the emotional risks that high-value strategies may bring, achieves the maximization of long-term comprehensive value under the premise of controllable risk, avoiding the situation where the user experience is seriously deteriorated due to the adoption of high-risk strategies.

[0009] Optionally, the method further includes: if the second auxiliary information is selected as the final auxiliary strategy, generating a strategy explanation text based on its corresponding task-related value and user experience-related value to explain the reasons for the strategy selection, and presenting the strategy explanation text together with the final auxiliary strategy. This strategy explanation function informs the agent of the value trade-offs behind the AI ​​decision in a concise way, solving the "black box" problem and enhancing the agent's understanding and trust in AI suggestions.

[0010] This application also provides an intelligent agent assistance system for performing the method described in any of the preceding claims, comprising: a monitoring module for real-time monitoring of the dialogue process between the agent and the customer, and calculating a dialogue state metric representing the dialogue state; a prediction module, electrically connected to the monitoring module, for receiving the dialogue state metric, and for generating rule-based first auxiliary information through an auxiliary model, and, when the dialogue state metric exceeds a preset trigger threshold, generating an evaluation result based on a dialogue strategy evaluation model for at least one candidate action, the evaluation result including task-related value and user experience-related value, and selecting a candidate action as second auxiliary information based on the evaluation result; and an arbitration module, electrically connected to the prediction module, for receiving... The system describes the first auxiliary information and the second auxiliary information, and is used for: when the type of agent action suggested by the first auxiliary information is different from the type of agent action suggested by the second auxiliary information, determining the final auxiliary strategy by comparing the user experience-related value corresponding to the second auxiliary information with a preset risk threshold; when the type of agent action suggested by the first auxiliary information is the same as the type of agent action suggested by the second auxiliary information, selecting the first auxiliary information as the final auxiliary strategy; and generating a source identifier indicating whether the final auxiliary strategy originates from the first auxiliary information or the second auxiliary information; a presentation module, electrically connected to the arbitration module, is used to receive and present the final auxiliary strategy to the agent.

[0011] Optionally, the monitoring module is used to calculate a multidimensional dialogue deadlock entropy as a measure of dialogue state. The calculation of the multidimensional dialogue deadlock entropy integrates at least two of the following features: intent information entropy calculated based on the dialogue intent distribution, the proportion of abnormal silence duration on the client side, and the variance of the client's speech rate change.

[0012] Optionally, the prediction module includes a bi-objective dialogue value network, which is used to generate the task-related value and the user experience-related value.

[0013] Optionally, the arbitration module is configured to: select the first auxiliary information as the final auxiliary strategy when the user experience-related value corresponding to the second auxiliary information is lower than the risk threshold; and select the second auxiliary information as the final auxiliary strategy when the user experience-related value corresponding to the second auxiliary information is equal to or higher than the risk threshold.

[0014] Optionally, the presentation module is further configured to: receive the source identifier, and when the source identifier indicates that the final assistance strategy originates from the second assistance information, generate a strategy explanation text based on its corresponding task-related value and user experience-related value, and present the strategy explanation text together with the final assistance strategy.

[0015] Compared with the prior art, the technical solution provided in this application has the following beneficial effects:

[0016] 1. Achieved decision-making that balances task and emotion: The system quantifies the assessment of task and emotion through a dual-objective dialogue value network and introduces a conflict arbitration mechanism based on risk thresholds. This enables the system to recommend the strategy with the best long-term comprehensive value under the premise that the assessment of emotional risk is controllable, thus solving the problem of poor communication effect caused by the single objective or simple decision-making of existing technologies.

[0017] 2. Achieved accurate and early warning for complex scenarios: By using multi-dimensional dialogue deadlock entropy, the vague concept of "deadlock" is transformed into a quantifiable multi-dimensional indicator, which integrates subtle features such as semantics, acoustics, and behavior. It can identify signs of dialogue deadlock earlier and more accurately than existing technologies, improving the timeliness and effectiveness of intervention.

[0018] 3. Achieved trustworthy human-machine collaboration: Through conflict arbitration and strategy interpretation generation functions, the "black box" problem of AI suggestions has been solved. It not only resolves internal strategy conflicts through intelligent risk assessment, but its strategy interpretation function also clearly and concisely informs agents of the value trade-offs behind AI decisions, enhancing agents' understanding and trust in AI suggestions and improving the efficiency of human-machine collaboration. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the architecture of the intelligent agent assistance system provided in the embodiments of this application.

[0021] Figure 2 This is a flowchart illustrating the intelligent agent assistance method provided in an embodiment of this application.

[0022] Figure 3 This is a schematic diagram of the structure of the dual-objective dialogue value network (DTD-VN) provided in an embodiment of this application.

[0023] Figure 4 This is a timing diagram illustrating the signaling interaction between the modules in the embodiments of this application.

[0024] Explanation of reference numerals in the attached figures:

[0025] 10 - Training Module; 20 - Monitoring Module; 30 - Prediction Module; 31 - Fast Response Layer; 32 - Deep Inference Layer; 40 - Arbitration Module; 50 - Presentation Module; 60 - Data Storage; 100 - Intelligent Agent Assistance System; 321 - Session State Vector; 322 - Shared Encoding Layer; 323 - Task Completion Value Header; 324 - Future Emotional Value Header; 325 - Q_task Output; 326 - E_Value Output. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] Unless otherwise defined, the technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0028] Before providing a further detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.

[0029] (1) Dialogue state measurement index: refers to at least one numerical value used to quantify the current state of the dialogue between the agent and the customer. This index can be calculated based on at least one dimension such as the text content, acoustic features, and interactive behavior of the dialogue. Its core function is to serve as a basis for judging whether the dialogue has reached a deadlock or an anomaly has occurred, thereby triggering the subsequent auxiliary strategy generation process.

[0030] (2) First auxiliary information: refers to the immediate auxiliary suggestions generated by modules with fast response speed and low computational overhead (such as engines or lightweight models based on preset rules). It is usually judged based on the direct information of the current dialogue round (such as customer intent, keywords), and aims to provide agents with quick and routine scripts or operational guidance.

[0031] (3) Secondary auxiliary information: refers to auxiliary suggestions generated by modules with more complex computations that are capable of in-depth analysis and long-term value assessment under specific conditions (such as when the conversation reaches a deadlock). This information not only considers the current situation, but also focuses on the final completion effect of the entire conversation task and its impact on future customer relationships, and therefore usually has higher strategic value.

[0032] (4) Task-related value: This refers to the quantitative assessment value of long-term returns directly related to the completion of core business objectives. For example, in a customer service scenario, it can be the prediction of the long-term value brought by tasks such as successfully resolving customer problems, completing an order, or effectively retaining a customer. This value focuses on whether the "task" can be accomplished.

[0033] (5) User experience-related value: This refers to the quantitative assessment value of the customer's future emotional state or overall service experience after an agent takes a certain action. This value focuses on the customer's feelings during the interaction process and subsequent processes, that is, whether the "human" experience is good. A positive value indicates that the customer's mood may improve, while a negative value indicates the risk of a decline in the customer experience.

[0034] like Figures 1 to 4 As shown, the embodiments of this application aim to provide an intelligent agent assistance method and system to solve the technical problems existing in the prior art, such as the single objective of dialogue strategies, inaccurate judgment of intervention timing in complex dialogue scenarios, and the inability to perform intelligent and secure arbitration when multiple module strategy suggestions conflict. This solution, by introducing a dual value assessment system and a dynamic conflict arbitration mechanism, achieves both ensuring task completion efficiency and taking into account the emotional experience of customers, and can enhance the trust level of human-machine collaboration by generating strategy explanations.

[0035] Reference Figure 1 This application provides an architecture for an intelligent agent assistance system 100 (hereinafter referred to as "System 100"). System 100 can be deployed in scenarios such as call centers and online customer service platforms, providing intelligent decision support to agents by analyzing the interaction process between agents and customers. System 100 mainly includes a monitoring module 20, a prediction module 30, an arbitration module 40, and a presentation module 50. Furthermore, System 100 may also include a training module 10 for offline model training and a data storage module 60 for data persistence.

[0036] In one specific implementation, the intelligent agent assistance method provided in this application begins with real-time monitoring of the conversation between the agent and the customer. For example... Figure 1 and Figure 2 As shown, the monitoring module 20 is responsible for capturing dialogue data in real time. This data may include speech streams, text converted from speech, and various accompanying signals during the call. The monitoring module 20 processes and analyzes this raw data to calculate at least one dialogue state metric that can characterize the current dialogue state. This dialogue state metric is the starting point for all subsequent decision-making processes, and its accuracy is directly related to the effectiveness of the system's intervention.

[0037] After acquiring the dialogue data, the prediction module 30 begins its operation. For example... Figure 1 As shown, the prediction module 30 can be designed with two parallel layers: a fast-response layer 31 and a deep inference layer 32. The fast-response layer 31 is responsible for generating rule-based first auxiliary information. Here, "rules" are generalized; they can be a series of predefined "if-then" logic statements or a lightweight classification model. For example, when a customer inquires about "price," the fast-response layer 31 can quickly generate a suggested sales pitch for the agent—this is the first auxiliary information. This process has low latency, ensuring efficient response to routine questions.

[0038] Meanwhile, the monitoring module 20 continuously sends the calculated dialogue state metrics to the deep inference layer 32 of the prediction module 30. The deep inference layer 32 has a preset trigger threshold. It compares the received dialogue state metrics with this threshold in real time. If the metrics do not exceed the threshold, it indicates that the dialogue is progressing smoothly, and the deep inference layer 32 remains silent to conserve computing resources. This design demonstrates efficient use of computing resources and avoids unnecessary complex analysis in simple dialogue scenarios.

[0039] When the dialogue state metrics exceed a preset trigger threshold, it indicates that the dialogue may have reached a stalemate, there is a risk of escalating negative emotions from the customer, or it has entered a critical juncture requiring careful decision-making. At this point, the deep inference layer 32 is activated and begins to execute its core functions. Once activated, the deep inference layer 32 evaluates a series of candidate actions for the current dialogue state. These candidate actions can be a predefined set of strategies, such as "appease," "reject," "transfer," and "offer compensation."

[0040] The deep reasoning layer 32 evaluates each candidate action, generating an evaluation result that includes both task-related value and user experience-related value. Task-related value aims to predict the action's contribution to achieving the final business goal, while user experience-related value aims to predict the action's impact on customer sentiment and long-term satisfaction. Through this dual-value evaluation system, the system 100 must simultaneously consider both task completion effectiveness and user interaction experience when making decisions.

[0041] After evaluating all candidate actions, the deep inference layer 32 selects an optimal candidate action based on these evaluation results. The selection can be based on a scoring function that combines task-related value and user experience-related value. For example, it can select the action that maximizes task-related value, or the action with the highest task-related value while ensuring that the user experience-related value does not fall below a certain threshold. This selected optimal action constitutes the second auxiliary information.

[0042] Both the first auxiliary information and the second auxiliary information (generated upon triggering) are sent to the arbitration module 40. The core responsibility of the arbitration module 40 is to process these two types of auxiliary information, which may have different sources, logics, and even objectives, to form a unified, secure, and effective final recommendation. This process is crucial for ensuring the consistency and reliability of the system output.

[0043] The arbitration module 40 first determines whether the type of action suggested by the first auxiliary information is the same as the type of action suggested by the second auxiliary information. The action type can be divided according to its underlying intent. For example, "soothing emotions" and "providing solutions" are different types of actions, while "suggesting that the customer restart the device" and "guiding the customer to check the network connection" are different in wording, but both belong to the "providing solutions" type.

[0044] When the action type suggested by the first auxiliary information is the same as that suggested by the second auxiliary information, it means that the immediate judgment of the rapid response layer 31 is consistent with the deliberated conclusion of the deep deduction layer 32. In this case, following the principle of simplicity and efficiency, the arbitration module 40 will select the first auxiliary information as the final auxiliary strategy. This can be understood as a kind of "confirmation" mode, where the result of the deep analysis verifies the correctness of the rapid response.

[0045] However, in more complex scenarios, the types of actions suggested by the two types of information often differ, resulting in a strategy conflict. For example, the rapid response layer 31 might suggest "immediate reassurance," while the deep deduction layer 32 might suggest "firmly rejecting but politely offering alternatives." This conflict is a key technical problem that this application needs to address. In this case, the arbitration module 40 will initiate its arbitration mechanism.

[0046] During the arbitration process, the arbitration module 40 compares the user experience-related value corresponding to the second auxiliary information with a preset risk threshold. This risk threshold can be understood as the maximum acceptable short-term decline in customer experience that the system 100 is willing to bear in order to achieve long-term task goals, i.e., the "emotional safety bottom line." Through this comparison, the arbitration module 40 can determine whether the emotional risk brought about by adopting this seemingly radical second auxiliary information is within a controllable range.

[0047] Based on the comparison results above, the arbitration module 40 will determine the final auxiliary strategy. For example, if the user experience-related value of the second auxiliary information is higher than or equal to the risk threshold, it indicates that the risk is controllable, and the second auxiliary information may be selected; conversely, if the user experience-related value of the second auxiliary information is lower than the risk threshold, it indicates that the risk is too high, and the more conservative first auxiliary information may be selected as a second choice.

[0048] Finally, the final support strategy determined by the arbitration module 40 is sent to the presentation module 50 for display. The presentation module 50 presents the final support strategy to the agents in a clear and intuitive way (such as text prompts, operation buttons, etc.). Agents can refer to this strategy to respond to customers. The entire process forms a closed loop from data monitoring, hierarchical prediction, conflict arbitration to final presentation, aiming to provide agents with high-quality real-time support.

[0049] like Figure 4 As shown, in a typical interaction, this method involves clear signaling communication between modules. The conversation between the customer and the agent is continuously analyzed by the monitoring module. When the metrics calculated by the monitoring module (MDDE in the figure) exceed the threshold, an activation signal is sent to the deep inference layer. The rapid response layer and the deep inference layer submit strategies to the arbitration module (CASE module in the figure). After completing the arbitration, the CASE module presents the final strategy to the agent, who then communicates with the customer based on the strategy.

[0050] The aforementioned solution achieves on-demand allocation of computing resources by setting up a dynamically triggered deep inference layer, initiating complex analysis only at critical moments. Simultaneously, by establishing an evaluation system that incorporates both task and experience value, the system can generate more comprehensive and optimized strategy recommendations. Furthermore, by introducing a conflict arbitration mechanism based on risk assessment, contradictions between different decision-making modules are effectively resolved, ensuring the rationality and security of the final output strategy, thereby enhancing the overall performance and practical value of the intelligent agent assistance system.

[0051] Furthermore, in a preferred embodiment, the dialogue state measurement metric is specifically defined. Specifically, the dialogue state measurement metric is Multidimensional Dialogue Deadlock Entropy (MDDE). Unlike traditional triggering methods that rely on only a single feature (such as silence duration or negative words), MDDE integrates features from multiple dimensions for comprehensive judgment, thereby enabling earlier and more accurate identification of the risk of a dialogue deadlock.

[0052] The calculation of this multidimensional dialogue deadlock entropy incorporates at least two of the following features: intent information entropy calculated based on the dialogue intent distribution, the proportion of abnormal silence duration on the client side, and the variance of the client's speech rate variation. For example, its calculation formula can be expressed as:

[0053]

[0054] in, , , These are the weighting coefficients for each feature, which can be adjusted according to the business scenario; This represents the entropy of intent information; this value increases when a customer repeatedly jumps between several intents or hesitates. This indicates the percentage of abnormal silence time on the customer side in the total call duration. This value increases when the customer remains silent for an extended period due to dissatisfaction or confusion. This represents the variance of the rate of change in the customer's speaking speed. This value increases when the customer's emotional state causes their speaking speed to fluctuate.

[0055] The technical advantage of using multidimensional dialogue deadlock entropy as a trigger indicator lies in its quantification and multi-dimensional deconstruction of the vague concept of "deadlock." Compared to triggers that rely on isolated events, MDDE can capture complex dialogue dilemmas formed by the combined effects of multiple factors such as semantic ambiguity, acoustic anomalies, and behavioral hesitation. For example, when a customer hasn't uttered any negative words, but their intent becomes confused and their speech rate fluctuates more, MDDE can sensitively detect this potential deadlock trend, thereby activating the deep inference layer in advance. This buys the system more valuable intervention time, improving the first-time problem resolution rate and customer satisfaction.

[0056] In a preferred embodiment, the method for generating evaluation results of task-related value and user experience-related value is specifically described. For example... Figure 3 As shown, the evaluation result was generated by the Bi-Objective Dialogue Value Network (DTD-VN). This network adopts a "shared-separated" structure, which can efficiently decode the value judgments of two different objectives from a unified input.

[0057] Specifically, the bi-objective dialogue value network takes a session state vector 321 containing multimodal information of the current dialogue as input. This session state vector 321 can be concatenated with word embeddings of the text, acoustic features such as customer tone and speech rate, and user profile information. The input vector first passes through at least one shared coding layer 322, which is responsible for extracting general feature representations useful for both objectives. Subsequently, these general feature representations are fed into two independent output heads: a task completion value head 323 and a future sentiment value head 324.

[0058] The task completion value head 323 is specifically used to predict the long-term reward associated with task completion. Its Q_task output 325 is the task-related value, specifically the long-term reward Q-value (Q_task) in this embodiment. The future emotion value head 324 is specifically used to predict the impact of the action on the customer's future emotions. Its E_Value output 326 is the user experience-related value, specifically the expected future emotion value E-Value in this embodiment. By performing forward propagation on multiple candidate actions, this bi-objective dialogue value network can output a pair of values ​​containing {Q_task, E-Value} for each action.

[0059] Based on this design, this application achieves decoupled evaluation of the two core objectives: "task" and "emotion." Unlike traditional reinforcement learning methods that weight and fuse multiple objectives into a single reward value, this scheme preserves the independent value information of each objective through separate output heads. This allows the subsequent arbitration module to perform more refined risk assessment based on the original, undamaged value data. For example, the arbitration module can clearly know that a strategy has a high Q_task but a low E-value, thus making a more prudent decision. This decision-making capability is key to solving the problem of poor communication effectiveness caused by the singularity of objectives in existing technologies.

[0060] Furthermore, the specific logic of the arbitration module 40 in handling strategy conflicts is further refined. When the type of agent action suggested by the first auxiliary information is different from the type of agent action suggested by the second auxiliary information, the arbitration module 40 compares the user experience-related value (i.e., E-Value) corresponding to the second auxiliary information with a preset risk threshold (i.e., emotional safety baseline threshold) to determine the final auxiliary strategy.

[0061] The logic for comparison and selection specifically includes: if the user experience-related value corresponding to the second auxiliary information is lower than the risk threshold, then the first auxiliary information is selected as the final auxiliary strategy; otherwise, when the user experience-related value is higher than or equal to the risk threshold, the second auxiliary information is selected as the final auxiliary strategy. Here, "lower than" means that implementing the second auxiliary information may cause the customer's emotional state to deteriorate beyond an acceptable range, constituting a high-emotional-risk event.

[0062] The technical advantage of this arbitration logic lies in introducing a clear risk control mechanism for System 100. In pursuing long-term task value (high Q_task), System 100 might recommend counterintuitive strategies that could cause short-term customer dissatisfaction (e.g., refusing even when strongly requested by the customer). While such strategies may be optimal in the long run, they carry high short-term risk. This arbitration logic, by setting an emotional safety threshold, ensures that System 100 will only adopt such high-return, high-risk depth strategies when emotional risk is controllable (E-Value not lower than the risk threshold). This effectively prevents System 100 from recommending "cold" or "rigid" strategies simply to complete tasks, maintaining customer experience while pursuing efficiency, thus achieving robustness and security in decision-making.

[0063] In another preferred embodiment, the application further includes a strategy explanation generation step linked to the arbitration result. Specifically, if the second auxiliary information is selected as the final auxiliary strategy after arbitration, the system 100 will automatically generate a strategy explanation text to explain the reasons for the strategy selection based on its corresponding task-related value (Q_task) and user experience-related value (E-Value), and present the strategy explanation text together with the final auxiliary strategy to the agent.

[0064] For example, when System 100 recommends a strategy of "rejecting unreasonable customer requests and providing voucher compensation," its corresponding Q_task might be +0.7 (representing that in the long run, it can maintain company policy and has a high customer retention rate), and its E-Value might be -0.2 (representing that customer sentiment may slightly decline in the short term). Based on these two values, System 100 can generate explanatory text such as "This solution may cause temporary dissatisfaction, but it can maintain company policy, and the long-term customer retention rate is estimated to be higher."

[0065] Based on this, this application addresses the "black box" problem prevalent in AI-assisted systems, enhancing trust in human-machine collaboration. When an agent faces an AI suggestion that contradicts their intuition, the strategy explanation text clearly explains "why" the system makes that suggestion—the long-term value trade-offs behind the decision. This allows agents to move beyond passively executing commands and instead understand the AI's "thought process," making them more willing and confident to adopt these high-quality suggestions. This establishment of trust transforms the human-machine relationship from a simple "command-execution" model to a proactive "cooperation-collaboration" model, which is crucial for improving the adoption rate and ultimate effectiveness of AI-assisted systems in practical work.

[0066] The following detailed description of a specific embodiment integrating the above-mentioned technical solutions will illustrate this point. This embodiment will demonstrate how the system and method provided in this application work together in a typical complex customer interaction scenario to achieve intelligent decision-making that balances task requirements and emotional considerations.

[0067] Imagine a call center scenario where a customer calls, demanding a refund for an item that has exceeded the 30-day no-reason return period, and is quite agitated. (Refer to...) Figure 1 , Figure 2 and Figure 4 The intelligent agent assistance system 100 of this application will operate according to the following steps.

[0068] First, after the conversation begins, monitoring module 20 acquires real-time voice and text data between the customer and the agent. The agent attempts to explain the return policy to the customer, but the customer repeatedly insists, "I don't care, you must give me a refund." Monitoring module 20 begins calculating the Multidimensional Dialogue Impasse Entropy (MDDE). At this point, the customer's intent is very stubborn, remaining fixed on "requesting a refund," leading to an intent information entropy... The variance is relatively low (e.g., 0.1); however, the customer is emotionally agitated, speaks quickly and with significant fluctuations, resulting in a higher variance in speech rate. The percentage of abnormal silence time is relatively high (e.g., 0.6); at the same time, customers interrupt and pause briefly in anger while the agent is explaining, which increases the percentage of abnormal silence time. It is at a moderate level (e.g., 0.3).

[0069] Monitoring module 20 according to the formula Perform the calculation, assuming the weights are... ,but The MDDE value exceeded the system's preset deadlock entropy threshold. Therefore, the monitoring module 20 sends an activation signal to the deep inference layer 32 of the prediction module 30.

[0070] In the prediction module 30, the rapid response layer 31 and the deep inference layer 32 work in parallel. Based on the detected customer agitation and the keyword "return," the rapid response layer 31 quickly generates the first auxiliary information (strategy A), which reads: "Suggested response: Please calm the customer down first. You could say, 'Sir, I understand how you feel. Please don't worry.'" This is a conventional, low-risk response strategy.

[0071] Meanwhile, the activated deep inference layer 32 inputs the current conversation state vector 321 (containing dialogue history, customer emotional characteristics, and user profile information indicating the customer is a high-value customer) into the bi-objective dialogue value network (DTD-VN). The bi-objective dialogue value network evaluates a series of candidate actions. For example, for "Action 1: Agree to return the goods," the network predicts that it will immediately satisfy the customer, but violates company policy and has very low long-term value; therefore, it outputs a bi-value pair as { Regarding "Action 2: Refuse return and proactively offer an 80% discount voucher as compensation," the bi-objective dialogue value network predicts that it will cause customer dissatisfaction in the short term, but it maintains the policy, and the compensation measure may retain customers in the long term. Therefore, the output bi-value pair is { }

[0072] The deep inference layer 32 selects logic based on a preset strategy (e.g., maximizing). It was found that action 2 had a higher overall score, so action 2 was selected as the second auxiliary information (strategy B).

[0073] Subsequently, Strategy A (appeasement) and Strategy B (rejection + compensation) were sent together to Arbitration Module 40. Arbitration Module 40 determined that the underlying intentions of the two strategies (one was purely emotional reassurance, and the other was a concrete action to solve the problem in opposite directions) were conflicting, and therefore initiated the arbitration process.

[0074] During the arbitration process, the arbitration module 40 obtained the emotional risk of strategy B, i.e., its user experience-related value. The system has a preset emotional safety threshold. The value is -0.4, representing the maximum decline in customer sentiment the company is willing to accept in order to maintain its policies. Arbitration module 40 is used for comparison: According to the arbitration logic, since the emotional risk of strategy B is within an acceptable range (i.e., the E-Value is not below the risk threshold), the arbitration decides to adopt this second auxiliary information (strategy B), which has a better long-term value, as the final auxiliary strategy.

[0075] Since the second auxiliary information was ultimately adopted, system 100 then triggered the policy interpretation generation function. Based on the value pair of policy B, { The system automatically generated a line of policy explanation text: "Although this solution may cause temporary dissatisfaction, it can maintain company policy and the long-term customer retention rate is estimated to be higher."

[0076] Finally, in presentation module 50, the system presents the final decision to the agent. The interface may display the following: - Final Strategy: "Recommendation: Reject the return request and proactively inform the customer that they can apply for an 80% discount voucher as compensation." - Strategy Explanation: "Reason: Although this solution carries short-term emotional risks, it is expected to maintain company policy and achieve higher long-term customer retention value."

[0077] Upon seeing the suggestion, and especially after reading the explanation behind it, the agent understood the underlying logic of AI decision-making. He adopted the solution and communicated it to the customer. Ultimately, although the customer was unable to return the product, they received compensation, their emotions were somewhat calmed, and a potentially serious customer complaint was averted, achieving a balance between task objectives and long-term customer relationships.

[0078] This embodiment fully demonstrates the synergistic effect of various technical features: Multidimensional Dialogue Impasse Entropy (MDDE) ensures that the system only utilizes deep analysis capabilities at critical moments (rather than all moments); Dual-Objective Dialogue Value Network (DTD-VN) provides "high emotional intelligence" strategy options that balance task and emotion; the conflict arbitration mechanism based on future emotional risk ensures that System 100 does not act recklessly while pursuing high returns, thus maintaining a safety baseline; and the strategy interpretation generation function linked to the arbitration result helps build agent trust in System 100. The organic combination of these features enables this application to provide an efficient, accurate, secure, and reliable intelligent agent assistance solution.

[0079] The application scenarios of this application are not limited to the after-sales service mentioned above. For example, in a support scenario for a complex technical problem, when the communication between the customer and the agent gets stuck in an "ineffective loop" (the customer repeatedly says "I've tried it, but it doesn't work"), the intent information entropy... This will increase and can also trigger MDDE. The Deep Analytical Model (DTD-VN) layer may identify this as a rare system defect based on massive amounts of historical data and recommend an unconventional solution (such as modifying configuration files). Even if this solution conflicts with the rapid response layer's suggestion to "transfer to a second-line expert," as long as the emotional risk is controllable, the arbitration module will adopt it and provide an explanation (e.g., "This is a special solution for this rare fault, with an estimated problem-solving rate of 95%)," thereby helping agents quickly resolve the problem and avoid unnecessary transfers.

[0080] Furthermore, the framework presented in this application exhibits good scalability. In some business scenarios that are particularly sensitive to emotional fluctuations (such as financial debt collection), the calculation formula for Multidimensional Dialogue Impasse Entropy (MDDE) can be further extended. For example, the variance of real-time emotion polarity scores can be incorporated. and the customer's historical complaint frequency As a new feature dimension. Adjusted formula. It can more accurately capture risk signals in specific scenarios and achieve flexible adaptation to different business needs.

[0081] This application also covers a collaborative enhancement working mode when there is no strategic conflict. When customers are simply making simple information inquiries, the MDDE value is typically low, and the deep deduction layer is not activated. The quick response layer generates standard script suggestions. In some cases, even if a brief fluctuation in MDDE triggers deep deduction, the generated strategy may still conflict with the suggestions from the quick response layer. Figure 1 At this point, the arbitration module 40 determines that there is no conflict in the strategy and does not proceed with arbitration, but instead performs collaborative enhancement. For example, it can combine standard rhetoric with the high-value assessment given by the deep network (such as { This information, combined with other factors, generates an enhanced explanation: "The current strategy is the optimal path, and the task is expected to be completed smoothly with the customer in a good mood." Providing this "certainty" information enhances the agent's confidence, enabling them to execute standard procedures more decisively and smoothly, thereby improving overall service efficiency.

[0082] In summary, the intelligent agent assistance method and system provided in this application, by constructing a complete technical closed loop integrating dynamic triggering, dual-objective value assessment, risk arbitration, and interpretable presentation, helps to solve some of the problems existing in the prior art. It not only enables smarter and more balanced decision-making but also strives to build a more transparent and trustworthy human-machine collaboration relationship, providing technical support for improving the operational efficiency and service quality of modern customer service centers.

[0083] Those skilled in the art will readily recognize that various modifications, combinations, and variations can be made to the above embodiments of this application without departing from the spirit and scope of this application. Therefore, although this specification describes this application with reference to specific preferred embodiments, these descriptions are not intended to limit the scope of this application. The scope of this application is defined only by the appended claims. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the claims of this application should be included within the protection scope of the claims of this application.

Claims

1. A method for assisting intelligent agents, characterized in that, include: Real-time monitoring of the conversation between agents and customers, and calculation of conversation state metrics that characterize the conversation state. Generate rule-based first auxiliary information; When the dialogue state metric exceeds a preset trigger threshold, an evaluation result is generated for at least one candidate action. The evaluation result includes task-related value and user experience-related value. Based on the evaluation result, one of the candidate actions is selected as second auxiliary information. When the type of agent action suggested by the first assistance information is different from the type of agent action suggested by the second assistance information, the user experience-related value corresponding to the second assistance information is compared with a preset risk threshold to determine the final assistance strategy. When the type of agent action suggested by the first auxiliary information is the same as the type of agent action suggested by the second auxiliary information, the first auxiliary information is selected as the final auxiliary strategy. The final assistance strategy is presented to the agent.

2. The method according to claim 1, characterized in that, The dialogue state metric is a multidimensional dialogue deadlock entropy, the calculation of which integrates at least two of the following features: intent information entropy calculated based on the dialogue intent distribution, the proportion of abnormal silence duration on the client side, and the variance of the client's speech rate variation.

3. The method according to claim 1, characterized in that, The task-related value and user experience-related value are defined by the dual-objective dialogue value network; the task-related value is the long-term reward Q value for task completion, and the user experience-related value is the expected future emotional value E-Value.

4. The method according to claim 1, characterized in that, The step of comparing the user experience-related value corresponding to the second assistance information with a preset risk threshold to determine the final assistance strategy further includes: If the user experience-related value corresponding to the second auxiliary information is lower than the risk threshold, then the first auxiliary information is selected as the final auxiliary strategy. If the user experience-related value corresponding to the second auxiliary information is equal to or higher than the risk threshold, the second auxiliary information is selected as the final auxiliary strategy.

5. The method according to any one of claims 1 to 4, characterized in that, Also includes: If the second auxiliary information is selected as the final auxiliary strategy, a strategy explanation text is generated based on its corresponding task-related value and user experience-related value to explain the reason for the strategy selection, and the strategy explanation text is presented together with the final auxiliary strategy.

6. An intelligent agent assistance system, characterized in that, For performing the method according to any one of claims 1-5, comprising: The monitoring module is used to monitor the dialogue process between agents and customers in real time and calculate dialogue status metrics that characterize the dialogue status. The prediction module is electrically connected to the monitoring module and is used to receive the dialogue state measurement index and generate rule-based first auxiliary information through the auxiliary model. When the dialogue state measurement index exceeds a preset trigger threshold, it generates an evaluation result for at least one candidate action based on the dialogue strategy evaluation model. The evaluation result includes task-related value and user experience-related value. Based on the evaluation result, a candidate action is selected as the second auxiliary information. An arbitration module, electrically connected to the prediction module, is used to receive the first auxiliary information and the second auxiliary information, and is used to: determine the final auxiliary strategy by comparing the user experience-related value corresponding to the second auxiliary information with a preset risk threshold when the type of agent action suggested by the first auxiliary information is different from the type of agent action suggested by the second auxiliary information; select the first auxiliary information as the final auxiliary strategy when the type of agent action suggested by the first auxiliary information is the same as the type of agent action suggested by the second auxiliary information; and generate a source identifier indicating whether the final auxiliary strategy originates from the first auxiliary information or the second auxiliary information. The presentation module, which is telecommunication connected to the arbitration module, is used to receive and present the final assistance strategy to the agent.

7. The system according to claim 6, characterized in that, The monitoring module is used to calculate the multidimensional dialogue deadlock entropy as a metric for dialogue state. The calculation of the multidimensional dialogue deadlock entropy integrates at least two of the following features: intent information entropy calculated based on the dialogue intent distribution, the proportion of abnormal silence duration on the client side, and the variance of the client's speech rate change.

8. The system according to claim 6, characterized in that, The prediction module includes a bi-objective dialogue value network, which is used to generate the task-related value and the user experience-related value.

9. The system according to claim 6, characterized in that, The arbitration module is configured as follows: When the user experience-related value corresponding to the second auxiliary information is lower than the risk threshold, the first auxiliary information is selected as the final auxiliary strategy. When the user experience-related value corresponding to the second auxiliary information is equal to or higher than the risk threshold, the second auxiliary information is selected as the final auxiliary strategy.

10. The system according to any one of claims 6 to 9, characterized in that, The presentation module is also used for: Upon receiving the source identifier, and when the source identifier indicates that the final assistance strategy originates from the second assistance information, a strategy explanation text is generated based on its corresponding task-related value and user experience-related value, and the strategy explanation text is presented together with the final assistance strategy.