Multi-agent dialogue training method and system based on semantic stage recognition
Patent Information
- Application Number
- CN202610988559.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]针对现有技术的不足,本发明提供了一种基于语义阶段识别的多智能体对话训练方法和系统,解决了对话训练中无法根据客户特征生成针对性建议,质检和推荐功能无法根据对话阶段动态切换的问题
[0016]本发明提供了一种基于语义阶段识别的多智能体对话训练方法和系统,通过引入语义阶段识别与规则匹配的双层级联架构,提高了对话训练中阶段识别的准确性和鲁棒性。根据语义阶段变化触发事件信号,实现了多智能体之间的解耦协作与实时调度,使得模拟用户智能体、质检智能体和推荐智能体能够协同执行对话训练内容。能够为对话训练提供更具针对性的实时反馈和辅助,提升了销售培训的效率和效果。
Smart Images

Figure CN122819243A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large language model technology, specifically to a multi-agent dialogue training method and system based on semantic stage recognition. Background Technology
[0002] Existing dialogue training systems suffer from significant shortcomings in real-time performance, architectural design, and functional coverage. Traditional customer service quality inspection systems focus solely on post-dialogue quality assessment, employing machine learning models to analyze complete dialogue records, but fail to provide immediate feedback during the dialogue. These systems feature tightly coupled architectures with interdependent functional modules, making it difficult to expand with new features and unsuitable for sales training scenarios, lacking crucial functions such as customer simulation, script recommendation, and training review. Furthermore, dialogue flow analysis techniques based on large language models are primarily used for offline historical data analysis, discovering general flow patterns from numerous dialogue samples through clustering and classification algorithms, but they cannot detect changes in dialogue stages in real-time or trigger corresponding actions during the dialogue.
[0003] AI sales training products on the market generally suffer from architectural flaws, lacking multi-agent event-driven collaboration mechanisms. This results in tightly coupled functional modules that cannot be independently expanded. Furthermore, the system lacks the ability to automatically identify semantic stages based on a large language model, making it difficult to accurately understand the evolution of dialogue context and leading to low stage recognition accuracy.
[0004] Furthermore, existing technologies lack a user profile-driven personalized script recommendation mechanism, making it impossible to generate targeted suggestions based on customer characteristics. The quality inspection and recommendation functions cannot be dynamically switched according to the dialogue stage, resulting in the inability to trigger script recommendations at the beginning of the training stage and trigger quality inspection evaluation at the end of the stage, thus affecting the training effect. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a multi-agent dialogue training method and system based on semantic stage recognition, which solves the problems in dialogue training where targeted suggestions cannot be generated based on customer characteristics, and the quality inspection and recommendation functions cannot be dynamically switched according to the dialogue stage.
[0006] To achieve the above objectives, this invention provides a multi-agent dialogue training method based on semantic phase recognition, comprising: Based on a two-layer cascaded architecture of semantic recognition and rule matching, the semantic stage of the current dialogue training is determined. Based on changes in the semantic stage, corresponding event signals are triggered; Based on the receipt of subscribed event signals by multiple intelligent agents, corresponding dialogue training content is executed. The intelligent agents include a simulated user intelligent agent for responding to the dialogue, a quality control intelligent agent for evaluating the dialogue, a recommendation intelligent agent for engaging in dialogue with the simulated user intelligent agent, and a debriefing intelligent agent for analyzing the dialogue. Based on the process and result data of the intelligent agent analyzing the dialogue content after the dialogue training is completed, the user profile information associated with the simulated user intelligent agent is updated. Based on the updated user profile information, the recommendation dialogue strategy of the recommending agent in subsequent dialogue training is adjusted to form a cyclical iteration of dialogue training.
[0007] In one embodiment of the present invention, determining the semantic stage of the current dialogue training based on a two-layer cascaded architecture of semantic recognition and rule matching includes: Semantic analysis is performed on prompt words containing the definition of the dialogue target stage and the dialogue context using a large language model, outputting the name of the semantic stage, and the validity is judged by a whitelist check. If the whitelist verification fails, the keyword mapping rule matching will be used to output the semantic stage name of the corresponding whitelist.
[0008] In one embodiment of the present invention, triggering a corresponding event signal based on a determined semantic stage change includes: Based on the phase start event at the beginning of each new phase of dialogue training, the recommending agent is triggered to generate dialogue training content; Based on the phase end event at the end of each dialogue training phase, the quality inspection agent is triggered to evaluate the dialogue training content.
[0009] In one embodiment of the present invention, based on multiple agents receiving subscribed event signals, corresponding dialogue training content is executed, including: Based on the user profile configured with the simulated user agent, which includes the dialogue response style, the user profile is injected into the prompt words of the simulated user agent to generate dialogue response content that matches the user profile.
[0010] In one embodiment of the present invention, based on multiple agents receiving subscribed event signals, corresponding dialogue training content is executed, including: Based on the multi-dimensional driving dialogue configured by the recommending agent, dialogue recommendation content is generated that is adapted to the current semantic stage and user profile.
[0011] In one embodiment of the present invention, generating dialogue recommendation content adapted to the current semantic stage and user profile includes: Based on constraints, the system identifies rejection intentions in dialogue responses and filters or adjusts the generated dialogue recommendations accordingly.
[0012] In one embodiment of the present invention, based on the process and result data of the review agent analyzing the dialogue content after the dialogue training is completed, the user profile information associated with the simulated user agent is updated, including: Based on the review of the agent's response to event signals at the end of the dialogue training phase, the execution status of each stage of the dialogue training recorded in the event history list is traced back to analyze the stage completion rate. Based on the dialogue training content, strengths, areas for improvement, and weaknesses were stratified and marked, and a scoring range analysis was conducted. Based on the analysis results, update the contact history, interest tags, and secondary marketing opportunities in the user profile information, and persist the updated user profile information. The analysis results include stage completion analysis and rating interval analysis.
[0013] This application also proposes a multi-agent dialogue training system based on semantic stage recognition, including: The central control brain module is configured to identify the semantic stages of dialogue training through a two-layer cascade method and issue event signals based on the identification results; The event bus module is configured to receive and distribute event signals; The intelligent agent module includes a simulated user intelligent agent, a quality inspection intelligent agent, a recommendation intelligent agent, and a review intelligent agent; each intelligent agent module pre-registers the event signals it responds to in the event bus module and is configured to asynchronously execute its function when the corresponding event signal is received.
[0014] This application also proposes a storage medium storing a computer program, wherein the computer program enables a computer to execute the above-described multi-agent dialogue training method based on semantic phase recognition.
[0015] This application also proposes an electronic device, including: a processor and a memory; The memory stores program instructions; The processor is used to run program instructions to execute the multi-agent dialogue training method based on semantic phase recognition described above.
[0016] This invention provides a multi-agent dialogue training method and system based on semantic stage recognition. By introducing a two-layer cascaded architecture of semantic stage recognition and rule matching, the accuracy and robustness of stage recognition in dialogue training are improved. Based on event signals triggered by semantic stage changes, decoupled collaboration and real-time scheduling among multiple agents are achieved, enabling the simulated user agent, quality inspection agent, and recommendation agent to collaboratively execute dialogue training content. This provides more targeted real-time feedback and assistance for dialogue training, improving the efficiency and effectiveness of sales training. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a multi-agent dialogue training method based on semantic stage recognition provided in the embodiments of this application.
[0019] Figure 2 This is a flowchart of the AI sales coaching system provided in the embodiments of this application.
[0020] Figure 3 This is a block diagram of the device structure of the AI sales coach system provided in the embodiments of this application.
[0021] Figure 4 This is a system architecture diagram of the AI sales coach system provided in the embodiments of this application.
[0022] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Existing customer service quality inspection systems or dialogue flow training methods have many limitations in sales training scenarios. These systems are functionally limited, lacking core functions required for sales training such as customer simulation, script recommendation, and training loops. Their evaluation models are outdated, unable to provide real-time feedback, and have limited semantic understanding capabilities, resulting in low accuracy in stage identification. Furthermore, the existing architecture is tightly coupled, has poor scalability, lacks multi-agent collaboration mechanisms and event-driven scheduling capabilities, and also lacks degradation strategies and training data accumulation mechanisms, leading to insufficient system availability, personalization, and continuous optimization capabilities.
[0025] In this regard, such as Figure 1 As shown, this application proposes a multi-agent dialogue training method based on semantic stage recognition, including: S10. A two-layer cascaded architecture based on semantic recognition and rule matching is adopted to determine the semantic stage of the current dialogue training.
[0026] Specifically, a two-tier cascaded architecture refers to a hierarchical processing mechanism containing at least two processing levels. When one level cannot effectively complete the task, it can automatically switch to or degrade to another level for processing, ensuring the robustness and availability of the task. Semantic stage recognition refers to identifying and determining the specific stage of a dialogue. This is achieved by manually defining a series of keywords or phrases and associating them with specific semantic stages. When these keywords or phrases appear in the dialogue content, they are identified as the corresponding semantic stage. For example, in a sales dialogue scenario, stages include the opening, value demonstration, needs assessment, action facilitation, or conclusion. By identifying these stages, subsequent processing can be targeted. As one implementation method, a classification model can be trained by manually annotating a large amount of historical dialogue data. This model can output the semantic stage to which the input dialogue text belongs.
[0027] S20. Trigger the corresponding event signal based on the determined semantic stage change.
[0028] Specifically, this can be achieved by periodically checking whether the current semantic stage has changed from the previous semantic stage. Once a change is detected, a general event signal is issued. The event signal can be a notification mechanism issued by one module when a specific condition is met or a state change occurs, and can be subscribed to and received by other modules. This signal can be used to drive asynchronous collaboration and function execution between different intelligent agent modules.
[0029] S30. Based on the received subscribed event signals by multiple agents, execute the corresponding dialogue training content. These agents include a simulated user agent for responding to the dialogue, a quality control agent for evaluating the dialogue, a recommendation agent for engaging in dialogue with the simulated user agent, and a debriefing agent for analyzing the dialogue.
[0030] Specifically, dialogue training content refers to the specific tasks performed or information generated by each agent according to its functional responsibilities during the dialogue training process, such as simulated user responses, evaluation results of quality control agents, or script recommendations from recommendation agents.
[0031] The simulated user agent can generate dialogue responses based on a preset fixed response template or a random generator. When it receives an event signal, the simulated user agent can simulate the dialogue behavior and response style of a real user to provide a realistic dialogue interaction environment for the training subjects, and randomly select a response from a predefined response list.
[0032] The function of the quality control agent is to evaluate and analyze the dialogue training process or results according to preset standards and rules, in order to provide quality feedback. It can initiate a preset evaluation process based on received event signals, performing simple grammar checks or length assessments on the current dialogue turn.
[0033] The function of a recommendation agent is to generate or recommend relevant dialogue content or strategies based on the context and goal of the current conversation, guiding the dialogue towards a preset objective. In a sales training scenario, this agent can be understood as a marketing assistant agent. It can select a general script from a pre-set script library based on received event signals, and recommend a general opening line related to the start of a new phase.
[0034] The debriefing agent can collect, process, and understand all relevant data generated during dialogue training, including dialogue text, agent behavior, user feedback, stage transition records, and the final training results. Its role is to provide a closed-loop feedback mechanism, transforming lessons learned during training into actionable knowledge to guide future training optimization.
[0035] S40. Based on the process and result data of the dialogue content analysis by the review agent after the dialogue training is completed, update the user profile information associated with the simulated user agent.
[0036] Specifically, the debriefing agent can be built based on machine learning models (such as natural language processing models and behavior analysis models). Through preset evaluation indicators and rules, it can conduct in-depth mining of dialogue data, such as identifying key turning points in the dialogue, changes in user emotions, the effectiveness of the agent's responses, and the achievement of goals.
[0037] After dialogue training, the review agent analyzes the process and outcome data of the dialogue content. The "process data" can include each round of interaction, the agent's responses at different semantic stages, the triggering of event signals, changes in the user's emotions or intentions during the dialogue, dialogue duration, and the frequency of key phrase usage. The "outcome data" may include whether the dialogue goal was achieved, the quality control agent's evaluation score, user satisfaction with the training, and the completion status of specific tasks. The review agent can use Natural Language Understanding (NLU) technology to parse the dialogue text and extract key information; use sentiment analysis models to identify user emotions; use behavioral sequence analysis models to evaluate the agent's strategy execution effectiveness; and use pre-set business rules or expert knowledge bases to quantitatively evaluate the results.
[0038] Furthermore, by analyzing the dialogue process and outcome data, user profiles can be dynamically revised and enriched to more accurately reflect the user's true situation. User profile information can include the user's interests, communication style, sensitivity to specific topics, historical rejected or accepted recommendations, learning progress, knowledge mastery, and emotional tendencies, allowing the simulated user agent to correlate their dialogue responses. The review agent identifies dimensions in the user profile that need adjustment based on the analysis results. For example, if a user consistently shows a rejection intention towards a certain type of recommended content during multiple training sessions, this preference can be marked in the user profile.
[0039] S50. Based on the updated user profile information, adjust the recommendation dialogue strategy of the recommending agent in subsequent dialogue training to form a cyclical iteration of dialogue training. Specifically, the recommendation dialogue strategy of the recommendation agent aims to provide customized guidance for subsequent dialogue training, thereby improving training efficiency and effectiveness. For example, it can adjust the response style of the simulated user agent based on user preferences, or adjust the content and order of the recommended dialogue; guide the quality control agent to focus more on the weaknesses marked in the user profile, or guide the recommendation agent to avoid repeating recommendation strategies that the user has already rejected. These decisions can be made through rule engines, recommendation systems, or reinforcement learning.
[0040] Furthermore, by reviewing the entire process and final results of each dialogue training session, the system conducts in-depth and comprehensive analysis. It can learn from historical data and dynamically update and improve user profile information. Based on more accurate user profiles, highly personalized decisions are made for subsequent dialogue training, such as adjusting the response style of the simulated user agent, optimizing the recommendation strategy of the recommendation agent, or adjusting the training difficulty according to the user's learning progress. This achieves a closed-loop iteration of "training → evaluation → profile update → strategy adjustment → retraining," continuously improving the personalization and targeting of dialogue training. This enhances the adaptability and intelligence of the multi-agent dialogue training system.
[0041] In this embodiment, by introducing a two-layer cascaded architecture of semantic stage recognition and rule matching, the accuracy and robustness of stage recognition in dialogue training are improved. By triggering event signals based on changes in semantic stages, decoupled collaboration and real-time scheduling among multiple agents are achieved, enabling the simulated user agent, quality inspection agent, and recommendation agent to collaboratively execute dialogue training content. Therefore, this embodiment can provide more targeted real-time feedback and assistance for sales training, effectively overcoming the limitations of traditional sales training systems such as single functionality, delayed feedback, limited semantic understanding capabilities, and tightly coupled architecture, thus improving the efficiency and effectiveness of sales training.
[0042] To better understand the above technical solution, the following will refer to the appendix to the instruction manual. Figure 1 and Figure 2 The specific implementation methods are described in detail below for the above technical solutions.
[0043] In step S10, a two-layer cascaded architecture based on semantic recognition and rule matching is used to determine the semantic stage of the current dialogue training.
[0044] Step S10 includes the following specific details.
[0045] S11. Perform semantic analysis on prompt words containing the definition of the dialogue target stage and the dialogue context using a large language model, output the name of the semantic stage, and determine its validity through whitelist verification.
[0046] Specifically, the large language model can employ pre-trained language models such as the GPT series. The construction of prompt words can include pre-defined dialogue target stages, such as "the opening stage aims to establish initial contact" or "the needs exploration stage aims to understand the user's specific needs," providing the large language model with a basis for judgment. It can also include the current dialogue context, i.e., historical dialogue records and the latest user input, ensuring the model can analyze based on the complete dialogue context. After receiving these prompt words, the large language model determines which predefined semantic stage the current dialogue conforms to and outputs the corresponding semantic stage name. To ensure that the output semantic stage name is recognizable and processable by the system, a whitelist verification mechanism is introduced. The whitelist is a pre-defined list of valid semantic stage names, such as "opening stage," "needs exploration stage," "solution introduction stage," "objection handling stage," and "closing stage." The semantic stage name output by the large language model is compared with this whitelist; if the output semantic stage name is not in the whitelist, the semantic recognition is considered invalid.
[0047] S12. If the whitelist verification is invalid, then use the keyword mapping rule matching to output the semantic stage name of the corresponding whitelist.
[0048] Specifically, when the large language model fails to provide a valid or accurate semantic stage, a rule matching module is activated as a supplement and fallback solution to the large language model's semantic analysis. This module is pre-configured with mapping relationships between keywords and semantic stages. For example, keywords such as "hello" and "consultation" are mapped to the "opening stage," and keywords such as "need" and "question" are mapped to the "demand mining stage." Keywords are extracted from the current dialogue context and compared with the pre-defined mapping table. Once a match is successful, the corresponding semantic stage name that is in the whitelist is output.
[0049] Thus, this application's embodiments construct a two-layer cascaded semantic stage recognition architecture that combines deep semantic understanding with high robustness. The large language model can handle complex and varied dialogue contexts, perform semantic analysis, and output more accurate semantic stage names. Simultaneously, a whitelist verification mechanism ensures the standardization and validity of the recognition results. When the large language model's recognition results are unclear or invalid, the keyword mapping rule matching mechanism can serve as an effective supplement, providing reliable alternatives and avoiding system interruptions or incorrect judgments due to the failure of a single recognition method. By combining advanced semantic models with traditional rule matching strategies, the accuracy and stability of semantic stage recognition during dialogue training are improved. This allows the simulated user agent, quality inspection agent, and recommendation agent to receive more accurate event signals, thereby executing training content more consistent with the current dialogue state and improving the efficiency and effectiveness of multi-agent dialogue training.
[0050] In step S20, the corresponding event signal is triggered based on the determined semantic stage change.
[0051] Step S20 includes the following specific details.
[0052] S21. Based on the phase start event at the beginning of each new phase of dialogue training, trigger the recommending agent to generate dialogue training content. Specifically, a phase start event refers to a specific signal indicating that a new dialogue training phase is about to begin or has already begun. This event is typically triggered automatically by the system after confirming that the previous phase has successfully concluded and may have passed the evaluation of the quality control agent, or according to a pre-set training process, to instruct the relevant agents to prepare for the new phase. Upon receiving a phase start event, the recommendation agent immediately executes the task of generating dialogue training content. This may include generating highly adaptable and guiding dialogue recommendations based on the current semantic stage, user profile, and pre-set multi-dimensional driving dialogue, ensuring that the dialogue in the new phase progresses efficiently in the expected direction.
[0053] S22. Based on the phase end event at the end of each dialogue training phase, trigger the quality inspection agent to perform an evaluation of the dialogue training content.
[0054] Specifically, a phase end event refers to a specific signal indicating the completion of a dialogue training phase. This event can be automatically generated and published by the system when it detects that the preset goal of the current phase has been achieved, a specific dialogue round has ended, or other preset phase end conditions are met. When the quality control agent receives the phase end event, it will activate its evaluation function to comprehensively analyze the completed dialogue interactions between the simulated user agent and the recommendation agent. This evaluation can be based on a preset evaluation model, such as using semantic analysis, sentiment recognition, and intent matching technologies, to determine the fluency, logic, goal achievement, and the presence of non-compliant content in the dialogue, and generate a corresponding evaluation report or score.
[0055] Thus, by explicitly defining "phase start events" and "phase end events" as trigger points for agent behavior, the quality control agent can intervene promptly and accurately during dialogue phase transitions, evaluating completed dialogue content to identify and correct training issues in a timely manner, ensuring training quality. The recommendation agent can immediately generate and provide suitable dialogue recommendations at the start of a new phase, effectively guiding the dialogue direction and improving the efficiency and relevance of dialogue training. This enhances the automation level and overall effectiveness of multi-agent dialogue training, avoiding lag or confusion in agent behavior, making the entire training process smoother and more efficient.
[0056] In step S30, based on the received subscribed event signals by multiple agents, corresponding dialogue training content is executed. These agents include a simulated user agent for responding to dialogues, a quality control agent for evaluating dialogues, and a recommendation agent for recommending dialogues.
[0057] Step S30 includes the following specific details.
[0058] S31. Based on the simulated user agent, a user profile containing dialogue response style is configured. The user profile is injected into the prompt words of the simulated user agent to generate dialogue response content that matches the user profile.
[0059] Specifically, the simulated user agent is configured with user profiles that include dialogue response styles. A user profile is an abstract and generalized description of a target user. It constructs a representative virtual user model by collecting and analyzing user behavioral data, preferences, and characteristics. A user profile specifically refers to a set of features that include a user's response style, tone, frequently used vocabulary, emotional tendencies, interests, and background information in a dialogue. Its purpose is to ensure that the simulated user agent can exhibit language habits and behavioral patterns consistent with a specific user group or individual in dialogue, thereby improving the realism and effectiveness of the simulated dialogue. User profiles can be constructed in various ways. For example, machine learning analysis can be used based on historical dialogue data to extract user language patterns, emotional tendencies, and topic preferences; different types of user profiles can also be defined through preset templates, manual annotation, or expert experience, such as "impatient sub-users," "hesitant users," and "professional domain users."
[0060] Furthermore, injecting user profiles into prompts for simulated user agents means integrating key information from the user profile (such as response style, background, preferences, etc.) as part of the context or instructions into the prompts sent to the simulated user agent. This allows the large language model to fully consider and adopt the features defined by the user profile when generating responses, thereby generating more personalized and style-appropriate dialogue content. When constructing prompts, for example, the prompts could be designed as: "You are now a [user profile description, such as: price-sensitive customer who likes to haggle], please reply according to the following dialogue: [dialogue context]". Alternatively, user profile information can be passed to the large language model as independent system instructions or role settings, allowing it to maintain a specific persona throughout the dialogue.
[0061] Building upon this foundation, generating dialogue responses that align with the user profile refers to the text output by the simulated user agent based on prompts infused with the user profile. This reflects the response style, emotion, preferences, and other characteristics defined in the user profile. Its purpose is to ensure the simulated user agent's responses are highly realistic and personalized, allowing the other party training the agent (such as a sales agent) to face challenges and feedback closer to real-world scenarios, thereby improving training effectiveness. Upon receiving prompts containing user profile information, the large language model leverages its powerful language generation capabilities, combining user profile constraints with the dialogue context, to generate appropriate responses. For example, if the user profile is set to "polite tone, detailed questions," the model will tend to generate polite and specific responses.
[0062] In this way, by simulating user agents to respond to dialogues, pre-configured user profiles containing dialogue response styles can be effectively injected into their prompts. This enables the simulated user agents to generate dialogue responses that highly match specific user profiles, thereby enhancing the realism and personalization of the simulated dialogues. This allows other agents trained with the simulated user agents to face more challenging and realistic dialogue scenarios. Consequently, the effectiveness and relevance of dialogue training are improved, and the trained agents can better adapt to diverse user needs and interaction styles in real-world applications, thus enhancing the overall quality and efficiency of dialogue training.
[0063] S32. Based on the multi-dimensional driving dialogue configuration of the recommendation agent, generate dialogue recommendation content that is adapted to the current semantic stage and user profile.
[0064] Specifically, the multi-dimensional driving scripts configured by the recommendation agent refer to a set of script strategies preset or dynamically generated within the recommendation agent. These strategies are rules or templates for generating recommendation content that can be dynamically adjusted and optimized based on multiple influencing factors (dimensions). These dimensions can include dialogue goals, user emotions, historical interactions, product features, marketing strategies, etc. For example, a series of predefined rules can be used, each containing triggering conditions (such as specific semantic stages or user profile features) and corresponding script templates or generation logic. Alternatively, a knowledge graph containing information such as dialogue goals, product knowledge, and user preferences can be constructed, and the recommendation agent can generate multi-dimensional scripts by querying and reasoning through the knowledge graph. In addition, the large language model can be fine-tuned so that it can understand and integrate multi-dimensional information when generating recommendation content, and output scripts that conform to specific strategies. For example, the model can be trained to generate scripts with different emphases (such as emphasizing product advantages, solving user pain points, and facilitating transactions) for different user profiles at different semantic stages.
[0065] Furthermore, generating dialogue recommendation content adapted to the current semantic stage and user profile aims to ensure that the recommended content output by the recommendation agent is highly context-relevant and personalized, thereby improving the effectiveness of dialogue training. When generating recommendation content, the recommendation agent first obtains information about the current semantic stage of the dialogue training. For example, if the current semantic stage is "demand exploration," the recommended content will focus on asking questions and understanding user pain points; if the semantic stage is "product introduction," the recommended content will focus on explaining product functions and advantages. This can be achieved by setting stage-based trigger conditions or weights in the multi-dimensional driving dialogue. Simultaneously, the recommendation agent also obtains user profile information configured by the simulated user agent. For example, if the user profile shows that the user prefers a concise and clear communication style, the recommended content will avoid lengthy and complex expressions; if the user profile shows that the user is price-sensitive, the recommended content may include discount information. This can be achieved by using user profile features as input parameters for generating dialogue or as a basis for selecting dialogue templates. By using the semantic stage and user profile as joint inputs, the logic of multi-dimensional driving dialogue is comprehensively judged to generate recommended content that both conforms to the current dialogue progress and fits the user's characteristics. For example, in the "demand exploration" stage, when facing "price-sensitive" users, the recommended content may subtly introduce cost-effective products or services while exploring their needs.
[0066] In this way, by adapting to the semantic stage, the recommended content can accurately match the current progress of the dialogue, avoiding inappropriate recommendations; by adapting to the user profile, the recommended content can fully consider the preferences and characteristics of the simulated user, improving the acceptance and effectiveness of the recommendations. This refined recommendation mechanism enhances the realism and effectiveness of dialogue training, enabling the trained agent to better learn how to communicate efficiently based on actual situations and user characteristics, thereby optimizing the overall dialogue training effect.
[0067] Step S32 includes the following specific details.
[0068] S321. Identify rejection intent in dialogue responses based on constraints, and filter or adjust the generated dialogue recommendation content accordingly.
[0069] Specifically, by using constraints such as pre-defined rules or models, the dialogue responses generated by the simulated user agent are analyzed to determine whether they contain attitudes or tendencies of non-acceptance, non-adoption, non-compliance, or disinterest. These constraints can include a predefined list of rejection keywords (e.g., "not needed," "not interested"), specific phrase patterns, syntactic structure rules, or intent recognition models built based on Natural Language Processing (NLP) techniques. For example, sentiment analysis models or deep learning classifiers (such as models based on the Transformer architecture) can be used to perform semantic analysis on the response text to identify the rejection sentiments or intentions contained within. Through these constraints, negative feedback from the simulated user agent can be accurately captured, providing crucial information for subsequent adjustments to recommended content.
[0070] Furthermore, after recognizing the simulated user agent's intention to refuse, the recommendation agent will revise the initially generated dialogue recommendations. Filtering operations can manifest as directly removing recommendations directly related to the refusal intention from the list of recommended content. For example, if the simulated user agent explicitly refuses the suggestion to "buy product A," the system will filter out all recommendations related to "buying product A." Adjustment operations can modify the wording and focus of existing recommendations to make them more attractive or more acceptable to the user, such as changing "buy" to "learn more." It can also switch recommendation strategies, shifting from direct sales to providing information, addressing concerns, or recommending related but not directly conflicting products or services. It can also lower the priority of recommendations related to the refusal intention and increase the priority of other irrelevant or potentially interesting content. Furthermore, after recognizing a refusal intention, the system can re-invoke the recommendation agent and inject the constraint "the user has refused X," causing it to generate entirely new recommendations that circumvent X.
[0071] Thus, by identifying rejection intentions in dialogue responses based on constraints, negative feedback from simulated users can be captured in a timely manner. The recommendation agent can filter or adjust the initially generated dialogue recommendations to avoid repeatedly recommending content that the user is not interested in or has explicitly rejected, thereby ensuring that the provided recommendations are more targeted and effective. This not only enhances the realism of dialogue training and the experience of the simulated user agent but also enables the trained agent to better handle rejection situations that may occur in real dialogues, improving the adaptability of its dialogue strategies.
[0072] In step S40, based on the process and result data of the agent analyzing the dialogue content after the dialogue training is completed, the user profile information associated with the simulated user agent is updated.
[0073] Step S40 includes the following specific details.
[0074] S41. Based on the review of the agent's response to the event signals at the end of the dialogue training phase, trace back the execution status of each stage of the dialogue training recorded in the event history list, and perform a stage completion analysis.
[0075] Specifically, the review agent only subscribes to event signals indicating the end of the dialogue training phase, such as the end event for the "End" phase. Upon receiving this event signal, the review agent initiates the review analysis process, which includes tracing back the event history list. The event bus maintains an event history list that records all published events in chronological order. The review agent extracts the execution records for each phase of the entire dialogue training process from this list, including the start time, end time, number of rounds participated in, and corresponding quality control scores for each phase. This information is used to analyze phase completion, statistically analyzing whether each dialogue target phase was fully executed and the distribution of execution rounds.
[0076] S42. Based on the dialogue training content, stratify and label the strengths, improvement areas, and weaknesses, and conduct a scoring interval analysis.
[0077] Specifically, based on preset scoring thresholds, the dialogue training content at each stage is stratified and labeled as, for example, strengths (score ≥ 90), areas for improvement (70 ≤ score < 90), or weaknesses (score < 70), forming a visualization capability and providing a quantitative basis for subsequent strategy adjustments.
[0078] S43. Update the contact history, interest tags, and secondary marketing opportunities in the user profile information based on the analysis results, and persist the updated user profile information. The analysis results include stage completion analysis and rating interval analysis.
[0079] Specifically, based on the results of stage completion analysis and rating interval analysis, the user profile information associated with the simulated user agent is automatically updated. The user profile information includes three dimensions: contact history (adding a summary and key interaction records of the current conversation), interest tags (dynamically adding or strengthening tags based on the user's expressed concerns during the conversation, such as "price sensitive," "time-sensitive," etc.), and secondary marketing opportunities (marking whether it is suitable to follow up again and suggesting an interval). The updated user profile information is persistently stored in the user profile storage module in JSON format.
[0080] In this way, the updated user profile information is loaded into subsequent dialogue training. The recommendation agent (marketing assistant) adjusts the recommendation dialogue strategy based on data such as interest tags, contact history, and secondary marketing opportunities in the profile. For example, it prioritizes using concise and efficient language for users marked as "time-sensitive" and emphasizes cost-effectiveness information for users marked as "price-sensitive". This achieves a closed-loop iteration of "training → evaluation → profile update → strategy adjustment → retraining", continuously improving the personalization and targeting of dialogue training.
[0081] For ease of understanding, the following is combined with Figures 2 to 4 The overall system operation process, device structure, and architecture hierarchy are described in a centralized manner.
[0082] Figure 2 This is a flowchart illustrating the methodology of an AI sales coaching system (i.e., a multi-agent dialogue training system based on semantic stage recognition). It details the complete business logic loop for processing a single conversation, primarily divided into four logical sections: The first step is the stage identification area. The process begins with the agent sending a message in S101, which is received by the central control unit in S102. In S103, the stage identification and confidence level are verified using LLM via the main path. If the whitelist verification passes, S105 is executed to determine the current marketing stage; if the verification fails or the identification is abnormal, the S104 downgrade path is followed, and rules are used as a fallback to ensure the success rate.
[0083] Next is the event distribution area. S106 determines whether the stage has changed. If "no", only the S107 user_session user session event is published; if "yes", the S111 three-channel event signal is published, synchronously triggering the S108 stage_start stage start event and the S109 stage_end stage end event. Subsequently, all paths converge to the S112 event bus (i.e., the event bus module), and concurrent distribution is performed through the asyncio mechanism.
[0084] The third is the parallel processing area for intelligent agents. After receiving and distributing instructions, the system executes three core tasks in parallel: S113 simulates a user intelligent agent that combines LLM and the five personality models (i.e., five personality models) to perform streaming optimization processing and outputs it to S116 three-way channel A (user message flow); S114 the quality inspector (i.e., the quality inspection intelligent agent) performs a preset quality inspection assessment and broadcasts the results after the structured assessment in S117; S115 the marketing assistant (i.e., the recommendation intelligent agent) performs six-dimensional context injection and five-layer constraint filtering and outputs it to S118 three-way channel B (marketing suggestion flow).
[0085] Finally, in the debriefing loop, S119 determines whether the current stage is the closing stage. If "no," it waits for the next round of dialogue; if "yes" (entering the closing stage), S120 triggers the debriefing agent (i.e., the debriefing agent) to conduct the debriefing, sequentially executing S122 stage completion analysis (reviewing event history) and S123 scoring item breakdown analysis (strengths / improvements / reductions, where "reductions" refers to "weaknesses"). Finally, S124 updates the three-dimensional profile (interest tags, marketing opportunities), and S125 completes the persistent storage of the debriefing.
[0086] Figure 3 This is a block diagram of the AI sales coach system, which shows the static architecture of the system from the perspective of module interaction. The core lies in the data flow and the coupling relationship between modules: The intelligent agent layer, located at the top of the architecture, comprises five key modules. The central control brain (module 102) acts as the hub, responsible for LLM stage identification and three-channel event publishing. The simulated user (automated user agent) (module 103) handles role reversal and streaming responses; the marketing assistant (recommendation agent) (module 108) and the quality inspector (quality inspection agent) (module 104) handle business strategy and quality checks respectively; and the review agent (review agent) (module 107) focuses on profile updates. The outputs of these modules are uniformly fed into the output management layer (module 109), which then distributes them to different communication channels.
[0087] The communication and coordination layer, located in the middle, serves as the hub connecting the upper and lower layers. The event bus module 101, at its core, maintains the subscriber dictionary and event history list, and is responsible for message publishing / subscription and asynchronous concurrent distribution. The dual-mode switching module 106 is responsible for switching communication modes according to the scenario and establishing a bidirectional data stream through the WebSocket communication module 110.
[0088] The data persistence layer is located at the bottom and consists of user profile storage 111. It centrally stores core data such as user interest tags, purchased products, contact history, and marketing opportunities, providing data support for the upper-layer intelligent agents.
[0089] Figure 4 This is the system architecture diagram of the AI sales coach system, which divides the system into a five-layer structure from bottom to top based on the technology stack: The infrastructure layer provides underlying resources, including LLM servers that execute non-streaming and streaming calls, configuration file storage for storing contract standards, personality models, script templates and phased strategies, and persistent media for storing user profile data.
[0090] The intelligent agent layer is built on top of the infrastructure and includes Agent Control, Agent User, Agent Marketing, Agent Quality, and Agent Review. Each intelligent agent processes specific business logic by loading the underlying models and policies.
[0091] The coordination layer (orange area) is responsible for decoupling and scheduling within the system. Its core is the Event Bus, which uses the asyncio mechanism to handle concurrent tasks. This layer also includes a front-end bridging and adaptation module to ensure smooth message flow between different layers.
[0092] The communication layer uses a WebSocket bidirectional channel to ensure real-time data interaction between the front-end and back-end.
[0093] The presentation layer is user-facing and includes the browser front-end interface and mode switching dashboard. To enhance user experience, the presentation layer is designed with Channel A (TextNode rendering the user message stream word by word) and Channel B (independent streaming rendering of the marketing suggestion stream), enabling differentiated display of different information streams.
[0094] This application also proposes a multi-agent dialogue training system based on semantic stage recognition, including a central control brain module configured to identify the semantic stages of dialogue training through a two-layer cascade method and to issue event signals based on the recognition results, an event bus module configured to receive and distribute event signals, and agent modules including simulated user agents, quality inspection agents, recommendation agents, and review agents. Each agent module pre-registers the event signals it responds to in the event bus module and is configured to asynchronously execute its function when the corresponding event signal is received.
[0095] Specifically, please refer to the appendix. Figure 3 By combining the central control brain module, event bus module, and intelligent agent module in an event-driven manner, complete decoupling and collaboration among multiple intelligent agents are achieved, enabling real-time stage recognition, accurate event triggering, and personalized dialogue training in sales training scenarios. Specifically, the central control brain module employs a two-layer cascaded architecture for semantic stage recognition. The main path constructs structured prompts containing marketing stage definitions and the entire dialogue history, using a non-streaming mode to call a large language model for semantic analysis, outputting stage names and ensuring validity through whitelist verification. When the main path fails, it automatically switches to a fallback path, using a keyword mapping table to match high-frequency words to determine the stage, ensuring valid results are returned. The central control brain module maintains a set of published stages and releases three-channel event signals when a new stage is identified: user session events are triggered each round, stage start events are triggered in the new stage, and stage end events are triggered when the previous stage is completed.
[0096] The event bus module acts as the communication hub, maintaining a subscriber dictionary with event type as the key and a chronologically recorded list of event history. Its publishing method creates an independent asynchronous task for each callback function, executing concurrently through a pooled waiting mechanism to ensure that downstream agents do not block each other. This module supports dynamic subscription and unsubscription mechanisms, providing infrastructure support for the system's dual-mode switching. The intelligent agent module achieves precise collaboration through an event bus: the simulated user intelligent agent module generates differentiated responses based on five personality models (consultant, resistant, hesitant, straightforward, and friendly), injecting personality trait descriptions and response styles into prompts in the large language model system, and ensuring the inherent consistency of the generated content through role reversal mapping; the quality inspection intelligent agent module subscribes to the end of the subscription phase, loads key checkpoints, weights, and passing standards from the JSON configuration file, injects them into the large language model evaluation prompts to obtain structured results, and employs a three-level fault tolerance strategy of regular expression extraction, direct parsing, and default value fallback to ensure evaluation continuity; the recommendation intelligent agent module subscribes to the start of the subscription phase, generates personalized scripts by integrating six dimensions of context including product information, phase strategies, script templates, user profiles, historical dialogues, and the latest messages, and outputs them after being filtered through a five-layer constraint architecture of decision layer, style layer, rejection rule layer, call termination hard rule layer, and word count control layer.
[0097] In this embodiment, please refer to the appendix. Figure 4 The overall system architecture is divided into five layers: presentation layer (browser front end, including dual-channel independent streaming rendering), communication layer (WebSocket bidirectional channel), coordination layer (EventBus), intelligent agent layer (five professional intelligent agents) and infrastructure layer (LLM server + configuration file storage + user profile storage).
[0098] In this embodiment, the event bus module, or EventBus, is the communication hub, maintaining a subscriber dictionary (with event type as the key and a list of callback functions as the value) and an event history list (recording all published events in chronological order). An independent asynchronous task (asyncio.create_task) is created for each callback function, and they are executed concurrently through a gathering and waiting mechanism (asyncio.gather) to ensure that downstream agents do not block each other. Dynamic subscription and unsubscription methods are supported, providing a foundation for dual-mode switching.
[0099] The central control module, or AgentControl, undertakes two core responsibilities: marketing stage identification and event dissemination. Stage identification employs a two-tiered cascading architecture: the main path constructs structured prompts containing five marketing stage definitions (opening, value presentation, needs discovery, action facilitation, and closing), all historical dialogues, and the current round of messages. It calls the LLM in non-streaming mode, requiring only the stage name to be output, and its validity is ensured through whitelist validation. The fallback path is automatically triggered when the main path fails, using a keyword mapping table to match high-frequency words to determine the stage, ensuring valid results are returned. AgentControl maintains a set of published stages (published_stages) and releases three-channel event signals when a new stage is identified: user session event (user_session, triggered per round), stage start event (stage_start, triggered when the new stage is triggered), and stage end event (stage_end, triggered when the previous stage is completed).
[0100] The simulated user agent (AgentUser) employs a passive invocation model, generating differentiated customer responses based on five personality models. Each personality model is defined by three dimensions: name, feature description, and response style. These are then integrated into the LLM system's prompts to generate consistent, personalized responses. Role-reversal mapping ensures the LLM correctly understands role relationships. The LLM is invoked in a streaming mode (temperature 0.8, representing parameters controlling the randomness and creativity of the model's output), and responses are pushed through a three-stage streaming output protocol.
[0101] The AgentQuality intelligent agent subscribes to the end of the phase event, loads the quality inspection standards (key checkpoints, weights, and passing scores) from the JSON configuration file, injects LLM evaluation prompts, and obtains the structured JSON evaluation results. A three-level fault-tolerant parsing strategy is employed: the first level extracts the JSON string using regular expressions, the second level directly parses the raw output, and the third level returns the default evaluation results, ensuring zero-interruption operation.
[0102] The Marketing Assistant (AgentMarketing) subscription phase begins, generating user profiles and driving the sales script through six-dimensional context injection: product information, phase strategy, script template, user profile (interest tags + purchased products + contact history), historical conversations, and latest messages. The generated script is filtered through a five-layer constraint architecture: decision layer → style layer → rejection rule layer → call termination hard rule layer → word count control layer.
[0103] The AgentReview intelligent agent only subscribes to the end event (stage_end_closing) of the closing stage and performs a three-step review: the first step is stage completion analysis (backtracking the event history to statistically analyze the execution of each stage), the second step is score interval analysis (stratified marking of strengths / improvements / weaknesses), and the third step is dynamic updating of the three-dimensional user profile (contact history + interest tags + secondary marketing opportunities). The updated profile is persistently stored in JSON format.
[0104] The following examples will provide a more detailed explanation of the above technical solutions.
[0105] Example 1 This embodiment describes how to use a multi-agent dialogue training method based on semantic stage recognition to improve the training effect of sales representatives in a financial product marketing training scenario.
[0106] First, the system initializes. A sales representative (i.e., the trainee) selects "Quality Inspection Mode" training on the training interface and specifies the simulated customer's personality model as "Consultant". The system loads preset financial product information, marketing stage strategies, and quality inspection standard configurations.
[0107] When a sales representative initiates a conversation and sends the message, "Hello, I am a sales representative from XX Financial. Today I would like to introduce our latest wealth management product to you," the system starts the conversation training process.
[0108] In the semantic stage of recognition and rule matching steps.
[0109] After receiving the message from the sales representative, the central control brain module immediately activates a two-layer cascaded architecture to determine the semantic stage of the current dialogue training.
[0110] 1) Large Language Model Semantic Analysis (Main Path): The central control module constructs a prompt word containing the definitions of five predefined marketing stages (e.g., opening, value demonstration, demand mining, action facilitation, and closing), as well as the complete historical dialogue context and the message of the current round. This prompt word is sent to the large language model for semantic analysis. Based on the semantic analysis results, the large language model outputs the name of the current stage, such as "opening." Subsequently, the system uses a whitelist verification mechanism to determine the validity of this output, ensuring that the recognition result conforms to the preset stage list.
[0111] 2) Keyword Mapping Rule Matching (Degradation Path): If the stage name returned by the large language model is invalid, or the large language model service is temporarily unavailable, the system will automatically switch to the degradation path. In this case, the system uses keyword mapping rule matching to identify high-frequency words in the current message (such as "introduction" or "product") to determine the corresponding semantic stage name and outputs the "opening" stage from the whitelist. This two-layer cascading mechanism ensures the robustness and high availability of semantic stage recognition, overcomes the limitations of single-model solutions in existing technologies, and avoids system interruptions due to model failure.
[0112] In the event signal triggering step.
[0113] Once the central control module confirms that the current semantic stage is "opening," it checks the set of published stages. Since "opening" is a new stage, the central control module triggers corresponding event signals based on the change in semantic stage. Specifically, it publishes a "stage start event" to notify the initiation of the "opening" stage, and simultaneously publishes a "user conversation event" to drive the simulated user agent to generate a response. This event-driven scheduling mechanism enables the system to respond to dialogue progress in real time, solving the problems of delayed quality control feedback and the inability to provide immediate feedback during dialogue in existing technologies.
[0114] In the step of executing dialogue training content by multiple agents.
[0115] After receiving event signals from the central control brain module, the event bus module will asynchronously and concurrently distribute these event signals to the corresponding intelligent agent modules according to the pre-registered subscription relationship.
[0116] 1) Simulated User Agent Reply Generation: After receiving a "user conversation event," the simulated user agent module, based on its configured user profile (in this example, "consultant") including dialogue response style, injects the user profile's characteristics (such as curiosity, initiative, rationality) and response style (such as politeness, clear information needs) into the prompts of its large language model. The large language model then generates dialogue responses that match the "consultant" user profile, such as: "Oh, really? What are the features of this product?" This response is pushed to the front-end interface through the streaming output management module, simulating a realistic typing rhythm. This differentiated customer simulation based on user profiles overcomes the lack of differentiation in existing customer simulation technologies, enabling sales representatives to interact with simulated customers of different personality types and improving their responsiveness.
[0117] 2) Quality Inspection Agent Execution Evaluation: When the dialogue transitions from the "opening" phase to the "value demonstration" phase, the central control module issues a "phase end event" to notify of the completion of the "opening" phase. Upon receiving this "phase end event," the quality inspection agent module immediately performs an evaluation of the dialogue training content. It loads the quality inspection criteria for the "opening" phase (including key checkpoints, scoring weights, and passing scores) from the configuration file and injects these criteria into the evaluation prompts of the large language model to obtain structured evaluation results. This real-time, event-driven quality inspection evaluation significantly reduces feedback latency from the post-event evaluation model of existing technologies, enabling sales representatives to receive improvement guidance at the optimal time for error correction.
[0118] 3) Recommendation Agent Generates Recommended Content: If the system is in "Recommendation Mode," the recommendation agent (i.e., the marketing assistant agent) will subscribe to "Phase Start Events." When it receives a "Phase Start Event" for the "Value Demonstration" phase, it will generate recommended dialogue content adapted to the current "Value Demonstration" semantic phase and user profile based on the configured multi-dimensional driving script. For example, it will integrate six dimensions of context, including product information, phase strategy, script template, user profile (such as interest tags, purchased products, contact history), historical conversations, and latest news, to generate personalized script suggestions: "You can emphasize the product's robustness and explain it in conjunction with the customer's investment goals."
[0119] 4) Rejection Intent Recognition and Adjustment: In subsequent conversations, if the simulated user agent's response contains a clear rejection intent (e.g., "I don't need it right now, please don't introduce it further"), the marketing assistant agent will identify this rejection intent based on constraints. Accordingly, it will filter or adjust its generated conversation recommendations; for example, it will no longer recommend further marketing pitches, but instead suggest that the sales representative politely end the conversation. This ensures the compliance of marketing, avoids harassment complaints, and provides a personalized and security mechanism lacking in existing technologies.
[0120] In the training closed loop and user profile update steps.
[0121] When the entire dialogue training is completed, for example, when the dialogue enters the "end" phase and triggers the "phase end event", the debriefing agent module will be activated.
[0122] 1) Analyze the dialogue content and results data: The debriefing agent module analyzes the process and results data of this dialogue, including the completion status of each stage and quality inspection scores.
[0123] 2) Update User Profile Information: The debriefing agent module dynamically updates user profile information based on the analysis results. For example, if the simulated user shows interest in long-term investment during the conversation, the debriefing agent module will add "long-term investment" to the user profile's interest tags; if the simulated user repeatedly mentions "time pressure," its "time-sensitive" tag will be updated. Simultaneously, a summary of the conversation will be added to the contact history.
[0124] 3) Determining Subsequent Personalized Decisions: Based on the updated user profile information, the system can determine personalized decisions for subsequent dialogue training. For example, in the next training session, the system can automatically adjust the simulated user's response rhythm or recommend more concise and efficient scripts based on the updated "time-sensitive" tag, thus forming a complete closed loop from training and evaluation to data accumulation and strategy adjustment. This overcomes the shortcomings of existing technologies, such as the inability to accumulate training data, the lack of user profile management, and dynamic update mechanisms.
[0125] Example 2 This embodiment demonstrates the application of an AI sales coaching system in telecommunications package marketing training, and represents the preferred implementation plan.
[0126] System initialization: Agents select "Quality Inspection Mode" and "Consultative" personality models in the browser frontend. The frontend sends the configuration to the backend via WebSocket. The central control system loads telecom package product information, five-stage marketing strategies, and quality inspection standard configurations. The event bus registers quality inspectors to subscribe to the stage end event.
[0127] The training process is as follows: The agent sends the message, "Hello, this is Xiao Wang from XX Telecom. Today, I'd like to introduce our new package to you." Upon receiving the message, the central control system constructs a prompt containing five phase definitions, the 10 most recent historical dialogues, and the current message. It then calls the LLM (Limited Module Management) for phase identification. The LLM returns "opening," which passes whitelist verification. The central control system checks the set of published phases; "opening" indicates a new phase, and it sequentially publishes three-channel event signals.
[0128] The event bus asynchronously and concurrently distributes events. A simulated user (consultant type, passively invoked) loads a personality configuration, injecting the characteristics of the "consultant type" (curious, proactive, rational) and response style (polite, clear information needs) into the LLM system prompts. The most recent 10 dialogues are extracted and role-reversed, and the LLM is invoked in streaming mode to generate a response: "Oh, really? What promotional offers are available now?" This response is pushed to frontend channel A via a streaming output protocol. The frontend uses TextNode for word-by-word rendering, simulating the rhythm of real typing.
[0129] Upon receiving the phase completion event, the quality inspector loads the opening phase quality inspection standards (key checkpoints: proactive greeting, self-introduction, product introduction, friendly tone, moderate speaking speed, weights [15,20,20,25,20], passing score 70 points) from the configuration file, injects LLM evaluation prompts, and obtains a score. Three-level fault-tolerant parsing ensures the validity of the evaluation results. The evaluation results are pushed to the front-end for display via front-end bridging.
[0130] The agent continues the conversation into the value demonstration phase. The central control system recognizes the change from "opening" to "value" and issues a stage start event (stage_start_value) and a stage end event (stage_end_opening). Upon receiving the stage start event, the marketing assistant (if in recommendation mode) constructs a six-dimensional context to generate suggested sales scripts.
[0131] Training continues, progressing through the needs assessment and action facilitation phases before entering the closing phase. The reviewer receives the closing event (stage_end_closing), analyzes the completion status of the five phases from the event history list, performs interval analysis based on quality control scores, updates the three-dimensional user profile data, and persists it to a JSON file, thus forming a complete training loop.
[0132] Example 3 This embodiment demonstrates the application of an AI sales coaching system in insurance product marketing training. The difference between this embodiment and the first embodiment is that the product type is insurance; the training subjects are experienced agents; a recommendation mode is used; and existing user profile data is loaded.
[0133] During system initialization, the target user profile is loaded: interest tags include "health protection," purchased products include "medical insurance," and contact history includes two previous contact records. Agents select the "resistant" personality model and "recommendation mode." The event bus registration marketing assistant subscription phase begins.
[0134] The agent sends an opening message: "Hello, I previously recommended medical insurance to you. Now, we have an accident insurance plan that complements your existing coverage." The central control system identifies this as the opening phase through a two-layer cascading mechanism, initiating the event. Upon receiving the event, the marketing assistant constructs a six-dimensional context: product information is the accident insurance details; the phase strategy is to build trust and introduce the product; and the user profile shows that the user has already purchased medical insurance and is interested in health protection. Based on this six-dimensional integration, LLM generates a personalized sales pitch suggestion: "Understanding your busy schedule, we can introduce this accident insurance plan that complements your existing medical insurance in just one minute." This suggestion explicitly links to the user's existing products, reflecting a profile-driven personalization approach.
[0135] The simulated user (resistant) replies, "No, I'm busy," in a cold and direct tone. The agent continues the conversation as suggested in the script, and after several rounds, the simulated user repeatedly expresses refusal. The marketing assistant detects the continued refusal intent at the third level of constraint (rejection rule) and determines that the termination condition has been triggered at the fourth level of constraint (hard call termination rule), suggesting the agent politely end the call.
[0136] The debriefinger triggers a debriefing after the conversation ends. Phase completion analysis shows only the opening and partial value demonstration phases were completed. The scoring range analysis identifies needs assessment and action facilitation as weaknesses. Profile update: A summary of this conversation is added to the contact history; the interest tag "time-sensitive" (based on repeated mentions of "busy") is added; and secondary marketing opportunities are marked "no." After the profile is persistently stored, the system can automatically adjust its strategy based on the "time-sensitive" tag during subsequent training, prioritizing concise and efficient communication.
[0137] The profile is continuously updated through multiple training sessions, forming a closed-loop iteration of "training → evaluation → profile update → strategy adjustment → retraining".
[0138] Through the above examples, this method addresses the problems of existing sales training systems, such as limited functionality, delayed feedback, limited semantic understanding, lack of multi-agent collaboration, and lack of training loops, by employing two-layer cascaded semantic stage recognition, event-driven multi-agent collaboration, personalized customer simulation and script recommendation, and debriefing-driven user profile updates. This enables the training process to be intelligent, real-time, and personalized.
[0139] This invention also provides a computer-readable storage medium storing a computer program that enables a computer to execute the above-described multi-agent dialogue training method based on semantic stage recognition.
[0140] Specifically, the computer-readable storage medium can be any combination of one or more storage media, such as flash memory, hard disk, optical disk, or cloud storage media. As one specific implementation, the computer-readable storage medium can be deployed on a server or cloud platform, and when the processor loads and executes the program, it can automatically complete the entire process from acquiring dialogue text to training multi-agent dialogue based on semantic stage recognition.
[0141] This application also proposes an electronic device, Figure 5 The diagram shown illustrates the structure of an electronic device 200 provided in an embodiment of the present invention. In some embodiments, the electronic device may be a terminal device such as a mobile phone, computer, or personal computer. Furthermore, the multi-agent dialogue training method based on semantic stage recognition provided in this embodiment can also be applied to service response systems based on terminal artificial intelligence. This embodiment does not limit the specific application scenarios of the multi-agent dialogue training method based on semantic stage recognition.
[0142] like Figure 5 As shown, the electronic device 200 provided in this embodiment of the invention includes a memory 201 and a processor 202.
[0143] Specifically, memory 201 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. Electronic device 200 may further include other removable / non-removable, volatile / non-volatile computer system storage media. Memory 201 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0144] The processor 202 is connected to the memory 201 and is used to execute the computer program stored in the memory 201 so that the electronic device 200 executes the multi-agent dialogue training method based on semantic stage recognition provided in any embodiment of the present invention.
[0145] In an alternative implementation, the processor 202 may be a general-purpose processor such as a central processing unit (CPU); it may also be other programmable logic devices such as a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA).
[0146] In an optional embodiment, the electronic device 200 of this invention may further include a display 203. The display 203 is communicatively connected to the memory 201 and the processor 202, and is used to display the relevant GUI interactive interface for multi-agent dialogue training based on semantic stage recognition.
[0147] It is understood that the multi-agent dialogue training system, electronic device and storage medium based on semantic stage recognition provided in the embodiments of the present invention correspond to the multi-agent dialogue training method based on semantic stage recognition provided in the embodiments of the present invention. The explanation, examples and beneficial effects of the relevant contents can be referred to the corresponding parts of the method, and will not be repeated here.
[0148] In summary, compared with existing technologies, it has the following beneficial effects: 1. Improved recognition accuracy during the dialogue phase.
[0149] By using LLM to fuse the entire dialogue history for semantic-level stage recognition as the main path, combined with keyword matching as a downgrade path and a whitelist verification mechanism, the accuracy of stage recognition is improved.
[0150] 2. Quality inspection feedback delays have been reduced.
[0151] By injecting configurable quality control standards into LLM prompts to enable real-time online assessment, coupled with an event-driven stage completion trigger mechanism, the quality control feedback delay has been reduced from hours of manual review to less than seconds. Agents can receive improvement guidance at the optimal time for error correction, significantly improving the effectiveness of dialogue training.
[0152] 3. Improved consistency in assessments.
[0153] By replacing manual review with LLM-based configurable assessment, subjective differences among different reviewers are eliminated, and dialogues of the same quality receive consistent scores at different times.
[0154] 4. Expanded training scenario coverage.
[0155] LLM-driven customer simulation, based on five personality models (counseling, resistant, hesitant, straightforward, and friendly), expands the training scenario coverage to more diverse scenarios (combinations of the five personality types and five marketing stages). Agents can interact with simulated customers of different personality types, improving their responsiveness to real customers.
[0156] 5. The system architecture achieves decoupling.
[0157] Through the publish / subscribe pattern of the event bus and the asymmetric event subscription design, collaboration between different agents can be achieved with zero direct references. New agents only need to register their event subscriptions with the event bus; no modifications to the code of any existing agents are required, thus improving scalability.
[0158] 6. Achieve closed-loop training data.
[0159] By analyzing the completion status of each stage based on the event history by the reviewers, the three-dimensional data of the user profile are automatically updated and persistently stored, and the training data is reused, providing data support for the personalized optimization of subsequent training.
[0160] 7. Training costs are reduced.
[0161] By replacing manual quality inspection and simulation training with AI systems, the average training cost per person (including the cost of manual quality inspection, the cost of trainers, and the cost of recorded learning) is significantly reduced.
[0162] 8. Improved accuracy in rejecting intent recognition.
[0163] The five-layer constraint architecture employs a dual safeguard of a rejection rule layer and a call termination hard rule layer to ensure that the system will not continue to push conversational marketing content after the user has explicitly rejected it, thus guaranteeing the compliance of the conversation and effectively avoiding harassment complaints.
[0164] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-agent dialogue training method based on semantic stage recognition, characterized in that, include: Based on a two-layer cascaded architecture of semantic recognition and rule matching, the semantic stage of the current dialogue training is determined. Based on the changes in the semantic phase, corresponding event signals are triggered; Based on multiple agents receiving subscribed event signals, corresponding dialogue training content is executed. The agents include a simulated user agent for responding to the dialogue, a quality control agent for evaluating the dialogue, a recommendation agent for engaging in dialogue with the simulated user agent, and a debriefing agent for analyzing the dialogue. Based on the process and result data of the intelligent agent analyzing the dialogue content after the dialogue training is completed, the user profile information associated with the simulated user intelligent agent is updated. Based on the updated user profile information, the recommendation dialogue strategy of the recommending agent in subsequent dialogue training is adjusted to form a cyclical iteration of dialogue training.
2. The method according to claim 1, characterized in that, The two-layer cascaded architecture based on semantic recognition and rule matching determines the semantic stage of the current dialogue training, including: Semantic analysis is performed on prompt words containing the definition of the dialogue target stage and the dialogue context using a large language model, outputting the name of the semantic stage, and the validity is judged by a whitelist check. If the whitelist verification is invalid, then the keyword mapping rule matching is used to output the semantic stage name of the corresponding whitelist.
3. The method according to claim 1, characterized in that, The step of triggering a corresponding event signal based on the determined semantic stage change includes: Based on the phase start event at the beginning of each new phase of dialogue training, the recommending agent is triggered to generate dialogue training content; Based on the phase end event at the end of each dialogue training phase, the quality inspection agent is triggered to perform an evaluation of the dialogue training content.
4. The method according to claim 1, characterized in that, The process of executing corresponding dialogue training content based on multiple agents receiving subscribed event signals includes: Based on a simulated user agent configured with a user profile that includes dialogue response style, the user profile is injected into the prompt words of the simulated user agent to generate dialogue response content that conforms to the user profile.
5. The method according to claim 1 or 4, characterized in that, The process of executing corresponding dialogue training content based on multiple agents receiving subscribed event signals includes: Based on the multi-dimensional driving dialogue configured by the recommending agent, dialogue recommendation content is generated that is adapted to the current semantic stage and user profile.
6. The method according to claim 5, characterized in that, The generation of dialogue recommendation content adapted to the current semantic stage and user profile includes: Based on constraints, the system identifies rejection intentions in dialogue responses and filters or adjusts the generated dialogue recommendations accordingly.
7. The method according to claim 1, characterized in that, The process and results data of analyzing dialogue content by the review agent after dialogue training are used to update the user profile information associated with the simulated user agent, including: Based on the review of the agent's response to event signals at the end of the dialogue training phase, the execution status of each stage of the dialogue training recorded in the event history list is traced back to analyze the stage completion rate. Based on the dialogue training content, strengths, areas for improvement, and weaknesses were stratified and marked, and a scoring range analysis was conducted. Based on the analysis results, update the contact history, interest tags, and secondary marketing opportunities in the user profile information, and persist the updated user profile information. The analysis results include stage completion analysis and rating interval analysis.
8. A multi-agent dialogue training system based on semantic stage recognition, characterized in that, include: The central control brain module is configured to identify the semantic stages of dialogue training through a two-layer cascade method and issue event signals based on the identification results; An event bus module is configured to receive and distribute the event signals; The intelligent agent module includes a simulated user intelligent agent, a quality inspection intelligent agent, a recommendation intelligent agent, and a review intelligent agent; each intelligent agent module pre-registers the event signals it responds to in the event bus module and is configured to asynchronously execute its function when the corresponding event signal is received.
9. A computer-readable storage medium, characterized in that, It stores a computer program, wherein the computer program causes the computer to execute the multi-agent dialogue training method based on semantic phase recognition as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, The electronic device includes: Processor and memory; The memory stores program instructions; The processor is configured to run the program instructions to execute the multi-agent dialogue training method based on semantic phase recognition as described in any one of claims 1 to 7.