Scenarized AI partner training self-adaptive interaction and multi-dimensional evaluation method, system and equipment

By breaking down real business scenarios into multi-stage nodes and building adaptive interaction and multi-dimensional evaluation models, the problems of rigid processes, singular roles, and ambiguous evaluation in scenario-based simulation training are solved, achieving a highly realistic and accurate training experience and improving practical capabilities.

CN121073718APending Publication Date: 2025-12-05SHANGHAI JIAO ZHIYAN (SUZHOU) TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511055490.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing scenario-based simulation training technologies suffer from rigid scenario processes, limited role simulation, and one-sided evaluation dimensions, making it difficult to achieve dynamic interaction and accurate feedback in real-world scenarios.

Method used

By breaking down real-world business scenarios into multi-stage nodes, configuring detection points and jump rules, and building an adaptive interaction model, the AI ​​role can dynamically adjust its response style based on user performance. Furthermore, through multi-dimensional evaluation models, quantitative analysis is conducted to generate accurate feedback.

Benefits of technology

It achieves a highly realistic training experience, precise capability diagnosis and improvement, personalized adaptation, significantly improved efficiency in converting training content into practical skills, and enhanced matching degree between training content and real business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121073718A_ABST
    Figure CN121073718A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a scene-based AI partner training self-adaptive interaction and multi-dimensional evaluation method. The method comprises the steps that S1, a system scene library and a role library are constructed to configure real service scene requirements, and a structured scene model is obtained and stored; s2, automatically matching and initializing AI agent roles according to scene nodes; s3, constructing an interaction entrance of a user and an AI agent role, and inputting a training demand and interaction content by the user; s4, performing real-time scene adaptation detection according to the interaction content; and S5, performing multi-dimensional data acquisition on the interactive content of the user to obtain a multi-dimensional feature data set. And S6, constructing an evaluation model based on the multi-dimensional feature data set and outputting a multi-dimensional evaluation report. According to the invention, the scene process can be flexibly adjusted along with the user performance through the nodal disassembly and jump rule; the AI role adjusts response logic and mood to improve interaction authenticity according to scene characteristics; and specific improvement suggestions are generated through multi-dimensional data acquisition and analysis, so that comprehensive evaluation and closed-loop promotion are realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of artificial intelligence, and in particular to a scene-based AI accompanying training adaptive interaction and multi-dimensional evaluation method, system and device. BACKGROUND

[0002] With the increasing demand for practicality and individualization in the field of vocational skill training, scene-based simulation training has become a core way to improve the practical ability of users. Whether it is a job interview, customer communication, business negotiation or professional skill demonstration (such as PPT report, product explanation), it is necessary to consolidate the skills through high-frequency and high-imitation interactive drills. In the traditional training mode, offline role-playing relies on human accompanying training, which has problems such as high cost, limited scene coverage, and delayed feedback. Although digital training tools have gradually become popular, they are mostly limited to static content presentation or simple question-and-answer level, and it is difficult to reproduce the dynamic interaction logic in real scenes, resulting in a gap between learning and practice for users and low skill conversion rate.

[0003] Under this background, scene-based AI accompanying training technology has gradually become a research focus. Its core demand is to simulate real scene interactions through artificial intelligence to achieve a training closed loop of "anytime practice, on-demand adjustment, and precise feedback". This technology needs to solve three major problems: first, how to decompose complex business scenarios into quantifiable and configurable interaction processes; second, how to enable AI characters to have response styles and logic that adapt to scenario characteristics; and third, how to build a multi-dimensional evaluation system to achieve objective quantification and targeted guidance of user performance.

[0004] Currently, scene-based simulation training related technologies mainly present the following implementation directions: Fixed script-based scenario simulation: training scenarios are built through pre-set dialogue processes and response content, and users complete interactions according to fixed paths. For example, in a customer visit simulation, the system pushes dialogue nodes in a fixed order of "greeting -> product introduction -> objection handling", and the AI character can only match pre-set answers according to user input, and cannot respond to expressions not included in the script (such as the user suddenly mentioning competitor information). The limitation of this solution is that the scene is not flexible and it is difficult to cover the diversified situations in real business, and the training effect is close to "rote practice".

[0005] Single-dimensional evaluation system: focusing on evaluating local features of user expression, such as speech speed, fluency, or keyword matching. For example, in a speech training, the system only detects whether "the number of words per minute meets the standard" and "whether it contains pre-set professional terms", but ignores core indicators such as content logic coherence (such as the relevance of viewpoints and arguments) and scene adaptability (such as the appropriateness of tone in formal situations). The evaluation result is one-sided and cannot support users in accurately improving their skills.

[0006] Unified role interaction model: A unified AI response style is used to deal with different scenarios, such as handling interview questions and customer negotiations with the same tone and logic. For example, when simulating an "interviewer" and a "customer procurement manager", the AI uses neutral statement responses, without reflecting the rigorous questioning characteristics of the interviewer or the interest-oriented questioning of the procurement manager, resulting in a lack of role authenticity and weak training immersion.

[0007] Unstructured scenario management: Real business scenarios are not broken down into processes, and training is carried out only through "free dialogue". For example, in sales training, users can randomly jump to "negotiation" and "after-sales" links, and the system cannot identify the logical loopholes of "negotiating directly when the demand is not clear", nor can it guide the dialogue back to the reasonable process, resulting in a lack of goal-oriented training and difficulty in forming a structured path for improving skills.

[0008] In combination with existing scenario simulation training related technologies (including virtual environment interaction, generative AI coaching, etc.), the core shortcomings can be summarized as the following three points: Scenario process is rigid and lacks dynamic guidance mechanism: Existing technologies simulate business scenarios using fixed scripts or single processes (such as fixed order dialogue nodes), and cannot dynamically adjust the interaction logic based on user performance. For example, when the user has not completed key steps such as "demand mining", the system cannot guide the dialogue back to the reasonable process through active questioning, resulting in training that deviates from real business logic and users being easily trapped in ineffective interactions.

[0009] Single role simulation, insufficient adaptability: The response style of AI coaching roles is fixed and cannot adapt to the characteristics of different scenarios. For example, when simulating "customer visits" and "interviews", the AI uses a neutral tone, without reflecting the interest-oriented questioning of the customer role (such as "can the cost be reduced by 10%") or the rigorous logical verification of the interviewer (such as "please specify the quantitative data of the project results"), resulting in weak role authenticity and insufficient training immersion.

[0010] Evaluation dimension is one-sided and feedback is vague: Existing evaluations focus on a single dimension (such as speech fluency or keyword matching), lacking coverage of key business scenario indicators. For example, in business negotiation simulation, only "speech speed is up to standard" is detected, while ignoring "logical integrity of objection handling" and "core interest point coverage", resulting in users being unable to clearly identify skill gaps and difficulty in targeted improvement. SUMMARY

[0011] Explanations of terms involved in the present invention: 1. Scenario-based AI coaching: refers to simulating real business scenarios (such as customer visits, interviews, etc.) through artificial intelligence technology, allowing users to interact in a virtual environment to improve their practical skills.

[0012] 2. Node disassembly: Split the complete business scenario into multiple continuous stages according to the process sequence (such as "opening → demand mining → objection handling"), and each stage is a "node".

[0013] 3. Detection point: The key verification item preset in each scene node, used to judge the integrity and accuracy of user interaction content (such as the detection points of "demand mining node" include "budget range inquiry" and "core pain point identification").

[0014] 4. Jump rule: The switching logic between scene nodes, when the user meets the detection point requirements of the current node, it automatically enters the next node, and when it does not meet, it triggers AI guidance or repeats the current node.

[0015] 5. AI agent role: A virtual interactive object simulated by the system (such as customers, interviewers, etc.), with response style and interaction logic adapted to the characteristics of the scene.

[0016] 6. Multi-dimensional evaluation: Quantitative analysis of user performance from multiple dimensions such as speech characteristics (speech rate, fluency), semantic content (keyword hit, logical integrity), and scene adaptability (node completion, detection point coverage).

[0017] 7. Adaptive interaction: The AI role dynamically adjusts the response content and style according to the user's real-time performance, such as adding follow-up questions when the user's expression is incomplete, and reducing the intensity of questioning when the user's emotion is tense.

[0018] 8. BiLSTM-CRF: A deep learning model that combines bidirectional long short-term memory network (BiLSTM) and conditional random field (CRF), commonly used for sequence labeling tasks (such as automatically identifying scene node turning points).

[0019] 9. BERT: A pre-trained language model based on Transformer, which can be used for semantic understanding, text classification and other natural language processing tasks, and is used in this invention to improve the accuracy of the evaluation model.

[0020] The technical problem to be solved by the present invention is to design a scenario-based AI practice adaptive interaction and multi-dimensional evaluation method, system and device, and the specific purposes include: Implementing structured disassembly and dynamic process management of scenarios: Disassemble real business scenarios into multi-stage dialogue nodes, configure detection points and jump rules, and enable AI to actively guide interactions based on user performance (such as automatically asking follow-up questions when the "demand mining" detection point is not hit), solving the problem of "scenario process rigidity".

[0021] Constructing an adaptive simulation mechanism for roles: enabling AI companion roles to dynamically adjust their response style according to the characteristics of the scene (such as customer role simulating interest-oriented questioning, and interviewer role strengthening logical verification), improving the authenticity of interaction and the adaptability of the scene.

[0022] Establishing a multi-dimensional real-time evaluation model: integrating speech characteristics (speech rate, fluency), semantic analysis (keyword hit, logical integrity), etc. to generate sub-item scores and targeted improvement suggestions (such as "core interest point coverage rate is only 40%, suggest adding 'cost advantage' related expressions"), achieving a "practice-feedback-improvement" closed loop, and solving the problem of "fuzzy evaluation".

[0023] To solve the above technical problems, the present application provides a method for adaptive interaction and multi-dimensional evaluation of scenario-based AI coaching, specifically including the following steps: Step S1: Constructing a system scenario library and a role library: configuring real business scenario requirements, obtaining scenario nodes, defining node jump rules (if the user does not mention "budget range", automatically trigger AI follow-up question: "What is the budget range for your procurement / project?"), role behavior rules (such as "customer role needs to prioritize questioning 'cost-effectiveness' in the 'dispute handling' node"), forming a dynamically adjustable scenario logic network, obtaining a structured scenario model, and storing it.

[0024] Step S2: Automatically matching and initializing AI agent roles (including role identity setting, response style rules) according to scenario nodes, serving as "smart opponents" in the interaction link.

[0025] Step S3: Constructing an interaction portal for users and AI agent roles, users input training requirements (selecting scenarios, starting training instructions) and interaction content (voice / text responses).

[0026] Step S4: Real-time scenario adaptation detection based on interaction content: if the current node detection point is hit, proceed to the next node; if the current node detection point is not hit, the AI agent role will guide the user to ask follow-up questions until the detection node is hit.

[0027] Step S5: Multi-dimensional data collection on user interaction content, obtaining multi-dimensional feature data set.

[0028] Step S6: Building an evaluation model based on the multi-dimensional feature data set, realizing quantitative analysis of user performance, and outputting a multi-dimensional evaluation report, the evaluation results including scores, ratings and improvement suggestions.

[0029] Further, in step S1, the scene nodes are obtained by scene structured decomposition, and a real business scene (such as "sales customer visit" and "job interview defense") is subdivided into a node process (for example, the customer visit scene is decomposed into "opening greeting -> demand mining -> scheme recommendation -> objection handling -> cooperation conclusion"), and each node is configured with a detection point rule (for example, the "demand mining node needs to cover three key information of the customer, including "budget range", "core pain point" and "decision maker information").

[0030] Further, in step S2, a preset role template is called according to the scene type, and the professionalism and tone intensity of the role response are adjusted in combination with the scene industry characteristics (the financial scene focuses on the verification of standard expressions, and the Internet scene focuses on the verification of innovative ideas), so as to ensure that the role adapts to the scene demand.

[0031] Further, in step S3, the user input supports multi-modal input such as voice dialogue, text input and micro-expression collection combined with a camera, and real-time user interaction data is captured, including voice signal, text content, interaction time length and emotion word frequency.

[0032] Further, in step S5, the data collection includes: Voice dimension: extract voice signal features, including speech rate (words / minute) and fluency (number of pauses / total time length).

[0033] Semantic dimension: analyze text content by NLP technology, including keyword hit number (such as "demand mining node needs to hit 3 keywords 'budget, pain point, decision maker'"), logical correlation (whether the derivation of viewpoint -> argument is complete) and professional term coverage.

[0034] Scene dimension: record node progress (currently in "demand mining" node), detection point hit state (has hit "budget", has not hit "decision maker"), and interaction time length (whether it is overtime), to provide multi-angle basis for evaluation.

[0035] Further, in step S6, the quantitative analysis of user performance includes checking the completion degree of scene detection points (such as "demand mining node needs to hit 3 detection points, actually hits 2 -> compliance deduction 20%") and evaluating semantic logic (whether the viewpoint is clear and the argument is sufficient), voice expressiveness (whether the speech rate is suitable for the scene demand, such as "180-220 words / minute" in the interview scene), and industry knowledge application (whether professional terms are misused, whether industry compliance expressions are covered).

[0036] Further, step S7 is also included: based on the evaluation result, a personalized training path scheme is planned for the user, specifically including: Short board positioning: identify the dimension with the lowest evaluation score (e.g. "semantic logic score 50, current short board"), associate the corresponding training resources (logical deduction dialogue template, industry case library).

[0037] Scenario reinforcement: for the failed scenario node (e.g. "demand mining node detection point not completely hit"), automatically generate a special training task (force to repeat the node 3 times, randomly adjust the AI character questioning direction each time).

[0038] Further, it also includes step S8: returning user training data and evaluation results to the scene library and character library, optimizing scene detection point rules (e.g. finding that 80% of users miss the "decision maker information" detection point → increasing the weight of the detection point) and AI character response strategies (e.g. high error rate of users for high-frequency follow-up dialogue → adjusting the follow-up logic to "guided questioning") through machine learning, realizing system "self-evolution".

[0039] The application also provides a self-adaptive interaction and multi-dimensional evaluation system for scenario-based AI training, which executes the aforementioned self-adaptive interaction and multi-dimensional evaluation method for scenario-based AI training, comprising: Scenario disassembly and configuration module: used for scenario structural disassembly and interaction logic configuration, obtaining a structured scenario model (including node flow, detection point, interaction rule), and storing in the system scene library and character library for subsequent module calling.

[0040] AI agent character initialization module: based on the scenario disassembly result, automatically matching and initializing the AI agent character.

[0041] User interaction module: used for building an interaction entrance for users and AI characters, and capturing user interaction data in real time.

[0042] AI adaptive response module: based on user interaction content, scene node rules and AI character features, dynamically generating AI response content (text / voice form) and feeding back to the user in real time, while recording interaction process data (user input timing, AI response strategy) for intelligent evaluation and calculation module analysis.

[0043] Multi-dimensional data acquisition module: used for synchronously acquiring full-quantity data in the interaction process, including voice dimension, semantic dimension and scene dimension, outputting multi-dimensional feature data set (structured voice parameters, semantic label, scene progress label), and transmitting to the intelligent evaluation and calculation module in real time.

[0044] Intelligent evaluation and calculation module: based on the data of the multi-dimensional acquisition module, constructing an evaluation model to realize quantitative analysis of user performance, and finally outputting a multi-dimensional evaluation report and pushing it to the result feedback and optimization module.

[0045] Adaptive learning path module: based on the evaluation results, a personalized training path scheme is planned for the user, and is pushed to the user end and the scene disassembly and configuration module.

[0046] Result feedback and optimization module: based on the evaluation report of the evaluation calculation module and the user historical training data, a user feedback interface (including report and training entry) and system optimization instruction (scene rule update and role strategy adjustment) are output, and the "training-evaluation-optimization" closed loop is completed.

[0047] The application also provides an adaptive interaction and multi-dimensional evaluation device for scenario-based AI accompanying training, comprising: At least one processor; and At least one memory in communication connection with the processor; Wherein the memory stores instructions executable by the processor, and the instructions are executed by the processor to enable the device to perform the adaptive interaction and multi-dimensional evaluation method of the scenario-based AI accompanying training.

[0048] The application relates to natural language processing, sentiment computing, human-computer interaction evaluation and intelligent accompanying training system technology, and can be applied to the fields of scenario-based training such as professional skill training, interview simulation and customer communication. The application has the following beneficial effects: 1. High-simulation scene training experience Compared with the fixed virtual scene in some prior art, the application disassembles a real business scene into multiple stage nodes through "scene disassembly and process modeling", and combines dynamic jump rules (such as automatic follow-up questions of AI when the user does not hit the "demand mining" detection point), so that the training scene can flexibly adapt to the user's performance and reproduce the complex interaction logic in the real business (such as sudden questioning of the customer and high-frequency follow-up questions of the interviewer). The user can experience the complete business process from the beginning to the end in the training, and the interaction of each link is consistent with the scene characteristics (such as the customer role focusing on "cost-benefit" and the interviewer role focusing on "logic-data"), and the immersion and real scene consistency are significantly improved.

[0049] 2. Precise ability diagnosis and improvement Unlike some prior art which focuses on single-dimensional evaluation, the application combines voice features (speech rate, fluency), semantic analysis (keyword hit, logical integrity) and scene adaptability (node progress, detection point coverage) through "multi-dimensional real-time evaluation" to generate quantitative scores and specific improvement suggestions. The user can clearly locate the weaknesses (such as "fast speech rate leading to unclear information transmission" and "logical discontinuity lacking supporting arguments"), and perform targeted intensive training (such as repeating weak nodes and adjusting the difficulty of the AI role) through the "adaptive learning path", so as to realize the closed loop of "practice-feedback-improvement", and the ability improvement efficiency is improved by more than 40% compared with the traditional training mode.

[0050] 3. Personalization and scalability Compared with the partial basis training function, the present application supports non-technical personnel to configure scene nodes, detection points and role characteristics (such as the professional term library of "academic customers" and the follow-up logic of "pressure interviewers") through a visual interface, which can quickly adapt to the training needs of different industries and different positions. At the same time, the system continuously accumulates user data through the "result feedback and optimization module", dynamically optimizes scene rules and role strategies (such as automatically increasing the weight of high-frequency missed detection points), realizes the two-way iteration of "user ability improvement" and "system self-evolution", and can make the matching degree of training content and real business improve to more than 90% in the long run.

[0051] 4. Real combat ability conversion efficiency improvement Through high-simulation scene training and precise feedback, the performance of users in actual business is significantly improved: for example, after the "customer visit scene" training, the demand mining accuracy of sales position users is improved by 50%, and the success rate of objection handling is improved by 35%; after the "pressure interview scene" training, the logical expression clarity of interviewers is improved by 45%, and the core problem answer completeness is improved by 40%. The system effectively solves the problems of "learning and training disconnection" and "difficulty in real combat ability conversion" in traditional training, and realizes efficient connection from "theoretical learning" to "real combat application". BRIEF DESCRIPTION OF DRAWINGS

[0052] The specific embodiments of the present application will be further illustrated below in conjunction with the drawings.

[0053] Figure 1 The flowchart of the adaptive interaction and multi-dimensional evaluation method of the scene-based AI training of the present application.

[0054] Figure 2 The system block diagram of the adaptive interaction and multi-dimensional evaluation system of the scene-based AI training of the present application.

[0055] Figure 3 The execution flow schematic diagram of the adaptive interaction and multi-dimensional evaluation system of the scene-based AI training of the present application. DETAILED DESCRIPTION Embodiment 1

[0056] In combination Figure 1 , the adaptive interaction and multi-dimensional evaluation method of the scene-based AI training of the present embodiment specifically includes the following steps: Step S1: Constructing system scenario library and character library: configure real business scenario requirements, obtain scenario nodes, define node jump rules (if the user does not mention "budget range", automatically trigger AI to ask: "What is the budget range of your procurement / project?"), character behavior rules (such as "customer character in 'objection handling' node needs to prioritize questioning 'cost effectiveness'"), form a dynamically adjustable scenario logic network, obtain a structured scenario model, and store it.

[0057] In this embodiment, preferably, in step S1, scenario nodes are obtained by scenario structured decomposition, and real business scenarios (such as "sales customer visit" and "job interview defense") are subdivided into nodal processes (for example, the customer visit scenario is decomposed into "opening greeting → demand mining → solution recommendation → objection handling → cooperation conclusion"). Each node is configured with a detection point rule (for example, the "demand mining node needs to cover three key information of the customer: "budget range", "core pain point", and "decision maker information").

[0058] In this embodiment, specifically, scenario nodes can also be automatically generated based on machine learning, without relying on manual decomposition of scenario nodes. Instead, a large amount of real business scenario dialogue data (such as historical interview recordings and customer communication records) is collected, and unsupervised learning algorithms (such as clustering analysis) are used to automatically identify key turning points in the dialogue (such as "switching from greeting to demand discussion" and "objection raising moment"), which are used as scenario nodes. Sequence labeling models (such as BiLSTM-CRF) are used to automatically extract the detection points corresponding to the nodes (such as "budget", "pain point", and other key words frequently appearing in the "demand discussion node"). Without manually configuring node logic, scenario processes are automatically generated through data-driven methods, still achieving "scenario dynamic jump" (node conversion rules based on model recognition), and ultimately achieving the purpose of scenario-based interaction. Existing mature natural language processing tools (such as spaCy and NLTK) can support text clustering and sequence labeling, and those skilled in the art can achieve model training through parameter tuning without breaking through the technology.

[0059] Step S2: Automatically match and initialize AI agent characters according to scenario nodes, including character identity setting and response style rules, as "smart opponents" in the interaction link.

[0060] Preferably in step S2, a preset role template is called according to the scene type, and the professionalism and tone intensity of the role response are adjusted in combination with the scene industry characteristics to ensure that the role adapts to the scene requirements. Specifically in this embodiment, if it is a financial scene, the scene industry characteristics need to focus on the verification of compliance expressions, and the tone intensity needs to be “precise statement type response”; if it is an Internet scene, the scene industry characteristics need to focus on the verification of innovative ideas, and the tone intensity needs to support “flexible humorous interaction”. Through the context understanding ability of the large language model, the role self-adaptive response is simulated, and the implementation mode of “role feature library + rule matching” is replaced, and the purpose of “role real-time interaction” can still be achieved. The existing open source large language model supports custom prompt word engineering, and technical personnel can adapt to the scene requirements by fine-tuning the model parameters without the need to redevelop the underlying model.

[0061] Specifically in this embodiment, the role response can also be generated based on the large language model. Instead of presetting a role feature library and response rules, the large language model is input with a role description (such as “you are a strict interviewer who is good at asking details about career planning”) and current scene information (such as “interview scene - self-introduction stage”), and the model directly generates response content that conforms to the role style; by embedding detection point requirements (such as “if the user does not mention past project experience, ask ‘please give an example of the core project you are responsible for’”) in the model prompt word, guided interaction is achieved.

[0062] Step S3: An interaction entrance for the user and the AI agent role is constructed, the user inputs training requirements (selects a scene and starts a training instruction) and interaction content (voice / text response). After each round of response of the user, scene rule verification and AI response generation are triggered immediately to avoid “lag interaction” that destroys the training immersion.

[0063] Preferably in step S3, the user input supports multi-modal input such as voice dialogue, text input, and micro-expression collection combined with a camera, and real-time user interaction data is captured, including voice signals, text content, interaction duration, and emotion word frequency.

[0064] Step S4: Real-time scene adaptation detection is performed according to the interaction content: if the current node detection point is hit, the next node is promoted; if the current node detection point is not hit, the AI agent role guides the user for further questioning until the detection node is hit.

[0065] In this embodiment, specifically, the priority check is performed on whether the user interaction hits the current node detection point (such as "the user mentions 'budget 500-800 thousand', which hits the 'demand mining' node detection point"), and the response direction is determined (if it hits, it proceeds to the next node, and if it does not hit, it triggers a guided follow-up question); the response style is adjusted in combination with the scene node characteristics (the "demand mining" node requires an "open guided response": "You just mentioned efficiency improvement, what specific problems do you want to solve?"; the "objection handling" node requires a "questioning and challenging response": "The scheme you mentioned sounds good, but can it really achieve this effect in practice?").

[0066] Step S5: Multi-dimensional data collection is performed on the interaction content of the user to obtain a multi-dimensional feature data set.

[0067] In this embodiment, preferably, in step S5, the data collection synchronously collects all data in the interaction process, including: Voice dimension: voice signal features are extracted, including speech rate (words / minute) and fluency (number of stalls / total duration).

[0068] Semantic dimension: text content is analyzed through NLP technology, including keyword hit number (such as "the 'demand mining' node needs to hit the three keywords 'budget, pain point, and decision maker"), logical correlation (whether the derivation of viewpoint→argument is complete), and professional term coverage.

[0069] Scene dimension: node progress (currently in the "demand mining" node), detection point hit status (has hit "budget" and has not hit "decision maker"), and interaction duration (whether it is overtime) are recorded, providing multi-perspective basis for evaluation.

[0070] Step S6: An evaluation model is constructed based on the multi-dimensional feature data set to realize quantitative analysis of the user's performance and output a multi-dimensional evaluation report, and the evaluation result includes score, rating, and improvement suggestion.

[0071] In this embodiment, specifically, the score items include voice, semantic, scene adaptation, etc., the rating is a comprehensive ability rating, presented in the form of a radar chart, and improvement suggestions are generated (such as "semantic logic is insufficient, it is suggested to supplement 'argument data expression: this scheme has been implemented in 3 peer companies, reducing costs by 20%'").

[0072] In this embodiment, preferably, in step S6, the quantitative analysis of the user's performance includes checking the completion degree of the scene detection point (such as "the 'demand mining' node needs to hit 3 detection points, and actually hits 2→20% of compliance deduction") and evaluating semantic logic (whether the viewpoint is clear and the argument is sufficient), voice expressiveness (whether the speech rate is adapted to the scene requirements, such as "180-220 words / minute" in an interview scene), and industry knowledge application (whether professional terms are misused, and whether industry compliance expressions are covered).

[0073] In this embodiment, the evaluation model can also be trained based on expert annotation data. Instead of calculating the evaluation result by preset scoring dimensions and weights, the scores (such as "voice fluency 80 points, logical integrity 60 points") and comments of field experts on the user practice data are collected as training data to train an end-to-end evaluation model (such as a BERT-based regression model). The model input is the user's voice transcription text and voice features, and the output is multi-dimensional evaluation scores and improvement suggestions (generated by the model learning from expert comments). Through data-driven machine learning models to replace "standardized scoring", "multi-dimensional accurate evaluation" can still be achieved, and the evaluation result is more in line with the actual business standards. Existing deep learning frameworks (such as TensorFlow, PyTorch) support the training of such models, and technical personnel can verify the model effect through public data sets (such as speech evaluation corpus).

[0074] The embodiment preferably further comprises a step S7 of planning a personalized training path scheme for the user based on the evaluation result, specifically including: Short board positioning: identify the dimension with the lowest evaluation score (such as "semantic logic score 50 points, which is the current short board"), and associate the corresponding training resources (logical deduction dialogue template, industry case library).

[0075] Scenario reinforcement: for the failed scenario nodes (such as "demand mining node detection point not completely hit"), automatically generate special training tasks (force to repeat the node 3 times, and randomly adjust the AI role questioning direction each time).

[0076] The embodiment preferably further comprises a step S8 of returning the user training data and evaluation result to the scenario library and role library to optimize the scenario detection point rules (such as finding that 80% of users miss the "decision maker information" detection point → increasing the weight of the detection point) and AI role response strategies (such as high error rate for users of high-frequency follow-up dialogue → adjusting the follow-up logic to "guided questioning") through machine learning, realizing the "self-evolution" of the system.

[0077] The adaptive interaction and multi-dimensional evaluation method of the scenario-based AI coaching in this embodiment is applied to "sales new employee training - customer visit scenario", and the completion process timing of this embodiment is as follows: Scenario construction phase: The administrator disassembles "customer visit" into a 5-node process through the "scenario disassembly and configuration module", configures "demand mining node needs to hit 3 detection points (budget, pain point, decision maker)", and outputs the structured scenario model to the scenario library.

[0078] Training start phase: User selects "Client Visit Training" -> "AI Character Initialization Module" calls "Business Client Character Template" (with "Cost-Oriented Questioning Style") -> initializes AI character, ready for interaction.

[0079] Interaction Training Phase: User says: "Mr. Wang, today I mainly want to talk about our new product, which is especially suitable for your company!" (voice input) The "User Interaction Module" captures the content and triggers the "Multi-Dimensional Data Collection Module" to extract voice features (speech rate 150 words per minute, fluency 90%) and semantic features (only hit "product introduction", missed "demand mining" detection points).

[0080] The "AI Adaptive Response Module" determines that the "missed detection points" and calls the scene rule to generate a response: "Thank you for the introduction, but I am more concerned about how much money this solution can save us? If the budget exceeds the limit, it will be difficult for the leader to approve (ask for 'budget range' detection point)" The interaction continues until all nodes are completed or the user actively ends.

[0081] Evaluation Optimization Phase: The "Intelligent Evaluation Calculation Module" outputs the report: "Semantic logic score 60 points (did not cover 'decision maker, core pain points'), scene adaptation score 70 points (promoted cooperation intention)".

[0082] The "Result Feedback and Optimization Module" pushes improvement suggestions: "In the 'demand mining' node, add: 'Mr. Wang, besides cost, what other aspects of the current solution do you think need to be optimized? Also, please tell me who the project decision maker is?'"

[0083] The "Adaptive Learning Path Module" generates reinforcement training: forces repetition of the "demand mining" node training, and the AI character randomly switches questioning direction (such as "Does your solution's delivery period match our project progress?").

[0084] The adaptive interaction and multi-dimensional evaluation method of the scenario-based AI coaching of this embodiment disassembles real business scenarios (such as interviews, client visits) into multiple stages of node disassembly (such as "opening speech -> demand mining -> objection handling -> conclusion") through scene node disassembly and dynamic process configuration methods, presets detection points for each node (such as "opening speech must include 'past cooperation' and 'purpose of visit'"), and configures node jump rules through a visual interface (such as the AI automatically asking for guidance when missing detection points). This method solves the problem of rigid process in traditional scenario training and is the core foundation of realizing dynamic interaction.

[0085] The adaptive interaction and multi-dimensional evaluation method of the scenario-based AI coaching of this embodiment adjusts the response style and interaction logic of the AI coaching role dynamically based on the characteristics of the scene through an AI intelligent body role adaptive simulation mechanism: for example, the "interview intelligent body" simulates a rigorous logical verification tone and conducts in-depth verification on "career planning" content; the "customer intelligent body" simulates the questioning manner of an academic expert (such as asking professional questions about parameters) according to the "product advantages" expressed by the user. This mechanism improves the authenticity of interaction through precise adaptation of the role and the scene.

[0086] The adaptive interaction and multi-dimensional evaluation method of the scenario-based AI coaching of this embodiment fuses voice signal processing (speech rate, fluency) and semantic analysis (keyword hit, logical integrity) through a multi-dimensional real-time evaluation model, configures differential scoring weights for different scenes (for example, "core detection point coverage" accounts for 70% of the weight and "speech rate + fluency" accounts for 30% of the weight in the PPT presentation scene), and generates an evaluation report containing sub-item scores, problem annotations, and improvement suggestions. This model realizes comprehensive quantification and accurate feedback on user performance.

[0087] The adaptive interaction and multi-dimensional evaluation method of the scenario-based AI coaching of this embodiment supports non-technical personnel to complete scene node disassembly, detection point configuration, and AI role feature setting (such as response style, professional terminology library) through a visual interface using a low-code scene and role configuration platform, without the need for code to quickly build exclusive training scenes and intelligent bodies, thereby reducing the technical threshold for enterprises to customize AI coaching systems. Embodiment 2

[0088] In combination with Figure 2 and Figure 3 , the adaptive interaction and multi-dimensional evaluation system of the scenario-based AI coaching of this embodiment executes the adaptive interaction and multi-dimensional evaluation method of the scenario-based AI coaching of embodiment 1, including: A scene disassembly and configuration module: used for scene structured disassembly and interaction logic configuration, to obtain a structured scene model (including node flow, detection point, and interaction rule) and store it in the system scene library and role library for subsequent module calling.

[0089] In this embodiment, as the "scene engine" of the system, the administrator supports two core operations through a visual interface: ① Scene structured disassembly: real business scenarios (such as "sales customer visit" and "job interview defense") are subdivided into nodal processes (for example, the customer visit scene is disassembled into "opening greeting → demand excavation → solution introduction → objection handling → cooperation conclusion"), and each node is configured with detection point rules (for example, the "demand excavation node" needs to cover three key information items of the customer, namely "budget range", "core pain points", and "decision maker information").

[0090] ②Interaction logic configuration: Define node jump rules (if the user does not mention the "budget range", automatically trigger AI to ask: "What is the budget range of your procurement / project?"), role behavior rules (such as "the customer role needs to prioritize questioning 'cost effectiveness' in the 'dispute handling' node"), and form a dynamically adjustable scene logic network.

[0091] The data input and output of this module is: The input is the scene requirements configured by the administrator (business scene description, industry rule document), and the output is the structured scene model (including node flow, detection point, interaction rule), which is stored in the system scene library for subsequent module calling.

[0092] AI agent role initialization module: Based on the scene disassembly result, automatically match and initialize the AI agent role, specifically: ①Role feature loading: According to the scene type, call the preset role template (such as "academic customer role" associated with professional term library, rigorous questioning logic; "pressure interviewer role" associated with high-frequency questioning skills, negative response style).

[0093] ②Dynamic parameter calibration: Combined with the scene industry characteristics (financial scene focuses on rule statement verification, Internet scene focuses on innovative idea verification), adjust the professionalism and tone intensity of the role response (such as "rigorous statement response" in financial scene, "flexible humorous interaction" in Internet scene), to ensure that the role adapts to the scene requirements.

[0094] The data input and output of this module is: The input is the scene type output by the scene disassembly module, and the output is the AI role instance (including role identity setting, response style rule, ), which serves as the "smart opponent" in the interaction link.

[0095] User interaction module: Used to build the interaction entrance between the user and the AI role, supporting multi-modal input (voice dialogue, text input, even combining camera micro-expression collection), and real-time capturing of user interaction data (voice signal, text content, interaction duration, emotion word frequency). The core feature of this module is real-time interaction: after each round of user response, the scene rule verification and AI response generation are triggered immediately to avoid "lagging interaction" that destroys the training immersion.

[0096] The data input and output of this module is: The input is the user's training requirements (selecting scenes, starting training instructions) and interaction content (voice / text response), and the output is the user interaction data sequence (including content text, voice feature parameters, interaction timing), which is pushed to the AI adaptive response module and multi-dimensional data collection module.

[0097] AI self-adaptive response module: dynamically generate AI response content (text / audio form) based on user interaction content, scene node rules and AI character characteristics, real-time feedback to users, while recording interaction process data (user input timing, AI response strategy) for intelligent evaluation and calculation module analysis.

[0098] In this embodiment, the module is the "intelligent interaction hub" of the system, which dynamically generates AI responses based on three core logics: ① Rule matching layer: priority check whether the user interaction hits the current node detection point (such as "the user mentions 'budget 50-80 million', hits the 'demand mining' node detection point"), decide the response direction (hit the next node, not hit the trigger guided follow-up question).

[0099] ② Scene adaptation layer: adjust the response style according to the characteristics of the scene node ("demand mining" node needs "open guided response": "You just mentioned efficiency improvement, which specific link do you want to solve the problem?"; "objection handling" node needs "question challenging response": "The scheme you mentioned sounds good, but can it really achieve this effect in practice?").

[0100] Data input and output of this module: The input is "user interaction content (voice / text), scene node rules, AI character characteristics", the output is AI response content (text / audio form), real-time feedback to users, while recording interaction process data (user input timing, AI response strategy) for evaluation module analysis.

[0101] Multi-dimensional data acquisition module: used for synchronous acquisition of full-quantity data in the interaction process, including voice dimension, semantic dimension, scene dimension, output multi-dimensional feature data set (structured annotated voice parameters, semantic labels, scene progress labels), real-time transmission to intelligent evaluation and calculation module.

[0102] In this embodiment, the module acts as the "data sensor" of the system, synchronously acquiring full-quantity data in the interaction process, covering three dimensions: ① Voice dimension: extract voice signal characteristics (speech rate (words / minute), fluency (number of stalls / total duration).

[0103] ② Semantic dimension: analyze text content through NLP technology (number of keyword hits (such as "demand mining node needs to hit 'budget, pain point, decision maker' three keywords"), logical correlation (whether the deduction from viewpoint to argument is complete), professional term coverage).

[0104] ③ Scene dimension: Record node progress (currently in the "demand mining" node), detection point hit state (hit "budget", not hit "decision maker"), interaction duration (whether timeout), provide multi-angle basis for evaluation.

[0105] Data input and output of this module: The input is the original interactive data of the user interaction module (voice stream, text stream), and the output is a multi-dimensional feature data set (structured annotated voice parameters, semantic labels, scene progress labels), which is transmitted to the intelligent evaluation calculation module in real time.

[0106] Intelligent evaluation calculation module: Based on the data of the multi-dimensional acquisition module, build an evaluation model to realize quantitative analysis of user performance, and finally output a multi-dimensional evaluation report to the result feedback and optimization module.

[0107] In this embodiment, the quantitative analysis of user performance is specifically: ① Basic compliance layer: Check the completion degree of scene detection points (such as "the "demand mining node" needs to hit 3 detection points, and actually hits 2 → deduct 20% of compliance points").

[0108] ② Professional ability layer: Evaluate semantic logic (whether the viewpoint is clear, whether the argument is sufficient), voice expressiveness (whether the speech speed is suitable for the scene requirements, such as "180-220 words per minute" in an interview scene), and industry knowledge application (whether professional terms are misused, whether industry compliance expressions are covered).

[0109] Finally, output multi-dimensional evaluation results (radar chart form presents voice, semantic, scene adaptation and other sub-item scores, and comprehensive ability rating), and generate improvement suggestions (such as "insufficient semantic logic, suggest adding 'argument data expression: this scheme has been implemented in 3 peer companies, reducing cost by 20%'").

[0110] Data input and output of this module: The input is the feature data set of the multi-dimensional data acquisition module, and the output is the evaluation report (including score, rating, and improvement suggestions), which is pushed to the result feedback and optimization module.

[0111] Adaptive learning path module: Based on the evaluation results, plan a personalized training path scheme for the user, and push it to the user end and scene disassembly and configuration module.

[0112] In this embodiment, specifically, the personalized training path includes: ① Shortboard positioning: Identify the lowest dimension of the evaluation score (such as "semantic logic score 50 points, which is the current short board"), and associate the corresponding training resources (logical deduction dialogue template, industry case library).

[0113] ② Scene reinforcement: For the failed scene nodes (such as "demand mining node detection point not fully hit"), automatically generate special training tasks (force to repeat the node 3 times, and randomly adjust the AI character questioning direction each time).

[0114] The data input and output of this module: The input is the evaluation report of the evaluation calculation module, and the output is the personalized training plan (including training nodes, resource recommendation, difficulty parameters), which is pushed to the user end and the scene disassembly and configuration module (used to adjust the subsequent training logic).

[0115] The result feedback and optimization module: based on the evaluation report of the evaluation calculation module and the user historical training data, output the user feedback interface (including report, training entrance) and system optimization instruction (scene rule update, role strategy adjustment), complete the "training-evaluation-optimization" closed loop.

[0116] In this embodiment, the module specifically constructs the "user touch layer" and "system optimization layer" of the evaluation result: ① User feedback: present the evaluation result in the form of a visual report (radar chart shows sub-item scores, text annotation key problems, pop-up window pushes improvement suggestions), support user "one-key review interaction process" (locate problem occurrence node), "directly start reinforcement training" (jump to adaptive learning path module).

[0117] ② System optimization: return user training data and evaluation results to scene library and role library, optimize scene detection point rules (such as finding that 80% of users miss "decision maker information" detection point → increase the weight of this detection point) and AI character response strategies (such as high-frequency asking users with high error rate → adjust the asking logic to "guided questioning") through machine learning, realize system "self-evolution".

[0118] The data input and output of this module: The input is the evaluation report of the evaluation calculation module and the user historical training data, and the output is the user feedback interface (including report, training entrance) and the system optimization instruction (scene rule update, role strategy adjustment), which completes the "training-evaluation-optimization" closed loop. Embodiment 3

[0119] The adaptive interaction and multi-dimensional evaluation device of the scene-based AI coaching in this embodiment includes: At least one processor; and At least one memory in communication connection with the processor; Wherein, the memory stores instructions executable by the processor, and the instructions are executed by the processor to make the device execute the adaptive interaction and multi-dimensional evaluation method of the scene-based AI coaching in embodiment 1.

[0120] To ensure the efficient operation of the device in this embodiment, the device relies on the following core hardware and architecture: Hardware layer: ① Multi-modal interaction terminal: supports voice collection (microphone array to ensure clear far-field sound pickup), text input (touch screen / keyboard), and optional video collection (camera for micro-expression analysis extension).

[0121] ② GPU server cluster: carries AI roles of natural language processing (NLP) and sentiment computing models (such as BERT model for semantic analysis and WaveNet model for speech synthesis), ensuring real-time interaction (response delay < 500 ms).

[0122] ③ Distributed database: stores scenario library (structured scenario model), role library (AI role template), and user training data (multi-dimensional evaluation results and interaction logs), supporting concurrent read and write of billions of data.

[0123] Software layer: ① Scenario configuration platform: based on a Web visual interface, administrators complete scenario disassembly and rule configuration, using a low-code engine (such as Retool) to reduce the operation threshold.

[0124] ② AI interaction engine: integrates open-source / self-developed NLP frameworks (such as TensorFlow and PyTorch), deploys dialogue management models (DialogueStateTracking), and realizes intelligent agent dynamic response.

[0125] ③ Evaluation analysis background: builds evaluation models through Python data analysis libraries (Pandas and Matplotlib) and outputs visual reports.

[0126] In the above description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the above description is only a preferred embodiment of the present application, and the present application can be implemented in many other ways different from those described herein, therefore the present application is not limited by the specific implementation disclosed above. Meanwhile, any person skilled in the art can make many possible changes and modifications to the technical solutions disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the technical solutions of the present application. Any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application, without departing from the scope of the technical solutions of the present application, are still within the scope of protection of the present application.

Claims

1. An adaptive interaction and multi-dimensional evaluation method for scenario-based AI coaching, characterized in that: Comprise the following steps: Step S1: Construct system scenario library and role library: configure real business scenario demand, obtain scenario node, define node jump rule, role behavior rule, form dynamically adjustable scene logic net, get structured scenario model, and store; Step S2: According to the scene node, the AI intelligent agent role is automatically matched and initialized; Step S3: Build the interaction entrance of user and AI intelligent agent role, user input training demand and interaction content; Step S4: Real-time scene adaptation detection is carried out according to the interaction content: if the current node detection point is hit, it is promoted to the next node; if the current node detection point is not hit, the AI intelligent agent role guides the user to ask until the detection node is hit; Step S5: Multi-dimensional data collection is carried out on the user's interaction content, and a multi-dimensional feature data set is obtained; Step S6: Based on the multi-dimensional feature data set, an evaluation model is constructed, the quantitative analysis of user performance is realized, and a multi-dimensional evaluation report is output, and the evaluation results include score, rating and improvement suggestion.

2. The adaptive interaction and multi-dimensional evaluation method of the scenario AI coaching according to claim 1, characterized in that: In step S1, the scene node is obtained by scene structured decomposition, the real business scenario is subdivided into nodal process, and each node is configured with detection point rule.

3. The adaptive interaction and multi-dimensional evaluation method of the scenario AI coaching according to claim 1, characterized in that: In step S2, according to the scene type, the preset role template is called, and the professional degree and tone intensity of the role response are adjusted combined with the scene industry characteristics to ensure that the role adapts to the scene demand.

4. The adaptive interaction and multi-dimensional evaluation method of the scenario AI coaching according to claim 1, characterized in that: In step S3, the user input supports multi-modal input such as voice dialogue, text input and micro-expression collection combined with camera, and real-time user interaction data is captured, including voice signal, text content, interaction time and emotion word frequency.

5. The adaptive interaction and multi-dimensional evaluation method of the scenario AI coaching according to claim 1, characterized in that: In step S5, the data collection includes: Voice dimension: extract voice signal features, including speech rate and fluency; Semantic dimension: analyze text content through NLP technology, including keyword hit number, logical correlation and professional term coverage; Scene dimension: record node progress, detection point hit state and interaction time.

6. The adaptive interaction and multi-dimensional evaluation method of the scenario AI coaching according to claim 1, characterized in that: In step S6, the quantitative analysis of user performance includes checking the completion degree of scene detection point and evaluating semantic logic, voice expressiveness and industry knowledge application.

7. The adaptive interaction and multi-dimensional evaluation method of the scenario AI coaching according to claim 1, characterized in that: It also includes step S7: based on the evaluation results, a personalized training path scheme is planned for the user, specifically including: Short board positioning: identify the lowest score dimension, and associate the corresponding training resources; Scene reinforcement: for the scene nodes that do not pass, automatically generate special training tasks.

8. The adaptive interaction and multi-dimensional evaluation method of the scenarized AI coaching according to claim 7, characterized in that: It also includes step S8: user training data and evaluation results are fed back to the scenario library and role library, and scene detection point rules and AI role response strategies are optimized through machine learning.

9. An adaptive interaction and multi-dimensional evaluation system for scenario-based AI coaching, characterized in that: The system performs the adaptive interaction and multi-dimensional evaluation method of the scenario AI training of any one of claims 1-8, comprising: Scene decomposition and configuration module: used for scene structured decomposition and interaction logic configuration, to obtain structured scenario model and store in system scenario library and role library for subsequent module calling; AI intelligent agent role initialization module: based on the scene decomposition result, the AI intelligent agent role is automatically matched and initialized; User interaction module: used for building the interaction entrance of user and AI role, and capturing user interaction data in real time; An AI adaptive response module: based on user interaction content, scene node rules and AI character characteristics, dynamically generate AI response content (text / voice form), real-time feedback to users, while recording interaction process data (user input timing, AI response strategy) for intelligent evaluation and calculation module analysis; Multi-dimensional data acquisition module: used to synchronously collect full-quantity data in the interaction process, including voice dimension, semantic dimension, scene dimension, output multi-dimensional feature data set, real-time transmission to intelligent evaluation and calculation module; Intelligent evaluation and calculation module: based on the data of multi-dimensional acquisition module, build evaluation model to realize quantitative analysis of user performance, finally output multi-dimensional evaluation report, push to result feedback and optimization module; Adaptive learning path module: based on the evaluation results, plan personalized training path scheme for users, and push to user end and scene disassembly and configuration module; Result feedback and optimization module: based on the evaluation report of evaluation and calculation module and user historical training data, output user feedback interface and system optimization instruction.

10. An adaptive interaction and multi-dimensional evaluation device for scenario-based AI coaching, characterized in that it comprises: Comprise: At least one processor; And At least one memory connected with the processor in communication; Wherein, the memory stores instructions executable by the processor, the instructions are executed by the processor to make the device execute the adaptive interaction and multi-dimensional evaluation method of the scene AI training of any one of claims 1-8.

Citation Information

Cited By

  • Intelligent deep questioning method and system

    CN121542399A