Training method, system and equipment based on capability vector and dynamic game, and medium

By employing a training method based on capability vectors and dynamic game theory, the system updates trainees' capability status in real time and modulates AI personality parameters to generate dynamic interactive content. This solves the rigidity problem of existing training systems, provides a personalized and realistic training experience, and improves training efficiency and effectiveness.

CN122050211APending Publication Date: 2026-05-15SHANGHAI HAOYI INFORMATION SCI & TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI HAOYI INFORMATION SCI & TECH CO LTD
Filing Date
2026-01-27
Publication Date
2026-05-15

Smart Images

  • Figure CN122050211A_ABST
    Figure CN122050211A_ABST
Patent Text Reader

Abstract

The invention provides a training method, system and device based on a capability vector and a dynamic game, and a medium, belongs to the technical field of artificial intelligence and education, and aims at solving the problems that an existing training system is rigid in content and insufficient in interaction reality sense. The method comprises the following steps: receiving interaction data of a student, and updating a multi-dimensional capability vector representing the capability state of the student in real time; based on the capability vector, personality parameters of the AI personality model are dynamically modulated; and on the basis of the modulated personality parameters and the interaction context, response content is generated in real time by utilizing a generative model to continue interaction. According to the method, a closed-loop feedback mechanism of the dynamic game of the student behaviors and the AI personality is constructed, so that highly-personalized and immersive training experience is realized, and intensive training can be carried out aiming at the defects of the students.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the intersection of artificial intelligence and educational technology, and in particular to a training method, system, device and medium based on ability vectors and dynamic game theory. Background Technology

[0002] In professional fields such as corporate training and medical education, simulation training systems are widely used to improve trainees' practical skills. In existing technologies, a common approach is based on pre-designed tree-like branching scenarios, where trainee selections trigger the system to jump to the next preset scenario. While this method achieves basic interaction, all interaction paths and results are pre-defined, making the content rigid and predictable. Trainees are prone to rote memorization after repeated training, making it difficult to improve their real-world adaptability. Furthermore, designing numerous branching paths to increase realism leads to extremely high development and maintenance costs.

[0003] To overcome these shortcomings, other technical solutions attempt to introduce artificial intelligence models for adaptive adjustments. For example, they might adjust the difficulty of subsequent training tasks or recommend supplementary training content by evaluating user behavior in real time, and utilize generative models to generate interactive content. However, the adaptability of these solutions typically remains at the level of adjusting the external stimuli provided to the user. Their core flaw lies in the fact that the AI ​​coach's own behavioral strategies or interaction style are fixed; it cannot dynamically change its interaction style based on the student's real-time performance, for example, changing from "patient" to "impatient." This results in an interactive environment that still lacks the dynamism and immersion of real interpersonal interaction, making it impossible to provide targeted, in-depth, and intensive training to address the student's specific skill weaknesses. Summary of the Invention

[0004] The purpose of this application is to provide a training method, system, device and medium based on capability vectors and dynamic game theory to solve the problems of rigid content, insufficient realism of interaction and weak training targeting in existing simulation training systems.

[0005] To address the aforementioned technical problems, this application provides a training method based on ability vectors and dynamic game theory, comprising the following steps: a) In an interactive session, receiving student interaction data via a user terminal and recording the interaction data and historical interactions to form the current interaction context; b) Based on the interaction data, updating the multidimensional ability vector representing the student's ability state in real time; c) Based on the updated multidimensional ability vector, dynamically modulating at least one personality parameter of a preset AI personality model; d) Based on the modulated personality parameter and the current interaction context, generating response content in real time using a generative model; and e) Presenting the response content to the student via the user terminal to continue the interactive session.

[0006] Optionally, the interactive data includes at least one of the student's text input, acoustic features of the voice data, and decision response time.

[0007] Furthermore, in step b), the multidimensional capability vector is updated using a nonlinear update algorithm. The nonlinear update algorithm reflects the cumulative effect of continuous errors by assigning increasing penalty weights to consecutive errors, or it uses a recurrent neural network structure to handle the temporal relationship of decision-making behavior.

[0008] Optionally, the personality parameters include at least one of patience, aggression, questioning, emotionality, and depth of questioning.

[0009] Optionally, the generative model is a large language model.

[0010] Optionally, the method further includes the following steps: after the interactive session ends, recording the time series changes of the multidimensional capability vector during the session; analyzing the time series changes to identify at least one key event point in which the rate of change of the capability vector value exceeds a preset threshold and the sign of its second derivative changes; and generating an evaluation report based on the key event point, wherein the evaluation report lists the key event point, as well as the interaction data of students who are temporally adjacent to the key event point and the system's response content.

[0011] This application also provides a training system based on ability vectors and dynamic game theory, including a user interaction module for interacting with trainees, collecting their interaction data, and forming a current interaction context based on the interaction data and historical interactions; a multi-dimensional ability vector evaluation engine connected to the user interaction module, configured to update the multi-dimensional ability vector representing the trainee's ability state in real time according to the interaction data; an AI personality modulation engine connected to the multi-dimensional ability vector evaluation engine, configured to dynamically modulate at least one personality parameter of the AI ​​personality model according to the updated multi-dimensional ability vector; and a dynamic response generator connected to the AI ​​personality modulation engine and the user interaction module, configured to generate response content in real time based on the modulated personality parameters and the current interaction context obtained from the user interaction module; wherein, the user interaction module is further configured to present the response content to the trainee.

[0012] Optionally, the personality parameters include at least one of patience, aggression, questioning, emotionality, and depth of questioning.

[0013] Optionally, the dynamic response generator includes a large language model.

[0014] Furthermore, it also includes a report generation module, which is connected to the multidimensional capability vector assessment engine, the user interaction module, and the dynamic response generator, and is configured to: record the capability vectors output by the multidimensional capability vector assessment engine to form a time series; analyze the time series to identify at least one key event point in which the rate of change of the capability vector value exceeds a preset threshold and the sign of its second derivative changes; and generate an assessment report based on the key event point, the assessment report showing the key event point, as well as the interaction data of trainees who are temporally adjacent to the key event point and the system's response content.

[0015] This application also provides an electronic device, including: a processor; and a memory storing a computer program; wherein, when the computer program is executed by the processor, it implements any of the methods described above.

[0016] This application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any of the methods described above.

[0017] Compared with the prior art, the technical solution provided in this application has the following beneficial effects:

[0018] 1. Achieving highly personalized training: By directly linking trainees' ability assessment results (multi-dimensional ability vectors) with AI's behavioral strategies (personality parameters), the training system based on ability vectors and dynamic game theory in this application can perceive trainees' unique ability weaknesses in real time and dynamically adjust the AI ​​opponent's strategies and difficulty for targeted reinforcement training, providing a customized training experience that is "one face for a thousand people", which greatly improves training efficiency and effectiveness.

[0019] 2. Enhanced Realism and Immersion: AI is no longer a sparring partner with fixed behavioral strategies, but an intelligent adversary capable of perceiving learners' weaknesses and dynamically adjusting its own "personality" and strategies. By dynamically generating unpredictable interactive content based on modulated personality parameters, it simulates the complexity and variability of real-world interpersonal interactions, effectively training learners' on-the-spot adaptability and stress management skills.

[0020] 3. Fundamentally reduce content development costs: Since interactive content is generated in real time, rather than relying on a large library of manually written scripts, this invention shifts the burden of content creation from "exhaustively exploring all possibilities" to "defining capability dimensions and modulation rules," resulting in an order-of-magnitude reduction in development and update costs.

[0021] 4. Provides in-depth competency diagnosis: By recording and analyzing the time series changes of competency vectors and identifying key event points and their contexts, it is possible to trace the trajectory and root causes of changes in trainees' competency performance in key situations. This transforms the assessment results from simple score evaluations into diagnostic feedback with in-depth causal analysis, providing highly instructive guidance. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of the structure of an adaptive training system provided in an embodiment of this application;

[0024] Figure 2 A flowchart illustrating an adaptive training method provided in an embodiment of this application;

[0025] Figure 3 This is a schematic diagram illustrating the internal working principle of the AI ​​personality modulation engine provided in the embodiments of this application;

[0026] Figure 4 This is a schematic diagram illustrating capability vector time series and critical event analysis provided in an embodiment of this application.

[0027] The main reference numerals in the attached diagrams are explained as follows: 10 - User interaction module; 20 - Multidimensional ability vector assessment engine; 30 - AI personality modulation engine; 31 - Modulation rule library; 40 - Dynamic response generator; 50 - Personality prototype library; V - Multidimensional ability vector. Detailed Implementation

[0028] To make the implementation methods, technical solutions, and beneficial effects of this application clearer, the following detailed description will be provided in conjunction with the accompanying drawings and specific embodiments. It should be noted that the specific embodiments described herein should not be considered as limitations on this application, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0029] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings.

[0030] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0031] Before providing a further detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.

[0032] A multidimensional competency vector V is a data structure used to quantitatively represent a trainee's overall competency in a specific training scenario. It is typically represented as an N-dimensional floating-point vector, such as V = [c1, c2, …, cn], where each vector ci represents a predefined key competency dimension that needs to be evaluated. For example, in sales training, these dimensions might include "product knowledge," "relationship building," and "objection handling"; in medical consultation training, they might include "history taking," "differential diagnosis," and "communication and empathy." The value of this vector changes dynamically, reflecting the trainee's performance level in real time during the interaction. The competency vector value is the value of that specific vector within the multidimensional competency vector.

[0033] AI personality model: This refers to a computational model used to define the behavioral style, strategic tendencies, and interaction patterns of an artificial intelligence (AI) tutor. It does not refer to a specific person, but rather to characterize their "personality" traits through a set of quantifiable personality parameters. The system can preset one or more personality archetypes, such as a "picky customer," a "patient and professional mentor," or a "nervous new colleague," each archetype corresponding to an initial set of personality parameters.

[0034] Personality parameters refer to a set of variables used to quantitatively describe specific traits of the AI ​​personality model. For example, they may include "patience," "aggression," and "skepticism." The base personality parameter P_base refers to the default parameter values ​​of the AI ​​personality before the interaction begins or when it is not subjected to specific stimuli; it is usually stored in the personality prototype library 50. The modulated personality parameter P_modulated refers to the result obtained by the system of this application dynamically adjusting the base personality parameters according to the trainee's multidimensional ability vector (V) during the interaction process. This set of modulated personality parameters will directly guide the generation style and strategy of subsequent response content.

[0035] Current interaction context refers to all relevant information accumulated up to the current point in time during the interaction session. This includes not only the direct input from the learner in this round and the direct output from the system, but also the dialogue records throughout the entire session history, the learner's key decision sequences, and the historical trajectory of the multidimensional capability vector (V). The current interaction context provides the generative model with the necessary background information to ensure that the generated responses are logically coherent and contextually appropriate.

[0036] Generative models are artificial intelligence models that can automatically generate natural language text that conforms to grammar and logic based on input instructions and contextual information. In the scenario of this application, the core task of the generative model is to receive the modulated personality parameters P_modulated and the current interaction context, and to create the next dialogue or the next behavioral description for the AI ​​tutor in real time based on these inputs, thereby driving the continuous interaction.

[0037] Based on the technical problem to be solved by this application and the foregoing content, this application proposes a training method based on ability vectors and dynamic game theory. This method constructs a closed-loop feedback control mechanism of "student behavior → real-time ability assessment → AI personality modulation → dynamic content generation". The method includes the following steps: First, in the interactive session, the user terminal receives the student's interaction data and records the interaction data and historical interactions to form the current interaction context; then, based on the interaction data, the multi-dimensional ability vector representing the student's ability state is updated in real time; then, based on the updated multi-dimensional ability vector, at least one personality parameter of the preset AI personality model is dynamically modulated; subsequently, based on the modulated personality parameter and the current interaction context, response content is generated in real time using a generative model; finally, the response content is presented to the student through the user terminal to continue the interactive session. Through the cyclical execution of the above steps, a training environment in which the student and the AI ​​coach influence each other and engage in dynamic game theory is formed.

[0038] In one possible implementation, the interaction data includes at least one of the student's text input, acoustic features of speech data, and decision response time. By collecting multimodal interaction data, the system can capture the student's behavior and state more comprehensively and accurately, thereby enabling more effective evaluation.

[0039] In one possible implementation, the updated multidimensional ability vector employs a nonlinear update algorithm. This algorithm assigns increasing penalty weights to consecutive errors to reflect the cumulative effect of consecutive mistakes, or uses a recurrent neural network structure to handle the temporal relationships of decision-making behaviors. This allows ability assessment to capture the complex dynamics of learners' performance, rather than simple linear scoring, resulting in assessment results that more closely reflect reality.

[0040] In one possible implementation, the personality parameters include at least one of patience, aggression, skepticism, emotionality, and questioning depth. By modulating these specific personality parameters, AI coaching can exhibit a rich variety of interaction styles and strategic intentions, thereby simulating a more realistic opponent.

[0041] In one possible implementation, the generative model is a large language model. Leveraging the powerful natural language generation capabilities of a large language model ensures that the generated response content achieves a high level of syntactic, logical, and contextual coherence, thereby eliminating reliance on pre-defined script libraries and enabling truly dynamic and fluid interaction.

[0042] In one possible implementation, the method further includes the following steps: after the interactive session ends, recording the time-series changes of the multidimensional capability vector during the session; analyzing the time-series changes to identify at least one key event point where the rate of change of the capability vector value exceeds a preset threshold and the sign of its second derivative changes; and generating an evaluation report based on the key event point, the evaluation report listing the key event point, as well as student interaction data and system response content that are temporally adjacent to the key event point. This method can provide students with in-depth diagnostic feedback including scenario replay and causal analysis, enabling them to accurately understand their performance at critical moments and its consequences, and has a strong guiding role.

[0043] To achieve the above method, another aspect of the present invention provides a training system based on ability vectors and dynamic game theory, comprising: a user interaction module 10, used to interact with trainees, collect their interaction data, and form a current interaction context based on the interaction data and historical interactions; a multi-dimensional ability vector evaluation engine 20, connected to the user interaction module 10, configured to update the multi-dimensional ability vector V representing the trainee's ability state in real time according to the interaction data; an AI personality modulation engine 30, connected to the multi-dimensional ability vector evaluation engine 20, configured to dynamically modulate at least one personality parameter of the AI ​​personality model according to the updated multi-dimensional ability vector V; and a dynamic response generator 40, connected to the AI ​​personality modulation engine 30 and the user interaction module 10, configured to generate response content in real time based on the modulated personality parameters and the current interaction context obtained from the user interaction module 10; wherein, the user interaction module 10 is further configured to present the response content to the trainee. Optionally, the personality parameters include at least one of patience, aggression, skepticism, emotionality, and questioning depth.

[0044] Optionally, the dynamic response generator 40 includes a large language model.

[0045] Optionally, the training system based on capability vectors and dynamic game theory further includes a report generation module. This module is connected to the multidimensional capability vector evaluation engine, the user interaction module, and the dynamic response generator, and is configured to: record the capability vectors output by the multidimensional capability vector evaluation engine to form a time series; analyze the time series to identify at least one key event point where the rate of change of the capability vector value exceeds a preset threshold, or where the sign of its second derivative changes; and generate an evaluation report based on the key event point, the report listing the key event point, as well as the interaction data of trainees temporally adjacent to the key event point and the system's response content.

[0046] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects:

[0047] 1. Enable personalized training: By directly linking trainees' ability assessment results (multi-dimensional ability vectors) with AI's behavioral strategies (personality parameters), the system can perceive trainees' unique ability weaknesses in real time and dynamically adjust the AI ​​opponent's strategies and difficulty for targeted reinforcement training, providing a customized training experience and improving training efficiency and effectiveness.

[0048] 2. Enhance realism and immersion: By dynamically modulating the AI ​​personality and generating unpredictable interactive content in real time, the AI ​​is no longer a fixed question setter, but an intelligent opponent that can perceive the weaknesses of trainees and adjust its own "personality" and strategies. This simulates the complexity and variability of real-world interpersonal interactions and can effectively train trainees' on-the-spot adaptability and stress management skills.

[0049] 3. Reduced content development costs: Since interactive content is generated in real time, rather than relying on a large library of manually written scripts, this invention shifts the burden of content creation from "exhaustively exploring all possibilities" to "defining capability dimensions and modulation rules," resulting in an order-of-magnitude reduction in development and update costs.

[0050] 4. Provide in-depth competency diagnosis: By recording and analyzing the time series changes of competency vectors and identifying key event points, it can provide trainees with in-depth diagnostic reports that include scenario replay and causal analysis, transforming the assessment results from simple "judgment" into providing guidance for improvement, thus having a strong guiding role.

[0051] The working process of this application will be described in detail below through examples.

[0052] This embodiment provides a training system and method based on capability vectors and dynamic game theory. In one embodiment of this application, taking a simulated training scenario of a new salesperson dealing with demanding customers as an example, the closed-loop feedback control mechanism of "trainee behavior → real-time capability assessment → AI personality modulation → dynamic content generation" proposed in this application is specifically illustrated.

[0053] Reference Figure 1 and Figure 2 This application provides an adaptive training method and system based on ability vectors and dynamic game theory, aiming to solve the technical problems of rigid content, lack of realism and relevance in interaction in existing simulation training systems. This solution constructs a closed-loop feedback control mechanism of "student behavior → real-time ability assessment → AI personality modulation → dynamic content generation" to achieve real-time adaptive adjustment of AI coaching behavior strategies, thereby creating a highly personalized and immersive dynamic game theory training environment for students.

[0054] like Figure 1 As shown in the embodiments of this application, an adaptive training system (hereinafter referred to as the "system") is provided, which may include: a user interaction module 10, a multidimensional capability vector evaluation engine 20 (MCVEE), an AI personality modulation engine 30 (APME), and a dynamic response generator 40 (DRG). These modules, engines, and generators work together to form a complete data processing and feedback closed loop.

[0055] User interaction module 10 serves as the interface for direct communication between the system and the learner. On one hand, it receives various interactive data generated by the learner during the interactive session through a user terminal (such as an application on a personal computer, tablet, or smartphone), including text input, recorded voice clips, and selections made on the interface. On the other hand, it presents the system-generated responses to the learner in text, voice, or graphical form, thereby continuing the interactive session. Furthermore, user interaction module 10 records the interaction history, which, together with the learner's current input, constitutes the current interaction context for use by other modules of the system.

[0056] The multi-dimensional capability vector assessment engine 20 is connected to the user interaction module 10, and its core function is to assess the trainee's capability status in real time. After receiving the latest interaction data from the user interaction module 10, it analyzes the data according to preset assessment rules and updates a multi-dimensional capability vector V to represent the trainee's capability status. This multi-dimensional capability vector V is a core data structure, with each dimension corresponding to a key capability that needs to be trained. For example, this multi-dimensional capability vector is a five-dimensional vector, and the five dimensions may represent "product knowledge," "communication skills," "adaptability," "emotional control," and "decision-making efficiency," respectively.

[0057] The update process for the multidimensional capability vector V is real-time and continuous. In each round of the interactive session, the multidimensional capability vector assessment engine 20 adjusts the corresponding vector of multidimensional capability vector V upwards or downwards based on the learner's latest performance. For example, if a learner accurately answers a difficult question about product specifications, the "product knowledge" vector may be increased; conversely, if the learner hesitates or gives an incorrect answer when faced with a challenge, the corresponding capability vector may be decreased. This real-time update mechanism helps maintain the timeliness and accuracy of the learner's capability profile.

[0058] The AI ​​personality modulation engine 30 is connected to the multi-dimensional ability vector evaluation engine 20. It receives the updated multi-dimensional ability vector V output by the multi-dimensional ability vector evaluation engine 20. Simultaneously, the system of this application has a pre-set personality prototype library 50, which stores the basic personality parameters P_base of various AI personalities. The function of the AI ​​personality modulation engine 30 is to dynamically modulate the basic personality parameters P_base of the selected personality prototype based on the current value of the multi-dimensional ability vector V, thereby generating a set of modulated personality parameters P_modulated.

[0059] like Figure 3 As shown, the internal operation of the AI ​​personality modulation engine 30 relies on a modulation rule base 31. This rule base 31 defines a series of "IF-THEN" mapping rules that map different states of the multidimensional ability vector V (e.g., an ability vector below a certain threshold) to specific adjustment operations on a certain personality parameter (e.g., increasing the "questioning" parameter by 50%). For example, a rule could be: "If the trainee's 'objection handling' ability vector is below 0.3, then multiply the AI ​​customer's 'price sensitivity' parameter value by 1.5." In this way, the trainee's ability shortcomings are directly translated into targeted changes in the AI ​​coach's behavioral strategies.

[0060] The dynamic response generator 40 maintains connections with both the AI ​​personality modulation engine 30 and the user interaction module 10. It receives two key inputs: the modulated personality parameter P_modulated from the AI ​​personality modulation engine 30 and the current interaction context from the user interaction module 10. Its task is to generate, based on these two inputs, a sentence or paragraph of grammatically, logically, and emotionally coherent response content in real time using a generative model.

[0061] In this process, the modulated personality parameter P_modulated acts as a "director" or "instructor." It tells the generative model what style, tone, and strategic intent to use when generating responses. For example, a combination of high "aggression" and high "questioning" parameters will guide the generative model to produce sharp, challenging questions; while a combination of high "patience" and low "aggression" parameters will guide the model to generate persuasive and encouraging remarks. The current interaction context ensures that the generated content seamlessly connects with the previous dialogue.

[0062] Once generated, the dynamic response generator 40 outputs the response content to the user interaction module 10. The user interaction module 10 then presents it to the trainee. After seeing or hearing the AI's response, the trainee will perform new interactive behaviors, which will be captured by the user interaction module 10, thus initiating the next "evaluation-modulation-generation" cycle. This continuous closed-loop process constitutes the core of the training of dynamic game between the trainee and the AI.

[0063] Accordingly, embodiments of this application also provide a training method based on capability vectors and dynamic game theory (hereinafter referred to as the "method"), such as Figure 2 As shown, this method corresponds to the operation flow of the above system, and is characterized by including the following steps:

[0064] Step S101: Receive interaction data. During the interaction session, the user terminal receives the student's interaction data and records the interaction data and historical interactions to form the current interaction context. This step is executed by the user interaction module 10, which is the starting point of the entire feedback loop and is responsible for collecting raw student behavior data.

[0065] Step S102: Update the multidimensional capability vector. Based on the interaction data, the multidimensional capability vector V, which represents the trainee's capability state, is updated in real time. This step is performed by the multidimensional capability vector assessment engine 20, which transforms the trainee's raw behavior into a structured, quantifiable capability assessment result.

[0066] Step S103: Modulate AI personality parameters. Based on the updated multidimensional ability vector V, dynamically modulate one or more personality parameters of a preset AI personality model to obtain the modulated personality parameters P_modulated. This step is executed by the AI ​​personality modulation engine 30, which directly maps the trainee's ability state to the AI ​​opponent's behavioral strategy adjustment.

[0067] Step S104: Generate a dynamic response. Based on the modulated personality parameter P_modulated and the current interaction context, response content is generated in real time using a generative model. This step is executed by the dynamic response generator 40, which eliminates the dependence on preset scripts and creates interactive content in real time according to dynamic instructions.

[0068] Step S105: Present the response content. The response content is presented to the student through the user terminal to continue the interactive session. This step is executed by the user interaction module 10, completing the closure of the feedback loop.

[0069] Step S106: Determine if the session has ended. During the interactive session, the system will determine whether the session has ended through step S106. If the session has not ended (for example, the student continues to input or the training has not reached the preset end conditions), the process will return to step S101 and start a new round of interaction. If the session has ended, the process will terminate. This cyclical process makes the behavior of the AI ​​tutor no longer static and predictable, but evolves in real time with the fluctuations in the student's performance, thus simulating the complexity and dynamism of interpersonal interaction in the real world.

[0070] This method and system achieve personalized training through the aforementioned closed-loop feedback mechanism. The system can perceive the unique skill gaps of trainees in real time and dynamically adjust the strategies and difficulty of the AI ​​opponent for targeted reinforcement training, providing a customized training experience. Compared with traditional training systems based on fixed branch scripts, this improves training efficiency and effectiveness.

[0071] Meanwhile, because AI tutors can adjust their "personality" and strategies based on the learner's weaknesses and dynamically generate unpredictable interactive content, this enhances the realism and immersion of the training. Learners feel as if they are playing against a real, intelligent opponent, effectively improving their on-the-spot adaptability and stress management skills.

[0072] Furthermore, this invention shifts the burden of content creation from script writing that "exhausts all possibilities" to framework design that "defines capability dimensions and modulation rules." Since interactive content is generated in real-time, rather than relying on a large, manually written script library, it helps reduce the development and maintenance costs of training content, making it possible to rapidly deploy and iterate new training scenarios.

[0073] In a preferred embodiment, to make the capability assessment more comprehensive and accurate, the interactive data received in step S101 and the method of the system may include at least one of the student's text input, acoustic features of voice data, and decision response time. This means that the user interaction module 10 and the multi-dimensional capability vector assessment engine 20 have the ability to process multimodal data.

[0074] Specifically, text input is the most direct form of interaction. The Multidimensional Capability Vector Assessment Engine 20 can perform natural language understanding on it, analyzing its semantic content, logical structure, and keyword usage to assess the trainee's knowledge mastery and logical thinking ability. For example, in sales training, it can assess whether the trainee used the correct product terminology, or whether their arguments were clear in a debate.

[0075] The acoustic features of the speech data provide non-linguistic information. The user interaction module 10 can collect the learner's speech and extract its acoustic features such as speech rate, pitch, volume, and pauses. The multi-dimensional ability vector assessment engine 20 can analyze these features to assess the learner's emotional state, confidence, and communication persuasiveness. For example, in a simulated management communication, a learner's excessively fast speech rate and high pitch may indicate impatience, and the system can accordingly lower their "emotional control" ability vector.

[0076] Decision response time refers to the time it takes for a learner to make a decision or provide an answer from the moment they receive a question. The user interaction module 10 can accurately record this time. The multidimensional competency vector assessment engine 20 can use this data to evaluate the learner's decision-making decisiveness and knowledge proficiency. For example, for a routine question, an excessively long decision response time may indicate that the learner is unfamiliar with the relevant knowledge or lacks confidence.

[0077] The technical advantage of the above solution lies in its ability to construct a far more comprehensive and accurate profile of learners' abilities than through the integration of multimodal interactive data. It not only focuses on "what" learners said, but also on "how" they said it and "how quickly" they reacted. This enables effective assessment and training of advanced soft skills such as communication skills, emotional management, and adaptability, making training more in-depth and comprehensive.

[0078] Furthermore, in another preferred embodiment, in step S102 and the method of the system, updating the multidimensional ability vector V can employ a nonlinear update algorithm. This algorithm is designed to more realistically simulate the complex dynamics of changes in human abilities, rather than simple linear addition and subtraction. For example, the nonlinear update algorithm can reflect the cumulative effect of consecutive errors by assigning increasing penalty weights to consecutive mistakes.

[0079] Specifically, when the multidimensional ability vector assessment engine 20 detects that a learner has made a specific type of error for the first time, it may deduct a base score, such as 0.05, from the corresponding ability vector. However, if the learner makes the same mistake again in a subsequent interaction, the algorithm will impose a larger penalty, such as deducting 0.1. The third time, it may deduct 0.2. This escalating penalty weight can effectively capture the learner's "repeatedly failing to correct" or "getting stuck in a fixed mindset" on a certain knowledge point, thus more significantly reflected in the sharp drop in the multidimensional ability vector V, thereby triggering a stronger targeted adjustment of the AI ​​personality.

[0080] Furthermore, the nonlinear update algorithm can also employ recurrent neural networks (RNNs) or their variants (such as LSTM (Long Short-Term Memory) and GRU (Gated Recurrent Unit)) to handle the temporal relationships of decision-making behaviors. The multidimensional ability vector evaluation engine 20 can take the trainee's interaction behavior sequence as input and provide it to a pre-trained RNN model. The RNN's inherent memory mechanism can capture the temporal dependencies in the behavior sequence; for example, a correct decision made after a series of incorrect attempts clearly represents a lower level of ability than a decision made correctly on the first attempt. The RNN can understand this contextual relationship and output a more accurate ability vector update value.

[0081] The advantage of employing a nonlinear update algorithm lies in its ability to move beyond simple, mechanical scoring methods in competency assessment models, enabling a deeper understanding of the underlying logic and dynamic trends of learners' performance. Whether it's the "collapse" effect of consecutive mistakes or the thought processes revealed in complex decision-making sequences, both can be quantified more accurately. This provides higher-quality input for downstream AI personality modulation, resulting in a more precise and effective response from the entire adaptive system.

[0082] In one specific implementation, to make the AI ​​tutor's behavior richer and more human-like, the personality parameters of the AI ​​personality model may include at least one of patience, aggression, skepticism, emotionality, and questioning depth. These parameters collectively define the AI ​​tutor's interaction style.

[0083] The patience parameter determines how tolerant the AI ​​is of learners' mistakes or repeated questions. High patience will cause the AI ​​to generate encouraging and guiding messages; while low patience may cause the AI ​​to show impatience, such as generating responses like "We've already discussed this issue," to train learners to perform under pressure.

[0084] The aggression parameter controls the challenging and adversarial intensity of AI speech. High aggression makes the AI ​​more inclined to directly refute the learner's views and take the initiative to challenge them; low aggression makes the AI ​​more cooperative and compliant.

[0085] The questioning parameter focuses on the depth to which the AI ​​delves into the learner's statements. A high questioning level will prompt the AI ​​to continuously ask "why," "what exactly does it mean," and "are you sure?" to test the learner's logical rigor and depth of knowledge.

[0086] The emotionality level parameter determines how much emotion AI adds to its expression. AI with a high emotionality level may exhibit emotions such as happiness, disappointment, and anger, making the interaction more like real interpersonal communication; while AI with a low emotionality level makes it behave like an objective and calm machine.

[0087] The question depth parameter controls the complexity and level of abstraction of the questions posed by the AI. When the learner's ability is low, the AI ​​can ask some basic factual questions; as the learner's ability improves, the AI ​​can gradually increase the depth of the questions, asking complex questions that require analysis, synthesis, and evaluation.

[0088] The technical advantage of the above solution lies in providing a clear control mechanism for the AI ​​personality modulation engine 30 by defining these specific and operable personality parameters. The system can combine and adjust these parameters to create a myriad of AI "personalities," thereby simulating various complex real-world interaction scenarios. Whether it's a customer requiring patient reassurance or an aggressive negotiating opponent, both can be vividly simulated, expanding the application scope and training depth of this invention.

[0089] Furthermore, to achieve high-quality dynamic content generation and reduce development costs, in a preferred embodiment, the generative model used in the dynamic response generator 40 can be a Large Language Model (LLM). Large language models, such as the Generative Pre-trained Transformer (GPT) series or similar models, possess powerful natural language understanding and generation capabilities.

[0090] Using the large language model as a generative model, the process is as follows: The dynamic response generator 40 packages the current interaction context and the modulated personality parameters P_modulated generated by the AI ​​personality modulation engine 30 into a structured prompt and sends it to the large language model. This prompt explicitly indicates the role the model should play, the current dialogue history, and most importantly—the current behavioral guidelines (i.e., personality parameters). For example, the prompt might contain instructions like: "You are now playing the role of a customer with a 'price sensitivity' of 0.9 and a 'skepticism' of 0.8. Please respond to the user's latest reply, 'Our product quality is better,' based on the following dialogue history."

[0091] Upon receiving this prompt, the large language model will utilize its generative capabilities to create a response that is both logically consistent with the context and precisely reflects the required "high price sensitivity" and "high skepticism" style. For example: "'Better quality' is too broad. Can you provide specific third-party test reports to prove it? Otherwise, why should I pay double the price for this so-called 'better'?"

[0092] First, it doesn't rely on pre-set script libraries; every sentence the AI ​​utters is generated in real time. This makes the interactive content virtually unlimited, preventing learners from passing levels by memorizing scripts and forcing them to rely on their actual abilities. Second, it reduces content development costs. Developers no longer need to write thousands of dialogue branches, but only need to design the ability dimensions and modulation rules. Finally, the fluency and creativity of the large language model make the AI ​​tutor's language expression very natural and vivid, enhancing the immersion and realism of the training.

[0093] In another preferred embodiment, the method of the present invention can also provide in-depth diagnostic analysis functions after the interactive session ends. For example... Figure 4 As shown, the method may further include the following steps: after the interactive session ends, record the time series changes of the multidimensional capability vector V during the session; analyze the time series changes to identify at least one key event point E1; and generate an evaluation report based on the key event point E1, wherein the evaluation report lists the key event point, as well as the interaction data of students who are temporally adjacent to the key event point and the system's response content.

[0094] Specifically, throughout the interaction, the system records the instantaneous values ​​of the multidimensional capability vector V at fixed time intervals (e.g., per second or per round), thus forming one or more curves showing the capability vector changing with time t, such as... Figure 4 The curve is shown in the diagram. After the session ends, the report generation module (which can be an independent functional module of the system or an extension implemented by modules such as the multidimensional capability vector assessment engine 20) will analyze these time series data.

[0095] Identifying the critical event point E1 is the core of the analysis. A critical event point E1 can be defined as the point where the rate of change of the capability vector value exceeds a preset threshold (i.e., a very steep rise or fall on the curve, representing a "sudden realization" or "collapse" of the capability), or the point where the sign of its second derivative changes (i.e., an inflection point on the curve, representing a reversal of the direction of acceleration of capability change, such as changing from accelerated decline to decelerated decline). Figure 4 As shown, on the time axis t, the system identified a key performance event E1, which corresponds to a significant change on the capability value axis c.

[0096] After identifying the critical event point E1, the report generation module traces back the interaction records around that time. It extracts the student interaction data that led to the event, as well as the system's subsequent response. Finally, the system generates a visually appealing evaluation report, which displays the critical event point E1 (e.g., at time 3 minutes and 15 seconds, the 'communication empathy' vector drops sharply from 0.7 to 0.3), student interaction data that is temporally close to the event point (e.g., a student says, "You must understand, this is company policy."), and the system's response (e.g., an AI-controlled employee replies, "Rules are rules, people are flexible; you don't care about my situation at all!"). The report can also include supplementary information such as... Figure 4 The event analysis annotation box A1 shown provides a textual description of the causal chain.

[0097] The technical advantage of this solution lies in its ability to provide trainees with in-depth diagnostic feedback. Trainees no longer receive a cold, impersonal final score, but rather a traceable and analyzable "retrospective report." By replaying key events, trainees can clearly see which specific behaviors, at what point in time, led to an improvement or decline in their abilities, and how these behaviors affected the AI ​​opponent's reactions. This causal-chain-based feedback has strong guiding value, helping trainees understand their behavioral patterns and skill gaps, thereby achieving more effective learning and improvement.

[0098] The present application will now be described in detail through a specific embodiment. This embodiment aims to simulate a scenario in which a medical student (trainee) conducts a simulated consultation with an AI-controlled patient complaining of abdominal pain, in order to train their clinical diagnostic and doctor-patient communication skills.

[0099] In this embodiment, the adaptive training system is as follows: Figure 1Configure as shown. The user interaction module 10 supports voice and text input and can record decision response time. The multi-dimensional ability vector assessment engine 20 uses a non-linear update algorithm based on LSTM (a type of recurrent neural network). The AI ​​personality model is set as an "anxious patient," and its moduloizable personality parameters include "patience," "skepticism," and "emotionality." The dynamic response generator 40 embeds a large language model. The system also includes a report generation module for generating post-meeting reports.

[0100] The multidimensional capability vector V is defined as six dimensions: V = [History taking, physical examination recommendations, differential diagnosis ability, clinical knowledge application, communication empathy, emotional stability]. All vectors are initialized to 0.6.

[0101] The basic personality parameter P_base for AI patients is set as follows: [Patience: 0.7, Questioning: 0.4, Emotionality: 0.5].

[0102] Modulation rule base 31 contains the following rules (examples): 1. IF V. Communication Empathy < 0.4 THEN P. Emotional Level = P_base. Emotional Level * 1.6 AND P. Questioning = P_base. Questioning * 1.4. 2. IFV. Differential Diagnosis Ability > 0.8 THEN P. Patience = P_base. Patience * 1.2.

[0103] The complete workflow of this embodiment is as follows:

[0104] The interaction begins when the student initiates the conversation via voice input. The user interaction module 10 receives the voice data, converts it into text, and extracts acoustic features such as speech rate and pitch. Step S101 is executed.

[0105] The trainee begins the consultation: "Hello, where does it hurt?" The multidimensional capability vector assessment engine 20 analyzes this opening and deems it conventional, so capability vector V will not be significantly adjusted for the time being. Step S102 is executed.

[0106] The AI ​​patient (generated by Dynamic Response Generator 40) replied: "Doctor, my stomach hurts terribly."

[0107] After asking several questions about the location and nature of the pain, the trainee, without fully collecting information on accompanying symptoms and past medical history, directly used a large amount of technical jargon to explain the possible causes to the AI ​​patient: "Your condition may be an acute attack of cholecystitis, or it may be referred pain from pancreatitis. We need to rule out the possibility of gastric perforation." At the same time, the trainee spoke quickly and in a high-pitched tone.

[0108] At this point, the multidimensional capability vector assessment engine 20 (step S102) captures information from multiple aspects: First, the LSTM model analyzes the consultation sequence and concludes that necessary data collection steps were skipped, thus lowering the vectors for "medical history collection" and "differential diagnosis ability." Second, by analyzing the text content, it detects that the trainee used professional terminology that was difficult for the other party to understand, which violates the basic principles of doctor-patient communication. Third, by analyzing acoustic features, it is determined that the trainee may be experiencing tension or an eagerness to perform. Based on the above information, the engine significantly lowers the "communication empathy" vector from 0.6 to 0.3 and slightly lowers the "emotional stability" vector.

[0109] The AI ​​personality modulation engine 30 (step S103) receives the updated ability vector V and detects that the "communication empathy" vector (0.3) has fallen below the threshold of 0.4 set in rule 1. Therefore, it triggers the rule to modulate the AI ​​patient's basic personality parameter P_base, generating the modulated personality parameter P_modulated: P_modulated. emotionality is increased to 0.5 * 1.6 = 0.8, and P_modulated. skepticism is increased to 0.4 * 1.4 = 0.56.

[0110] The dynamic response generator 40 (step S104) receives the set of parameters P_modulated, which are highly "emotional" and "skeptical," along with the interaction context containing the previously mentioned technical terms. Based on this, the large language model generates a challenging response with anxiety and confusion: "Doctor, I don't understand any of what you're saying... Is it very serious? Are you sure you know what my problem is?"

[0111] User interaction module 10 presents this audio and text to the learner (step S105). The learner now faces a more difficult situation caused by their own inappropriate communication: a distrustful and anxious patient. Subsequent interactions will proceed from this new and more challenging baseline. If the learner can adjust their communication style, reassure the patient, and explain in simple language, their "communication empathy" vector may rebound, thereby reducing the AI ​​patient's "emotional level" and bringing the interaction back on track.

[0112] After the session ended, the report generation module started. It analyzed the time series curve of the "Communication Empathy" vector and found a "cliff-like drop" at the aforementioned time point, marking it as the critical event point E1. Figure 4As shown in the report, the event analysis annotation box A1 will be generated as follows: "At [time point], due to the use of a large number of professional terms in the student's reply 'Your condition may be cholecystitis…' and the excessive speaking speed, the student's 'communication empathy' vector dropped sharply. This behavior triggered a shift in the AI ​​patient personality towards 'anxiety' and 'questioning,' and generated a challenging question: 'Doctor, I don't understand, is it very serious?', increasing communication barriers. It is recommended that in subsequent communications, open-ended questions be used to fully gather information first, and the condition should be explained to the patient in plain and easy-to-understand language."

[0113] It is evident that the synergistic effect of various technical features enhances the overall performance of the present invention, achieves the training objective of in-depth diagnostic capabilities, and helps to solve the problems of rigid interaction, superficial evaluation, and weak feedback guidance in the prior art.

[0114] The training method and system based on capability vectors and dynamic game theory provided by this invention have a wide range of applications. Besides the aforementioned sales training, management communication, and medical consultation scenarios, it can also be applied to various professional fields requiring high-intensity interpersonal interaction and on-the-spot decision-making abilities. For example, in emergency response training, it can simulate a fire commander (trainee) collaborating with multiple AI-controlled firefighters with different personalities (such as an experienced veteran and a nervous rookie), training the trainee's command, coordination, and clear communication abilities under chaotic information and high-pressure environments.

[0115] In customer service training, various types of "difficult" customers can be simulated, such as angry complainers, hesitant customers, and customers who constantly challenge the rules. The system dynamically adjusts the AI ​​customer's personality parameters, such as "aggression," "patience," and "stubbornness," to provide a safe yet challenging environment for new customer service personnel to hone their emotional control, problem-solving, and reassurance skills.

[0116] In negotiation training, this invention can simulate a negotiation opponent. The AI's behavioral strategy will be adjusted in real time according to the trainee's (negotiation expert's) questioning skills, logical pressure, and psychological insight. For example, when the trainee's "logical pressure" ability vector increases, the AI's "defensiveness" and "lying tendency" parameters may also increase, thereby creating a more realistic high-IQ game environment.

[0117] This application also provides an electronic device, which may be a server, a personal computer (PC), a tablet computer, or a smartphone. The electronic device includes a processor and a memory, on which a computer program is stored. When the computer program is executed by the processor, it can implement any of the training methods based on ability vectors and dynamic game theory described above. For example, the processor can execute instructions to run the logic of the user interaction module 10, the multi-dimensional ability vector assessment engine 20, the AI ​​personality modulation engine 30, and the dynamic response generator 40, while the memory is used to store the multi-dimensional ability vector V, the personality prototype library 50, interaction data, and the generation model itself.

[0118] This application also provides a computer-readable storage medium storing a computer program thereon. The computer-readable storage medium can be volatile or non-volatile, such as a read-only memory (ROM), random access memory (RAM), solid-state drive (SSD), or optical disc. When the computer program is executed by a processor, it can implement any of the training methods based on capability vectors and dynamic game theory described above. For example, the program code embedded on the storage medium, when loaded into the processor of an electronic device and run, will cause the device to execute the complete interactive loop of steps S101 to S105, and optionally implement functions such as multimodal input processing, nonlinear capability updates, and post-meeting diagnostic report generation.

[0119] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A training method based on capability vectors and dynamic game theory, characterized in that, Includes the following steps: a) During the interactive session, the user terminal receives the student's interactive data and records the interactive data and historical interactions to form the current interactive context; b) Based on the interaction data, update the multidimensional ability vector used to characterize the trainee's ability status in real time; c) Based on the updated multidimensional ability vector, dynamically modulate at least one personality parameter of the preset AI personality model; d) Based on the modulated personality parameters and the current interaction context, generate response content in real time using a generative model; as well as e) Present the response content to the student through the user terminal to continue the interactive session.

2. The method according to claim 1, characterized in that, It also includes the following steps: After the interactive session ends, record the time-series changes of the multidimensional capability vector during the session; Analyze the time series changes to identify at least one key event point: either the rate of change of the capability vector value exceeds a preset threshold, or the sign of its second derivative changes. as well as Based on the key event points, an evaluation report is generated, which lists the key event points, as well as the interaction data of students who are temporally close to the key event points and the system's response content.

3. The method according to claim 1 or 2, characterized in that, The interactive data includes at least one of the following: student text input, acoustic features of voice data, and decision response time; and / or, The personality parameters include at least one of patience, aggression, questioning, emotionality, and depth of questioning.

4. The method according to claim 1 or 2, characterized in that, In step b), the multidimensional capability vector is updated using a nonlinear update algorithm. The nonlinear update algorithm reflects the cumulative effect of continuous errors by assigning increasing penalty weights to consecutive errors, or by using a recurrent neural network structure to handle the temporal relationship of decision-making behavior.

5. The method according to claim 1 or 2, characterized in that, The generative model is a large language model.

6. A training system based on capability vectors and dynamic game theory, characterized in that, include: The user interaction module is used to interact with students, collect their interaction data, and form the current interaction context based on the interaction data and historical interactions. A multidimensional ability vector assessment engine, connected to the user interaction module, is configured to update the multidimensional ability vector used to characterize the trainee's ability status in real time based on the interaction data. The AI ​​personality modulation engine, connected to the multidimensional ability vector evaluation engine, is configured to dynamically modulate at least one personality parameter of the AI ​​personality model based on the updated multidimensional ability vector. A dynamic response generator, connected to the AI ​​personality modulation engine and the user interaction module, is configured to generate response content in real time based on the modulated personality parameters and the current interaction context obtained from the user interaction module. The user interaction module is also configured to present the response content to the student.

7. The system according to claim 6, characterized in that, It also includes a report generation module, which is connected to the multidimensional capability vector assessment engine, the user interaction module, and the dynamic response generator, and is configured to: Record the capability vectors output by the multidimensional capability vector evaluation engine to form a time series; The time series is analyzed to identify at least one key event point in which the rate of change of the capability vector value exceeds a preset threshold and the sign of its second derivative changes; as well as Based on the key event points, an evaluation report is generated, which lists the key event points, as well as the interaction data of students who are temporally close to the key event points and the system's response content.

8. The system according to claim 6 or 7, characterized in that, The personality parameters include at least one of patience, aggression, questioning, emotionality, and depth of questioning; and / or, The dynamic response generator contains a large language model.

9. An electronic device, comprising: processor; as well as A memory, on which computer programs are stored; When the computer program is executed by the processor, it implements the method as described in any one of claims 1-5.

10. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5.