Scenarized critical thinking evaluation system and method, electronic device and storage medium
By using a scenario-based critical thinking assessment system, which leverages coordination, questioning, and evidence curation to dynamically control the dialogue process, the system addresses the issues of insufficient interpretability and behavioral evidence capture in AI scoring, thus achieving efficient and transparent critical thinking assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for AI scoring lack interpretability, fail to capture behavioral evidence, and lack quality control in the dialogue process, resulting in low assessment reliability and user fatigue.
A scenario-based critical thinking assessment system is adopted. This system coordinates agents to generate micro-context texts, asks questions to interact with agents, uses evidence curation agents to monitor evidence saturation, and uses reasoning and arbitration agents to determine the assessment thought chain. The system dynamically controls the dialogue length and generates transparent and traceable assessment results.
Significantly improves the reliability and efficiency of assessments, meets the requirements of educational assessments for transparency and auditability, and provides highly interpretable assessment results.
Smart Images

Figure CN121808022A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a scenario-based critical thinking assessment system, method, electronic device, and storage medium. Background Technology
[0002] Critical thinking (CT) is one of the core training objectives of higher education. Traditional assessment methods mainly rely on standardized psychometric scales, such as the California Critical Thinking Disposition Inventory (CCTDI). These scales assess thinking tendencies by having test takers answer abstract Likert scale questions such as "I always try to seek the truth."
[0003] With the development of Large Language Models (LLMs), the education field has begun to explore the use of AI (Artificial Intelligence) for automatic grading. Current attempts mainly fall into two categories: one is using LLMs to automatically grade users' static texts such as academic papers; the other is using LLMs to generate question-and-answer questions. Multi-agent systems, due to their ability to simulate complex role interactions, are also increasingly being applied to simulated teaching scenarios.
[0004] However, the above-mentioned method of using LLM for automatic scoring still has the following drawbacks.
[0005] (1) The direct scoring method based on LLM is like a black box, which only outputs scores and cannot provide evidence-based scoring basis. This does not meet the requirements of transparency and auditability for educational assessment. Therefore, the direct scoring method based on LLM lacks interpretability.
[0006] (2) In human-computer dialogue evaluation, existing methods usually adopt fixed dialogue rounds or simple question-and-answer logic, often rushing to score before the user has fully demonstrated their thought process, or making ineffective follow-up questions after obtaining sufficient information, resulting in low evaluation reliability or user fatigue.
[0007] Therefore, how to provide a critical thinking assessment method that is highly interpretable, can capture behavioral evidence, and has controllable dialogue quality is a technical problem that urgently needs to be solved. Summary of the Invention
[0008] This invention provides a scenario-based critical thinking assessment system, method, electronic device, and storage medium to address the shortcomings of existing technologies, such as the lack of interpretability in AI scoring, the inability to capture behavioral evidence, and the lack of quality control in the dialogue process.
[0009] This invention provides a scenario-based critical thinking assessment system, comprising the following intelligent agents.
[0010] A coordinating agent is used to traverse all dimensions corresponding to critical thinking and generate micro-contextual text corresponding to the current dimension. The questioning agent is used to interact with the user to be evaluated based on the micro-context text of the current dimension, and obtain the dialogue text of the user to be evaluated in all rounds. An evidence curation agent is used to monitor evidence saturation based on the current dialogue text of the user to be evaluated in the current round, and obtain indication information corresponding to the target agent; the target agent includes at least one of a coordinating agent, a questioning agent, and a reasoning arbitration agent; The reasoning arbitration agent is used to determine the evaluation thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be evaluated in all rounds, upon receiving the instruction information sent by the evidence curation agent; the evaluation thought chain includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimension score.
[0011] According to the scenario-based critical thinking assessment system provided by the present invention, the evidence curation agent is specifically used for: Obtain the dimension operation definition from the preset definition information of the current dimension; the dimension operation definition is used to characterize the behavioral meaning of the current dimension; According to the priority of the matching steps from high to low, the current dialogue text is matched with the dimension operation definition step by step to obtain the evidence saturation state corresponding to the current dialogue text. Based on the evidence saturation state, determine the indication information corresponding to the target intelligent agent; The matching steps, in descending order of priority, are keyword matching, semantic equivalence matching, and dialogue turn matching.
[0012] According to the scenario-based critical thinking assessment system provided by the present invention, the evidence curation agent is specifically used for: Perform keyword matching between the current dialogue text and the dimension operation definition to determine the keyword matching result; If the keyword matching result is successful, the evidence saturation state corresponding to the current dialogue text is determined to be saturated. If the keyword matching result fails, determine the semantic similarity between the current dialogue text and the dimension operation definition; If the semantic similarity is greater than or equal to a first preset threshold, the evidence saturation state corresponding to the current dialogue text is determined to be saturated. If the semantic similarity is less than the first preset threshold and greater than or equal to the second preset threshold, the evidence saturation state corresponding to the current dialogue text is determined to be partially saturated; the first preset threshold is greater than the second preset threshold. If the semantic similarity is less than the second preset threshold, the evidence saturation state corresponding to the current dialogue text is determined to be unsaturated. If the evidence saturation state corresponding to the current dialogue text is unsaturated or partially saturated, an intervention attribute is added to the evidence saturation state based on the comparison result between the dialogue round corresponding to the current dialogue text and the round intervention threshold.
[0013] According to the scenario-based critical thinking assessment system provided by the present invention, the instruction information includes at least one of dimension switching information, scoring trigger information, and continued detection information; The evidence curation agent is specifically used for: When the evidence saturation state is saturated, dimension switching information corresponding to the coordinating agent and scoring trigger information corresponding to the reasoning arbitration agent are generated; the dimension switching information is used to instruct the coordinating agent to generate micro-context text for the next dimension; the scoring trigger information is used to instruct the reasoning arbitration agent to determine the evaluation thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be evaluated in all rounds. When the evidence saturation state is unsaturated or partially saturated, the questioning agent generates continued detection information based on the intervention attribute; the continued detection information is used to instruct the questioning agent to generate the evaluation question text for the next round.
[0014] According to the scenario-based critical thinking assessment system provided by the present invention, the reasoning arbitration agent is specifically used for: The dialogue text of each round is mapped to the preset behavior indicators in the preset definition information to obtain the mapping relationship between the matching behavior indicators and the dialogue text, as well as the number of evidence and the confidence level of the evidence corresponding to the matching behavior indicators. Based on the mapping relationship between the matching behavior indicators and the dialogue text, the current behavior pattern of the user to be evaluated in the current dimension is determined; Based on the number of pieces of evidence and the confidence level of the evidence corresponding to the matching behavior indicators, the level score corresponding to the current behavior pattern is determined; Based on the grade score, the number of pieces of evidence and the confidence level of the evidence corresponding to the matching behavior indicator, the dimension score corresponding to the current dimension is determined; Based on the mapping relationship between the matching behavior indicators and the dialogue text, the current behavior pattern, the dimension score, the number of evidences, and the confidence level of the evidence, the evaluation thought chain corresponding to the current dimension is determined.
[0015] According to the scenario-based critical thinking assessment system provided by the present invention, the reasoning arbitration agent is specifically used for: Based on the number of evidence corresponding to the matching behavior indicators and the number of preset behavior indicators, the evidence coverage of the current dimension is determined; Based on the evidence coverage and the evidence confidence of the current dimension, the level score corresponding to the current behavior pattern is determined.
[0016] According to the scenario-based critical thinking assessment system provided by the present invention, the questioning agent is specifically used for: Upon receiving a continue detection message or a forced selection message from the evidence curation agent, the attribute indicators corresponding to the current dialogue text are determined; based on the attribute indicators, the questioning strategy for the next round is determined; the attribute indicators include at least one of content completeness, depth of thought, and expression certainty.
[0017] The present invention also provides a scenario-based critical thinking assessment method, which is applied to the scenario-based critical thinking assessment system described in any of the above claims, and the method includes the following steps.
[0018] By coordinating the intelligent agent to traverse all dimensions corresponding to critical thinking, micro-contextual text corresponding to the current dimension is generated; By having the intelligent agent interact with the user to be evaluated based on the micro-context text of the current dimension, the dialogue text of the user to be evaluated in all rounds is obtained. The evidence curation agent monitors the evidence saturation based on the dialogue text of the user to be evaluated in the current round, and obtains the indication information corresponding to the target agent; the target agent includes at least one of the coordinating agent, the questioning agent, and the reasoning arbitration agent; Upon receiving the instruction information sent by the evidence curation agent, the reasoning arbitration agent determines the evaluation thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be evaluated in all rounds. The evaluation thought chain includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimension score.
[0019] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the scenario-based critical thinking assessment method described above.
[0020] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the scenario-based critical thinking assessment method as described above.
[0021] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the scenario-based critical thinking assessment method described above.
[0022] This invention provides a scenario-based critical thinking assessment system, method, electronic device, and storage medium. It involves a coordinating agent that transforms abstract mental constructs of the current dimension into scenario-based micro-context text. The questioning agent interacts with the user to be assessed based on the micro-context text, obtaining the dialogue text of all rounds of the user's responses to the current dimension. An evidence curating agent monitors the evidence saturation of the current dialogue text in the current round, obtaining the corresponding instruction information from the target agent. Upon receiving the instruction information from the evidence curating agent, the reasoning arbitration agent, based on the preset definition information of the current dimension and the dialogue text of all rounds of the user to be assessed, determines the mapping relationship between at least all rounds of dialogue text in the current dimension and the preset definition information, as well as the assessment thought chain for dimension scoring. In this invention, the evidence curation agent monitors the evidence saturation in real time to verify the sufficiency of evidence, and dynamically controls the dialogue length between the questioning agent and the user to be evaluated, avoiding insufficient evidence or invalid follow-up questions caused by fixed dialogue rounds, thus significantly improving the reliability and efficiency of the evaluation. At the same time, the evaluation thought chain generated by the reasoning arbitration agent includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimensional scores. The scoring process is completely transparent and traceable, improving the interpretability of the dimensional scores and meeting the requirements of educational evaluation for transparency and auditability. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0024] Figure 1 This is a schematic diagram of the structure of the scenario-based critical thinking assessment system provided in an embodiment of the present invention.
[0025] Figure 2 This is a flowchart illustrating the scenario-based critical thinking assessment method provided in this embodiment of the invention.
[0026] Figure 3This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0028] To address the problems of existing AI (Artificial Intelligence) scoring methods, such as lack of interpretability, inability to capture behavioral evidence, and lack of quality control in the dialogue process, this invention provides a scenario-based critical thinking assessment system. Figure 1 This is a schematic diagram of the structure of the scenario-based critical thinking assessment system provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the system includes: a coordinating agent, a questioning agent, an evidence curating agent, and a reasoning arbitration agent.
[0029] A coordinating agent is used to traverse all dimensions corresponding to critical thinking and generate micro-contextual text corresponding to the current dimension.
[0030] The questioning agent is used to interact with the user to be evaluated based on the micro-context text of the current dimension, and obtain the dialogue text of the user to be evaluated in all rounds.
[0031] An evidence curation agent is used to monitor evidence saturation based on the current dialogue text of the user to be evaluated in the current round, and obtain indication information corresponding to the target agent; the target agent includes at least one of a coordinating agent, a questioning agent, and a reasoning arbitration agent.
[0032] The reasoning arbitration agent is used to determine the evaluation thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be evaluated in all rounds, upon receiving the instruction information sent by the evidence curation agent; the evaluation thought chain includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimension score.
[0033] Specifically, in this system, the questioning agent communicates with the coordinating agent and the evidence curating agent, which in turn communicates with the coordinating agent and the reasoning arbitration agent. The four agents work collaboratively.
[0034] The coordinating agent is located at the top level of the system's architecture. It is used to traverse the seven dimensions corresponding to critical thinking, generate the micro-context text corresponding to the current dimension, and send the micro-context text of the current dimension to the questioning agent.
[0035] It should be noted that the critical thinking assessment includes seven dimensions: Dimension 1, Truth-seeking; Dimension 2, Open-mindedness; Dimension 3, Analyticity; Dimension 4, Systematicity; Dimension 5, Self-confidence (CT); Dimension 6, Inquisitiveness; and Dimension 7, Maturity.
[0036] It should be noted that the pre-built CCTDI dimension library stores the preset definition information and prompt word templates for each dimension. The prompt word templates for each dimension include elements such as situation type constraints, decision point requirements, and role settings. The preset definition information for each dimension includes the dimension operation definition and preset behavioral indicators. The dimension operation definition is used to characterize the behavioral meaning of the corresponding dimension. The dimension operation definitions for each dimension include: Dimension 1, Seeking Truth: Delaying judgment and actively seeking evidence when there is insufficient evidence; Dimension 2, Open-mindedness: Remaining open to different viewpoints and willing to consider alternative solutions; Dimension 3, Analytical: Tendency to use reasoning and evidence to solve problems; Dimension 4, Systematic: Handling problems in an organized and planned manner; Dimension 5, Confidence: Trust in one's reasoning ability; Dimension 6, Curiosity: A desire to learn new things; Dimension 7, Cognitive Maturity: Understanding that there may be multiple reasonable solutions to a problem. The preset behavioral indicators for each dimension are shown in Table 1.
[0037] Table 1 In existing technologies, traditional self-report scales such as the CCTDI rely on the test-taker's self-report. Test-takers often choose high-scoring options based on social desirability, resulting in measurement results reflecting "the test-taker's perceived abilities" rather than "actual problem-solving performance." Therefore, in this embodiment of the invention, a coordinating agent is constructed to transform abstract psychological constructs of each dimension into scenario-based micro-context texts guided by specific behaviors. This forces the system to operate according to the predetermined psychological measurement logic of each dimension, capturing the actual behavior of the user being evaluated when solving problems. This effectively overcomes the social desirability error of traditional self-report scales and solves the problem of the separation of knowledge and action. The coordinating agent generates the micro-context text for the current dimension through the following steps.
[0038] (1) Based on the dimension identifier (Identity Document, ID) corresponding to the current dimension, obtain the preset definition information and prompt word template corresponding to the current dimension from the pre-built CCTDI dimension library.
[0039] (2) Embed the dimension operation definition in the preset definition information of the current dimension into the prompt word template to obtain the prompt word corresponding to the current dimension.
[0040] (3) Input the prompt words corresponding to the current dimension into the large language model to obtain the scene description text output by the large language model. Perform context validity verification on the scene description text. If the scene description text meets the context conditions, determine the scene description text as the micro-context text corresponding to the current dimension. The context conditions include that the context corresponding to the scene description text contains a clear decision point, the scene description text is related to the dimension operation definition, and the scene description text has sufficient openness to allow the user to be evaluated to demonstrate their thought process.
[0041] For example, taking the truth-seeking dimension as an example, the coordinating agent is input with the dimension identifier dim_id=1. The coordinating agent loads the dimension operation definition corresponding to the truth-seeking dimension from the CCTDI dimension library, which is defined as "operational_definition="Truth seekers tend to postpone judgments and actively seek evidence when there is insufficient evidence." The prompt word constructed based on the dimension operation definition is "Generate a micro-context for the truth-seeking dimension. The context should include a decision point, allowing the user to be evaluated to demonstrate whether they would postpone judgments and actively seek evidence when there is insufficient evidence. The context type is a daily decision-making scenario." Inputting this prompt word into a large language model, the resulting micro-context text, output by the large language model and validated for context effectiveness, is: "Suppose you are reading a popular online article that claims a new data analysis technique 'X-Algo' is 50% more efficient than traditional methods. Your team members are excited and want to adopt it in the project immediately. What would you do in this situation?" For example, taking the open-mindedness dimension as an example, the coordinating agent is input with the dimension identifier dim_id=2. The coordinating agent loads the dimension operation definition corresponding to the open-mindedness dimension from the CCTDI dimension library, which is defined as "operational_definition="open to different viewpoints and willing to consider alternatives". The prompt word constructed based on the dimension operation definition is "The scenario type is a daily decision-making scenario involving differences of opinion or conflict of choice; it contains a clear decision point, requiring the user to be evaluated to demonstrate whether they are willing to openly consider different viewpoints; the roles are set to include the user to be evaluated and at least one other role, such as a friend or colleague, who express different viewpoints." Inputting this prompt word into the large language model yields the micro-context text output by the large language model, which has been validated for contextual effectiveness: "Suppose your good friend proposes a travel plan you've never considered—to a destination you've always thought 'uninteresting'. Your friend enthusiastically shares many unique aspects of that place. How would you respond?" The questioning agent resides in the middle layer of the system, interacting with the user through dialogue based on micro-contextual text, and sending the dialogue text from each round to the evidence curation agent in real time. Unlike traditional question-and-answer machines, the questioning agent in this embodiment employs a Socratic questioning strategy to interact with the user being evaluated. This agent dynamically adjusts its questioning strategy for the next round based on the current dialogue text responded to by the user. This strategy includes conventional follow-up questioning, in-depth follow-up questioning, and strategies that encourage expression. By dynamically adjusting the questioning strategy, the agent guides the user to externalize their implicit thought processes, thereby obtaining high-quality behavioral evidence (i.e., dialogue text).
[0042] The evidence curation agent, acting as an observer in the system, monitors the evidence saturation in real time based on the current dialogue text in the current round, controlling the next steps of the coordinating agent, the questioning agent, and the reasoning arbitration agent. For example, when the evidence saturation state corresponding to the current dialogue text is unsaturated, the agent is controlled to continue asking questions; when the evidence saturation state corresponding to the current dialogue text is saturated, the agent is controlled to switch to the next dimension, generating the micro-contextual text for the next dimension. Simultaneously, the reasoning arbitration agent is triggered to score the current dimension based on the preset definition information of the current dimension and the dialogue texts of all rounds, obtaining the evaluation thought chain corresponding to the current dimension. This evaluation thought chain includes at least the mapping relationship between the dialogue texts of all rounds and the preset definition information, as well as the dimension score.
[0043] Optionally, all four agents are constructed based on a Large Language Model (LLM) and structured prompts. For example, all four agents are based on the multimodal language model GPT-4o and constructed through prompt engineering; this embodiment of the invention does not limit this approach.
[0044] This invention provides a scenario-based critical thinking assessment system. A coordinating agent transforms abstract mental constructs of the current dimension into scenario-based micro-context text. A questioning agent interacts with the user to be assessed based on the micro-context text, obtaining the dialogue text of all rounds of the user's responses to the current dimension. An evidence curating agent monitors the evidence saturation of the current dialogue text in the current round, obtaining the corresponding instruction information from the target agent. Upon receiving the instruction information from the evidence curating agent, the reasoning arbitration agent, based on the preset definition information of the current dimension and the dialogue text of all rounds of the user to be assessed, determines the mapping relationship between at least all rounds of dialogue text and the preset definition information in the current dimension, as well as the assessment thought chain for dimension scoring. In this embodiment of the invention, the evidence curation agent monitors the evidence saturation in real time to verify the sufficiency of evidence, and dynamically controls the dialogue length between the questioning agent and the user to be evaluated, avoiding insufficient evidence or invalid follow-up questions caused by fixed dialogue rounds, thus significantly improving the reliability and efficiency of the evaluation. At the same time, the evaluation thought chain generated by the reasoning arbitration agent includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimension score. The scoring process is completely transparent and traceable, improving the interpretability of the dimension score and meeting the requirements of educational evaluation for transparency and auditability.
[0045] After receiving the micro-context text corresponding to the current dimension sent by the coordinating agent, the questioning agent displays the micro-context text to the user to be evaluated on the interactive interface. After receiving the current dialogue text of the reply to the micro-context text input by the user to be evaluated, the current dialogue text is sent to the evidence curation agent in real time.
[0046] In one embodiment, the evidence curation agent is specifically used for: Obtain the dimension operation definition from the preset definition information of the current dimension; the dimension operation definition is used to characterize the behavioral meaning of the current dimension; According to the priority of the matching steps from high to low, the current dialogue text is matched with the dimension operation definition step by step to obtain the evidence saturation state corresponding to the current dialogue text. Based on the evidence saturation state, determine the indication information corresponding to the target intelligent agent; The matching steps, in descending order of priority, are keyword matching, semantic equivalence matching, and dialogue turn matching.
[0047] Specifically, after receiving the current dialogue text for the current round, the evidence curation agent obtains the dimension operation definition corresponding to the current dimension. Following the matching steps, it performs a three-tiered evidence saturation judgment based on the current dialogue text and the dimension operation definition: first, keyword matching; then, semantic equivalence matching; and finally, dialogue round matching, to determine the evidence saturation state corresponding to the current dialogue text. This evidence saturation state includes a saturated state, a partially saturated state, or an unsaturated state. Subsequently, based on this evidence saturation state, it determines the instruction information corresponding to the target agent. This instruction information is used to guide the next actions of the coordinating agent, the questioning agent, and the reasoning arbitration agent.
[0048] It should be noted that evidence saturation refers to whether the behavioral evidence (i.e., dialogue text) collected by the system during the dialogue process is sufficient to reliably score the current dimension. In this embodiment of the invention, the concept of theoretical saturation in qualitative research is adopted, and the system is considered to be in a saturated state when the current dialogue text no longer generates new categories of evidence.
[0049] In one embodiment, the evidence curation agent is specifically used for: Perform keyword matching between the current dialogue text and the dimension operation definition to determine the keyword matching result; If the keyword matching result is successful, the evidence saturation state corresponding to the current dialogue text is determined to be saturated. If the keyword matching result fails, determine the semantic similarity between the current dialogue text and the dimension operation definition; If the semantic similarity is greater than or equal to a first preset threshold, the evidence saturation state corresponding to the current dialogue text is determined to be saturated. If the semantic similarity is less than the first preset threshold and greater than or equal to the second preset threshold, the evidence saturation state corresponding to the current dialogue text is determined to be partially saturated; the first preset threshold is greater than the second preset threshold. If the semantic similarity is less than the second preset threshold, the evidence saturation state corresponding to the current dialogue text is determined to be unsaturated. If the evidence saturation state corresponding to the current dialogue text is unsaturated or partially saturated, an intervention attribute is added to the evidence saturation state based on the comparison result between the dialogue round corresponding to the current dialogue text and the round intervention threshold.
[0050] Specifically, after obtaining the current dialogue text, a three-layer evidence saturation judgment mechanism is adopted, which includes the following steps.
[0051] In the first layer of this invention, to quickly identify high-quality behavioral evidence, improve evaluation efficiency, and ensure scoring reliability, the current dialogue text and the dimensional operation definition are input into a large language model for keyword matching. The model detects whether the current dialogue text directly contains the core behavioral description keywords of the dimensional operation definition. If the current dialogue text contains these keywords, the keyword matching result is a successful match between the current dialogue text and the dimensional operation definition. The large language model outputs an evidence saturation state of SUFFICIENT, and the current dialogue text corresponding to this saturation state is considered high-confidence evidence. If the current dialogue text does not contain the core behavioral description keywords of the dimensional operation definition, the keyword matching result is a failed match between the current dialogue text and the dimensional operation definition, and the process proceeds to the second layer.
[0052] In the second layer of this invention, to expand the scope of evidence recognition, improve the inclusiveness of evaluation, and reduce misjudgments due to different wording that may result in the omission of valid evidence, a large language model is used to calculate the semantic similarity between the current dialogue text and the dimension operation definition. Then, this semantic similarity is compared with a first preset threshold to determine whether the current dialogue text is semantically equivalent to the dimension operation definition. If the semantic similarity is greater than or equal to the first preset threshold, it indicates that the current dialogue text is semantically equivalent to the dimension operation definition, and the evidence saturation state corresponding to the current dialogue text is saturated. If the semantic similarity is less than the first preset threshold but greater than or equal to the second preset threshold, it indicates that the current dialogue text is semantically partially equivalent to the dimension operation definition, and the evidence saturation state corresponding to the current dialogue text is partially saturated (PENDING). If the semantic similarity is less than the second preset threshold, it indicates that the current dialogue text is semantically inequivalent to the dimension operation definition, and the evidence saturation state corresponding to the current dialogue text is unsaturated (INSUFFICIENT), and the process proceeds to the third layer.
[0053] In the third layer of this invention, to prevent inefficient evaluation due to endless follow-up questions and ensure user experience, when the evidence saturation state determined in the first two layers is unsaturated or partially saturated, the dialogue round of the current dialogue text is compared with the round intervention threshold to determine whether the dialogue round satisfies the evidence saturation theory. If the comparison result is that the dialogue round is less than the round intervention threshold, it indicates that the dialogue round does not satisfy the evidence saturation theory, and the intervention attribute is determined to be non-intervention, and this intervention attribute is added to the evidence saturation state. If the comparison result is that the dialogue round is greater than or equal to the round intervention threshold, it indicates that the dialogue round satisfies the evidence saturation theory, and the intervention attribute is determined to be mandatory intervention, and this intervention attribute is added to the evidence saturation state. This mandatory intervention is used to adjust the questioning method in the next round.
[0054] It should be noted that the evidence saturation theory states that 3 to 5 rounds of dialogue are generally sufficient to reach theoretical saturation, and experimental evidence shows that 87% of valid behavioral evidence is captured in the first 3 rounds of dialogue. Therefore, the intervention threshold for this round is determined based on the evidence saturation theory, and this intervention threshold can be 3 or 4, etc., and the embodiments of the present invention do not limit it in this way.
[0055] Furthermore, if the number of dialogue rounds exceeds the round intervention threshold, the number of dialogue rounds is compared with the forced round threshold to determine whether the dialogue rounds satisfy the cognitive load theory. If the number of dialogue rounds is greater than or equal to the forced round threshold, it indicates that the dialogue rounds satisfy any load theory; that is, subsequent dialogues will increase the cognitive burden on the user being evaluated, and the evaluation efficiency will be low. Multiple rounds of dialogue may still fail to obtain valid behavioral evidence. In this case, the dialogue with the user being evaluated can be forcibly terminated, and the current dimension can be marked as insufficient evidence. The forced round threshold is determined based on the cognitive load theory and can be 5 or 6, etc., ensuring that the forced round threshold is greater than the round intervention threshold. This embodiment of the invention does not limit this.
[0056] It should be noted that the first preset threshold and the second preset threshold can be determined based on historical experience. The first preset threshold serves as a threshold between the saturated state and the partially saturated state, and the second preset threshold serves as a threshold between the partially saturated state and the unsaturated state. For example, the first preset threshold can be 0.75, 0.8, or 0.85, and the second preset threshold can be 0.5, 0.55, or 0.6, etc. The embodiments of the present invention do not limit this.
[0057] In one embodiment, the indication information includes at least one of dimension switching information, scoring trigger information, and continued detection information; The evidence curation agent is specifically used for: When the evidence saturation state is saturated, dimension switching information corresponding to the coordinating agent and scoring trigger information corresponding to the reasoning arbitration agent are generated; the dimension switching information is used to instruct the coordinating agent to generate micro-context text for the next dimension; the scoring trigger information is used to instruct the reasoning arbitration agent to determine the evaluation thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be evaluated in all rounds. When the evidence saturation state is unsaturated or partially saturated, the questioning agent generates continued detection information based on the intervention attribute; the continued detection information is used to instruct the questioning agent to generate the evaluation question text for the next round.
[0058] Specifically, if the evidence saturation state is saturated, it indicates that sufficient behavioral evidence corresponding to the current dimension has been obtained. At this time, dimension switching information can be generated to control the coordinating agent to switch to the evaluation of the next dimension. That is, the micro-context text corresponding to the next dimension is generated and sent to the questioning agent. At the same time, scoring trigger information is generated to trigger the reasoning arbitration agent to determine the evaluation thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be evaluated in all rounds.
[0059] If the evidence saturation state is unsaturated or partially saturated, it indicates that the acquired dialogue text is insufficient. In this case, the intervention attribute in the evidence saturation state can be further obtained. If the intervention attribute is uninterrupted, the generated continuation detection information will not intervene in the questioning method of the next round. If the intervention attribute is forced intervention, the generated continuation detection information includes forced selection information. Through this forced selection information, the questioning agent is instructed to provide binary options for the user to be evaluated to choose from in the generated evaluation question text of the next round, so that the user to be evaluated can express their opinion clearly.
[0060] For example, taking the current dimension as the dimension of seeking truth, with a first preset threshold of 0.7, an iterative preset threshold of 0.4, and a round intervention threshold of 3, the evidence curation agent makes a judgment on the evidence saturation state by including the following steps.
[0061] Round 1: The dialogue text of the user's response to be evaluated is "This technology sounds interesting".
[0062] First-level check: 0 keyword matches, no keyword matching result found.
[0063] Second layer check: semantic similarity 0.15, semantic inequivalence, indicating an unsaturated state.
[0064] Third-level check: Current round 1, intervention attribute is no intervention.
[0065] Evidence saturation status determination: unsaturated state; indication information: continue detection.
[0066] Round 2: The dialogue text of the user's response to be evaluated is "I might think about it, but I'm not sure".
[0067] First-level check: 0 keyword matches, no keyword matching result found.
[0068] Second layer check: semantic similarity 0.32, semantic inequivalence, indicating an unsaturated state.
[0069] Third-level check: Current round 2, intervention attribute is no intervention.
[0070] Evidence saturation status determination: unsaturated state; indication information: continue detection.
[0071] Round 3: The dialogue text of the user's response to be evaluated is "Hmm...I'll probably look into the relevant information?" First-level check: 0 keyword matches, no keyword matching result found.
[0072] Second layer check: semantic similarity 0.45, semantic inequivalence, indicating partial saturation.
[0073] Third-level check: Current round 3, intervention attribute is mandatory intervention.
[0074] Evidence saturation status determination: Partially saturated state, indicating information: Continued detection information including forced selection information. The questioning agent generates the following evaluation question text for the next round based on this continue detection information: "Faced with this new information, are you more inclined to 'postpone the judgment and look for more evidence' or 'believe what the article says and adopt it directly'?" Round 4: The dialogue text of the user's response to be evaluated is "I will definitely postpone my judgment for now. I will check the source of this article to see if it is published by an authoritative source, and I will also search for other independent studies that have verified this result."
[0075] First-level check: The keywords "delay judgment", "check", "source", and "verify" match the dimension operation definition. The keyword matching count is 4, and the match is successful.
[0076] Evidence saturation status determination: saturated state; evidence quality: direct match; evidence confidence: high; indication information: dimension switching information and scoring trigger information.
[0077] The valid behavioral evidence that can be extracted from round 4 is ["Deferred judgment", "Investigate source", "Search independent research", "Verification results"].
[0078] In one embodiment, the questioning agent is specifically used for: Upon receiving the continued detection information sent by the evidence curation agent, the attribute indicators corresponding to the current dialogue text are determined; based on the attribute indicators, the questioning strategy for the next round is determined; the attribute indicators include at least one of content completeness, depth of thought, and expression certainty.
[0079] Specifically, when the questioning agent receives a "continue detection" message from the evidence curating agent, indicating that the current dialogue text is insufficient, it can input the current dialogue text into a large language model to extract attribute indicators. These indicators include at least one of content completeness, thought depth, and expression certainty. Content completeness is obtained through semantic analysis of the current dialogue text by the large language model, and it characterizes whether the user being evaluated directly responded to the evaluation question text in the current round. Thought depth is obtained through semantic analysis of preset keywords, including "because," "therefore," and "considering," and it characterizes whether the user being evaluated demonstrated a reasoning process. Expression certainty is obtained through frequency detection of preset fuzzy words, including "possibly," "maybe," and "probably," and it characterizes whether the user being evaluated used ambiguous language.
[0080] Next, based on this attribute indicator, the questioning strategy for the next round is determined. This strategy includes a standard follow-up questioning strategy, an in-depth follow-up questioning strategy, and a strategy that encourages expression. When the completeness and depth of thought of the current dialogue text are both high, it indicates that the user's answer is complete and clear, and more relevant behavioral evidence can be collected. In this case, the standard follow-up questioning strategy can be determined for the next round. For example, if the user's answer to the current dialogue text is "I will first check the source of this article," then the evaluation question text generated according to the standard follow-up questioning strategy for the next round would be "Great, besides checking the source, what other steps would you take to verify this information?" When the completeness of the content of the current dialogue text is medium and the depth of thought is low, it indicates that the user's answer is partially relevant but lacks depth. The user can be guided to elaborate on their thought process. In this case, the in-depth follow-up questioning strategy can be determined for the next round. For example, if the user's answer to the current dialogue text is "I will consider it before deciding," then the evaluation question text generated according to the in-depth follow-up questioning strategy would be "You mentioned you would 'consider,' could you elaborate on what aspects you would consider? How would you weigh the different factors?" When the completeness, depth of thought, and certainty of expression in the current dialogue text are all low, it indicates that the user being evaluated is vague, uses a lot of vague words, or has difficulty expressing themselves. It's necessary to reduce the pressure on the user to express themselves and guide them to approach the issue from a more concrete perspective. At this point, the questioning strategy for the next round can be determined to be an encouragement-based approach. For example, if the user's current dialogue text is "Hmm...I'm not sure, maybe...I'll take a look," the evaluation question text generated based on the encouragement-based approach could be: "That's okay, this is indeed a question that needs thought. Let's look at it from a different angle—suppose your most trusted friend asks you for advice on this matter, how would you answer them?"
[0081] In one embodiment, the reasoning arbitration agent is specifically used for: The dialogue text of each round is mapped to the preset behavior indicators in the preset definition information to obtain the mapping relationship between the matching behavior indicators and the dialogue text, as well as the number of evidence and the confidence level of the evidence corresponding to the matching behavior indicators. Based on the mapping relationship between the matching behavior indicators and the dialogue text, the current behavior pattern of the user to be evaluated in the current dimension is determined; Based on the number of pieces of evidence and the confidence level of the evidence corresponding to the matching behavior indicators, the level score corresponding to the current behavior pattern is determined; Based on the grade score, the number of pieces of evidence and the confidence level of the evidence corresponding to the matching behavior indicator, the dimension score corresponding to the current dimension is determined; Based on the mapping relationship between the matching behavior indicators and the dialogue text, the current behavior pattern, the dimension score, the number of evidences, and the confidence level of the evidence, the evaluation thought chain corresponding to the current dimension is determined.
[0082] Specifically, upon receiving the scoring trigger information from the evidence curation agent, the reasoning arbitration agent indicates that the collected dialogue text from all rounds is sufficient. Then, based on the dimensional representation corresponding to the current dimension, a preset behavioral indicator for that dimension can be obtained. Next, the dialogue text from all rounds and the preset behavioral indicator are input into a large language model. The large language model identifies dialogue fragments in the dialogue text that match the preset behavioral indicator and identifies these fragments as matching behavioral indicators, thus mapping the natural language of the user to be evaluated to specific dimensional operation definitions. After feature mapping, the large language model obtains the mapping relationship between the matching behavioral indicator and the dialogue text, as well as the evidence confidence level. Based on this mapping relationship, the amount of evidence corresponding to the matching behavioral indicator can be determined.
[0083] For example, the mapping relationship between the matching behavior metric and the dialogue text is as follows.
[0084] The dialogue text "I will definitely pause first" is mapped to the matching behavior indicator "postpone drawing conclusions", with a high confidence level of evidence and a source round of 2.
[0085] The dialogue text "I will check the source of this article" is mapped to the matching behavior indicator "questioning the source of information", with a high confidence level of evidence and a source round of 2.
[0086] The dialogue text "Search for other independent studies" maps to the matching behavior indicator "Seeking multi-party verification", with a high confidence level of evidence and a source round of 3.
[0087] The dialogue text "I won't easily believe it if I can't find corroborating evidence" maps to the matching behavior indicator "requires evidence", with a high confidence level for the evidence and a source round of 3.
[0088] As can be seen from the above mapping relationship, the number of pieces of evidence corresponding to each matching behavior indicator is 1.
[0089] After determining the mapping relationship between the matching behavior indicators and the dialogue text, the current behavior pattern of the user to be evaluated in the current dimension is determined based on this mapping relationship. This current behavior pattern is used to determine whether the user's behavior pattern meets the high-score criteria for the current dimension. For example, the current behavior pattern could be: "The user to be evaluated demonstrates a strong awareness of information verification. When faced with new information, the user first chooses to postpone judgment, and then actively proposes verification strategies: including verifying the source, assessing authority, and seeking independent verification. The user also explicitly states that they will not readily believe information when the evidence is insufficient." This current behavior pattern indicates that the user's behavior pattern meets the high-score criteria for the current dimension.
[0090] After determining the amount and confidence level of evidence corresponding to the matching behavioral indicators, the rating of the current behavioral pattern of the user to be evaluated is determined based on the amount and confidence level of evidence.
[0091] In one embodiment, the reasoning arbitration agent is specifically used for: Based on the number of evidence corresponding to the matching behavior indicators and the number of preset behavior indicators, the evidence coverage of the current dimension is determined; Based on the evidence coverage and the evidence confidence of the current dimension, the level score corresponding to the current behavior pattern is determined.
[0092] Specifically, based on the number of pieces of evidence corresponding to the matched behavioral indicators, the number of indicators with evidence is determined. The ratio of the number of indicators with evidence to the total number of indicators corresponding to the preset behavioral indicators is calculated. This ratio is the evidence coverage, and the value range of evidence coverage is [0,1]. Next, the confidence level of the evidence is determined as the strength of the evidence. Then, the evidence coverage rate is compared with the coverage scoring threshold in the high-score standard judgment rule to determine the first interval of the evidence coverage. Simultaneously, the evidence strength is compared with the strength averaging threshold in the high-score standard judgment rule to determine the second interval of the evidence strength. Based on the first and second intervals, the level score corresponding to the current behavioral pattern is determined from the high-score standard judgment rule.
[0093] For example, taking the coverage scoring threshold as having an upper limit of 75% and a lower limit of 50%, and the intensity scoring threshold as having an upper limit of 75% and a lower limit of 50%, the high score standard judgment rule is as follows.
[0094] High level (50-60 points): First interval: evidence coverage ≥75%, second interval: evidence strength ≥75%.
[0095] Medium level (30-49 points): First interval: evidence coverage ≥50%, second interval: evidence strength ≥50%.
[0096] Low level (10-29 points): The first interval is: evidence coverage <50%, or the second interval is: evidence strength <50%.
[0097] When the current behavioral pattern is determined to be at a high level based on the first and second intervals, the corresponding level score can be determined from the score range corresponding to the high level based on the average of evidence coverage and evidence strength.
[0098] After determining the level score corresponding to the current behavioral pattern, this level score is used as the initial score. Based on the amount of evidence and the confidence level of the evidence, this initial score is calibrated to obtain the calibrated dimension score for the current dimension. For example, the evidence coverage is determined based on the amount of evidence. If the evidence coverage is 100% and the confidence level of the evidence corresponding to each matching behavioral indicator is high, 2 points are added to the initial score. Conversely, if the evidence coverage is less than 50%, 3 points are subtracted from the initial score. It should be noted that the value range of this dimension score can be [10, 60].
[0099] After obtaining the dimensional scores, the mapping relationship between matching behavioral indicators and dialogue text, the current behavioral pattern, dimensional scores, the amount of evidence, and the confidence level of evidence are summarized to generate the evaluation thought chain corresponding to the current dimension. This evaluation thought chain specifically includes: using the mapping relationship as an evidence summary table, and determining the feature mapping summary based on the mapping relationship and the confidence level of evidence; summarizing the current behavioral pattern as the current dimension, and determining the score derivation for the current dimension based on the amount of evidence, the confidence level of evidence, and the dimensional scores; and determining the audit trail based on the mapping relationship, evidence coverage, evidence confidence level, current behavioral pattern, and dimensional scores.
[0100] For example, the thought process for this assessment is as follows: Evidence Summary Table: E1 "I will definitely pause first" Source round 2, mapping indicator "postpone conclusion".
[0101] E2 "Check the source of this article" Source round 2, mapping indicator "Question the source of information".
[0102] Is E3 published in a top-tier journal, or just a company promotional blog? (Source Round 2, Mapping Indicator) Questions the source of information.
[0103] E4 "Search if other independent studies or benchmarks have replicated this result" Source Round 3, Mapping Metrics "Seek multi-party verification".
[0104] E5 "If I can't find supporting evidence, I won't easily believe the source round 3, mapping indicator" which requires evidence.
[0105] Feature mapping summary: Five high-confidence pieces of evidence were extracted from the three rounds of dialogue, covering all four preset behavioral indicators of the "seeking the truth" dimension.
[0106] Dimensional Summary: The user being evaluated proactively proposed a verification strategy, demonstrating a strong and consistent initiative in information verification. The user's verification strategy, encompassing source verification, authoritative assessment, and independent verification, reflects a systematic approach. The user explicitly stated that they would not readily accept information without sufficient evidence, demonstrating steadfastness. Behavioral Pattern Judgment: Highly consistent with the operational definition of the "seeking truth" dimension.
[0107] Scoring Derivation: Based on the CCTDI scoring standard (10-60 points), with 100% evidence coverage and all evidence confidence being high, the high score threshold was met. A score of 50-60 corresponds to a "high level," which the user being evaluated fully meets. Considering the sufficiency and high quality of the evidence, the proactive rather than reactive behavior, and the clear and unambiguous expression, the final score is 52 points.
[0108] Audit trail: 5 pieces of evidence → full coverage of 4 indicators → highly consistent behavioral pattern → meets high score criteria → 52 points.
[0109] In addition, the assessment thought process includes an uncertainty statement: there is no significant uncertainty, all evidence is of high confidence, and the behavioral patterns are consistent.
[0110] This invention also provides a scenario-based critical thinking assessment method, applicable to the scenario-based critical thinking assessment system described in any of the above embodiments. Figure 2 This is a flowchart illustrating the scenario-based critical thinking assessment method provided in this embodiment of the invention, such as... Figure 2 As shown, the method includes steps 210 to 240.
[0111] Step 210: By coordinating the agent to traverse all dimensions corresponding to critical thinking, generate the micro-context text corresponding to the current dimension.
[0112] Step 220: The questioning agent interacts with the user to be evaluated based on the micro-context text of the current dimension to obtain the dialogue text of the user to be evaluated in all rounds.
[0113] Step 230: The evidence curation agent monitors the evidence saturation based on the dialogue text of the user to be evaluated in the current round to obtain the indication information corresponding to the target agent; the target agent includes at least one of the coordinating agent, the questioning agent, and the reasoning arbitration agent.
[0114] Step 240: Upon receiving the instruction information sent by the evidence curation agent, the reasoning arbitration agent determines the evaluation thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be evaluated in all rounds; the evaluation thought chain includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimension score.
[0115] This invention provides a scenario-based critical thinking assessment method. A coordinating agent transforms the abstract mental constructs of the current dimension into scenario-based micro-context text. The questioning agent interacts with the user to be assessed based on the micro-context text, obtaining the dialogue text of all rounds of the user's responses to the current dimension. An evidence curating agent monitors the evidence saturation of the current dialogue text in the current round, obtaining the corresponding instruction information from the target agent. Upon receiving the instruction information from the evidence curating agent, the reasoning arbitration agent, based on the preset definition information of the current dimension and the dialogue text of all rounds of the user to be assessed, determines the mapping relationship between at least all rounds of dialogue text in the current dimension and the preset definition information, as well as the assessment thought chain for dimension scoring. In this embodiment of the invention, the evidence curation agent monitors the evidence saturation in real time to verify the sufficiency of evidence, and dynamically controls the dialogue length between the questioning agent and the user to be evaluated, avoiding insufficient evidence or invalid follow-up questions caused by fixed dialogue rounds, thus significantly improving the reliability and efficiency of the evaluation. At the same time, the evaluation thought chain generated by the reasoning arbitration agent includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimension score. The scoring process is completely transparent and traceable, improving the interpretability of the dimension score and meeting the requirements of educational evaluation for transparency and auditability.
[0116] It should be noted that the specific implementation steps of the scenario-based critical thinking assessment method provided in the embodiments of the present invention can refer to the corresponding implementation steps of the scenario-based critical thinking assessment system described in any of the above embodiments, and will not be repeated here.
[0117] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present invention, such as... Figure 3As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communications bus 340. The processor 310 can call logical instructions in the memory 330 to execute a contextualized critical thinking assessment method. This method includes: coordinating an agent to traverse all dimensions corresponding to critical thinking and generate micro-context text corresponding to the current dimension; having an interrogating agent interact with the user to be assessed based on the micro-context text of the current dimension to obtain the dialogue text of the user in all rounds; having an evidence curation agent monitor evidence saturation based on the dialogue text of the user in the current round to obtain instruction information corresponding to a target agent; the target agent includes at least one of a coordinating agent, an interrogating agent, and a reasoning arbitration agent; and, upon receiving the instruction information sent by the evidence curation agent, determining the assessment thought chain corresponding to the current dimension based on preset definition information of the current dimension and the dialogue text of the user in all rounds; the assessment thought chain includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as a dimension score.
[0118] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0119] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the scenario-based critical thinking assessment method provided by the above methods. The method includes: traversing all dimensions corresponding to critical thinking by a coordinating agent to generate micro-context text corresponding to the current dimension; interacting with the user to be assessed by a questioning agent based on the micro-context text of the current dimension to obtain the dialogue text of the user to be assessed for all rounds; monitoring the evidence saturation by an evidence curation agent based on the dialogue text of the user to be assessed for the current round to obtain the instruction information corresponding to the target agent; the target agent includes at least one of a coordinating agent, a questioning agent, and a reasoning arbitration agent; and, upon receiving the instruction information sent by the evidence curation agent, determining the assessment thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be assessed for all rounds; the assessment thought chain includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimension score.
[0120] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the scenario-based critical thinking assessment method provided by the above methods. The method includes: traversing all dimensions corresponding to critical thinking by a coordinating agent to generate micro-context text corresponding to the current dimension; interacting with the user to be assessed by a questioning agent based on the micro-context text of the current dimension to obtain the dialogue text of the user to be assessed for all rounds; monitoring evidence saturation by an evidence curation agent based on the dialogue text of the user to be assessed for the current round to obtain instruction information corresponding to a target agent; the target agent includes at least one of a coordinating agent, a questioning agent, and a reasoning arbitration agent; upon receiving the instruction information sent by the evidence curation agent, the reasoning arbitration agent determines the assessment thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be assessed for all rounds; the assessment thought chain includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimension score.
[0121] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0122] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A scenario-based critical thinking assessment system, characterized in that, include: A coordinating agent is used to traverse all dimensions corresponding to critical thinking and generate micro-contextual text corresponding to the current dimension. The questioning agent is used to interact with the user to be evaluated based on the micro-context text of the current dimension, and obtain the dialogue text of the user to be evaluated in all rounds. An evidence curation agent is used to monitor evidence saturation based on the current dialogue text of the user to be evaluated in the current round, and obtain the indication information corresponding to the target agent. The target intelligent agent includes at least one of a coordinating intelligent agent, a questioning intelligent agent, and a reasoning arbitration intelligent agent; The reasoning arbitration agent is used to determine the evaluation thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be evaluated in all rounds, upon receiving the instruction information sent by the evidence curation agent. The evaluation thought chain includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimension score.
2. The scenario-based critical thinking assessment system according to claim 1, characterized in that, The evidence curation agent is specifically used for: Obtain the dimension operation definition from the preset definition information of the current dimension; the dimension operation definition is used to characterize the behavioral meaning of the current dimension; According to the priority of the matching steps from high to low, the current dialogue text is matched with the dimension operation definition step by step to obtain the evidence saturation state corresponding to the current dialogue text. Based on the evidence saturation state, determine the indication information corresponding to the target intelligent agent; The matching steps, in descending order of priority, are keyword matching, semantic equivalence matching, and dialogue turn matching.
3. The scenario-based critical thinking assessment system according to claim 2, characterized in that, The evidence curation agent is specifically used for: Perform keyword matching between the current dialogue text and the dimension operation definition to determine the keyword matching result; If the keyword matching result is successful, the evidence saturation state corresponding to the current dialogue text is determined to be saturated. If the keyword matching result fails, determine the semantic similarity between the current dialogue text and the dimension operation definition; If the semantic similarity is greater than or equal to a first preset threshold, the evidence saturation state corresponding to the current dialogue text is determined to be saturated. If the semantic similarity is less than the first preset threshold and greater than or equal to the second preset threshold, the evidence saturation state corresponding to the current dialogue text is determined to be a partially saturated state. The first preset threshold is greater than the second preset threshold; If the semantic similarity is less than the second preset threshold, the evidence saturation state corresponding to the current dialogue text is determined to be unsaturated. If the evidence saturation state corresponding to the current dialogue text is unsaturated or partially saturated, an intervention attribute is added to the evidence saturation state based on the comparison result between the dialogue round corresponding to the current dialogue text and the round intervention threshold.
4. The scenario-based critical thinking assessment system according to claim 3, characterized in that, The indication information includes at least one of dimension switching information, scoring trigger information, and continued detection information; The evidence curation agent is specifically used for: When the evidence saturation state is saturated, dimension switching information corresponding to the coordinating agent and scoring trigger information corresponding to the reasoning arbitration agent are generated; the dimension switching information is used to instruct the coordinating agent to generate micro-contextual text for the next dimension. The scoring trigger information is used to instruct the reasoning arbitration agent to determine the evaluation thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be evaluated in all rounds. When the evidence saturation state is unsaturated or partially saturated, the questioning agent generates continued detection information based on the intervention attribute; the continued detection information is used to instruct the questioning agent to generate the evaluation question text for the next round.
5. The scenario-based critical thinking assessment system according to claim 1, characterized in that, The reasoning and arbitration agent is specifically used for: The dialogue text of each round is mapped to the preset behavior indicators in the preset definition information to obtain the mapping relationship between the matching behavior indicators and the dialogue text, as well as the number of evidence and the confidence level of the evidence corresponding to the matching behavior indicators. Based on the mapping relationship between the matching behavior indicators and the dialogue text, the current behavior pattern of the user to be evaluated in the current dimension is determined; Based on the number of pieces of evidence and the confidence level of the evidence corresponding to the matching behavior indicators, the level score corresponding to the current behavior pattern is determined; Based on the grade score, the number of pieces of evidence and the confidence level of the evidence corresponding to the matching behavior indicator, the dimension score corresponding to the current dimension is determined; Based on the mapping relationship between the matching behavior indicators and the dialogue text, the current behavior pattern, the dimension score, the number of evidences, and the confidence level of the evidence, the evaluation thought chain corresponding to the current dimension is determined.
6. The scenario-based critical thinking assessment system according to claim 5, characterized in that, The reasoning and arbitration agent is specifically used for: Based on the number of evidence corresponding to the matching behavior indicators and the number of preset behavior indicators, the evidence coverage of the current dimension is determined; Based on the evidence coverage and the evidence confidence of the current dimension, the level score corresponding to the current behavior pattern is determined.
7. The scenario-based critical thinking assessment system according to claim 4, characterized in that, The questioning agent is specifically used for: Upon receiving a continue detection message or a forced selection message from the evidence curation agent, the attribute indicators corresponding to the current dialogue text are determined; based on the attribute indicators, the questioning strategy for the next round is determined; the attribute indicators include at least one of content completeness, depth of thought, and expression certainty.
8. A scenario-based critical thinking assessment method, characterized in that, The method applied to the scenario-based critical thinking assessment system according to any one of claims 1-7, the method comprising: By coordinating the intelligent agent to traverse all dimensions corresponding to critical thinking, micro-contextual text corresponding to the current dimension is generated; By having the intelligent agent interact with the user to be evaluated based on the micro-context text of the current dimension, the dialogue text of the user to be evaluated in all rounds is obtained. The evidence curation agent monitors the evidence saturation based on the dialogue text of the user to be evaluated in the current round, and obtains the indication information corresponding to the target agent; the target agent includes at least one of the coordinating agent, the questioning agent, and the reasoning arbitration agent; Upon receiving the instruction information sent by the evidence curation agent, the reasoning arbitration agent determines the evaluation thought chain corresponding to the current dimension based on the preset definition information of the current dimension and the dialogue text of the user to be evaluated in all rounds. The evaluation thought chain includes at least the mapping relationship between the dialogue text of all rounds and the preset definition information, as well as the dimension score.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the scenario-based critical thinking assessment method as described in claim 8.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the scenario-based critical thinking assessment method as described in claim 8.