Closed-loop learning control method and device based on artificial intelligence, equipment and storage medium
By using an AI-based closed-loop learning control method, guiding questions are generated and multi-dimensional evaluations are performed to detect semantic deviations and form a learning loop. This solves the problems of insufficient learning loop and feedback in AI education systems and improves the stability and security of the learning process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 成耿
- Filing Date
- 2026-01-23
- Publication Date
- 2026-04-21
AI Technical Summary
Existing AI education systems lack closed-loop control for learning, making it difficult to form a continuous learning loop of questioning, answering, and feedback. They are unable to detect and correct deviations from the dialogue, have insufficient feedback and evaluation dimensions, and lack the ability to organize interdisciplinary knowledge and assess learning levels.
By using an AI-based closed-loop learning control method, guiding questions are generated, semantic analysis is performed, multi-dimensional scores are calculated, semantic deviation is detected, and corrective feedback is generated, forming a closed loop of questioning, answering, and feedback. Combined with a strategy task library, multi-dimensional regulation is carried out, supporting interdisciplinary knowledge integration and learning path planning, and local storage and encryption mechanisms are used to protect learning records.
It enables multi-dimensional learning process assessment and feedback, automatically detects and regresses dialogue deviations, supports interdisciplinary knowledge integration, and improves the stability and security of the learning process.
Smart Images

Figure CN121903818A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a closed-loop learning control method, apparatus, device and storage medium based on artificial intelligence. Background Technology
[0002] With the development of generative AI and natural language processing technologies, intelligent teaching systems are able to interact with learners in multiple rounds and provide instant feedback. However, existing AI education systems or general conversational learning systems still have at least the following shortcomings in actual learning guidance: 1. Lack of learning loop control: Most systems only provide single-round questioning and feedback, making it difficult to form a continuous learning loop of questioning, answering, feedback, and re-questioning, and thus unable to continuously correct and advance on the same learning goal; 2. Difficulty in timely correction of deviations in dialogue: In multi-round dialogues, learners are prone to interrupting or changing the topic. Existing systems lack detection and correction mechanisms for deviations, which can easily lead to the loss of learning objectives or the learning process being led astray by irrelevant content. 3. Insufficient feedback and evaluation dimensions: Existing systems mostly evaluate based on the accuracy of answers or a few quantitative indicators, lacking comprehensive analysis and targeted correction of learning performance across multiple dimensions such as the completeness of the answer content, the logical structure, the standardization of language, and the clarity of expression.
[0003] Furthermore, with the increasing complexity of learning scenarios and the rising requirements for data security, some application scenarios also need to support interdisciplinary knowledge organization and learning level assessment, while taking into account the localized management and security protection capabilities of learning data in the process of personalized learning.
[0004] Therefore, there is an urgent need for a learning control scheme that can achieve closed-loop control of learning in multi-turn dialogues, detect and regress semantic deviations, and output corrective feedback based on multi-dimensional evaluation, so as to improve the stability of the learning process and the learning results. Summary of the Invention
[0005] This application provides an artificial intelligence-based closed-loop learning control method, apparatus, device, and storage medium to solve at least one problem existing in related technologies. The technical solution is as follows: In a first aspect, embodiments of this application provide a closed-loop learning control method based on artificial intelligence, comprising: The question generation submodule generates guiding questions based on the target topic and receives learners' answers to the guiding questions. The response analysis and feedback submodule performs semantic analysis on the response content and calculates content dimension scores, structure dimension scores, language dimension scores, and expression dimension scores respectively. The scoring dimension mapping submodule then calculates a comprehensive score based on these four dimensions and compares the comprehensive score with the target threshold. Compare with T1; The semantic deviation of the answer content relative to the target topic is analyzed through the interruption hosting and regression submodule, and the semantic deviation is compared with the deviation threshold T2. When the semantic deviation is greater than the deviation threshold T2 or the comprehensive score is lower than the achievement threshold T1, the regression count C is accumulated by the regression counter. If the regression count C is less than or equal to the round limit N, correction feedback content is generated and used as a new guiding question in the next round. The process returns to the learner's answer to the guiding question to form a closed loop of questioning, answering, feedback, and re-questioning until the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the achievement threshold T1, or the regression count C is greater than the round limit N. The corrective feedback includes: corrective follow-up questions to improve the overall score, or regression prompts to indicate the cause of semantic deviation and switch the conversation topic of the answer back to the target topic.
[0006] In one implementation, the step of analyzing the semantic deviation of the response content relative to the target topic through the interruption hosting and regression submodule includes: The first vector of the answer content and the second vector of the guiding question are generated by the natural language processing model in the interruption hosting and regression submodules, and the semantic embedding similarity between the first vector and the second vector is calculated. The keyword matching score and edit distance between the keywords of the answer content and the topic keywords corresponding to the target topic are determined through the interruption hosting and regression sub-modules. The BM25 score is determined by using the interruption hosting and regression submodules based on the content of the answer and the topic keywords corresponding to the target topic. The semantic embedding similarity, the keyword matching score, the edit distance, and the BM25 score are normalized respectively to obtain normalized semantic embedding similarity, normalized keyword matching score, normalized edit distance, and normalized BM25 score. Each preset value used for comparison is determined, and a first difference between each preset value and the normalized semantic embedding similarity, a second difference between each preset value and the normalized keyword matching score, and a third difference between each preset value and the normalized BM25 score are calculated. One of the first difference, the second difference, the third difference, and the normalized edit distance is used as the semantic deviation, or at least two of the first difference, the second difference, the third difference, and the normalized edit distance are used to calculate a comprehensive semantic deviation. The comprehensive calculation includes any one of weighted summation, logistic regression, and neural networks.
[0007] In one implementation, the regression hints include at least one of missing concepts, off-topic keywords, and topic difference summaries.
[0008] In one embodiment, the method further includes: When the number of regressions C exceeds the maximum number of rounds N, switch back to the main task, redirect the learner to the learning objective of the target topic, and reset the number of regressions C to zero.
[0009] In one implementation, the step of calculating the regression number C using a regression counter, generating corrective feedback content within rounds where C ≤ N, and using the corrective feedback content as a new guiding question includes: The scores for the four dimensions—content, structure, language, and expression—are compared with their corresponding scoring thresholds to identify the target dimensions that are below the thresholds. Based on the preset task mapping relationship between the target dimension and the scoring interval, at least one strategic task is triggered to generate targeted corrective follow-up questions corresponding to the target dimension.
[0010] In one embodiment, the method further includes: When the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the achievement threshold T1, a structured summary containing strengths and suggestions for optimization is output based on the learner's final answer.
[0011] In one implementation, the target threshold T1, the deviation threshold T2, and the round limit N are adaptively updated based on historical performance and dialogue stability; wherein the dialogue stability includes at least one of regression count, round score fluctuation, or round duration fluctuation; and the adaptive update method includes at least one of exponential smoothing or sliding window estimation.
[0012] Secondly, embodiments of this application provide a closed-loop learning control device based on artificial intelligence, comprising: The question generation submodule is used to generate guiding questions based on the target topic and receive learners' answers to the guiding questions. The answer analysis and feedback submodule is used to perform semantic analysis on the answer content and calculate the content dimension score, structure dimension score, language dimension score, and expression dimension score respectively. The scoring dimension mapping submodule calculates the comprehensive score based on the above four dimension scores and compares the comprehensive score with the pass threshold T1. The interruption hosting and regression submodule is used to analyze the semantic deviation of the answer content relative to the target topic and compare the semantic deviation with the deviation threshold T2; The scoring dimension mapping submodule is used to accumulate the number of regressions C by a regression counter when the semantic deviation is greater than the deviation threshold T2 or the comprehensive score is lower than the achievement threshold T1. If the number of regressions C is less than or equal to the upper limit of rounds N, corrective feedback content is generated and the corrective feedback content is used as a new guiding question to enter the next round. The steps of returning the learner's answer to the guiding question are used to form a closed loop of questioning, answering, feedback, and re-questioning until the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the achievement threshold T1, or the number of regressions C is greater than the upper limit of rounds N. The corrective feedback includes: corrective follow-up questions to improve the overall score, or regression prompts to indicate the cause of semantic deviation and switch the conversation topic of the answer back to the target topic.
[0013] In one implementation, the system further includes a policy task library and an output module, for at least one of the following: When the semantic deviation is greater than the deviation threshold T2, the reverse question generation task is invoked to generate explanatory questions for the missing key concepts of the target topic; When the content dimension score and / or the structure dimension score are lower than the corresponding score threshold, the semantic hierarchy decomposition task is invoked to generate guiding sub-questions corresponding to the content dimension and / or structure dimension. When the language dimension score and / or the expression dimension score are lower than the corresponding score threshold, the information density restatement task and / or the hierarchical summary output task are invoked. When the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the achievement threshold T1, the output module calls the hierarchical summary output task based on the learner's final answer and outputs a structured summary. Obtain information related to the learner's learning status, determine the learner's learning status based on the information related to the learning status, and adjust the pace and difficulty by calling the learning status detection task based on the learning status. Based on the learner's historical memory data, a memory scheduling task is invoked to arrange a review; Trigger the relationship extraction and comparison task, and feed the extraction results back into the question generation submodule and / or the strategy task library for the generation of new guiding questions in the next round.
[0014] In one embodiment, a STEM extension module is further included. This STEM extension module comprises a knowledge map generation submodule, a hierarchical growth assessment submodule, and a personalized learning path submodule. The knowledge map generation submodule generates a knowledge map, and its output is used to update the parameters of the question generation submodule and the personalized learning path submodule. The hierarchical growth assessment submodule calculates the learning level according to the understanding layer, application layer, and innovation layer. The personalized learning path submodule updates the question granularity and prerequisite dependencies based on the knowledge map.
[0015] In one implementation, an output module is also included, which supports local storage and clearing functions, and by default only saves learning records locally.
[0016] In one implementation, the output module also supports end-to-end encryption for encrypting and storing learning records generated during dialogue with the learner on the terminal side. The end-to-end encryption includes at least a key rotation and an offline erasure trigger mechanism.
[0017] In one implementation, the policy task library further includes at least one of the following learning regulation tasks: The retrieval practice task is used to select target knowledge points that need to be reviewed first based on the learner's historical practice records on the target knowledge points, and generate practice questions that encourage the learner to actively recall the target knowledge points without relying on external materials; the historical practice records include at least: the number of correct answers, the number of incorrect answers, and the time of the last practice; Alternating training tasks are used to alternate learning tasks between different themes or subjects. The proportion of questions is dynamically adjusted according to the correctness of answers and the answering time for each theme or subject, so as to improve the ability to transfer and generalize knowledge. Interval repetitive scheduling tasks are used to determine the next review interval for the target knowledge point based on the learner's performance in answering questions about the target knowledge point at different times, so that the review time of each knowledge point is adaptively shifted forward or backward during the learning process to enhance long-term retention. Metacognitive monitoring tasks are used to determine the learner's current comprehension and mastery level by comprehensively considering at least two of the following: response time, learner self-assessment confidence, and error type information. Based on this, the pace of subsequent feedback, the granularity of explanation, and the difficulty of questions are adjusted.
[0018] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory, wherein the memory stores instructions that are loaded and executed by the processor to implement the methods in any of the above-described embodiments.
[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed, implements the methods in any of the above-described embodiments.
[0020] The beneficial effects of the above technical solution include at least the following: The system generates guiding questions based on the target topic, receives learners' responses to these questions, and uses an answer analysis and feedback submodule to perform semantic analysis on the responses and calculate scores for content, structure, language, and expression dimensions. A scoring dimension mapping submodule calculates a comprehensive score for these four dimensions and compares it to a threshold T1 to achieve a multi-dimensional evaluation of the response performance. An interruption and regression submodule analyzes the semantic deviation of the response from the target topic and compares it to a deviation threshold T2. If the semantic deviation exceeds T2 or the comprehensive score falls below T1, a regression counter is used to adjust the response. The number of regressions C is accumulated. If the number of regressions C is less than or equal to the maximum number of rounds N, corrective feedback content is generated as a new guiding question, and the process proceeds to the next round. The process returns to the learner's answer to the corresponding guiding question, forming a closed loop of questioning, answering, feedback, and re-questioning, until the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the achievement threshold T1, or the number of regressions C is greater than the maximum number of rounds N. The corrective feedback content includes corrective follow-up questions to improve the comprehensive score, or regression prompts to indicate the cause of semantic deviation and switch the conversation topic of the answer back to the target topic, thereby realizing the detection of the cause of deviation and automatic regression.
[0021] Furthermore, by combining the scoring dimension mapping submodule with a preset strategy task library, this application can also automatically trigger various control tasks based on scoring results in different dimensions such as content, structure, language, and expression. These tasks include generating reverse questions, semantic hierarchy decomposition, information density restatement, hierarchical summary output, memory scheduling, learning state detection, and relation extraction and comparison, allowing for fine-tuning of the learning process in terms of time, content, and structure.
[0022] By setting up the STEM extension module and introducing sub-modules for knowledge mapping generation, hierarchical growth assessment, and personalized learning paths, it is possible to support interdisciplinary knowledge integration and continuously assess and plan the learner's growth process according to the understanding, application, and innovation layers, making the learning process more long-term and visible.
[0023] By employing a local storage and erasure mechanism in the output module, and supporting privacy protection mechanisms such as end-to-end encryption in one implementation, combined with key rotation and offline erasure triggering mechanisms, this application improves the security and controllability of learning records without sacrificing the utilization value of personalized learning records.
[0024] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, these aspects, embodiments, and features will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description
[0025] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in this application and should not be construed as limiting the scope of this application.
[0026] Figure 1 This is a flowchart illustrating the steps of an artificial intelligence-based closed-loop learning control method according to an embodiment of this application. Figure 2 This is a structural block diagram of an artificial intelligence-based closed-loop learning control device according to an embodiment of this application; Figure 3 This is a schematic diagram of the learning hierarchy in one embodiment of this application; Figure 4 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0027] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of this application. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0028] Reference Figure 1 The flowchart illustrates an embodiment of an artificial intelligence-based closed-loop learning control method according to this application. This artificial intelligence-based closed-loop learning control method may include at least steps S100-S400: S100: Generate guiding questions based on the target topic through the question generation submodule, and receive learners' answers to the guiding questions.
[0029] Specifically, the question generation submodule can include a topic parsing unit, a template retrieval unit, and a question generation unit. Upon receiving a target topic, the system first parses it, extracting the subject category, knowledge point tags, grade level, difficulty level, and scenario constraints. A strategy task library and a question template library are pre-established locally and / or on the server. Each template includes corresponding knowledge point tags and applicable conditions. Based on the target topic's tags, the template retrieval unit retrieves several highly matching question templates from the question template library, filling the template's placeholder slots with information such as entity names, scenario descriptions, and preconditions from the target topic, generating the first type of candidate guiding questions.
[0030] Simultaneously, the question generation unit can transform the target topic and its reference answer (if any) into vectors or other feature representations, inputting them into the text generation model, which automatically generates several candidate guidance questions of the second type. The quality assessment unit evaluates and filters these two types of candidate guidance questions based on dimensions such as relevance, comprehensibility, language standardization, and difficulty level, eliminating duplicate or obviously unsuitable questions for the teaching scenario. One or more questions that meet preset conditions are selected as the final guidance question output for the current round and recorded along with the target topic for subsequent updates to the template library and optimization of question generation strategies.
[0031] For example, when the target topic is "the formation of rainfall", the system can match templates such as "Please explain in your own words how ×× is formed" and "Please tell the whole process of ×× in order" in the template library, and fill them in to generate guiding questions such as "Please explain in your own words how rainfall is formed" and "Please tell the whole process from water evaporation to rainfall in order". When the number of templates is insufficient or the matching degree is low, the text generation model can also supplement the topic by generating guiding questions with different expressions but consistent teaching objectives.
[0032] S200. Through the answer analysis and feedback submodule, semantic analysis is performed on the answer content and the content dimension score, structure dimension score, language dimension score, and expression dimension score are calculated respectively. The scoring dimension mapping submodule calculates the comprehensive score based on the above four dimension scores and compares the comprehensive score with the pass threshold T1.
[0033] S300. Through the interruption hosting and regression sub-modules, analyze the semantic deviation of the answer content relative to the target topic, and compare the semantic deviation with the deviation threshold T2.
[0034] S400. When the semantic deviation is greater than the deviation threshold T2 or the comprehensive score is lower than the achievement threshold T1, the regression count C is accumulated by the regression counter. If the regression count C is less than or equal to the round limit N, correction feedback content is generated and used as a new guiding question in the next round. The process returns to step S100 to receive the learner's answer to the guiding question, thus forming a closed loop of questioning, answering, feedback, and re-questioning, until the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the achievement threshold T1, or the regression count C is greater than the round limit N.
[0035] In step S400, the interruption hosting and regression submodule works in conjunction with the regression counter, and is executed in conjunction with the comprehensive score output by the scoring dimension mapping submodule and the task trigger signal. The corrective feedback includes corrective follow-up questions to improve the comprehensive score, or regression prompts to indicate the cause of semantic deviation and switch the conversation topic of the answer back to the target topic.
[0036] The technical solution of this application embodiment generates guiding questions based on a target topic, receives learners' answers to the corresponding guiding questions, and uses an answer analysis and feedback submodule to perform semantic analysis on the answers and calculate content dimension scores, structure dimension scores, language dimension scores, and expression dimension scores respectively. A scoring dimension mapping submodule calculates a comprehensive score for the four dimensions and compares it with a passing threshold T1 to achieve a multi-dimensional evaluation of the answer performance. An interruption and regression submodule analyzes the semantic deviation of the answer content from the target topic and compares it with a deviation threshold T2. When the semantic deviation is greater than the deviation threshold T2 or the comprehensive score is lower than the passing threshold T1, a regression analysis is performed. The regression counter accumulates the regression count C. When the regression count C is less than or equal to the round limit N, corrective feedback content is generated as a new guiding question, and the process proceeds to the next round. The process then returns to the learner's answer to the corresponding guiding question, forming a closed loop of questioning, answering, feedback, and re-questioning, until the semantic deviation is less than or equal to the deviation threshold T2 and the overall score is higher than or equal to the pass threshold T1, or the regression count C is greater than the round limit N. The corrective feedback content includes corrective follow-up questions to improve the overall score, or regression prompts to indicate the cause of semantic deviation and switch the conversation topic of the answer back to the target topic, thereby realizing the detection of deviation causes and automatic regression.
[0037] In one implementation, learners utilize the AI-based closed-loop learning control device (system) of this application embodiment in 100 learner terminals to learn. The device includes: an input module 110, an AI learning engine 120 (including a question generation submodule 121, an answer analysis and feedback submodule 122, an interjection hosting and regression submodule 123, and a rating dimension mapping submodule 124), a strategy task library, a STEM extension module 130 (which may include at least one of a knowledge map generation submodule 131, a hierarchical growth assessment submodule 132, and a personalized learning path submodule 133), and an output module 140 for receiving the content generated by each module and outputting it to the 100 learner terminals for display.
[0038] The strategy task library can include reverse question generation, semantic hierarchy decomposition, information density restatement, hierarchical summary output, learning state detection, memory scheduling, and relation extraction and comparison. The 140 output module has privacy capabilities: an encryption module can be set locally on the terminal to encrypt the learner's original answers, profile features, and evaluation results before writing them to local storage. The encryption key used is stored in a secure area on the terminal. Learning records are stored locally by default without cloud authorization. When interaction with the server is required, anonymized statistical features or vector representations can be uploaded instead of data that can be directly restored to the original text. The output module also provides a clearing entry point. When the user triggers this entry point in the interface, learning data related to the device and encryption keys and other related information are cleared from the local storage space. An offline erasure trigger mechanism is also set up. For example, if the terminal is offline for more than a preset time, the account is logged out, or the application is uninstalled, the local erasure process is automatically initiated to reduce the risk of learning data residue when the device is lost or idle for a long time.
[0039] In one application scenario, to further support interdisciplinary knowledge integration and hierarchical growth assessment, a STEM extension module can be introduced, where STEM is an acronym for Science, Technology, Engineering, and Mathematics.
[0040] In this embodiment of the application, in step S100, the user can determine the subject and corresponding learning objective / topic (i.e., target topic) through the input module on the learner terminal (100). Then, the question generation submodule generates guiding questions based on the target topic for the learner to answer. The learner's answer to the guiding question (which can be text or voice) is input through the input module and received and processed by the question generation submodule and the answer analysis and feedback submodule. For example, when learning a science subject and the target topic is "precipitation", the guiding question could be: "What are the formation and destination of precipitation?" The learner provides the answer on the terminal for subsequent analysis and feedback.
[0041] In this embodiment of the application, in step S200, the answer analysis and feedback submodule performs semantic analysis on the answer content and calculates content dimension scores, structure dimension scores, language dimension scores, and expression dimension scores respectively. Specifically, it may include the following processing methods: (1) In terms of content dimension, learners’ answers are input into the text encoding model along with pre-set standard answers or reference points to identify the core concepts, key links and important relationships in the answers; the system calculates the coverage of the content dimension based on the number of knowledge points covered and the importance of each knowledge point, and maps it to the content score.
[0042] (2) In terms of structure, the answers are syntactically analyzed and segmented into paragraphs to extract structural information such as time sequence, causal relationship, and condition-result to form the "event chain" or "process chain" of the answer; then compared with the reference structure corresponding to the target topic, and the structural dimension score is obtained based on indicators such as whether the steps are complete, whether the order is reasonable, and whether a closed loop is formed.
[0043] (3) In terms of language dimension, the grammar detection and language quality assessment model is used to statistically analyze the features of grammatical errors, spelling errors, word repetition and sentences that are too long or too short in the answers, and combined with the overall fluency and standardization, to generate a score for the language dimension.
[0044] (4) In terms of expression dimension, the model evaluates the “comprehensibility” and “organization” of the answer by combining features such as paragraph division, transition connection, clarity of reference, and coherence of argumentation, and maps them to the score of expression dimension.
[0045] The scores for the four dimensions—content, structure, language, and expression—are weighted according to pre-defined weights and combined into a comprehensive score using the scoring dimension mapping submodule. This comprehensive score is then compared to the target threshold T1. When the comprehensive score is higher than or equal to the target threshold T1, the current answer meets the target requirements. When the comprehensive score is lower than the target threshold T1, it indicates that at least one dimension's score is below the corresponding threshold, requiring the next round of guided questioning and correction.
[0046] To facilitate understanding of the differences between various answers across different dimensions, this application embodiment uses the generation and destination of "precipitation" as an example to provide different scenarios for answer A, answer B, and answer C: 1. Answer A: Content: Complete steps (evaporation → condensation → precipitation → runoff / infiltration → return to the ocean), accurate concepts; Structure: In chronological / causal order with a clear closed loop; Language: Standard word choice with few ambiguities; Expression: Clear and easy to understand.
[0047] 2. Answer B: Content: Contains the main structure but lacks details; Structure: The connection is not tight enough and a complete convergence is not formed; Language: Slightly colloquial; Expression: Lacks organization.
[0048] Using the same scoring process, answer B scored lower than answer A in all dimensions, and in at least one dimension, it scored below the corresponding scoring threshold.
[0049] 3. Answer C: Content: Only describes "precipitation," lacking regression path and energy source; Structure: No closed loop formed; Language: The expression is largely omitting; Expression: Insufficient information density.
[0050] Therefore, answer C scores lower than answer B in all dimensions, and is below the corresponding score threshold in multiple dimensions.
[0051] For answers such as D (conceptual error) and E (off-topic), the system analyzes based on the above indicators and obtains lower scores in the content dimension, structure dimension, language dimension, and expression dimension, which can trigger stronger subsequent correction and regression strategies.
[0052] In this embodiment, a corresponding weight is set for the score of each dimension. The scoring dimension mapping submodule uses the weight and the score of each dimension to calculate the comprehensive score, and compares the comprehensive score with the target threshold T1. If the comprehensive score is higher than or equal to the target threshold T1, the summary or extension task can be entered. When the comprehensive score is lower than the target threshold T1, it means that the score of at least one dimension is lower than the corresponding scoring threshold, and targeted corrective follow-up questions need to be generated in the next round.
[0053] In one implementation, assessing semantic deviation may include the following aspects: Topic coverage: checking whether the answer contains the key concepts and relationships that should be present in the question; Expression consistency: checking whether the cause and effect and order of the answer are consistent with the expected path of the target topic; Intent and literal matching: checking whether the answer contains a large number of nouns, scenes, or tasks unrelated to the topic. These signals can be implemented using equivalent techniques such as semantic embedding similarity, keyword / phrase matching, edit distance, BM25, or a weighted combination of rule features, etc., and this application is not limited to any specific implementation method.
[0054] In one implementation, step S300, through the interjection hosting and regression submodule, analyzes the semantic deviation of the response content from the target topic, and may include steps S310-S350: S310. The natural language processing models in the interjection hosting and regression sub-modules generate the first vector of the answer content and the second vector of the guiding question, respectively, and calculate the semantic embedding similarity between the first vector and the second vector to reflect the degree of similarity between the two texts in overall meaning.
[0055] S320. Through the interjection hosting and regression sub-modules, determine the keyword matching score and edit distance between the keywords of the answer content and the topic keywords corresponding to the target topic.
[0056] Optionally, topic keywords can be pre-configured. For example, when the target topic is "rainfall formation process," topic keywords can include, but are not limited to, "evaporation," "precipitation," "runoff / infiltration," and "return to the ocean." The system analyzes whether these keywords appear in the responses, their frequency, and their location, and calculates the keyword matching score and edit distance accordingly.
[0057] S330. Through the interjection hosting and regression sub-module, the BM25 score is determined based on the answer content and the topic keywords corresponding to the target topic. This score is used to measure the relevance of the answer content to the target topic in terms of retrieval significance.
[0058] S340. Normalize the semantic embedding similarity, keyword matching score, edit distance, and BM25 score respectively to obtain the normalized semantic embedding similarity, normalized keyword matching score, normalized edit distance, and normalized BM25 score.
[0059] This allows metrics from different sources to be compared and combined on the same scale. Among the normalized metrics, higher semantic embedding similarity, keyword matching score, and BM25 score indicate closer proximity to the target topic; higher edit distance indicates more significant deviation.
[0060] S350. Determine the preset values for comparison respectively, and calculate the first difference between each preset value and the normalized semantic embedding similarity, the second difference between each preset value and the normalized keyword matching score, and the third difference between each preset value and the normalized BM25 score. Take one of the first difference, the second difference, the third difference and the normalized edit distance as the semantic deviation, or use at least two of the first difference, the second difference, the third difference and the normalized edit distance to calculate the comprehensive semantic deviation.
[0061] In addition to the normalization results, to facilitate the representation of the "degree of deviation", a certain upper limit value can be used as a benchmark. For example, "complete similarity or complete match" can be regarded as the benchmark value. Then, several "difference features" are obtained by "subtracting the actual score from the benchmark value". These difference features, together with the normalized edit distance, serve as candidate features for semantic deviation.
[0062] In one implementation, one of the features (e.g., the difference based on semantic embedding, or the normalized edit distance) can be directly selected as the semantic deviation.
[0063] In another implementation, the aforementioned multiple features can be input together into the comprehensive calculation module, and then fused to obtain a comprehensive semantic deviation. The comprehensive calculation module can adopt any of the following methods: 1. Weighted Summation Method: Each feature is assigned a weight, reflecting its importance in deviation judgment. The system sums the weights of all features to obtain a single deviation value; the larger the value, the more severe the deviation. The weights can be set empirically or optimized based on historical samples.
[0064] 2. Logistic Regression Approach: Using features as input, the logistic regression model learns a mapping from features to deviation values. This model first performs a linear combination of multiple features, then uses an sigmoid compression function to map the result to the 0-1 range. The closer the output value is to 1, the greater the deviation. Model parameters can be obtained offline through training on datasets labeled with "whether it's off-topic" or "degree of deviation."
[0065] 3. Feedforward Neural Network Approach: This approach uses a vector composed of various features as input and employs a feedforward neural network with at least one hidden layer for nonlinear mapping. The network outputs a deviation value in the range of 0 to 1. This network can include several hidden units and commonly used activation functions, and is trained offline on samples labeled with the deviation level using the backpropagation algorithm.
[0066] Finally, the obtained semantic deviation is compared with the deviation threshold T2: if the semantic deviation is greater than the deviation threshold T2, it means that the current answer deviates significantly from the target topic and needs to be guided by questions and answered again in the next round; when the semantic deviation is less than or equal to the deviation threshold T2, it means that the answer is still centered around the target topic and can no longer trigger regression based on "off-topic", but can enter the summary or other subsequent tasks.
[0067] In one implementation, in step S400, the upper limit N of the number of rounds can be preset according to the teaching scenario. When the semantic deviation is greater than the deviation threshold T2, or the comprehensive score is lower than the target threshold T1, the regression counter increments the regression count C by one. Within rounds where C≤N, corrective feedback content is generated as a new guiding question, and the process proceeds to the next round. Then, step S100 is returned to receive the learner's new answer to the new guiding question, thus forming a closed loop of questioning, answering, feedback, and re-questioning. The corrective feedback content includes: corrective follow-up questions to improve the comprehensive score, or regression prompts to indicate the cause of semantic deviation and switch the conversation topic of the answer back to the target topic.
[0068] Optionally, in step S400, the regression count C is accumulated using a regression counter, and corrective feedback is generated when the regression count C is less than or equal to the upper limit of the number of rounds N. This can be further subdivided into steps S410-S420: S410. Compare the content dimension score, structure dimension score, language dimension score, and expression dimension score with the corresponding score thresholds to determine the target dimensions that are below the score thresholds.
[0069] For example, when the overall score is lower than the target threshold T1, the system checks the scores of the four dimensions in turn. If the score of one or more dimensions is lower than their respective score thresholds, these dimensions are marked as "target dimensions".
[0070] S420. Based on the preset task mapping relationship between the target dimension and the scoring interval, trigger at least one strategic task to generate targeted corrective follow-up questions corresponding to the target dimension.
[0071] Specifically, the system can pre-configure different strategy tasks for different scoring ranges of each dimension: when the score of the "content" dimension is lower than the corresponding scoring threshold, a question generation task is triggered to supplement key concepts or evidence; when the score of the "structure" dimension is lower than the corresponding scoring threshold, a semantic hierarchical decomposition task is triggered to break down the original question into several sub-questions that can be completed sequentially; when the score of the "language" dimension is lower than the corresponding scoring threshold, an information density restatement task is triggered to deduplicate and disambiguate the original answer and provide a more standardized example expression; when the score of the "expression" dimension is lower than the corresponding scoring threshold, a hierarchical summary output task is triggered to generate a structured answer for readers as a reference template for learners to imitate.
[0072] In this embodiment, when the semantic deviation exceeds the deviation threshold T2, the interruption hosting and regression submodule will also analyze the cause of the deviation by combining multiple signals. The system can select indicators with larger value features or lower similarity / matching degree from the normalized semantic embedding similarity, keyword matching score, edit distance, and BM25 score as the main source of deviation, and generate a deviation cause explanation and regression prompt accordingly.
[0073] For example, when the main problem is the absence of a key step, the missing step is assigned to the "missing concept" category; when the main problem is the presence of a large number of words unrelated to the topic, it is attributed to the "deviation from keywords" and "topic differences" categories.
[0074] Optionally, regression suggestions may include at least one of the following three categories: missing concept suggestions, off-keyword suggestions, and topic-difference summary suggestions. For example, if the target topic is "rainfall" or "water cycle": 1. Missing concept hints: The standard answer should include key steps such as "evaporation", "rising, cooling, and condensing into clouds", "forming precipitation", and "runoff / infiltration back into the ocean".
[0075] If a learner answers "Seawater evaporates to form clouds and precipitation, and the water returns to the Earth's surface," the system, after comparison, finds that it lacks two key points: "How does surface water return to the ocean?" and "Solar radiation is the energy source." It will then generate a prompt similar to: "You have already described up to the 'precipitation' step. To form a complete closed loop, two more points can be added: first, how surface water returns to the ocean through runoff or infiltration; and second, how solar radiation provides energy for evaporation. Please add these two points and connect them into a complete process in one or two sentences." 2. Keyword deviation warning: In the current round of dialogue, the system will maintain a set of keywords that are highly related to the target topic (such as "water vapor", "cooling", "clouds", "precipitation", "runoff", "infiltration", etc.).
[0076] If a learner answers with "It will get colder next week, wear more clothes, and be careful on slippery roads in the rain," and the system detects that the answer contains a large number of words unrelated to the target topic, such as "wearing clothes" and "travel safety," while lacking core words like "evaporation" and "condensation," it will give a prompt similar to: "Current responses focus more on clothing and travel safety, which deviates from the task of 'explaining the water cycle / rainfall formation process.' Please return to the scientific process itself and explain in sequence 'evaporation—condensation—precipitation—runoff / infiltration—return to the ocean,' and point out the role of solar radiation in it." 3. Summary of Topic Differences: When overall semantic clustering shows that the current answer mainly revolves around life advice, while the reference points revolve around physical or geographical processes, the system will summarize this difference in concise natural language, for example: "Your answer so far focuses on life tips and hasn't explained the whole process of water evaporating from the sea surface, condensing into clouds in the air, and eventually returning to the ocean. Next, please try to explain this natural process in chronological order." If the answer fails to return to the target topic and meet the specified quality requirements within the preset number of rounds (i.e., the number of regressions C does not exceed N), the system will automatically switch back to the main task, re-explain the learning objectives for this round to the learner, and provide a structured example answer in an outline format so that the learner can answer again in the new round.
[0077] In this embodiment, the criteria for "returning to the target topic" include: covering the core concepts and key relationships corresponding to the question, forming the required causal chain or closed-loop process, and reaching a level that can be summarized into a clear and structured answer. When these conditions are met, the system enters the summarization and output stage; otherwise, it will continue to use strategies such as targeted follow-up questions, hierarchical decomposition, information restatement, and regression prompts until the conditions are met or the maximum number of rounds is reached.
[0078] In one embodiment, the artificial intelligence-based closed-loop learning control method of this application may further include step S510: S510. Optionally, when C > N, switch back to the main task to redirect the learner to the learning objective of the target topic, ensuring that the learning process does not get out of control due to continuous deviation, and reset the number of regressions C to zero.
[0079] In one implementation, the artificial intelligence-based closed-loop learning control method of this application embodiment may further include step S520: S520. When the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is much higher than or equal to the pass threshold T1, output a structured summary based on the learner's final answer.
[0080] Optionally, when the semantic deviation is less than or equal to the deviation threshold T2 and the overall score is higher than or equal to the achievement threshold T1, it indicates that the learner's overall answer meets the target requirements. In this case, the answer analysis and feedback submodule can first select features higher than the corresponding threshold as "advantageous items" based on the scores of the four dimensions of content, structure, language, and expression, and select features close to the threshold but still with room for improvement as "suggested items," forming an intermediate results list. Then, the scoring dimension mapping submodule, according to a pre-set template or summary generation model, transforms these advantageous and suggested items into structured summary text in natural language form.
[0081] A structured summary should include at least: key concepts and correct steps covered in the content dimension; the causal order or completeness of steps demonstrated in the structural dimension; the standardization of vocabulary in the language dimension; and the logical flow and comprehensibility in the expression dimension. It should also provide targeted suggestions, such as hints for details that can be added or suggestions for improved transitions. The generated structured summary can serve as a learning report excerpt presented to learners and parents, or as a reference for subsequent task scheduling.
[0082] In this embodiment, the target threshold T1, the deviation threshold T2, and the round limit N can be adaptively updated based on historical performance and dialogue stability.
[0083] Optionally, the system maintains statistical records of the most recent W rounds of question-and-answer interactions in the background. These records include at least the overall score, semantic deviation, and whether there were any obvious off-topic or premature interruptions in each round. Dialogue stability can be understood as follows: in the most recent W rounds of interaction, the fluctuation range of the overall score and semantic deviation is small, and the proportion of rounds marked as "obvious off-topic" or "abnormally interrupted" is lower than a preset proportion.
[0084] In one implementation, adaptive updates can be based on sliding window estimation: the system periodically takes the comprehensive score of the most recent W rounds of interaction, calculates the mean and median or quantile of the score, and compares it with the current threshold T1; when the comprehensive score of the most recent W rounds is consistently higher than the current T1 and the dialogue is stable, T1 is slightly increased from its original level; when the comprehensive score of the most recent W rounds is close to or lower than T1, or frequently approaches the upper limit N of rounds before returning to the target topic, T1 is appropriately reduced or the upper limit N of rounds is increased, so that the system can give a more lenient evaluation and more opportunities for correction at the current stage.
[0085] Similarly, the deviation threshold T2 can also be adjusted according to the moving average, median, or quantile of the semantic deviation within the sliding window: for example, when the semantic deviation in the most recent W rounds is generally small and the fluctuations converge, T2 can be slightly reduced to more sensitively detect new deviations; when moderate deviations frequently occur in the most recent W rounds but can be brought back to the topic with a few follow-up questions, T2 can be appropriately increased to avoid over-intervention in minor interruptions.
[0086] In another implementation, adaptive updates can also use exponential smoothing, which assigns higher weights to recent rounds of interaction and lower weights to earlier rounds, so that the threshold adjustment can reflect the recent learning state without changing frequently due to accidental fluctuations.
[0087] Reference Figure 2 The diagram illustrates a structural block diagram of an artificial intelligence-based closed-loop learning control device according to an embodiment of this application. The device may include: The question generation submodule is used to generate guiding questions based on the target topic and receive learners' answers to the guiding questions. The answer analysis and feedback submodule is used to perform semantic analysis on the answer content and calculate the content dimension score, structure dimension score, language dimension score, and expression dimension score respectively. The scoring dimension mapping submodule calculates the comprehensive score based on the above four dimensions and compares the comprehensive score with the pass threshold T1. The interruption hosting and regression submodule is used to analyze the semantic deviation of the answer content relative to the target topic and compare the semantic deviation with the deviation threshold T2; The scoring dimension mapping submodule is used to accumulate the number of regressions C by a regression counter when the semantic deviation is greater than the deviation threshold T2 or the comprehensive score is lower than the target threshold T1. If the number of regressions C is less than or equal to the upper limit of rounds N, corrective feedback content is generated and used as a new guiding question in the next round. The steps of receiving the learner's answer to the guiding question are returned to form a closed loop of questioning, answering, feedback, and re-questioning, until the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the target threshold T1, or the number of regressions C is greater than the upper limit of rounds N. The corrective feedback includes: corrective follow-up questions to improve the overall score, or regression prompts to indicate the reasons for semantic deviation and switch the conversation topic of the answer back to the target topic.
[0088] The functions of each module in the device of this application embodiment can be found in the corresponding description in the above method, and will not be repeated here.
[0089] The apparatus in this embodiment further includes a strategy task library and an output module, used for at least one of the following: When the semantic deviation exceeds the deviation threshold T2, a reverse question generation task is invoked to generate explanatory questions about the missing key concepts in the target topic. Specifically, the strategy task library can first compare the standard key points list corresponding to the target topic with the key points already covered in the current answer to obtain the missing concepts; then, the missing concepts are filled into a preset question template to form follow-up questions. For example, if the target topic is "water cycle", and "infiltration / runoff returns to the ocean" is not mentioned, an explanatory question is generated: "You have not explained how surface water returns to the ocean through infiltration or runoff. Please supplement this part."
[0090] When the content dimension score and / or structure dimension score are lower than the corresponding score threshold, the semantic level decomposition task is invoked to generate guiding sub-questions corresponding to the content dimension and / or structure dimension. Specifically, the strategy task library can break down the target topic into several sub-steps according to causal or chronological order, and combine the completed and incomplete parts in the current answer to generate step-by-step guiding questions, such as "Step 1: Explain the evaporation process," "Step 2: Explain the condensation to form clouds," and "Step 3: Explain the destination of precipitation," etc., so that learners can complete the overall structure by following the sub-steps.
[0091] When the language dimension score and / or expression dimension score are lower than the corresponding score threshold, the information density restatement task and / or hierarchical summary output task are invoked.
[0092] Specifically, the information density restatement task can deduplicate, de-colloquialize, and merge sentences in learners' original answers, condensing fragmented expressions into well-organized short paragraphs, and outputting them in the form of "system demonstration + learner restatement".
[0093] When the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the pass threshold T1, the output module calls the hierarchical summary output task based on the learner's final answer and outputs a structured summary.
[0094] Obtain information related to the learner's learning status, determine the learner's learning status based on this information, and then invoke the learning status detection task to adjust the pace and difficulty accordingly.
[0095] Learning status information can be inferred from interactive behavior characteristics, which include at least one of the following: pause duration, number of error corrections, answer duration, and number of repeated questions; when conditions permit, peripheral sensor data can also be used as an optional signal source.
[0096] Based on learners' historical memory data, a memory scheduling task is invoked to arrange review. Specifically, the memory scheduling task can record the time when each knowledge point was last answered correctly and its historical mastery, select knowledge points that need to be reviewed using a preset interval repetition strategy, and insert short review questions at appropriate rounds to consolidate long-term memory.
[0097] The process triggers a relation extraction and comparison task, feeding the extraction results back into the question generation submodule and / or the strategy task library. In one implementation, the extraction results are also fed back into the knowledge map generation submodule and / or the personalized learning path submodule in the STEM extension module for generating new guiding questions in the next round. Specifically, the relation extraction and comparison task extracts semantic triples of "concept-relation-concept" from learners' answers and compares them with the relations in the standard knowledge map, marking missing or incorrect connections; then, the comparison results are written into the learner's personalized knowledge map, updating the mastery level and prerequisite dependencies of each node. Based on the updated knowledge map, the question generation submodule will prioritize generating new guiding questions around nodes with "missing or confused relations" in the next round, achieving personalized reinforcement.
[0098] For example, when a learner is detected frequently confusing the relationship between "cause" and "effect" in multiple rounds of responses, the relation extraction and comparison task records this phenomenon on the corresponding edge of the knowledge map and triggers a dedicated relation comparison exercise. This could involve providing several candidate statements and asking the learner to determine "which are causes and which are effects," after which the system provides an analysis. This task-linked approach and knowledge map feedback improves the coverage of subsequent supplementary questions and helps the learning path gradually converge towards the target topic through multiple rounds of interaction.
[0099] In a series of controlled experiments and scenario simulations, it can be observed that compared with traditional learning systems that do not employ closed-loop control and policy task linkage, the system of this application embodiment exhibits superior performance in terms of learning topic regression rate, irrelevant interruption control, and learning summary generation efficiency. It shows trends such as learning topics being more easily recalled, irrelevant dialogues being significantly reduced, and structured summary generation being faster. The relevant numerical values are only used to illustrate the direction of improvement in technical effects and do not limit the specific range of indicators. The above results are based on system logs and test statistics and can be verified by comparing experimental logs and offline playback simulations.
[0100] The device in this embodiment further includes a STEM extension module, which comprises a knowledge map generation submodule, a hierarchical growth assessment submodule, and a personalized learning path submodule. The knowledge map generation submodule generates a knowledge map, and its output is used to update the parameters of the question generation submodule and the personalized learning path submodule. The hierarchical growth assessment submodule calculates the learning level according to the understanding layer, application layer, and innovation layer. The personalized learning path submodule updates the question granularity and prerequisite dependencies based on the knowledge map. The knowledge map generation submodule generates knowledge maps, and its output is used to feed back parameters to the question generation submodule and the personalized learning path submodule. Specifically, based on learners' multi-round responses and a pre-built subject knowledge base, the knowledge map generation submodule analyzes the semantic relationships and knowledge point dependencies in the response text. It uses the identified target topics, key concepts, and intermediate concepts as nodes, and relationships such as causality, inclusion, succession, and comparison as edges, generating a mind map represented by nodes and connections. For example, under the topic of "water cycle," nodes may include "evaporation," "condensation," "precipitation," "surface runoff," "ground infiltration," "returning to the ocean," and "solar radiation," while edges are used to represent causal chains such as "solar radiation → evaporation," "evaporation → condensation," "condensation → precipitation," and "precipitation → runoff / infiltration → returning to the ocean."
[0101] The knowledge map generation submodule also assigns a proficiency value to each node, representing the learner's mastery of the knowledge point. The proficiency value can be updated by weighting the score results of the current round in the relevant tasks of that knowledge point with the historical proficiency. That is, the "performance in this round" and the "historical average performance" are smoothly integrated according to the preset smoothing coefficient A, so as to avoid drastic fluctuations in proficiency caused by a single abnormal answer.
[0102] Optionally, the knowledge map generation submodule can generate knowledge maps and update proficiency levels according to the following process: (1) Node extraction: Concept identification is performed on learners' multi-round responses and pre-set subject knowledge base to extract key concepts under the target topic as nodes. Concept identification can use at least one of keyword / thesaurus matching, entity recognition model, and concept normalization based on vector similarity to unify expressions such as "evaporation / water evaporation / seawater evaporation" into the same concept node.
[0103] (2) Relation extraction: Extracting causal, inclusion, contrast, succession, and other relationships between concepts from the response text. Relation extraction can use at least one of the following: rule templates (such as “because…therefore…” corresponding to causal relationships), dependency syntax features, and relation extraction models.
[0104] (3) Mind map construction and layer overlay: Concepts are represented by nodes and relationships are represented by edges to generate directed graphs or multi-relationship graphs. At the same time, the concept mapping method can be extended to overlay logical relationship layers, comparison relationship layers and inclusion relationship layers on the mind map to enhance the structure and interpretability of the mind map.
[0105] For example, “convective rain / frontal rain / orographic rain” can be treated as subclasses of “precipitation” and connected through inclusion relationships; “correct examples” and “common misunderstandings” can be connected through comparison relationships; and “necessary conditions / sufficient and sufficient conditions” can be connected through logical relationships, thereby forming a multi-level conceptual structure and causal tracing path.
[0106] (4) Proficiency update and reinjection: The performance of the current round of answers on relevant node tasks is mapped to node observation values, and the new proficiency is smoothly integrated with the historical proficiency. The generated knowledge map and the proficiency of each node are used as input parameters for the subsequent question generation submodule and personalized learning path submodule, which are used to select the question focus, determine the split granularity, and control the order of first patching.
[0107] For example, under the theme of "water cycle", nodes can include "evaporation", "condensation", "precipitation", "surface runoff", "ground infiltration", "return to the ocean", "solar radiation", etc., and edges can represent relationships such as "solar radiation → provide energy → evaporation", "evaporation → condensation", "condensation → precipitation", "precipitation → runoff / infiltration → return to the ocean". If learners consistently omit the "return to the ocean" path in multiple rounds of responses, their proficiency in that node will remain below the threshold. After retraining, the system will prioritize generating supplementary guidance questions around that node.
[0108] like Figure 3 As shown, the hierarchical growth assessment submodule is used to calculate the learning level according to the understanding, application, and innovation levels, and to plan personalized learning paths accordingly. Specifically, the hierarchical growth assessment submodule categorizes historical dialogues and practice results into "understanding tasks," "application tasks," and "innovation / expansion tasks" based on task tags. The comprehension level focuses on whether learners can explain concepts and key relationships completely and accurately in their own words, such as "using continuous sentences to explain the whole process of water evaporating and returning to the ocean"; The application layer mainly focuses on whether learners can transfer and use learned concepts in new situations, such as "analogizing the water cycle to local rainfall and river replenishment phenomena" and "explaining the causes of rainstorms by combining real-life examples." The innovation layer focuses on whether learners can raise extended questions or rewrite the context, such as "designing a small experiment to verify the influencing factors of evaporation rate" or "proposing hypotheses to improve the existing context setting."
[0109] Optionally, the hierarchical growth assessment submodule can calculate the performance of each level and determine the dominant learning level in the following manner: (1) Task labeling: The guiding questions and learner answers in each round of interaction are labeled. The labels include at least the task type (understanding / application / innovation), the set of target knowledge points, and the difficulty and openness information. The labels can be obtained from the template library metadata or automatically identified by the classification model.
[0110] (2) Indicator statistics: Within each level, the stability indicators such as the number of times the target was met, the length of consecutive target met, and the range of fluctuations in the most recent W rounds or the most recent time window are statistically analyzed and given higher weight in combination with the performance of the most recent rounds to reflect the recent status.
[0111] (3) Level determination: Based on the above indicators, calculate the level results of the understanding layer, application layer and innovation layer, and determine the dominant learning level of the current topic according to the preset rules; for example, when the understanding layer is stably up to standard but the application layer is not stably up to standard, it is determined as "understanding consolidation stage"; when the application layer is stably up to standard and remains stable in open questions, it is determined as "application consolidation stage"; when reasonable hypotheses or verification steps can be proposed in the innovation layer tasks, it is determined as "innovation expansion stage".
[0112] The results at this level will be fed back to the personalized learning path submodule to determine whether subsequent tasks should focus on supplementing understanding or increasing cross-contextual application and open-ended inquiry.
[0113] The personalized learning path submodule is used to update the granularity of questions and prerequisite dependencies based on the knowledge map. Here, "question granularity" describes the scope of knowledge and the number of reasoning steps covered by a single guided question: coarser-granular questions focus on the complete process or comprehensive explanation, such as "Please explain a complete water cycle from beginning to end"; finer-granular questions focus on a specific sub-node or single step, such as "Explain only the destination of surface runoff after precipitation" or "Explain the role of solar radiation in the water cycle".
[0114] Optionally, the personalized learning path submodule can quantify the granularity of questions to include at least one of the following: the number of covered nodes, the number of reasoning steps, and the number of sub-questions. The granularity can be dynamically selected based on the proficiency of each node in the knowledge map and the results of hierarchical growth assessment: for nodes with lower proficiency or unstable understanding, questions with finer granularity and clearer step breakdowns are prioritized; for nodes with higher proficiency and that have entered the application / innovation layer, questions with coarser granularity and more open-ended contexts are arranged.
[0115] "Prerequisite dependencies" are used to indicate the prerequisites for a certain node in the knowledge map in the learning sequence. For example, "understanding the evaporation mechanism" is a prerequisite for "explaining the complete water cycle," and "mastering fraction addition and subtraction" is a prerequisite for "solving fraction word problems." Prerequisite dependencies can be pre-marked by subject matter experts or automatically supplemented by the system based on the course syllabus, knowledge base structure, and historical learning paths.
[0116] When generating the question sequence for the next stage, the personalized learning path submodule first checks whether the prerequisite nodes corresponding to the target node have reached the set proficiency threshold. If there are insufficient prerequisite nodes, guiding questions that review and consolidate the prerequisite knowledge are inserted first. When the prerequisite nodes meet the requirements, higher-level, coarser-grained comprehensive tasks are then arranged. By dynamically adjusting the granularity of questions and prerequisite dependencies, the system can generate differentiated learning paths for learners with different foundations and paces, ensuring that the target topic is not "skipped" while avoiding repetitive practice.
[0117] The device in this embodiment further includes an output module for receiving content generated by other modules and outputting it to 100 learner terminals for display; the output module supports local storage and clearing functions, and by default only saves learning records locally.
[0118] Optionally, when a user triggers the clearing function (e.g., one-click clearing), the output module can execute a clearing process that includes at least the following steps: stopping read, write, and index services related to learning records; deleting local learning record data files and index files; erasing or destroying key materials used to decrypt learning records and triggering key rotation so that historical keys can no longer be used to decrypt stored data; recording a clearing operation log for local auditing and display (excluding content that can be restored to its original form) to reduce the risk of data residue; the output module also supports controlled output of learning reports in shareable formats for results display and parental supervision.
[0119] The methods and apparatus of this application can achieve the following technical effects: (1) Forming a complete closed-loop control: Through multiple rounds of questioning, answering, feedback, and re-questioning, the learning process is dynamically adjusted and converged. For example, when the comprehensive score is lower than the threshold T1, the system automatically enters the next round of targeted follow-up questions; when the semantic deviation is greater than the deviation threshold T2, the system prioritizes generating regression prompts to bring the conversation topic back to the target task.
[0120] (2) It has an interruption and regression mechanism: it detects deviations from the topic and automatically returns to the original task within a preset number of rounds N (accumulated by counter C) to maintain learning focus. Specifically, when the semantic deviation is detected to be greater than the deviation threshold T2 for several consecutive rounds, the device triggers a regression prompt task, generates a combination of feedback of "pointing out the reason for the deviation + giving a prompt to return to the target chain", and switches back to the main task when C>N to reiterate the target topic and basic outline.
[0121] (3) Multi-dimensional feedback evaluation: Comprehensive analysis of four dimensions: content, structure, language, and expression, and form dimensional feedback. For example, when the score of the structure dimension is lower than the corresponding threshold, the system can point out the causal chain or steps that need to be filled in the feedback; when the score of the language dimension is lower than the corresponding threshold, it can provide a standard expression example and a disambiguated expression template.
[0122] (4) Introducing multi-strategy task scheduling: The system triggers algorithmic tasks such as reverse question generation, semantic hierarchy decomposition, information density restatement, hierarchical summary output, learning state detection, memory scheduling, relation extraction and comparison according to the "score-task mapping" to enhance the learning depth. For example, when the content dimension score is lower than the corresponding threshold, the reverse question generation task outputs follow-up questions such as "supplement key concepts / supplement evidence"; when the structure score is lower than the corresponding threshold, the semantic hierarchy decomposition task breaks down the original question into several sub-steps and guides the answer step by step; when the expression dimension score is lower than the corresponding threshold, the information density restatement task generates a more compact example expression based on the current answer for comparison and modification; when the comprehensive score reaches the threshold T1, the hierarchical summary output task generates a structured summary that can be used for report presentation based on multiple rounds of answers.
[0123] (5) STEM Expansion and Knowledge Map Visualization: Realize the generation of interdisciplinary knowledge maps, hierarchical assessment and personalized learning path recommendation, and feed the map output back to the question generation and personalized learning path (133) sub-module to adjust the question granularity and path parameters. For example, in the topic of "water cycle", the knowledge map generation sub-module encodes the concept nodes such as "evaporation → condensation → precipitation → runoff / infiltration → return to the ocean" and their relationships into a directed graph, and maintains a proficiency value for each node; the hierarchical growth assessment sub-module determines the dominant learning level based on the node according to the task label and stability statistics; the personalized learning path sub-module decides whether to continue to deepen the application tasks in the same chain or to expand to interdisciplinary topics such as energy conservation and climate change.
[0124] (6) Privacy-friendly and highly portable: Learning records can be stored locally only and support deletion and report export; edge encryption is used to protect the security of learning records on the terminal. In one implementation, edge encryption can adopt a hierarchical mechanism of session key + master key: each learning session generates a session key to encrypt incremental session data, and the master key is used to encrypt the session key; when the user triggers "offline erasure" or the device is lost and reported, the historical key is no longer usable through key rotation, revocation and expiration marking, so that the local record quickly becomes invalid in the sense of decryption, thereby improving controllability.
[0125] (7) The AI learning engine can combine various learning control algorithm tasks to enhance intelligence and adaptability, including but not limited to: Metacognitive monitoring task: Based on features such as answer time, confidence assessment, and error type, the system determines comprehension deviations and adaptively adjusts the feedback pace. Specifically, the system maintains baseline answer time and accuracy statistics for each question type. When patterns such as "answer time significantly shorter than the baseline and incorrect answers" are detected, a prompt is triggered requiring learners to supplement their reasoning steps. When "answer time significantly longer than the baseline and repeated corrections are detected," the system triggers a reduction in granularity or switches to a concept review task.
[0126] Retrieval Practice Task: Based on the forgetting curve, knowledge points are scheduled over time, and recall questions are generated periodically to strengthen long-term memory. For example, the results of the most recent recall for each knowledge point are recorded, and the risk of forgetting is estimated. When the risk exceeds a threshold, a retrieval question that "does not rely on external materials" is inserted, and the next review interval is updated based on the answers.
[0127] Alternating training tasks: Problems are generated alternately across different themes / disciplines to enhance transfer and generalization abilities. For example, after explaining the "water cycle" process, the system alternately presents application problems related to "climate zones" and "urban drainage design"; the task scheduler controls the alternation frequency based on historical performance and learning load.
[0128] Interval repetition scheduling: The system automatically calculates review intervals based on memory scheduling strategies, achieving adaptive repetition over time. It estimates memory strength based on the correctness and fluency of each answer and maps the next review time accordingly, automatically inserting relevant review questions.
[0129] Concept mapping / relationship extraction task: Extract entities and causal / inclusion / contrast relationships from learner responses, and add logical, causal, and inclusion relationships to the knowledge map to enhance hierarchy and interpretability. For example, when answering "Why is precipitation more likely to form in mountainous areas," the system extracts concepts such as "orographic uplift," "condensation nucleus," and "updraft" and their relationship chains, attaches them to the "precipitation formation" node, updates the knowledge map, and serves as the basis for subsequent transfer tasks.
[0130] Reverse question generation task: Combining mind maps and scoring results, learners are guided to restate key concepts in the role of an explainer, which is used to test the depth of understanding and logical consistency. For example, when the system determines that the learner has covered the main link of the "water cycle", it generates a reverse question, "Please explain the key processes of the water cycle in three steps", and corrects it according to whether the restatement covers the key nodes and their logical order.
[0131] In one implementation, the aforementioned end-side encryption may include a key rotation and offline erasure triggering mechanism, which is used to quickly invalidate local learning records when the terminal is lost or the user actively triggers privacy clearing.
[0132] Reference Figure 4The diagram illustrates a structural block diagram of an electronic device according to an embodiment of this application. The electronic device includes a memory 310 and a processor 320. The memory 310 stores instructions that can be executed on the processor 320. The processor 320 loads and executes these instructions to implement the artificial intelligence-based closed-loop learning control method described in the above embodiment. The number of memories 310 and processors 320 can be one or more.
[0133] In one embodiment, the electronic device further includes a communication interface 330 for communicating with external devices and exchanging data. If the memory 310, processor 320, and communication interface 330 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0134] Optionally, in a specific implementation, if the memory 310, processor 320 and communication interface 330 are integrated on a single chip, the memory 310, processor 320 and communication interface 330 can communicate with each other through an internal interface.
[0135] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the artificial intelligence-based closed-loop learning control method provided in the above embodiments.
[0136] This application also provides a chip including a processor for calling and executing instructions stored in a memory, causing a communication device with the chip installed to perform the method provided in this application.
[0137] This application also provides a chip. In one embodiment, the chip further includes an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in this application.
[0138] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting the Advanced Reduced Instruction Set Computing (RISC) machine (ARM) architecture.
[0139] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM).
[0140] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in this application can be implemented, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
[0141] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0142] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0143] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.
[0144] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).
[0145] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, performs any step of the above method embodiments or any combination thereof.
[0146] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.
[0147] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A closed-loop learning control method based on artificial intelligence, characterized in that, include: The question generation submodule generates guiding questions based on the target topic and receives learners' answers to the guiding questions. The answer analysis and feedback submodule performs semantic analysis on the answer content and calculates content dimension score, structure dimension score, language dimension score and expression dimension score respectively. The scoring dimension mapping submodule calculates a comprehensive score based on the above four dimension scores and compares the comprehensive score with the pass threshold T1. The semantic deviation of the answer content relative to the target topic is analyzed through the interruption hosting and regression submodule, and the semantic deviation is compared with the deviation threshold T2. When the semantic deviation is greater than the deviation threshold T2 or the comprehensive score is lower than the achievement threshold T1, the regression count C is accumulated by the regression counter. If the regression count C is less than or equal to the round limit N, correction feedback content is generated and used as a new guiding question in the next round. The process returns to the learner's answer to the guiding question to form a closed loop of questioning, answering, feedback, and re-questioning until the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the achievement threshold T1, or the regression count C is greater than the round limit N. The corrective feedback includes: corrective follow-up questions to improve the overall score, or regression prompts to indicate the cause of semantic deviation and switch the conversation topic of the answer back to the target topic.
2. The method according to claim 1, characterized in that: The analysis of the semantic deviation of the response content relative to the target topic through the interruption hosting and regression submodule includes: The first vector of the answer content and the second vector of the guiding question are generated by the natural language processing model in the interruption hosting and regression submodules, and the semantic embedding similarity between the first vector and the second vector is calculated. The keyword matching score and edit distance between the keywords of the answer content and the topic keywords corresponding to the target topic are determined through the interruption hosting and regression sub-modules. The BM25 score is determined by using the interruption hosting and regression submodules based on the content of the answer and the topic keywords corresponding to the target topic. The semantic embedding similarity, the keyword matching score, the edit distance, and the BM25 score are normalized respectively to obtain normalized semantic embedding similarity, normalized keyword matching score, normalized edit distance, and normalized BM25 score. Each preset value used for comparison is determined, and a first difference between each preset value and the normalized semantic embedding similarity, a second difference between each preset value and the normalized keyword matching score, and a third difference between each preset value and the normalized BM25 score are calculated. One of the first difference, the second difference, the third difference, and the normalized edit distance is used as the semantic deviation, or at least two of the first difference, the second difference, the third difference, and the normalized edit distance are used to calculate a comprehensive semantic deviation. The comprehensive calculation includes any one of weighted summation, logistic regression, and neural networks.
3. The method according to claim 1, characterized in that: The regression hints include at least one of missing concepts, off-keywords, and topic difference summaries.
4. The method according to claim 1, characterized in that: The method further includes: When the number of regressions C exceeds the maximum number of rounds N, switch back to the main task, redirect the learner to the learning objective of the target topic, and reset the number of regressions C to zero.
5. The method according to claim 1, characterized in that: The process involves calculating the regression count C using a regression counter, generating corrective feedback content within rounds where C ≤ N, and using this corrective feedback content as a new guiding question, including: The scores for the four dimensions—content, structure, language, and expression—are compared with their corresponding scoring thresholds to identify the target dimensions that are below the thresholds. Based on the preset task mapping relationship between the target dimension and the scoring interval, at least one strategic task is triggered to generate targeted corrective follow-up questions corresponding to the target dimension.
6. The method according to claim 1, characterized in that: The method further includes: When the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the achievement threshold T1, a structured summary containing strengths and suggestions for optimization is output based on the learner's final answer.
7. The method according to claim 1, characterized in that: The attainment threshold T1, the deviation threshold T2, and the round limit N are adaptively updated based on historical performance and dialogue stability; wherein the dialogue stability includes at least one of regression count, round score fluctuation, or round duration fluctuation; and the adaptive update method includes at least one of exponential smoothing or sliding window estimation.
8. A closed-loop learning control device based on artificial intelligence, characterized in that, include: The question generation submodule is used to generate guiding questions based on the target topic and receive learners' answers to the guiding questions. The answer analysis and feedback submodule is used to perform semantic analysis on the answer content and calculate the content dimension score, structure dimension score, language dimension score, and expression dimension score respectively. The scoring dimension mapping submodule calculates the comprehensive score based on the above four dimension scores and compares the comprehensive score with the pass threshold T1. The interruption hosting and regression submodule is used to analyze the semantic deviation of the answer content relative to the target topic and compare the semantic deviation with the deviation threshold T2; The scoring dimension mapping submodule is used to accumulate the number of regressions C by a regression counter when the semantic deviation is greater than the deviation threshold T2 or the comprehensive score is lower than the achievement threshold T1. If the number of regressions C is less than or equal to the upper limit of rounds N, corrective feedback content is generated and the corrective feedback content is used as a new guiding question to enter the next round. The steps of returning the learner's answer to the guiding question are used to form a closed loop of questioning, answering, feedback, and re-questioning until the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the achievement threshold T1, or the number of regressions C is greater than the upper limit of rounds N. The corrective feedback includes: corrective follow-up questions to improve the overall score, or regression prompts to indicate the cause of semantic deviation and switch the conversation topic of the answer back to the target topic.
9. The artificial intelligence-based closed-loop learning control device according to claim 8, characterized in that: It also includes a strategy task library and an output module for at least one of the following: When the semantic deviation is greater than the deviation threshold T2, the reverse question generation task is invoked to generate explanatory questions for the missing key concepts of the target topic; When the content dimension score and / or the structure dimension score are lower than the corresponding score threshold, the semantic hierarchy decomposition task is invoked to generate guiding sub-questions corresponding to the content dimension and / or structure dimension. When the language dimension score and / or the expression dimension score are lower than the corresponding score threshold, the information density restatement task and / or the hierarchical summary output task are invoked. When the semantic deviation is less than or equal to the deviation threshold T2 and the comprehensive score is higher than or equal to the achievement threshold T1, the output module calls the hierarchical summary output task based on the learner's final answer and outputs a structured summary. Obtain information related to the learner's learning status, determine the learner's learning status based on the information related to the learning status, and adjust the pace and difficulty by calling the learning status detection task based on the learning status. Based on the learner's historical memory data, a memory scheduling task is invoked to arrange a review; Trigger the relationship extraction and comparison task, and feed the extraction results back into the question generation submodule and / or the strategy task library for the generation of new guiding questions in the next round.
10. The artificial intelligence-based closed-loop learning control device according to claim 9, characterized in that: It also includes a STEM extension module, which comprises a knowledge map generation submodule, a hierarchical growth assessment submodule, and a personalized learning path submodule. The knowledge map generation submodule generates a knowledge map, and its output is used to update the parameters of the question generation submodule and the personalized learning path submodule. The hierarchical growth assessment submodule calculates the learning level according to the understanding layer, application layer, and innovation layer. The personalized learning path submodule updates the question granularity and prerequisite dependencies based on the knowledge map.
11. The artificial intelligence-based closed-loop learning control device according to claim 8, characterized in that: It also includes an output module that supports local storage and clearing functions, and by default only saves learning records locally.
12. The artificial intelligence-based closed-loop learning control device according to any one of claims 9-11, characterized in that: The output module also supports end-to-end encryption, which is used to encrypt and store learning records generated during dialogue with the learner on the terminal side. The end-to-end encryption includes at least a key rotation and an offline erasure trigger mechanism.
13. The artificial intelligence-based closed-loop learning control device according to claim 9, characterized in that: The strategy task library also includes at least one of the following learning regulation tasks: The retrieval practice task is used to select target knowledge points that need to be reviewed first based on the learner's historical practice records on target knowledge points, and generate practice questions that encourage learners to actively recall the target knowledge points without relying on external materials. The historical practice record shall include at least: the number of correct answers, the number of incorrect answers, and the time of the last practice session; Alternating training tasks are used to alternate learning tasks between different themes or subjects. The proportion of questions is dynamically adjusted according to the correctness of answers and the answering time for each theme or subject, so as to improve the ability to transfer and generalize knowledge. Interval repetitive scheduling tasks are used to determine the next review interval for the target knowledge point based on the learner's performance in answering questions about the target knowledge point at different times, so that the review time of each knowledge point is adaptively shifted forward or backward during the learning process to enhance long-term retention. Metacognitive monitoring tasks are used to determine the learner's current comprehension and mastery level by comprehensively considering at least two of the following: response time, learner self-assessment confidence, and error type information. Based on this, the pace of subsequent feedback, the granularity of explanation, and the difficulty of questions are adjusted.
14. An electronic device, characterized in that, include: A processor and a memory, wherein instructions are stored in the memory and loaded and executed by the processor to implement the method as claimed in any one of claims 1-7.
15. A computer-readable storage medium storing a computer program therein, which, when executed, implements the method as described in any one of claims 1-7.