An AI-based teaching decision-making method, system, device, and medium

By acquiring students' historical status data and learning feedback, and updating the teaching decision model by combining temporary and delayed reward signals, the problem of the inability to assess long-term internalization and transfer capabilities in existing technologies is solved, achieving dynamic balance of student learning and improving the accuracy of intelligent teaching.

CN121883211BActive Publication Date: 2026-07-31SEVEN (BEIJING) EDUCATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SEVEN (BEIJING) EDUCATION TECH CO LTD
Filing Date
2025-11-01
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing reinforcement learning models based on temporary feedback cannot effectively assess and optimize students' long-term internalization and transfer abilities in the field of education. This leads to teaching decision-making models focusing excessively on short-term performance, which limits the overall teaching effectiveness and decision-making intelligence level of intelligent teaching systems.

Method used

By acquiring students' historical status data, the first teaching intervention strategy is determined, and the learning effect is obtained within a first preset time period to determine the temporary reward signal. During a second preset time period, assessment content is pushed out and learning feedback is obtained to determine the delayed reward signal. The teaching decision model is updated based on the dual reward signals to balance students' short-term learning performance and long-term knowledge acquisition.

Benefits of technology

It achieves a dynamic balance between students' immediate engagement and long-term knowledge retention, improves the intelligence level of the teaching decision-making model and the accuracy of personalized teaching, and ensures the continuous self-evolution and optimization of the teaching decision-making model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883211B_ABST
    Figure CN121883211B_ABST
Patent Text Reader

Abstract

This application provides an artificial intelligence-based teaching decision-making method, system, device, and medium. The method includes the following steps: acquiring student status data; determining and executing a first teaching intervention strategy through a teaching decision-making model; acquiring student learning performance within a first preset time period and determining a temporary reward signal accordingly; generating assessment content and pushing it to students within a second preset time period after the end of the first preset time period, based on the first teaching intervention strategy; acquiring student learning feedback on the assessment content and determining a delayed reward signal accordingly; and updating the teaching decision-making model based on the temporary and delayed reward signals to determine subsequent teaching intervention strategies. Implementing the technical solution provided in this application, by constructing a dual feedback loop combining temporary and delayed rewards, significantly improves the intelligence level of teaching decision-making and long-term teaching effectiveness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of reinforcement learning technology, and in particular to an artificial intelligence-based teaching decision-making method, system, device, and medium. Background Technology

[0002] With the deepening application of artificial intelligence technology in education, personalized learning and intelligent tutoring systems (ITS) have become important directions for improving teaching efficiency and realizing individualized instruction. How to efficiently collect and analyze learners' real-time status and learning outcomes, and dynamically adjust teaching strategies based on cognitive levels, has become a core requirement for ensuring teaching quality and improving learning efficiency.

[0003] Existing technologies typically rely on reinforcement learning models based on ad-hoc feedback. For example, the system adjusts the difficulty of subsequent content in real time based on whether the student answers the current knowledge point correctly or not; when the student answers correctly, the system immediately gives a positive reward, and vice versa.

[0004] However, in practical applications, this optimization mechanism, which relies solely on temporary feedback, suffers from a fundamental shortsightedness. This mechanism leads instructional decision-making models to overemphasize short-term student performance (such as temporary correct answers or participation), failing to effectively assess and optimize students' long-term internalization, retention, and transfer of knowledge. This limits the overall teaching effectiveness and decision-making intelligence level of intelligent teaching systems. Summary of the Invention

[0005] In view of this, this application provides an artificial intelligence-based teaching decision-making method, system, device, and medium to solve the above problems.

[0006] Firstly, an artificial intelligence-based teaching decision-making method is provided, which includes: Acquire students' status data at historical moments, and determine and execute the first teaching intervention strategy through a pre-set teaching decision model, recording the decision moment when the first teaching intervention strategy is started; Acquire students' learning performance within the first preset time period after the decision-making moment, and determine temporary reward signals based on the learning performance. The learning performance includes answer results and interaction behavior data. During the second preset time period after the first preset time period ends, assessment content is generated according to the first teaching intervention strategy and pushed to students. Obtain student learning feedback on assessment content and determine delayed reward signals based on the learning feedback, which includes data on student interaction with assessment content. The teaching decision-making model is updated based on temporary and delayed reward signals; Based on the updated teaching decision-making model and students' current status data, a second teaching intervention strategy is determined and output.

[0007] The above technical solution provides a targeted basis for determining the first teaching intervention strategy by acquiring students' historical status data; within the first preset time period, it determines temporary reward signals by acquiring learning outcomes to capture students' immediate response to the intervention; within the second preset time period, it determines delayed reward signals by pushing relevant assessment content and acquiring learning feedback to examine students' long-term knowledge retention; the teaching decision model is updated based on both temporary rewards (short-term effects) and delayed rewards (long-term effects), which overcomes the one-sidedness of the model relying solely on a single feedback and makes the optimization more comprehensive; the second teaching intervention strategy is output based on the updated model, ensuring that the new strategy can simultaneously balance students' short-term learning performance and long-term knowledge mastery, and is more in line with students' actual learning needs.

[0008] Optionally, the learning performance of students within a first preset time period after the decision-making moment is obtained, and a temporary reward signal is determined based on the learning performance, specifically including: Based on the answer results and interaction behavior data, the cognitive load status of students in the first preset time period is assessed, and the cognitive load status is either focused cognitive state or unfocused cognitive state. When the cognitive load state is assessed as a focused cognitive state, the temporary reward signal is determined as the first preset reward value; When the cognitive load state is assessed as an unfocused cognitive state, the temporary reward signal is determined to be the second preset reward value, which is lower than the first preset reward value.

[0009] The above technical solution assesses students' focused or unfocused cognitive state within a first preset time period based on their answer results and interactive behavior data, and sets different temporary reward values ​​accordingly. This avoids the one-sidedness of judging learning effectiveness solely based on correct or incorrect answers, and makes the temporary reward signal more accurately reflect the student's learning engagement and true cognitive state within that time period, thereby improving the rationality of the temporary reward.

[0010] Optionally, based on the answer results and interaction behavior data, assess the student's cognitive load during the first preset time period, specifically including: When the answer is incorrect or no answer is given, determine whether the interactive behavior data should be displayed as the first preset behavior pattern or the second preset behavior pattern. If the interactive behavior data shows the first preset behavior pattern, then the cognitive load state is assessed as a focused cognitive state. The first preset behavior pattern includes answering questions for a duration greater than the preset long-term threshold or repeatedly viewing teaching content associated with the first teaching intervention strategy more than the preset content viewing frequency threshold. If the interactive behavior data shows a second preset behavior pattern, the cognitive load state is assessed as an unfocused cognitive state. The second preset behavior pattern includes a question-answering time shorter than a preset short-term threshold or a number of user interface click events within a preset statistical time greater than a preset interface click number threshold, wherein the short-term threshold is less than the long-term threshold.

[0011] The above technical solution, by setting specific behavioral indicators (such as long-term and short-term thresholds for answering questions, thresholds for the number of times content is viewed, and thresholds for the number of times the interface is clicked), assesses the cognitive state based on matching the first or second preset behavioral pattern with the interactive behavior data when students answer questions incorrectly or do not answer. This provides a clear and quantifiable standard for distinguishing between focused and unfocused cognitive states, improves the objectivity and accuracy of cognitive load assessment, and thus ensures that the determination of temporary reward signals is more reliable.

[0012] Optionally, during a second preset time period following the end of the first preset time period, assessment content is generated based on the first teaching intervention strategy and pushed to students, specifically including: Analyze the first teaching intervention strategy and identify the first knowledge point associated with the first teaching intervention strategy; Obtain multiple candidate knowledge points associated with the first knowledge point from the preset knowledge graph; Calculate the knowledge transfer cost from the first knowledge point to each candidate knowledge point; Candidate knowledge points whose knowledge transfer costs meet the preset cost range are selected as the second knowledge points, and the evaluation content is determined based on the second knowledge points.

[0013] The above technical solution identifies the first knowledge point associated with the first teaching intervention strategy by analyzing it, filters related candidate knowledge points from the knowledge graph and calculates the knowledge transfer cost, and selects candidate knowledge points with costs within a preset range to generate assessment content. This ensures that the assessment content is reasonably related to the knowledge involved in the first teaching intervention strategy and that the transfer difficulty is moderate, effectively testing students' understanding and ability to transfer and apply relevant knowledge, and providing a scientific basis for determining delayed rewards.

[0014] Optionally, calculate the knowledge transfer cost from the first knowledge point to each candidate knowledge point, specifically including: In the knowledge graph, the semantic distance between the first knowledge point and the target candidate knowledge point is determined, and the semantic distance is quantified as the contextual difference degree, where the target candidate knowledge point is any one of multiple candidate knowledge points; Obtain the first cognitive level code of the first knowledge point in the preset cognitive classification method and the second cognitive level code of the target candidate knowledge point in the cognitive classification method; The difference between the first cognitive level code and the second cognitive level code is calculated to obtain the cognitive level span; Multiply the situational difference by the first preset weight to obtain the first weighted value; Multiply the cognitive level span by the second preset weight to obtain the second weighted value. The sum of the first preset weight and the second preset weight is a preset constant. Add the first weighted value to the second weighted value to obtain the knowledge transfer cost from the first knowledge point to the target candidate knowledge point.

[0015] The above technical solution quantifies the semantic distance between the first knowledge point and the candidate knowledge point into the contextual difference degree, combines the difference in hierarchical encoding between the two in the cognitive classification method (cognitive hierarchical span), and calculates the knowledge transfer cost by weighted summation. This makes the calculation of transfer cost take into account both the semantic association of knowledge points and the difference in cognitive difficulty, more comprehensively reflects the transfer difficulty between knowledge points, and ensures that the selected assessment knowledge points are more in line with the needs of testing students' learning outcomes.

[0016] Optionally, based on learning feedback, a delayed reward signal is determined, specifically including: Determine the assessment results corresponding to the learning feedback, and determine the basic reward value based on the assessment results; The interaction process data is matched with a preset typical error pattern library, which stores the correspondence between interaction process data and attribution types. Identify the matched attribution types, which include knowledge-based attributions and non-knowledge-based attributions. Knowledge-based attributions are associated with the first teaching intervention strategy, while non-knowledge-based attributions are not associated with the first teaching intervention strategy. Retrieve the attribution weights corresponding to the attribution type from the preset attribution weight library. The attribution weight library stores the mapping relationship between attribution types and attribution weights. Multiply the base reward value by the attribution weight to obtain the delayed reward signal.

[0017] The above technical solution determines the basic reward value based on the evaluation results, matches the attribution type (knowledge-based or non-knowledge-based) with the interaction process data and typical error pattern library, and adjusts the basic reward value according to the attribution weight to obtain the delayed reward signal. This enables the delayed reward to be more accurately associated with the actual effect of the first teaching intervention strategy, reduces the reward bias caused by non-knowledge factors (such as accidental mistakes), and improves the targeting of the delayed reward.

[0018] Optionally, the teaching decision-making model can be updated based on temporary and delayed reward signals, specifically including: Construct a multi-objective reward vector from temporary reward signals and delayed reward signals; The expected Pareto frontier of the multi-objective reward vector is calculated using a teaching decision-making model. Based on the expected Pareto frontier, the teaching decision-making model is updated.

[0019] The above technical solution, by constructing temporary reward signals and delayed reward signals into a multi-objective reward vector, calculating its expected Pareto frontier and updating the teaching decision model accordingly, can avoid the model only optimizing a single short-term (temporary reward) or long-term (delayed reward) objective. After the model is updated, it can balance students' short-term learning effects and long-term knowledge acquisition, thereby improving the overall decision-making effectiveness of the teaching decision model.

[0020] Secondly, an artificial intelligence-based teaching decision-making system is provided, the system including: The strategy execution module is configured to acquire students' status data at historical moments, determine and execute the first teaching intervention strategy through a preset teaching decision model, and record the decision moment when the first teaching intervention strategy is started. The temporary reward determination module is configured to obtain the student's learning performance within the first preset time period after the decision-making moment, and determine the temporary reward signal based on the learning performance. The learning performance includes answer results and interaction behavior data. The assessment content generation module is configured to generate assessment content based on the first teaching intervention strategy within a second preset time period after the end of the first preset time period, and push the assessment content to students. The delayed reward determination module is configured to obtain students' learning feedback on the assessment content and determine the delayed reward signal based on the learning feedback. The learning feedback includes data on the students' interaction process with the assessment content. The model update module is configured to update the teaching decision model based on temporary reward signals and delayed reward signals; The strategy output module is configured to determine and output a second teaching intervention strategy based on the updated teaching decision model and the students' current state data.

[0021] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the foregoing.

[0022] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described in any of the preceding descriptions.

[0023] In summary, implementing one or more technical solutions provided in this application has at least the following technical effects or advantages: The invention constructs a complete automated closed loop from intervention execution and determination of dual rewards (temporary and delayed) to model updates. Utilizing mechanisms such as cognitive load assessment, knowledge transfer cost calculation, and error attribution analysis, it provides the instructional decision-making model with highly accurate and low-noise composite feedback signals. This high-fidelity feedback enables the instructional decision-making model to continuously evolve and optimize itself. While ensuring high automation and scalability, it achieves a dynamic balance and consideration of the two core objectives of "immediate student engagement" and "long-term knowledge retention," significantly improving the system's intelligence level and the accuracy of personalized instruction. Attached Figure Description

[0024] Figure 1 This is an exemplary system architecture diagram of an AI-based teaching decision-making method or an AI-based teaching decision-making system applied in this application; Figure 2 This is a flowchart illustrating an artificial intelligence-based teaching decision-making method in the embodiments of this application; Figure 3 This is a schematic diagram of a module of an artificial intelligence-based teaching decision system according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in the application implementation method.

[0025] Explanation of reference numerals in the attached diagram: 100, System architecture; 101, First terminal device; 102, Second terminal device; 103, Third terminal device; 104, Network; 105, Server; 301, Policy execution module; 302, Temporary reward determination module; 303, Evaluation content generation module; 304, Delayed reward determination module; 305, Model update module; 306, Policy output module; 401, Processor; 402, Communication bus; 403, User interface; 404, Network interface; 405, Memory. Detailed Implementation

[0026] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments.

[0027] In the description of embodiments in this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any implementation or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other implementations or design options. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0028] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0029] Figure 1 An exemplary system architecture diagram is shown, illustrating an implementation of an AI-based instructional decision-making method or an AI-based instructional decision-making system to which this application can be applied.

[0030] like Figure 1 As shown, the system architecture 100 may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium to provide communication links between the terminal devices 101, 102, 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0031] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as model training applications, video recognition applications, web browser applications, social platform software, etc.

[0032] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptops, and desktop computers, etc. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as multiple software programs or software modules (e.g., multiple software programs or software modules used to provide distributed services) or as a single software program or software module. No specific limitations are imposed here.

[0033] When terminals 101, 102, and 103 are hardware devices, video capture devices can also be installed on them. These video capture devices can be various devices capable of capturing video, such as cameras, sensors, etc. Users can use the video capture devices on terminals 101, 102, and 103 to capture video.

[0034] Server 105 can be a server that provides various services, such as a backend server for processing data displayed on terminal devices 101, 102, and 103. The backend server can analyze and process the received data and can feed back the processing results (such as recognition results) to the terminal devices.

[0035] It should be noted that a server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software programs or software modules (e.g., multiple software programs or software modules used to provide distributed services), or as a single software program or software module. No specific limitations are made here.

[0036] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. In particular, if the target data does not need to be obtained remotely, the above system architecture may exclude the network and include only terminal devices or servers.

[0037] Figure 2This is a flowchart illustrating an artificial intelligence-based teaching decision-making method according to an embodiment of this application. This method can be implemented using a computer program, a microcontroller, or run on an artificial intelligence-based teaching decision-making system. The computer program can be integrated into the application or run as a standalone tool application. The specific steps of an artificial intelligence-based teaching decision-making method are described in detail below.

[0038] S201: Obtain students' status data at historical moments, determine and execute the first teaching intervention strategy through a pre-set teaching decision model, and record the decision moment when the first teaching intervention strategy is started.

[0039] In some embodiments of this application, the "pre-setting" process (i.e., initial training) of the pre-defined teaching decision-making model can be implemented using offline reinforcement learning algorithms. A large amount of historical interaction logs can be collected from the teaching system and processed into a sequence of four tuples (state data, teaching intervention strategy, next state data, reward). Since the reward is dual (temporary reward and delayed reward) and the delayed reward is sparse within the framework of this application, a specific offline reinforcement learning algorithm (e.g., Conservative Q-Learning (CQL) or Implicit Q-Learning (IQL)) can be used to process this offline, sparse reward dataset, thereby training a pre-defined teaching decision-making model with basic decision-making capabilities. Alternatively, in another implementation, the initial training can also employ a combination of behavioral cloning and supervised learning. For example, data pairs (state data, teaching intervention strategies) executed by "expert teachers" or "high-performance" models are selected from historical data. These data pairs are then used to train a policy network through supervised learning, enabling it to mimic expert behavior (i.e., inputting state data and outputting the teaching intervention strategy closest to the expert's choice). This policy network, trained through behavioral cloning, can serve as the initial strategy for the pre-defined teaching decision-making model and be updated online through reinforcement learning in subsequent step S205.

[0040] The system retrieves student status data from a database or real-time data stream at historical moments. This data can include students' historical answer records, knowledge point mastery, learning preferences, and historical interaction logs. This historical status data is fed into a pre-defined teaching decision model. The model performs inference calculations based on the status data and selects the optimal primary teaching intervention strategy from multiple alternative strategies (e.g., pushing videos, pushing practice questions, providing hints, etc.). This primary intervention strategy is then executed (e.g., pushing the corresponding teaching content to the student's user terminal). At the moment the primary intervention strategy is determined and triggered, the system records the current timestamp as the decision moment to begin its execution. This decision moment is stored for subsequent immediate and delayed evaluation of the primary intervention strategy's effectiveness.

[0041] S202: Obtain the student's learning performance within the first preset time period after the decision-making moment, and determine a temporary reward signal based on the learning performance. The learning performance includes answer results and interaction behavior data.

[0042] For example, after the decision to begin implementing the first teaching intervention strategy, the student's learning performance (including the student's answer results and learning-related interactive behavior data) is collected within a set first preset time period. Then, the student's learning situation at this stage is judged by combining these two aspects of information, and a temporary reward signal that can reflect the student's short-term learning input and performance is determined.

[0043] In one possible implementation, the learning performance of students within a first preset time period after the decision-making moment is obtained, and a temporary reward signal is determined based on the learning performance. Specifically, this includes: assessing the cognitive load state of students within the first preset time period based on answer results and interaction behavior data, wherein the cognitive load state is either a focused cognitive state or a non-focused cognitive state; when the cognitive load state is assessed as a focused cognitive state, a temporary reward signal is determined as a first preset reward value; when the cognitive load state is assessed as a non-focused cognitive state, a temporary reward signal is determined as a second preset reward value, wherein the second preset reward value is lower than the first preset reward value.

[0044] In some embodiments of this application, cognitive load state refers to a technical label used to technically characterize the level of learning engagement and quality of cognitive processes of students during the implementation of instructional intervention strategies. The label is inferred from objective behavioral data of students and is designed to distinguish between two cognitive processes that have different meanings in pedagogy, such as a focused cognitive state representing that the student is actively thinking and investing cognitive resources, and a non-focused cognitive state representing that the student may have given up thinking or engaged in ineffective interaction.

[0045] Specifically, after the decision-making moment, the system begins to acquire the student's learning performance within a first preset time period (e.g., a 5-minute real-time observation window). Learning performance is a dataset that includes at least the student's submitted answers (e.g., answers "A", "B", "C", "D", or "no answer") and the student's interactive behavior data during this period (e.g., answer time, click logs, page browsing history, etc.). After acquiring the learning performance, the system does not immediately determine a reward based on the correctness of the answers. Instead, it first comprehensively assesses the student's cognitive load state within the first preset time period based on the answer results and interactive behavior data, categorizing the cognitive load state as either focused or unfocused. Subsequently, a temporary reward signal is determined based on the assessed cognitive load state category: when the cognitive load state is assessed as focused (e.g., the student is determined to be engaged in valuable thinking), the temporary reward signal is determined to be a first preset reward value (e.g., a high positive number, such as +10); conversely, when the cognitive load state is assessed as unfocused (e.g., the student is determined to be clicking randomly or has given up), the temporary reward signal is determined to be a second preset reward value. To differentiate incentives for different states, a second preset reward value (e.g., a lower value, such as -5 or 0) is set to be lower than the first preset reward value.

[0046] In one possible implementation, based on the answer results and interactive behavior data, the cognitive load status of students within a first preset time period is assessed. Specifically, this includes: when the answer result is incorrect or no answer is given, determining whether the interactive behavior data exhibits a first preset behavior pattern or a second preset behavior pattern; if the interactive behavior data exhibits the first preset behavior pattern, the cognitive load status is assessed as a focused cognitive state, where the first preset behavior pattern includes an answer duration exceeding a preset long-term threshold or a number of times the teaching content associated with the first teaching intervention strategy is repeatedly viewed exceeding a preset content viewing frequency threshold; if the interactive behavior data exhibits the second preset behavior pattern, the cognitive load status is assessed as a non-focused cognitive state, where the second preset behavior pattern includes an answer duration less than a preset short-term threshold or a number of times user interface click events are generated within a preset statistical time period exceeding a preset interface click frequency threshold, wherein the short-term threshold is less than the long-term threshold.

[0047] In this embodiment, a preset behavioral pattern refers to a set of predefined and quantified objective interactive behavioral data. The pattern is used to objectively infer the underlying cognitive load state by analyzing a student's interaction process when the student's answer is unsatisfactory (e.g., incorrect or unanswered). For example, a first preset behavioral pattern can be defined as a set of objective behavioral data representing "active thinking," while a second preset behavioral pattern can be defined as a set of objective behavioral data representing "ineffective interaction" or "abandonment." Specifically, the step of assessing the student's cognitive load state within a first preset time period is a refinement of the "assessment" action in the previous embodiment. This assessment checks the answer results. When the answer result is incorrect or unanswered (i.e., when the mastery cannot be directly judged from the result), the system initiates a deep analysis process, analyzing the interactive behavioral data. The system determines whether the interactive behavioral data exhibits the first or second preset behavioral pattern. If the interactive behavioral data is determined to exhibit the first preset behavioral pattern, the cognitive load state is assessed as a focused cognitive state. The criteria for determining the first preset behavioral pattern include (any one of the following is sufficient): the student's answering time (e.g., 3 minutes) is greater than the preset long-term threshold (e.g., 1 minute); or, the number of times the student repeatedly views the teaching content associated with the first teaching intervention strategy (e.g., returning to view videos or documents) (e.g., 3 times) is greater than the preset content viewing frequency threshold (e.g., 2 times). Conversely, if the interactive behavior data is determined to exhibit the second preset behavioral pattern, the cognitive load state is assessed as an inattentive cognitive state. The criteria for determining the second preset behavioral pattern include (any one of the following is sufficient): the student's answering time (e.g., 5 seconds) is less than the preset short-term threshold (e.g., 10 seconds); or, the number of user interface click events generated within the preset statistical time (e.g., within a 10-second statistical window) (e.g., 15 times) is greater than the preset interface click frequency threshold (e.g., 8 times). To ensure the effectiveness of the time determination, the short-term threshold (e.g., 10 seconds) is set to be less than the long-term threshold (e.g., 1 minute).

[0048] In some embodiments of this application, it should also be noted that the aforementioned judgment logic based on the first and second preset behavior patterns is initiated when the answer result is incorrect or no answer is given. In a preferred implementation, when the answer result is "correct," the system can be configured to directly assess the cognitive load state as a focused cognitive state and determine the temporary reward signal as the first preset reward value (high reward value). This is because a "correct" answer result is itself the most direct and effective technical representation of a focused cognitive state. Even when the answer result is incorrect or no answer is given, there may be "intermediate behaviors" that do not satisfy either the first or second preset behavior pattern. For example, the answer duration is greater than or equal to a preset short-term threshold and less than or equal to a preset long-term threshold, and the interaction behavior does not meet the conditions of the content viewing count threshold or the interface click count threshold. In the case of such "intermediate behaviors," the system is configured to similarly assess the cognitive load state as a non-focused cognitive state and determine the temporary reward signal as the second preset reward value (low reward value). In this application, the technical label of non-focused cognitive state includes not only behaviors characterized by a second preset behavioral pattern (such as "quick abandonment"), but also "intermediate behaviors" that do not exhibit a first preset behavioral pattern (such as "repeated review") and ultimately lead to errors or no response.

[0049] S203: During the second preset time period after the end of the first preset time period, generate assessment content according to the first teaching intervention strategy and push the assessment content to the students.

[0050] For example, after the learning effect is collected in the first preset time period, in the subsequent second preset time period, content with reasonable knowledge relevance to the core teaching direction focused on by the first teaching intervention strategy will be selected to construct assessment content. This assessment content will be pushed to students, thereby examining students' long-term mastery and transfer application of the corresponding knowledge through assessments related to the previous teaching intervention.

[0051] In one possible implementation, during a second preset time period after the end of the first preset time period, assessment content is generated based on the first teaching intervention strategy and pushed to students. Specifically, this includes: parsing the first teaching intervention strategy to determine the first knowledge point associated with the first teaching intervention strategy; obtaining multiple candidate knowledge points associated with the first knowledge point in a preset knowledge graph; calculating the knowledge transfer cost from the first knowledge point to each candidate knowledge point; selecting candidate knowledge points whose knowledge transfer cost meets a preset cost range as second knowledge points, and determining the assessment content based on the second knowledge points.

[0052] In some embodiments of this application, knowledge transfer cost refers to an evaluation metric used to quantify the cognitive effort required for a student to learn and understand another related knowledge point (i.e., a candidate knowledge point) from a known knowledge point (i.e., the first knowledge point). The cost can be a comprehensive value; for example, the cost can take into account both the content similarity and the cognitive difficulty span between two knowledge points. The higher the cost, the greater the difficulty for the student to achieve knowledge transfer.

[0053] Specifically, within a second preset time period following the end of the first preset time period (e.g., between the 24th and 48th hour after the decision-making time), the process for generating assessment content is initiated. The first teaching intervention strategy is analyzed, and the first knowledge point associated with the first teaching intervention strategy is identified from its content or metadata. From a preset knowledge graph, multiple candidate knowledge points that have direct or indirect connections (e.g., "preceding," "following," or "similar" relationships) with the first knowledge point are retrieved. The knowledge graph construction methods may include: 1. Construction based on expert experience, where educational experts and subject teachers manually define the knowledge points (i.e., nodes) constituting the graph according to the curriculum syllabus and manually establish the relationships (i.e., edges) between knowledge points, such as "preceding relationships," "inclusion relationships," or "similar relationships." 2. Automatic construction based on text mining, applying Natural Language Processing (NLP) technology to automatically extract data from a large corpus of teaching materials (e.g., textbooks, course handouts, academic papers). For example, a Named Entity Recognition (NER) model is used to identify and extract core knowledge points as nodes; subsequently, relationship extraction algorithms (e.g., rule-based, co-occurrence statistics-based, or deep learning model-based) are used to mine and label the relationships between nodes, thereby constructing a knowledge graph. After obtaining multiple candidate knowledge points, the knowledge transfer cost from the first knowledge point to each candidate knowledge point is calculated one by one. After completing the knowledge transfer cost calculation for all candidate knowledge points, candidate knowledge points whose knowledge transfer costs fall within a preset cost range are selected. The preset cost range can be preset in at least one of the following objective ways: 1. Based on a fixed numerical range, the preset cost range is set to a fixed numerical range. For example, based on the knowledge transfer cost calculation method in S203, the range can be objectively set to [1.0, 2.5] (where costs less than 1.0 are considered "simple repetitions" and costs greater than 2.5 are considered "too difficult"). 2. Based on a dynamic range of statistical percentiles, in order to make the range more adaptive, the preset cost range can be determined based on the statistical distribution of all currently calculated knowledge transfer costs. For example, the system can calculate the median (50th percentile) and standard deviation of the knowledge transfer cost for all candidate knowledge points, or calculate its complete percentile distribution. A preset cost interval is objectively defined as a specific percentile range of this distribution; for example, all cost values ​​falling between [30th percentile, 70th percentile] are selected. After determining the preset cost interval, the system identifies the selected candidate knowledge point as the second knowledge point.Based on the second knowledge point (rather than the first knowledge point), the assessment content is determined. For example, questions that are strongly related to the second knowledge point are extracted from the question bank, combined into an assessment test paper, and the assessment content is pushed to the student's user terminal.

[0054] In one possible implementation, calculating the knowledge transfer cost from the first knowledge point to each candidate knowledge point specifically includes: determining the semantic distance between the first knowledge point and the target candidate knowledge point in the knowledge graph, and quantifying the semantic distance into a contextual difference degree, wherein the target candidate knowledge point is any one of multiple candidate knowledge points; obtaining the first cognitive level code corresponding to the first knowledge point in a preset cognitive classification method and the second cognitive level code corresponding to the target candidate knowledge point in the cognitive classification method; calculating the difference between the first cognitive level code and the second cognitive level code to obtain the cognitive level span; multiplying the contextual difference degree by a first preset weight to obtain a first weighted value; multiplying the cognitive level span by a second preset weight to obtain a second weighted value, wherein the sum of the first preset weight and the second preset weight is a preset constant; and adding the first weighted value and the second weighted value to obtain the knowledge transfer cost from the first knowledge point to the target candidate knowledge point.

[0055] In some embodiments of this application, the predefined cognitive classification method refers to a standard framework for classifying the cognitive difficulty of knowledge points or learning objectives. The classification method systematically divides cognitive activities from low-level (e.g., "memory") to high-level (e.g., "analysis" or "creation"), and assigns a first cognitive level code or a second cognitive level code that can be recognized and processed by a computer to each level. For example, the "memory" level can be coded as 1, and the "application" level can be coded as 3.

[0056] Specifically, the step of calculating the knowledge transfer cost from the first knowledge point to each candidate knowledge point is performed iteratively for any one of the multiple candidate knowledge points (referred to as the target candidate knowledge point in this step). In the knowledge graph, the semantic distance between the first knowledge point and the target candidate knowledge point is queried and determined. The semantic distance can be a value calculated by the graph path length or the cosine distance of word vectors, and this semantic distance is further quantified into the contextual difference degree. The first cognitive level code (e.g., code 1, representing "memory") pre-bound to the first knowledge point and the second cognitive level code (e.g., code 3, representing "application") pre-bound to the target candidate knowledge point are obtained from a preset cognitive classification method (e.g., Bloom's classification method). The difference between the first cognitive level code (1) and the second cognitive level code (3) is calculated (e.g., the absolute value |1-3|=2), and this difference (2) is obtained as the cognitive level span. After obtaining the two core indicators, contextual difference and cognitive level span, the contextual difference (e.g., 0.8) is multiplied by a first preset weight (e.g., 0.4) to obtain a first weighted value (0.8 × 0.4 = 0.32). Then, the cognitive level span (2) is multiplied by a second preset weight (e.g., 0.6) to obtain a second weighted value (2 × 0.6 = 1.2). To ensure that the sum of the weights of the two dimensions is constant, the sum of the first preset weight (0.4) and the second preset weight (0.6) is set to a preset constant (e.g., 1). The first weighted value (0.32) and the second weighted value (1.2) are added together, and the final sum (1.52) is the knowledge transfer cost from the first knowledge point to the target candidate knowledge point.

[0057] In some other embodiments of this application, the step of mapping the first knowledge point and the target candidate knowledge point to a preset cognitive classification method to obtain the corresponding cognitive level code can be specifically implemented as follows: 1. Manual annotation: When constructing the preset knowledge graph, the education expert team manually assigns a corresponding cognitive level code (e.g., "memory" = 1, "application" = 3) to each knowledge point node in the graph according to the preset cognitive classification method (e.g., Bloom's Taxonomy), and stores this code as metadata along with the knowledge point. 2. Verb analysis based on learning objectives: In many teaching systems, each knowledge point (e.g., the first knowledge point) is associated with one or more learning objectives, and learning objectives are usually defined by specific verbs (e.g., "enumerate...", "calculate...", "design..."). The system can maintain a "verb-cognitive level" mapping table (e.g., "enumerate" maps to "memory" [code 1]; "calculate" maps to "application" [code 3]). By automatically parsing the verbs in the learning objectives associated with the knowledge point, the first cognitive level code or the second cognitive level code for the knowledge point can be automatically determined.

[0058] S204: Obtain student learning feedback on the assessment content and determine delayed reward signals based on the learning feedback, which includes data on student interaction with the assessment content.

[0059] For example, after obtaining students' learning feedback on the assessment content (including data on the interaction process with the assessment content), the learning performance reflected in the feedback is combined with the analysis of the underlying reasons reflected in the interaction process. Based on this, the rewards are adjusted to determine the delayed reward signal that can accurately reflect the long-term effect of the first teaching intervention strategy.

[0060] In one possible implementation, determining a delayed reward signal based on learning feedback specifically includes: determining the evaluation result corresponding to the learning feedback and determining a basic reward value based on the evaluation result; matching the interaction process data with a preset typical error pattern library, which stores the correspondence between interaction process data and attribution types; determining the matched attribution type, which includes knowledge-based attribution and non-knowledge-based attribution, where knowledge-based attribution is associated with the first teaching intervention strategy and non-knowledge-based attribution is not associated with the first teaching intervention strategy; retrieving the attribution weight corresponding to the attribution type from a preset attribution weight library, which stores the mapping relationship between attribution types and attribution weights; and multiplying the basic reward value by the attribution weight to obtain the delayed reward signal.

[0061] In some embodiments of this application, attribution type refers to a classification label used to technically qualitatively characterize the root causes of errors made by students during content evaluation interactions. The label is determined by analyzing interaction process data and matching it with a library of typical error patterns, aiming to distinguish whether the error truly stems from a lack of mastery of the knowledge imparted by the first instructional intervention strategy (i.e., knowledge-based attribution) or from other factors unrelated to the strategy, such as carelessness or guesswork (i.e., non-knowledge-based attribution).

[0062] Specifically, after receiving learning feedback, the system determines the corresponding assessment result (e.g., "correct," "incorrect," or a specific score) and assigns an initial base reward value based on the assessment result (e.g., +10 points for "correct," -5 points for "incorrect"). The system extracts interaction process data contained in the learning feedback (e.g., student click sequences, answer times, answer modification records, etc.) and matches this data with a pre-defined library of typical error patterns. This library stores the correspondence between various interaction pattern sequences and attribution types (e.g., the sequence "submitting an incorrect answer after rapid, continuous clicking" corresponds to "non-knowledge-based attribution"). The system determines the attribution type through matching. Attribution types are divided into two main categories: knowledge-based attribution and non-knowledge-based attribution. Knowledge-based attribution is defined as an attribution strongly associated with the teaching effectiveness of the first teaching intervention strategy, while non-knowledge-based attribution is defined as an attribution not associated with the first teaching intervention strategy (e.g., student carelessness). After determining the attribution type, the system retrieves the corresponding attribution weight from a pre-set attribution weight library. This library stores the mapping relationship between attribution types and attribution weights (e.g., "knowledge-based attribution" has a weight of 0.8, and "non-knowledge-based attribution" has a weight of 1.0). The base reward value (e.g., -5 points) is multiplied by the retrieved attribution weight (e.g., 0.8 for "knowledge-based attribution") to obtain the adjusted final value (e.g., -5 × 0.8 = -4), which is then used as the delayed reward signal.

[0063] In other embodiments of this application, the typical error pattern library can store richer correspondences. For example, the typical error pattern library can map interaction process data such as "submitting an incorrect answer in a very short time (e.g., less than 3 seconds)" or "frequent (e.g., more than 3 times per second) continuous user interface click events during the evaluation process" to non-knowledge-based attribution (representing "guessing" or "ineffective interaction"). The typical error pattern library can also map interaction process data such as "interaction duration significantly longer than average duration", "reviewing teaching content associated with the first teaching intervention strategy multiple times (e.g., more than 3 times)", and "multiple answer modifications (e.g., more than 2 times)" to knowledge-based attribution (representing that although the student has made an effort, the teaching effect of the first teaching intervention strategy has not been achieved, i.e., "teaching failure") when the final evaluation result is determined to be incorrect. Correspondingly, the configuration logic of the attribution weight library can also be varied. In a preferred implementation, the core function of the attribution weight is to filter out penalties unrelated to the first teaching intervention strategy. For example, when the base reward value is negative (e.g., -5 points, representing an "error"): if the attribution type is determined to be a knowledge-based attribution (i.e., "teaching failure"), the attribution weight can be set to 1.0. In this case, the delayed reward signal is -5 × 1.0 = -5, meaning the model fully bears the penalty for this teaching failure. If the attribution type is determined to be a non-knowledge-based attribution (e.g., a student's carelessness leading to an error, but the student has already mastered the knowledge), the attribution weight can be set to a value close to 0 (e.g., 0.1). In this case, the delayed reward signal is -5 × 0.1 = -0.5, meaning the model bears almost no penalty for this "non-teaching factor." Through this method, the delayed reward signal is refined, more accurately reflecting the effectiveness of the primary teaching intervention strategy itself. Accordingly, in a preferred implementation, when the base reward value is positive (e.g., +10 points, representing an evaluation result of "correct"), regardless of whether the attribution type is determined to be knowledge-based or non-knowledge-based attribution, the corresponding attribution weight can be set to 1.0 to ensure that the model can fully learn and consolidate the long-term positive effects brought about by this successful teaching intervention.

[0064] S205: Update the teaching decision-making model based on temporary reward signals and delayed reward signals.

[0065] For example, in order to enable the teaching decision-making model to simultaneously adapt to short-term teaching feedback (corresponding to temporary reward signals) and long-term teaching effects (corresponding to delayed reward signals), the two reward signals are integrated to form a comprehensive reference that reflects the dual-objective needs. Based on this reference, the optimization trade-off relationship between the two reward objectives is analyzed to ensure that a balanced optimization direction is found without weakening the effect of either objective. According to this trade-off result, the core parameters or decision logic of the teaching decision-making model are adjusted to complete the model update, so that the teaching intervention strategies subsequently output by the model can better balance short-term learning feedback and long-term knowledge acquisition effects.

[0066] In one possible implementation, the teaching decision model is updated based on the temporary reward signal and the delayed reward signal, specifically including: constructing the temporary reward signal and the delayed reward signal into a multi-objective reward vector; calculating the expected Pareto frontier of the multi-objective reward vector through the teaching decision model; and updating the teaching decision model based on the expected Pareto frontier.

[0067] In some embodiments of this application, the expected Pareto front refers to a set of non-dominated solutions in a multi-objective optimization problem. Each solution in the set represents a specific trade-off, under which it is impossible to improve the performance of another objective (e.g., delayed reward signal) without sacrificing the performance of at least one objective (e.g., a temporary reward signal). "Expectation" indicates that the front is obtained by statistically estimating future rewards using a teaching decision model.

[0068] Specifically, after obtaining the temporary reward signal (one numerical value) and the delayed reward signal (another numerical value), the update process of the instructional decision model is initiated. The temporary and delayed reward signals are constructed into a multi-objective reward vector. For example, if the temporary reward signal is +10 and the delayed reward signal is +50, the constructed multi-objective reward vector can be represented as [+10, +50]. The expected Pareto front of the multi-objective reward vector is calculated using the instructional decision model (e.g., a multi-objective reinforcement learning model). This calculation process does not maximize the weighted sum of the two reward values, but rather aims to estimate a set of non-dominated solutions that is optimal in both the temporary and delayed reward dimensions, representing different policy preferences. After obtaining the expected Pareto front (i.e., a set of non-dominated solutions), the instructional decision model is updated based on the expected Pareto front. For example, the update step may specifically include: selecting a target Q vector from the expected Pareto frontier set according to a pre-defined preference strategy (e.g., selecting the vector that best balances the two objectives or best enhances the delayed reward dimension); and then using the target Q vector as a temporal difference objective to update the value function or policy network parameters of the instructional decision model.

[0069] S206: Based on the updated teaching decision-making model and students' current status data, determine and output a second teaching intervention strategy.

[0070] In some embodiments of this application, the student's current state data refers to the latest information about the student that the system acquires or observes in real time after the teaching decision model has been updated and the next teaching action needs to be determined. The current state data corresponds to the state data at the historical moment in step S201 and is used to represent the student's immediate state after experiencing the first teaching intervention strategy and its assessment (steps S202-S204). For example, it may include the student's mastery of knowledge points after completing the assessment, current fatigue level, or attention level.

[0071] Specifically, after updating the teaching decision model in step S205, the updated model (which integrates short-term and long-term feedback represented by temporary and delayed reward signals) is used for a new round of teaching decisions. The system acquires the student's current state data, which can be the latest interaction information collected in real time from the student's terminal, or student profile data updated based on the student's learning feedback on the assessment content (step S204). The student's current state data is used as new input to the updated teaching decision model. The updated teaching decision model performs inference calculations based on the current state data to determine a new teaching action aimed at balancing short-term engagement and long-term retention, and outputs this action as a second teaching intervention strategy. The second teaching intervention strategy (e.g., pushing a micro-lesson video of a new knowledge point) will be executed, thus starting the next round of the "intervention-feedback-update" cycle.

[0072] Figure 3 This is a schematic diagram of a module of an artificial intelligence-based teaching decision-making system according to an embodiment of this application. This system can be implemented through software, hardware, or a combination of both, becoming all or part of the overall system. For example... Figure 3 As shown, the system includes: a policy execution module 301, a temporary reward determination module 302, an evaluation content generation module 303, a delayed reward determination module 304, a model update module 305, and a policy output module 306, wherein: The strategy execution module 301 is configured to acquire students' status data at historical moments, determine and execute the first teaching intervention strategy through a preset teaching decision model, and record the decision moment when the first teaching intervention strategy is started. The temporary reward determination module 302 is configured to acquire the student's learning performance within a first preset time period after the decision-making moment, and determine a temporary reward signal based on the learning performance. The learning performance includes answer results and interaction behavior data. The assessment content generation module 303 is configured to generate assessment content according to the first teaching intervention strategy during a second preset time period after the end of the first preset time period, and push the assessment content to students. The delayed reward determination module 304 is configured to obtain students' learning feedback on the assessment content and determine the delayed reward signal based on the learning feedback. The learning feedback includes data on the students' interaction process with the assessment content. Model update module 305 is configured to update the teaching decision model based on temporary reward signals and delayed reward signals; The strategy output module 306 is configured to determine and output a second teaching intervention strategy based on the updated teaching decision model and the student's current state data.

[0073] Based on the above implementation method, as an optional implementation method, the temporary reward determination module 302 is specifically used for: Based on the answer results and interaction behavior data, the cognitive load status of students in the first preset time period is assessed, and the cognitive load status is either focused cognitive state or unfocused cognitive state. When the cognitive load state is assessed as a focused cognitive state, the temporary reward signal is determined as the first preset reward value; When the cognitive load state is assessed as an unfocused cognitive state, the temporary reward signal is determined to be the second preset reward value, which is lower than the first preset reward value.

[0074] As an optional implementation, the temporary reward determination module 302 is specifically used for: When the answer is incorrect or no answer is given, determine whether the interactive behavior data should be displayed as the first preset behavior pattern or the second preset behavior pattern. If the interactive behavior data shows the first preset behavior pattern, then the cognitive load state is assessed as a focused cognitive state. The first preset behavior pattern includes answering questions for a duration greater than the preset long-term threshold or repeatedly viewing teaching content associated with the first teaching intervention strategy more than the preset content viewing frequency threshold. If the interactive behavior data shows a second preset behavior pattern, the cognitive load state is assessed as an unfocused cognitive state. The second preset behavior pattern includes a question-answering time shorter than a preset short-term threshold or a number of user interface click events within a preset statistical time greater than a preset interface click number threshold, wherein the short-term threshold is less than the long-term threshold.

[0075] Based on the above implementation method, as an optional implementation method, the evaluation content generation module 303 is specifically used for: Analyze the first teaching intervention strategy and identify the first knowledge point associated with the first teaching intervention strategy; Obtain multiple candidate knowledge points associated with the first knowledge point from the preset knowledge graph; Calculate the knowledge transfer cost from the first knowledge point to each candidate knowledge point; Candidate knowledge points whose knowledge transfer costs meet the preset cost range are selected as the second knowledge points, and the evaluation content is determined based on the second knowledge points.

[0076] As an optional implementation, the evaluation content generation module 303 is specifically used for: In the knowledge graph, the semantic distance between the first knowledge point and the target candidate knowledge point is determined, and the semantic distance is quantified as the contextual difference degree, where the target candidate knowledge point is any one of multiple candidate knowledge points; Obtain the first cognitive level code of the first knowledge point in the preset cognitive classification method and the second cognitive level code of the target candidate knowledge point in the cognitive classification method; The difference between the first cognitive level code and the second cognitive level code is calculated to obtain the cognitive level span; Multiply the situational difference by the first preset weight to obtain the first weighted value; Multiply the cognitive level span by the second preset weight to obtain the second weighted value. The sum of the first preset weight and the second preset weight is a preset constant. Add the first weighted value to the second weighted value to obtain the knowledge transfer cost from the first knowledge point to the target candidate knowledge point.

[0077] Based on the above implementation method, as an optional implementation method, the delayed reward determination module 304 is specifically used for: Determine the assessment results corresponding to the learning feedback, and determine the basic reward value based on the assessment results; The interaction process data is matched with a preset typical error pattern library, which stores the correspondence between interaction process data and attribution types. Identify the matched attribution types, which include knowledge-based attributions and non-knowledge-based attributions. Knowledge-based attributions are associated with the first teaching intervention strategy, while non-knowledge-based attributions are not associated with the first teaching intervention strategy. Retrieve the attribution weights corresponding to the attribution type from the preset attribution weight library. The attribution weight library stores the mapping relationship between attribution types and attribution weights. Multiply the base reward value by the attribution weight to obtain the delayed reward signal.

[0078] Based on the above implementation method, as an optional implementation method, the model update module 305 is specifically used for: Construct a multi-objective reward vector from temporary reward signals and delayed reward signals; The expected Pareto frontier of the multi-objective reward vector is calculated using a teaching decision-making model. Based on the expected Pareto frontier, the teaching decision-making model is updated.

[0079] It should be noted that the system provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0080] This embodiment also discloses an electronic device, with reference to... Figure 4 The electronic device may include: at least one processor 401, at least one communication bus 402, user interface 403, network interface 404, and at least one memory 405.

[0081] The communication bus 402 is used to enable communication between these components.

[0082] The user interface 403 may include a display screen and a camera. Optionally, the user interface 403 may also include a standard wired interface and a wireless interface.

[0083] The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0084] The processor 401 may include one or more processing cores. The processor 401 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 405, and by calling data stored in memory 405. Optionally, the processor 401 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 401 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor 401.

[0085] The memory 405 may include random access memory (RAM) or read-only memory. Optionally, the memory 405 may include a non-transitory computer-readable storage medium. The memory 405 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 405 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 405 may also be at least one storage device located remotely from the aforementioned processor 401. Figure 4 As shown, the memory 405, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for an artificial intelligence-based teaching decision-making method.

[0086] exist Figure 4 In the electronic device shown, the user interface 403 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 401 can be used to call an application program stored in the memory 405 that is an artificial intelligence-based teaching decision method. When executed by one or more processors 401, the electronic device executes one or more methods as described in the above embodiments.

[0087] It should be noted that, for the sake of simplicity, the aforementioned implementation methods are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the implementation methods described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0088] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0089] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.

[0090] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0091] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0092] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 405 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned memory 405 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.

[0093] The above description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the disclosure in this specification. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are considered exemplary only, and the scope of this application is defined by the claims.

Claims

1. An artificial intelligence-based teaching decision method, characterized by, The method includes: Acquire students' status data at historical moments, and determine and execute the first teaching intervention strategy through a pre-set teaching decision model, recording the decision moment when the first teaching intervention strategy is started; The learning performance of the student within a first preset time period after the decision time is obtained, and a temporary reward signal is determined based on the learning performance, wherein the learning performance includes answer results and interaction behavior data; During the second preset time period after the end of the first preset time period, assessment content is generated according to the first teaching intervention strategy, and the assessment content is pushed to the student; Obtain the student's learning feedback on the assessment content, and determine a delayed reward signal based on the learning feedback, wherein the learning feedback includes the student's interaction process data on the assessment content; The teaching decision model is updated based on the temporary reward signal and the delayed reward signal; Based on the updated teaching decision model and the student's current status data, a second teaching intervention strategy is determined and output. The step of acquiring the student's learning performance within a first preset time period after the decision-making time, and determining a temporary reward signal based on the learning performance, specifically includes: Based on the answer results and the interaction behavior data, the cognitive load status of the student during the first preset time period is evaluated, and the cognitive load status is either a focused cognitive state or a non-focused cognitive state. When the cognitive load state is assessed as the focused cognitive state, the temporary reward signal is determined to be the first preset reward value; When the cognitive load state is assessed as the non-focused cognitive state, the temporary reward signal is determined to be a second preset reward value, which is lower than the first preset reward value. The determination of the delayed reward signal based on the learning feedback specifically includes: Determine the evaluation result corresponding to the learning feedback, and determine the basic reward value based on the evaluation result; The interaction process data is matched with a preset typical error pattern library, which stores the correspondence between the interaction process data and the attribution type. The matched attribution types are determined, including knowledge-based attribution and non-knowledge-based attribution. The knowledge-based attribution is associated with the first teaching intervention strategy, and the non-knowledge-based attribution is not associated with the first teaching intervention strategy. Retrieve the attribution weight corresponding to the attribution type from the preset attribution weight library, wherein the attribution weight library stores the mapping relationship between attribution types and attribution weights; The delayed reward signal is obtained by multiplying the base reward value by the attribution weight.

2. The method of claim 1, wherein, The step of assessing the student's cognitive load during the first preset time period based on the answer results and the interaction behavior data specifically includes: When the answer result is incorrect or no answer is given, determine whether the interactive behavior data exhibits a first preset behavior pattern or a second preset behavior pattern. If the interactive behavior data shows a first preset behavior pattern, then the cognitive load state is evaluated as the focused cognitive state. The first preset behavior pattern includes answering questions for a duration greater than a preset long-term threshold or repeatedly viewing teaching content associated with the first teaching intervention strategy more than a preset content viewing frequency threshold. If the interactive behavior data exhibits a second preset behavior pattern, then the cognitive load state is assessed as the non-focused cognitive state. The second preset behavior pattern includes a question-answering time shorter than a preset short-term threshold or a number of user interface click events generated within a preset statistical time greater than a preset interface click number threshold, wherein the short-term threshold is less than the long-term threshold.

3. The method of claim 1, wherein, The step of generating assessment content based on the first teaching intervention strategy and pushing the assessment content to the student within a second preset time period after the end of the first preset time period specifically includes: Analyze the first teaching intervention strategy and identify the first knowledge point associated with the first teaching intervention strategy; Obtain multiple candidate knowledge points associated with the first knowledge point from the preset knowledge graph; Calculate the knowledge transfer cost from the first knowledge point to each of the candidate knowledge points; Candidate knowledge points whose knowledge transfer costs meet the preset cost range are selected as second knowledge points, and the evaluation content is determined based on the second knowledge points.

4. The method of claim 3, wherein, The calculation of the knowledge transfer cost from the first knowledge point to each of the candidate knowledge points specifically includes: In the knowledge graph, the semantic distance between the first knowledge point and the target candidate knowledge point is determined, and the semantic distance is quantified as a contextual difference degree, wherein the target candidate knowledge point is any one of the plurality of candidate knowledge points; Obtain the first cognitive level code corresponding to the first knowledge point in the preset cognitive classification method and the second cognitive level code corresponding to the target candidate knowledge point in the cognitive classification method; The difference between the first cognitive level code and the second cognitive level code is calculated to obtain the cognitive level span; Multiply the scenario difference by a first preset weight to obtain a first weighted value; The cognitive level span is multiplied by a second preset weight to obtain a second weighted value, and the sum of the first preset weight and the second preset weight is a preset constant. The first weighted value and the second weighted value are added together to obtain the knowledge transfer cost from the first knowledge point to the target candidate knowledge point.

5. The method according to claim 1, characterized in that, The step of updating the teaching decision model based on the temporary reward signal and the delayed reward signal specifically includes: The temporary reward signal and the delayed reward signal are used to construct a multi-objective reward vector; The expected Pareto frontier of the multi-objective reward vector is calculated using the teaching decision model. The teaching decision model is updated based on the expected Pareto front.

6. An artificial intelligence-based teaching decision-making system, characterized in that the system... include: The strategy execution module is configured to acquire students' status data at historical moments, determine and execute the first teaching intervention strategy through a preset teaching decision model, and record the decision moment when the first teaching intervention strategy is started. A temporary reward determination module is configured to acquire the student's learning performance within a first preset time period after the decision-making time, and determine a temporary reward signal based on the learning performance. The learning performance includes answer results and interaction behavior data. Specifically, acquiring the student's learning performance within the first preset time period after the decision-making time and determining the temporary reward signal based on the learning performance includes: Based on the answer results and the interaction behavior data, the cognitive load status of the student during the first preset time period is evaluated, and the cognitive load status is either a focused cognitive state or a non-focused cognitive state. When the cognitive load state is assessed as the focused cognitive state, the temporary reward signal is determined to be the first preset reward value; When the cognitive load state is assessed as the non-focused cognitive state, the temporary reward signal is determined to be a second preset reward value, which is lower than the first preset reward value. The assessment content generation module is configured to generate assessment content according to the first teaching intervention strategy within a second preset time period after the end of the first preset time period, and push the assessment content to the student. A delayed reward determination module is configured to acquire the student's learning feedback on the assessment content and determine a delayed reward signal based on the learning feedback. The learning feedback includes data on the student's interaction with the assessment content. Specifically, determining the delayed reward signal based on the learning feedback includes: Determine the evaluation result corresponding to the learning feedback, and determine the basic reward value based on the evaluation result; The interaction process data is matched with a preset typical error pattern library, which stores the correspondence between the interaction process data and the attribution type. The matched attribution types are determined, including knowledge-based attribution and non-knowledge-based attribution. The knowledge-based attribution is associated with the first teaching intervention strategy, and the non-knowledge-based attribution is not associated with the first teaching intervention strategy. Retrieve the attribution weight corresponding to the attribution type from the preset attribution weight library, wherein the attribution weight library stores the mapping relationship between attribution types and attribution weights; The base reward value is multiplied by the attribution weight to obtain the delayed reward signal; The model update module is configured to update the teaching decision model based on the temporary reward signal and the delayed reward signal. The strategy output module is configured to determine and output a second teaching intervention strategy based on the updated teaching decision model and the student's current state data.

7. An electronic device, characterized in that, The device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-5.