A teaching interaction engine system and method based on multi-modal perception and rule inference

CN122839221APending Publication Date: 2026-09-29XIAN WEIDU EDUCATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611203249.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-10
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0008]本发明提供一种基于多模态感知与规则推理的教学交互引擎系统及方法,以克服现有技术中感知维度单一、策略决策缺乏协同、回复缺乏安全合规保障、存储缺乏高可用性、模块之间缺乏闭环的技术缺陷

Benefits of technology

[0063]采用本发明所述技术方案,与现有技术相比,具有以下有益效果:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839221A_ABST
    Figure CN122839221A_ABST
Patent Text Reader

Abstract

The application discloses a kind of teaching interactive engine systems and methods based on multi-modal perception and rule inference.The system includes perception layer, decision layer, interactive layer and data layer.Perception layer includes emotion processor, cognitive load processor and context enhancer, emotion processor adopts pre-training model and keyword heuristic dual-mode degradation architecture, cognitive load processor is based on typing speed and pause length dual-index estimation load level, is aggregated into structured enhancement metadata by context enhancer.Decision layer includes priority rule engine and dynamic prompt word builder, rule engine merges multiple rule decisions according to priority, and dynamic prompt word builder selects prompt word from template library according to decision and supports value guiding language injection.Interactive layer includes dialogue state tracker and security check module.Data layer includes student portrait storage.The application is applicable to teaching interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of teaching interaction in artificial intelligence, specifically to a teaching interaction engine system and method based on multimodal perception and rule reasoning. Background Technology

[0002] With the deepening application of artificial intelligence technology in education, intelligent teaching interaction systems have become an important development direction for educational informatization. However, existing intelligent teaching interaction systems have the following technical shortcomings:

[0003] Deficiency 1: Limited Perceptual Dimension and Lack of Multimodal Learning Understanding. Existing intelligent teaching systems typically rely solely on students' answer accuracy to determine teaching strategies, neglecting students' emotional states and cognitive load levels during the learning process. When students are experiencing negative emotions such as anxiety or frustration, or when their cognitive load is too high, the system still pushes challenging content according to predetermined strategies, leading to a deterioration in the learning experience and a decrease in learning efficiency. This single-dimensional perception cannot comprehensively depict students' true learning state, making teaching interventions lack specificity and timeliness. While some systems incorporate affective computing, they employ a single affective analysis model, lacking an effective degradation mechanism when the model is unavailable or the inference is uncertain, resulting in insufficient system usability.

[0004] Deficiency 2: Lack of a flexible multi-rule collaborative mechanism for instructional strategy decision-making. Existing systems typically use hard-coded if-else logic for selecting instructional strategies. Strategy rules are deeply bound to system code; adding or adjusting strategies requires code modification and redeployment, failing to adapt to the continuous iteration needs of instructional strategies. Furthermore, the lack of a priority coordination mechanism between different strategies means that when multiple strategies simultaneously meet their triggering conditions, the system cannot determine which strategy should be prioritized and which should be secondary, leading to decision conflicts and inconsistent strategy behavior.

[0005] Deficiency 3: AI-generated responses lack compliance safeguards for educational scenarios. Existing systems lack dedicated filtering mechanisms after generating responses using large language models, or only perform simple sensitive word filtering. They haven't built a multi-level filtering system specifically for educational scenarios, including harmful content detection, academic misconduct detection, and value alignment checks. In educational settings, AI responses not only need to avoid harmful content but also prevent inducing academic misconduct (such as assignment writing or cheating) and align with moral education values. Existing systems cannot meet this requirement.

[0006] Defect 4: The student profile storage lacks a highly available two-tier architecture. Existing systems typically use only a single database to store student profile data, which easily leads to performance bottlenecks under high-concurrency read / write scenarios. When the database service is unavailable, the entire profile storage module becomes unavailable, causing the core functions of the system to fail. Although some systems have introduced caching, the lack of a consistent read / write strategy and data synchronization mechanism between the cache and the database can easily lead to data inconsistency.

[0007] Defect 5: The teaching interaction lacks a closed-loop architecture from perception to decision-making to feedback. The existing system treats emotion perception, strategy decision-making, dialogue management, and data storage as independent modules, lacking a unified data flow architecture between these modules. Perception results cannot directly drive decision-making, decision results cannot directly guide interaction, and interaction records cannot automatically flow back to the data layer to form a learning history. This results in data silos between different modules of the system, failing to achieve a complete closed loop from perception to decision-making to feedback. Summary of the Invention

[0008] This invention provides a teaching interaction engine system and method based on multimodal perception and rule reasoning, to overcome the technical defects of existing technologies such as single perception dimension, lack of coordination in strategy decision-making, lack of security and compliance guarantee for responses, lack of high availability of storage, and lack of closed loop between modules.

[0009] The present invention achieves the above objectives by adopting the following technical solution: Firstly, the present invention provides a teaching interaction engine system based on multimodal perception and rule reasoning, comprising:

[0010] The perception layer includes an emotion processor, a cognitive load processor, and a context enhancer;

[0011] The emotion processor employs a dual-mode degradation architecture combining a pre-trained sentiment analysis model and keyword heuristic matching to classify user input text by emotion. When the classification confidence of the pre-trained sentiment analysis model falls below a preset confidence threshold, it automatically downgrades to the keyword heuristic matching scheme. The cognitive load processor estimates the cognitive load level based on the user's typing speed and pause duration during input. The context enhancer aggregates the sentiment analysis results of the emotion processor and the cognitive load estimation results of the cognitive load processor to generate structured enhanced metadata.

[0012] The decision-making layer includes a rules engine and a dynamic prompt word builder;

[0013] The rule engine maintains several teaching strategy rules. Each rule includes a condition judgment function, an action execution function, and a priority value. The rule engine traverses all rules, collects decision results for rules that meet the conditions, sorts them by priority, and merges them to generate a teaching strategy decision. The dynamic prompt word builder selects the corresponding system prompt word template and user prompt word template from the predefined prompt word template library based on the teaching strategy decision, fills in the context variables, and generates a complete interactive prompt word.

[0014] The interaction layer includes a dialogue state tracker and a security verification module;

[0015] The dialogue state tracker maintains the current topic, a list of student confusion points, a set of discussed achievements, a history of dialogue rounds, and the number of attempts for the current question. When it detects that the AI's response contains keywords of a recorded confusion point, it automatically marks the confusion point as resolved. The security verification module performs multi-level compliance checks on the AI-generated response text and generates a safe alternative response for responses that fail the checks.

[0016] The data layer, including student profile storage, adopts a two-layer storage architecture of memory caching and database persistence. When reading, the memory cache is queried first, and if the cache is not hit, the database is retrieved. When writing, both the memory cache and the database are updated simultaneously.

[0017] The perception layer, decision layer, interaction layer, and data layer are connected in sequence. The enhanced metadata output by the perception layer serves as the input to the decision layer. The response generated by the interaction prompts produced by the decision layer after inference by the AI ​​model is filtered by the interaction layer and then output. The analysis results of the perception layer and the interaction records of the interaction layer are persisted to the data layer.

[0018] Furthermore, the dual-mode degradation architecture of the emotion processor is specifically as follows:

[0019] The emotion processor prioritizes loading a pre-trained Chinese sentiment analysis model to classify the input text into emotions, and outputs emotion labels and corresponding model confidence scores.

[0020] When the pre-trained sentiment analysis model is unavailable or the model confidence is lower than the preset confidence threshold, the sentiment processor automatically downgrades to the keyword heuristic matching scheme.

[0021] The keyword heuristic matching scheme maintains a positive keyword database and a negative keyword database, and counts the occurrence frequency of positive and negative keywords in the input text. If the number of positive keywords is greater than the number of negative keywords, it is determined to be a positive emotion; if the number of negative keywords is greater than the number of positive keywords, it is determined to be a negative emotion; if the two are equal, it is determined to be a neutral emotion.

[0022] Furthermore, the specific method by which the cognitive load processor estimates the cognitive load level is as follows:

[0023] Calculate typing speed, where typing speed is the number of characters in the input text divided by the input time;

[0024] When the typing speed is lower than a preset speed threshold, the speed score is marked as high load; otherwise, it is marked as low load.

[0025] The number of long pauses whose duration exceeds a preset pause threshold is counted, and the pause score is calculated by dividing the number of long pauses by the total number of pauses.

[0026] The comprehensive load score is the average of the speed score and the pause score. When the comprehensive load score is greater than a first threshold, it is determined to be a high load. When the comprehensive load score is greater than a second threshold and less than or equal to the first threshold, it is determined to be a medium load. When the comprehensive load score is less than or equal to the second threshold, it is determined to be a low load.

[0027] Furthermore, the rule merging method of the rule engine is specifically as follows:

[0028] The priority value for each rule is a preset non-negative integer, with higher values ​​indicating higher priority.

[0029] When the rule engine is executed, it traverses all rules and collects the decision results corresponding to the rules whose condition judgment function returns true into the candidate set.

[0030] The decision results in the candidate set are sorted in descending order of priority value. The highest priority decision result is the main decision, and the lower priority decision results that do not conflict with the main decision are merged into the final strategy as supplementary decisions.

[0031] The rules pre-set by the rule engine include: triggering a high scaffolding, low challenge strategy when the emotion label output by the emotion processor is negative; triggering a value-injection and emotional support strategy when the cognitive load is high and the number of attempts for the current question exceeds 2; triggering a Socratic extension strategy when the answer is correct; triggering a simplified explanation strategy when the cognitive load is high; and a default Socratic guidance strategy.

[0032] Furthermore, the dynamic prompt word builder includes a value injection module;

[0033] The value injection module maintains a value keyword mapping table, which contains the association between value categories and corresponding introductory templates;

[0034] When the teaching strategy decision output by the rule engine includes value objectives, the dynamic prompt word builder automatically appends the guiding statement corresponding to the value category to the end of the user prompt word template.

[0035] Furthermore, the multi-level compliance checks of the security verification module specifically include:

[0036] Level 1 Harmful Content Detection: Traverse the predefined harmful keyword library and detect whether the AI-generated response text contains violent, pornographic, politically sensitive, or discriminatory language keywords. If any keyword is found, it is judged as a violation.

[0037] Level 2 academic misconduct detection: Traverse the predefined academic misconduct keyword library and check whether the reply text contains keywords that guide the writing of assignments or cheating. If any keyword is hit, it is determined to be a violation.

[0038] Level 3 Value Alignment Check: When the teaching strategy decision output by the rule engine includes value objectives, verify whether the response text contains keywords of the corresponding value category. If not, it is determined to be a violation.

[0039] When any level determines a violation, the security verification module selects a security response corresponding to the violation type from a predefined security response template library to replace the original text output.

[0040] Furthermore, the two-tier storage architecture for storing student profiles is specifically as follows:

[0041] The memory cache stores information using student identifiers as keys and student profile data structures as values. The student profile data structure includes learning history, proficiency scores for each knowledge point, emotional history, and cognitive load history.

[0042] During a read operation, the memory cache is queried using the target student identifier. If a match is found, the student profile data structure is returned directly. If no match is found, the database is queried using the target student identifier, and the query result is written to the memory cache and then returned.

[0043] During a write operation, both the memory cache and the database are updated simultaneously. If the database write operation fails, the system automatically downgrades to pure memory mode, records the downgrade log, and automatically synchronizes it after the database is restored.

[0044] Furthermore, the list of confusion points maintained by the dialogue state tracker includes confusion point description text and timestamps added to the confusion points;

[0045] When the dialogue state tracker receives AI response text, it iterates through each confusion point description text in the confusion point list, checks whether the AI ​​response text contains keywords from the confusion point description text, and if it does, marks the corresponding confusion point as resolved and removes it from the confusion point list.

[0046] The dialogue state tracker supports exporting the session state as a context dictionary, which includes the current topic, a list of remaining points of confusion, the current number of attempts, and a dialogue round count, for the rule engine to consume during conditional judgment.

[0047] Secondly, the present invention provides a teaching interaction engine method based on multimodal perception and rule reasoning, applied to the aforementioned teaching interaction engine system based on multimodal perception and rule reasoning, comprising the following steps:

[0048] Step S1: Receive user input data, which includes text content, student identifier, input time, and a list of pause durations;

[0049] Step S2: Query the student profile storage based on the student identifier. If no corresponding profile exists, create a new student profile.

[0050] Step S3: The context enhancer calls the emotion processor to perform emotion analysis on the text content, calls the cognitive load processor to estimate the cognitive load level based on the input time and the pause duration list, and writes the emotion analysis result and the cognitive load estimation result into the enhanced metadata;

[0051] Step S4: Update the dialogue state tracker, record the current dialogue round, detect and remove resolved points of confusion;

[0052] Step S5: Merge the enhanced metadata with the dialogue state of the dialogue state tracker into a context dictionary, send it to the rule engine to perform rule matching and decision merging, and output the teaching strategy decision;

[0053] Step S6: The dynamic prompt word builder selects a template from the prompt word template library according to the teaching strategy decision, fills in the context variables, and adds value guidance when the teaching strategy decision includes value goals, thereby generating the interactive prompt word;

[0054] Step S7: Append the emotion analysis results and the cognitive load estimation results to the emotion history record and cognitive load history record of the student profile, and write them into the student profile storage;

[0055] Step S8: Send the response text generated by the AI ​​model based on the interaction prompts to the security verification module to perform multi-level compliance checks; if the response text passes all checks, output the response text.

[0056] If the response text fails any level of inspection, a security response template corresponding to the violation type will be selected from the security response template library for alternative output.

[0057] Furthermore, the specific steps for the rule engine to perform rule matching and decision merging in step S5 are as follows:

[0058] Step S51: Traverse all preset rules and collect the decision results of rules whose condition judgment functions return true on the context dictionary into the candidate set;

[0059] Step S52: Sort the decision results in the candidate set from high to low priority values;

[0060] Step S53: Select the decision with the highest priority as the main decision;

[0061] Step S54: Traverse the remaining decision results and merge the decision results that do not conflict with the main decision in terms of attributes and have a lower priority than the main decision into the final teaching strategy decision.

[0062] The beneficial effects of this invention are as follows:

[0063] Compared with the prior art, the technical solution described in this invention has the following advantages:

[0064] Multimodal fusion perception enhances the accuracy of learning comprehension. This invention achieves real-time and robust perception of students' emotional states and cognitive load levels through a dual-mode degradation architecture (model + keywords) of an emotion processor and a dual-index estimation of a cognitive load processor. The dual-mode degradation architecture ensures that the system can still provide reliable emotion judgments when the model is unavailable or the inference is uncertain, balancing accuracy and usability; the dual-index cognitive load estimation comprehensively determines the load level from two dimensions: speed and pauses, which is more accurate than a single index. Perceived information is uniformly aggregated into structured enhanced metadata by a context enhancer and then injected into the decision-making process, enabling teaching strategies to be dynamically adjusted according to students' actual psychological states, avoiding the drawback of continuing to impose high-difficulty content when students are emotionally depressed or cognitively overloaded.

[0065] Priority-based rules and collaborative decision-making enhance the flexibility and accuracy of strategy adaptation. The rule engine of this invention adopts a priority-based condition-action-priority triple rule architecture, decoupling teaching strategy rules from system code. It supports dynamic addition and deletion of rules at runtime, adapting to continuous iteration of teaching strategies. When multiple rules are triggered simultaneously, they are merged in descending priority order—high-priority decisions take precedence, while low-priority decisions supplement non-conflicting attributes, resolving the conflict problem in traditional hard-coded strategy decision-making. Six pre-defined core rules cover key teaching scenarios such as negative emotions, high cognitive load, correct answers, and historical topics, achieving refined strategy adaptation.

[0066] A three-tiered filtering system ensures the safety of educational content. The security verification module of this invention performs three levels of filtering on AI-generated responses: harmful content detection, academic misconduct detection, and value alignment checks. If any level fails, a safe alternative response is generated, ensuring that the AI-generated teaching content complies with educational standards and preventing harmful information, academic misconduct, and value deviations from negatively impacting students. The value alignment check is automatically triggered when the rule engine specifies value objectives, achieving a unity between knowledge transmission and character development.

[0067] A dual-layer storage architecture ensures high availability of profile data. The student profile storage of this invention employs a dual-layer architecture of memory caching and database persistence. During reads, the memory cache is queried first to improve performance; if the cache misses, the data is retrieved from the source database to ensure data integrity. During writes, both layers are updated simultaneously to ensure consistency. In the event of a database write failure, the system automatically degrades to pure memory mode, ensuring continued operation even in the event of a database failure. After database recovery, automatic synchronization significantly improves the system's fault tolerance and availability.

[0068] This invention achieves end-to-end collaboration between perception, decision-making, interaction, and data through a four-layer closed-loop architecture. It integrates emotion perception, cognitive load estimation, rule-based reasoning, dynamic prompt generation, dialogue state tracking, security and compliance filtering, and student profile storage into a four-layer architecture: perception layer, decision layer, interaction layer, and data layer. Each layer forms a complete closed loop through structured data transfer: enhanced metadata output from the perception layer directly drives rule-based reasoning in the decision layer; prompts generated by the decision layer are filtered by the interaction layer before being output; and interaction records automatically flow back to the data layer to form a learning history. The historical data in the data layer provides contextual reference for subsequent perception and decision-making, thus realizing a complete closed loop from perception to decision to feedback.

[0069] A progressive degradation design ensures system resilience. This invention implements progressive degradation at multiple levels: at the sentiment analysis level, it automatically degrades to a keyword-based solution when the model is unavailable or the confidence level falls below a set value; at the storage level, it automatically degrades to a pure memory mode when the database is unavailable. This multi-layered degradation mechanism ensures that the system can still provide basic services when resources are limited or components fail, avoiding complete system unavailability caused by a single point of failure. Attached Figure Description

[0070] Figure 1 This is a schematic diagram of the overall architecture of a teaching interaction engine system based on multimodal perception and rule reasoning provided in an embodiment of the present invention. Detailed Implementation

[0071] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0072] Example 1: System Overall Architecture

[0073] Reference Figure 1 The teaching interaction engine system based on multimodal perception and rule reasoning described in this invention includes a four-layer architecture: perception layer, decision layer, interaction layer, and data layer.

[0074] The perception layer consists of an emotion processor, a cognitive load processor, and a context enhancer. The emotion processor receives text input from the user, performs emotion analysis, and outputs an emotion label (POSITIVE / NEGATIVE / NEUTRAL) and a confidence score. The cognitive load processor receives behavioral data from the user's input process (input time, pause duration list), performs cognitive load estimation, and outputs a load level (HIGH / MODERATE / LOW) and a comprehensive score. The context enhancer aggregates the emotion analysis results and the cognitive load estimation results to generate structured enhanced metadata containing fields such as emotion_label, emotion_confidence, load_level, and load_score.

[0075] The decision layer consists of a rule engine and a dynamic prompt word builder. The rule engine receives enhanced metadata output by the context enhancer and a dialogue state context dictionary exported by the dialogue state tracker, performs rule matching and decision merging, and outputs a teaching strategy decision (including strategy name, matched rule_name, priority, and optional value_target). The dynamic prompt word builder selects a corresponding prompt word template from the template library based on the strategy decision, populates the context variables, and generates a complete interactive prompt word (including "system" and "user" parts).

[0076] The interaction layer consists of a dialogue state tracker and a security verification module. The dialogue state tracker maintains the current topic, a list of confusion points (including descriptions and timestamps), a set of discussed achievements, the history of dialogue rounds, and the number of attempts for the current question. When an AI response is received, the dialogue state tracker automatically detects confusion points in the confusion point list that have been covered by responses and marks them as resolved. The security verification module performs a three-level filter on the AI ​​response text, outputting either the original response that passes the check or a safe alternative response.

[0077] The data layer includes student profile storage, employing a two-tier architecture of in-memory caching and a database. The in-memory cache stores copies of the student profile data structure using the student identifier as the key, while the database serves as persistent storage. Read operations prioritize querying the in-memory cache; if a match is not found, the database is retrieved and the result is written back to the in-memory cache. Write operations simultaneously update both the in-memory cache and the database; if a database write fails, the system automatically downgrades to pure in-memory mode.

[0078] The AI ​​model is a large language model deployed locally or in the cloud. It receives interactive prompts generated by the dynamic prompt word builder and returns the generated response text. The security verification module filters the response text before outputting it to the user interface.

[0079] Example 2: Dual-mode Degradation Process of Emotion Processor

[0080] The workflow of the dual-mode degradation architecture of the emotion processor described in this invention includes the following steps:

[0081] Step A01: Receive the input text content.

[0082] Step A02: Attempt to load the pre-trained Chinese sentiment analysis model. The pre-trained sentiment analysis model is a Chinese sentiment analysis model based on the RoBERTa architecture, pre-trained on a large Chinese sentiment corpus. If model loading fails (e.g., insufficient memory, missing model files, or missing dependency libraries), proceed to step A06.

[0083] Step A03: Feed the input text into the model for inference to obtain the sentiment classification result and the corresponding model confidence. The model's output layer is a three-class softmax layer, corresponding to the probability distributions of positive, negative, and neutral categories, respectively.

[0084] Step A04: Determine whether the model confidence is greater than or equal to the preset confidence threshold (0.6). If it is greater than or equal to 0.6, proceed to step A05; otherwise, proceed to step A06.

[0085] Step A05: Use the emotion classification result output by the model as the final emotion label, output the emotion label (POSITIVE / NEGATIVE / NEUTRAL), and label the confidence as model_confidence = confidence and heuristic_confidence = 0.

[0086] Step A06: Downgrade to a keyword heuristic matching scheme. The keyword heuristic matching scheme maintains a positive keyword library and a negative keyword library. In this embodiment, the positive keyword library contains words such as ["happy", "like", "not bad", "good", "great", "understand", "got it", "haha", "thank you"]; the negative keyword library contains words such as ["annoying", "tired", "don't want to", "hate", "difficult", "collapse", "can't", "give up", "confused", "anxious", "terrible"].

[0087] Step A07: Count the number of positive keywords (denoted as pos_count) and the number of negative keywords (denoted as neg_count) in the input text.

[0088] Step A08: Determine the relationship between pos_count and neg_count. If pos_count > neg_count, the emotion is POSITIVE; if neg_count > pos_count, the emotion is NEGATIVE; if pos_count == neg_count, the emotion is NEUTRAL. At this point, the confidence level is set to model_confidence = 0 and heuristic_confidence = 1.0.

[0089] Step A09: Output the final emotion label.

[0090] In a specific application example of this invention, when a student inputs the text "This question is so hard, I'm about to collapse, I have no idea how to do it," the model's inference confidence is 0.73 (≥0.6), and the model outputs a negative emotion (NEGATIVE), which is directly taken as the model result. When the student inputs the text "Okay," the model's inference confidence is 0.52 (<0.6), and it is downgraded to the keyword scheme: the positive keyword "okay" appears once (pos_count=1), and the negative keyword appears 0 times (neg_count=0), which is judged as a positive emotion (POSITIVE).

[0091] Example 3: Cognitive Load Estimation Process

[0092] The cognitive load estimation method for the cognitive load processor of the present invention includes the following steps:

[0093] Step B01: Receive the input text content, the input time typing_time (in seconds), and the pause duration list pause_durations (in seconds).

[0094] Step B02: Calculate typing speed: typing_speed = len(content) / typing_time, in words per second.

[0095] Step B03: Determine if the typing speed (typing_speed) is less than the preset speed threshold (2.0 characters / second). If it is less than 2.0 characters / second, then the speed score (speed_score) = 1.0 (marked as high load); otherwise, the speed_score = 0.0 (marked as low load).

[0096] Step B04: Count the number of long pauses (long_pause_count) in the pause_durations list that exceed the preset pause threshold (5.0 seconds).

[0097] Step B05: Calculate the pause score: pause_score = long_pause_count / len(pause_durations). If len(pause_durations) = 0, then pause_score = 0.0.

[0098] Step B06: Calculate the overall load score load_score = (speed_score + pause_score) / 2.0, with a value range of [0.0, 1.0].

[0099] Step B07: Determine the load level. If load_score > 0.7, it is determined to be a high load (HIGH); otherwise, if load_score > 0.4, it is determined to be a medium load (MODERATE); otherwise, it is determined to be a low load (LOW).

[0100] Step B08: Output the overall load score (load_score) and load level (load_level).

[0101] In a specific application example of this invention, the student inputs 20 characters of text over a time of 15 seconds. The pause duration list is [3.2, 6.8, 8.1, 2.5, 12.3] (unit: seconds). Typing speed = 20 / 15 ≈ 1.33 characters / second < 2.0 characters / second, speed_score = 1.0; long pauses (>5.0 seconds) occur 3 times (6.8, 8.1, 12.3 seconds), for a total of 5 pauses, pause_score = 3 / 5 = 0.6; load_score = (1.0 + 0.6) / 2 = 0.8 > 0.7, which is determined to be high load (HIGH).

[0102] In another application example, a student inputs 30 characters of text over 8 seconds. The pause duration list is [1.2, 0.8, 2.1, 1.5]. The typing speed is 30 / 8 = 3.75 characters / second ≥ 2.0 characters / second, so speed_score = 0.0. There are no long pauses (all ≤ 5.0 seconds), so pause_score = 0.0. Load_score = 0.0, which is considered low load.

[0103] Example 4: Priority rule matching and decision merging in a rule engine

[0104] The priority rule matching and decision merging method of the rule engine described in this invention includes the following steps:

[0105] Step C01: Receive the context dictionary. The context dictionary contains fields such as student identifier, current knowledge point, emotion tag (from emotion processor), cognitive load level (from cognitive load processor), number of attempts for the current question (from dialogue state tracker), correctness of answer result (if any), and course focus tag.

[0106] Step C02: Initialize the candidate decision set candidates = [].

[0107] Step C03: Traverse all preset rules in the rule engine. The data structure of each rule is {name: rule name, condition: condition evaluation function, action: action execution function, priority: priority value}.

[0108] Step C04: For each rule, call its condition evaluation function, passing the context dictionary as a parameter. If the condition evaluation function returns True, call the action execution function to generate the decision result, and add the decision result (including fields such as strategy, rule_name, priority, value_target, etc.) to the candidate set candidates.

[0109] Step C05: Determine if all rules have been traversed. If not, return to step C03 to continue; if completed, proceed to step C06.

[0110] Step C06: Sort the decision results in the candidate set in descending order of priority value.

[0111] Step C07: Take the first decision result after sorting as the primary decision.

[0112] Step C08: Initialize the final strategy final_strategy = primary_decision.

[0113] Step C09: Iterate through the remaining decision results. For each candidate decision, determine whether it conflicts with the primary_decision in terms of attributes. If there is no conflict (e.g., the candidate's value_target is different from the primary_decision's value_target, or there are fields in the candidate's strategy that are not covered by primary_decision), then merge the candidate as a supplementary decision into the final_strategy.

[0114] Step C10: Output the final teaching strategy decision final_strategy.

[0115] In a specific application example of the present invention, the rule engine presets the following rules, as shown in Table 1 below.

[0116] Table 1. Pre-defined rules for the rule engine

[0117]

[0118] When the student's emotion is NEGATIVE, cognitive load is HIGH, and the current problem has been attempted 3 times, the rule matching results are: Emotional Support (priority 3), High Scaffolding for Negative Emotions (priority 2), Simplified Explanation for High Load (priority 1), and Default Socratic Guidance (priority 0). After sorting by priority, Emotional Support is the primary decision (highest priority), and High Scaffolding for Negative Emotions is the supplement (priority 2, no conflict with the primary decision, can be merged if the strategy names are different). The final strategy is a combination of strategy="emotional_support" + "high_scaffolding", with value_target="Emotional Support".

[0119] Example 5: Dynamic Prompt Construction and Value Injection

[0120] The dynamic prompt word construction method of the dynamic prompt word builder of the present invention includes the following steps:

[0121] Step D01: Receive the teaching strategy decision final_strategy output by the rule engine, which includes the strategy name (or the combined name if multiple strategies are merged) and value_target (optional).

[0122] Step D02: Parse the main policy identifier from the policy name. For example, "emotional_support+high_scaffolding" is parsed as the main policy "emotional_support".

[0123] Step D03: Based on the parsed policy identifier, select the corresponding system prompt word template and user prompt word template from the predefined prompt word template library.

[0124] In this embodiment, the prompt word template library includes:

[0125] High scaffolding, low challenge template (corresponding to high_scaffolding): system is set to "You are a patient and gentle math tutor", user is set to "Please guide students to think in a step-by-step manner, do not give the answer directly. Student's current state: {emotion}, {load}.";

[0126] Socratic extension template (corresponding to socratic_extension): system is set to "You are a Socratic mentor who is good at inspiring thinking", and user is set to "The student just answered this question correctly, please ask a deeper question";

[0127] The default Socratic guidance template (corresponding to socratic_default) has the following settings: system is set to "You are a mentor for guided learning", and user is set to "Please guide students to discover problem-solving strategies on their own by asking questions";

[0128] Emotional support template (corresponding to emotional_support): system is set to "You are an empathetic mentor", user is set to "Please first acknowledge and encourage the student's efforts, and then guide them in solving the problem";

[0129] Simplified explanation template (corresponding to simplified): system is set to "You are a mentor who is good at simplifying complex things", and user is set to "Please explain this concept in the simplest language and analogy".

[0130] Step D04: Fill the placeholders in the selected template with context variables (student input, current knowledge point, emotion tag, etc.).

[0131] Step D05: Determine whether the teaching strategy decision includes the value_target field (value objectives). If not, proceed to step D07.

[0132] Step D06: If a value objective is included, append a value guidance statement to the end of the user prompt template. The value injection module maintains a value keyword mapping table, which includes the association between value categories and corresponding guidance templates:

[0133] Patriotism: The guiding statement is "Please appropriately integrate the knowledge points of this question with the connection between national development and national spirit into your explanation";

[0134] The guiding principle is: "Please encourage students during your explanation, emphasizing the importance of perseverance and hard work."

[0135] Honesty and Integrity: The guiding text is "Please emphasize the importance of honest learning and completing assignments independently."

[0136] Cooperation and mutual assistance: The guiding words are "Please encourage students to actively seek help and cooperate when they encounter difficulties."

[0137] Step D07: Assemble the complete interactive prompt, which includes two parts: system (system role settings) and user (user prompt).

[0138] Step D08: Output the complete interactive prompts.

[0139] In a specific application example of this invention, a student inputs "I can't do this geometry problem," with the emotion tag NEGATIVE, cognitive load HIGH, and number of attempts 3. The rule engine outputs strategy="emotional_support+high_scaffolding" and value_target="perseverance". The dynamic prompt word builder selects the emotional support template and the high scaffolding low challenge template, fills in the context variables, and adds the perseverance prompt, ultimately generating:

[0140] System message: "You are an empathetic, patient, and gentle math tutor."

[0141] User: "Please first acknowledge and encourage the student's efforts, then guide the student's thinking step by step. The student's current state: low spirits, high cognitive load. Please integrate the importance of perseverance and effort into the explanation of this question."

[0142] Example 6: Dialogue State Tracking and Confusion Management

[0143] The dialogue state maintenance and confusion point management method of the dialogue state tracker of the present invention includes the following steps:

[0144] Step E01: Receive user input data, including the current topic and the input text content.

[0145] Step E02: Update the current topic to topic and increment the dialogue round count (turn_count).

[0146] Step E03: Invoke the automatic confusion point detection module to scan the input text and identify whether it contains new confusion point descriptions. The identification rule is: when the input text contains interrogative words such as "don't understand," "can't," "why," or "how," extract the name of the knowledge point in that sentence or context as a new confusion point, add it to the confusion point list, and record the timestamp of the addition.

[0147] Step E04: Receive the AI's response text.

[0148] Step E05: Iterate through the description text "confusion" for each confusion point in the confusion point list, and check if the response contains the keyword "confusion". If it does, mark the confusion point as resolved, remove it from the confusion point list, and record the resolution timestamp.

[0149] Step E06: If the current problem is not resolved after the number of attempts reaches the preset limit (e.g., 5 times), trigger the special processing rules in the rule engine (e.g., switch the explanation method).

[0150] Step E07: After the current round ends, export the dialogue state (including current_topic, remaining_confusion_points, turn_count, attempts_count, and history) as a context dictionary format.

[0151] Step E08: The exported context dictionary is consumed by the rule engine during the next rule matching.

[0152] In a specific application example of this invention, a student says, "I don't understand what logarithms are." The dialogue state tracker adds "logarithm" to the confusion list and records turn_count=1. After the AI ​​replies by explaining the definition of logarithms, the reply contains the keyword "logarithm," and the dialogue state tracker automatically removes "logarithm" from the confusion list and marks it as resolved.

[0153] Example 7: Three-level filtering process of security verification module

[0154] The three-level filtering method of the security verification module of the present invention includes the following steps:

[0155] Step F01: Receive the AI ​​response text and the teaching strategy decision final_strategy (including optional value_target).

[0156] Step F02: First-level filtering – Harmful content detection. Traverse a predefined harmful keyword database to detect whether the response contains violent, pornographic, politically sensitive, or discriminatory language keywords. The harmful keyword database includes, but is not limited to, words such as "violence," "suicide," "pornography," "reactionary," and "mentally retarded." If any keyword is matched, it is considered a violation, marked with `violation_type = "harmful"`, and the process proceeds to step F08.

[0157] Step F03: If the first-level filter passes, proceed to the second-level filter—academic misconduct detection. Iterate through a predefined database of academic misconduct keywords to check if the response contains keywords such as "assignment writing service," "plagiarism," or "cheating." This database includes, but is not limited to, terms like "I'll write your homework for you," "copy this directly," "cheating," and "assignment writing service." If any keyword is matched, it is considered a violation, and `violation_type = "academic"` is set, then proceed to step F08.

[0158] Step F04: If the second-level filter passes, determine whether the `final_strategy` contains the `value_target` field. If it does not, proceed to step F07 (all passed).

[0159] Step F05: If the `value_target` field is included, proceed to the third level of filtering—value alignment check. Retrieve the keyword list for the corresponding value category from the value keyword mapping table, and check if the response contains at least one keyword of that category. The value keyword mapping table includes:

[0160] Patriotism: ["Nation", "Ethnicity", "Struggle", "Rejuvenation", "Chinese Dream", "Dedication"];

[0161] The spirit of perseverance: ["persistence", "effort", "never give up", "perseverance", "persistence"];

[0162] Honesty and Integrity: ["Honesty", "Integrity", "Uprightness", "Trustworthiness"];

[0163] Cooperation and mutual assistance: ["cooperation", "mutual assistance", "unity", "sharing"].

[0164] Step F06: If the response does not contain any keyword of the corresponding value category, it is judged as a violation, marked as violation_type = "value_misalignment", and jumps to step F08.

[0165] Step F07: All three levels of filtering pass, output passed = True, and return the original response.

[0166] Step F08: Select a security response template corresponding to the violation type from the security response template library. The security response template library includes:

[0167] Harmful type: "Sorry, I can't answer that question. Let's get back to math. What topic have you been studying lately?"

[0168] Academic type: "Learning requires independent thinking and practice to improve. Let's analyze the solution to this problem together."

[0169] The `value_misalignment` type selects the corresponding template based on the missing value category. For example, if patriotism is missing, the response would be "Let's also think about this issue from the perspective of national development."

[0170] Step F09: Output passed = False, violations contains a list of violation types, and safe_response is the selected safe response.

[0171] Example 8: Two-layer read / write process for student portrait storage

[0172] The two-layer read / write method for storing student profiles as described in this invention includes the following steps:

[0173] Step G01: Receive the profile operation request, which includes the operation type (read / write), student identifier student_id, and profile data (when writing).

[0174] Step G02: If it is a read operation, proceed to step G03; if it is a write operation, proceed to step G07.

[0175] Step G03 (Read Path): Query the memory cache using student_id as the key.

[0176] Step G04: Determine if a corresponding record exists in the memory cache. If a cache hit occurs, proceed to step G05; if a cache miss occurs, proceed to step G06.

[0177] Step G05: Directly return the student profile data structure in the memory cache.

[0178] Step G06 (Cache Miss, Return to Origin): Query the database using student_id as the query condition. If a corresponding record exists in the database, write the query result to the memory cache and return it; if no corresponding record exists in the database, create a new student profile (including initial proficiency and empty historical records), write it to the memory cache and the database, and return it.

[0179] Step G07 (Write Path): Update both the memory cache and the database simultaneously.

[0180] Step G08: Determine if the database write was successful. If successful, proceed to step G10.

[0181] Step G09: If database write fails (e.g., database connection timeout, lock conflict, etc.), automatically downgrade to pure memory mode, only update the memory cache, and record the downgrade log (including timestamp, student_id, and reason for failure). At the same time, start a background retry thread, which attempts to reconnect to the database every 30 seconds. After a successful connection, automatically synchronize the changes in the memory cache to the database and clear the downgrade flag.

[0182] Step G10: Write success is returned.

[0183] In a specific application example of the present invention, the student profile data structure includes the following fields:

[0184] student_id: A unique identifier for the student;

[0185] learning_history: A list of learning history records, each record contains timestamp, topic, is_correct, emotion_label, and load_level;

[0186] proficiency_scores: Proficiency score for each knowledge point, a floating-point number between 0 and 1;

[0187] emotion_history: A list of emotion history records;

[0188] load_history: A list of historical cognitive load records;

[0189] total_attempts: Total number of attempts;

[0190] correct_count: Number of correct answers;

[0191] The proficiency score uses a moving average update strategy: when a student answers correctly, proficiency = 0.7 * old_proficiency + 0.3 * 1.0; when a student answers incorrectly, proficiency = 0.7 * old_proficiency + 0.3 * 0.0.

[0192] Example 9: Complete Methodology and Flow of the Teaching Interaction Engine

[0193] The complete process of the teaching interaction engine method described in this invention includes the following steps:

[0194] Step S1: Receive user input data. The input data includes text content, student identifier student_id, input time typing_time (the time interval from when the user starts typing to when the user submits the input, in seconds), and pause duration list pause_durations (a list of durations in seconds during which the user pauses for more than 0.5 seconds during the input process).

[0195] In this embodiment, suppose a student inputs "How do I draw the graph of this function?", student_id = "S001", typing_time = 12 seconds, pause_durations = [2.1, 4.3, 6.8, 1.2].

[0196] Step S2: Query the student profile storage based on student_id = "S001". The query path is as follows: first check the memory cache; if a match is found, return the result directly; if no match is found, query the database, write the result to the cache, and then return it. If the student profile does not exist, create a new student profile, initialize the proficiency of each knowledge point to 0.5, and set the learning history, emotional history, and cognitive load history to empty.

[0197] Step S3: The context enhancer calls the emotion processor to perform emotion analysis on the content. In this embodiment, the model inference confidence is 0.78 (≥0.6), and the model output emotion label is NEUTRAL. The context enhancer calls the cognitive load processor to estimate the cognitive load level: typing speed = len("How to draw the graph of this function?") / 12 = 9 / 12 = 0.75 words / second < 2.0 words / second, speed_score = 1.0; long pause (>5.0 seconds) occurs once at 6.8 seconds, with a total of 4 pauses, pause_score = 1 / 4 = 0.25; load_score = (1.0 + 0.25) / 2 = 0.625, which is judged as moderate load (MODERATE). The context enhancer 130 writes the emotion analysis result (NEUTRAL, confidence 0.78) and the cognitive load estimation result (MODERATE, overall score 0.625) into the enhanced metadata.

[0198] Step S4: Update the dialogue state tracker. Set the current topic to "Function Graphs", increment turn_count to 1, and initialize attempts_count to 1. Check if the input text contains new points of confusion. If the question "How do I draw this function graph?" contains the question word "how", add "How to draw a function graph" to the list of points of confusion.

[0199] Step S5: Merge the enhanced metadata with the dialogue state from the dialogue state tracker into a context dictionary, and send it to the rule engine for rule matching. The context dictionary contains: emotion_label="NEUTRAL", load_level="MODERATE", attempts_count=1, current_topic="function graph", confusion_points=["function graph drawing method"]. The rule engine iterates through all rules, matching only the default Socratic guide rule (priority 0), outputting strategy="socratic_default", with no value_target.

[0200] Step S6: The dynamic prompt word builder selects the default Socratic guidance template based on the strategy name "socratic_default", populates the context variables, and generates interactive prompt words: system = "You are a guided learning mentor", user = "A student asked a question about function graphs: 'How do you draw this function graph?'. Please guide the student to discover the solution themselves by asking questions, rather than giving a complete answer directly."

[0201] Step S7: Append the emotion tag (NEUTRAL) and cognitive load level (MODERATE) to the emotion history and cognitive load history of student profile S001, and write them to the student profile storage. The write operation simultaneously updates the memory cache and the database.

[0202] Step S8: Send the interactive prompt generated in Step S6 to the AI ​​model for inference to obtain the response text. Assume the AI ​​model returns the response: "Okay, let's think about it together. First, do you know what type of function this is? What is its domain?" Send the response text to the security verification module 320 for three-level filtering: Level 1 harmful content detection failed, Level 2 academic misconduct detection failed, and Level 3 value alignment check (no value_target) is skipped. If all checks pass, output the original response to the user interface.

[0203] After receiving the AI's response, the dialogue state tracker detects that the response contains relevant content such as "domain" and determines that the point of confusion, "how to draw a function graph", is being gradually resolved. It keeps the list of points of confusion unchanged and waits for the next round of interaction to determine whether it has been completely resolved.

[0204] Example 10: Dynamic triggering of rules in multi-turn interactions

[0205] This embodiment demonstrates a scenario where the rule engine dynamically adjusts its strategy based on state changes during multi-round interactions.

[0206] In the first round of interaction, the student's mood was NEGATIVE, cognitive load was MODERATE, and the number of attempts was 1. The rule engine matched the negative emotion high scaffolding rule (priority 2) and the default Socratic guidance rule (priority 0), and merged them into the strategy "high_scaffolding", generating high scaffolding low challenge cue words.

[0207] After the student completes the first round of questions and answers correctly, at the start of the second round of interaction, the rule engine detects that the answer is correct (is_correct = True), matching the Socratic extension rule (priority 2). Simultaneously, the student's emotion remains NEGATIVE (priority 2), and both rules have the same priority. The rule engine sorts the rules by priority, prioritizing the rule that matches first, with the other rule serving as a supplement. The merging strategy is "socratic_extension + high_scaffolding," achieving synergy between the Socratic extension and high scaffolding.

[0208] If a student answers incorrectly three times in a row and the cognitive load remains HIGH, the rule engine matches the perseverance injection rule (priority 3). This rule has a higher priority than other rules and serves as the main decision output, triggering emotional support and value guidance to prevent the student from giving up learning due to repeated setbacks.

[0209] Example 11: Fault Recovery Process for Database Degradation

[0210] This embodiment demonstrates the degradation and recovery process when student profiles are stored in the database in case of a failure.

[0211] Suppose the database is temporarily unavailable due to a network outage. When the system receives a write request, it attempts to update both the memory cache and the database simultaneously. If the database connection times out (exceeding 30 seconds), the write operation fails, and the system automatically enters degrade mode.

[0212] Step H01: Mark the degraded status as True and record the degrade log: {timestamp: "2026-07-20T10:30:00", student_id: "S001", operation: "write", reason: "DBconnection timeout"}.

[0213] Step H02: Update only the memory cache to ensure the write operation returns successfully.

[0214] Step H03: Start a background recovery thread that attempts to reconnect to the database every 30 seconds.

[0215] Step H04: When the database connection is restored (e.g., network recovery), the background thread detects a successful connection and triggers the data synchronization process.

[0216] Step H05: The synchronization process writes all the profile data that was updated during the downgrade from the memory cache to the database.

[0217] Step H06: After synchronization is complete, clear the degrade flag degraded = False and record the recovery log: {timestamp: "2026-07-20T10:31:30", status: "recovered", synced_count: 5}.

[0218] During the degradation period, system read operations always prioritize hitting the memory cache, ensuring that even if the database is unavailable, the system can still provide users with complete profile read and write services. The degradation flag only takes effect when the database is unavailable and is automatically removed after the database is recovered, making it completely transparent to upper-layer business processes.

[0219] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A teaching interaction engine system based on multimodal perception and rule-based reasoning, characterized in that, include: The perception layer includes an emotion processor, a cognitive load processor, and a context enhancer; The emotion processor adopts a dual-mode degradation architecture of pre-trained sentiment analysis model and keyword heuristic matching to classify the emotion of user input text. When the classification confidence of the pre-trained sentiment analysis model is lower than the preset confidence threshold, it automatically degrades to the keyword heuristic matching scheme. The cognitive load processor estimates the cognitive load level based on the typing speed and pause duration during user input; the context enhancer aggregates the emotion analysis results of the emotion processor and the cognitive load estimation results of the cognitive load processor to generate structured enhanced metadata. The decision-making layer includes a rules engine and a dynamic prompt word builder; The rule engine maintains several teaching strategy rules. Each rule includes a condition judgment function, an action execution function, and a priority value. The rule engine traverses all rules, collects decision results for rules that meet the conditions, sorts them by priority, and merges them to generate teaching strategy decisions. The dynamic prompt word builder, based on the teaching strategy decision, selects the corresponding system prompt word template and user prompt word template from the predefined prompt word template library, fills in the context variables, and generates a complete interactive prompt word; The interaction layer includes a dialogue state tracker and a security verification module; The dialogue state tracker maintains the current topic, a list of student confusion points, a set of discussed achievements, a history of dialogue rounds, and the number of attempts for the current question. When it detects that the AI's response contains keywords of a recorded confusion point, it automatically marks the confusion point as resolved. The security verification module performs multi-level compliance checks on the AI-generated response text and generates a safe alternative response for responses that fail the checks. The data layer, including student profile storage, adopts a two-layer storage architecture of memory caching and database persistence. When reading, the memory cache is queried first, and if the cache is not hit, the database is retrieved. When writing, both the memory cache and the database are updated simultaneously. The perception layer, decision layer, interaction layer, and data layer are connected in sequence. The enhanced metadata output by the perception layer serves as the input to the decision layer. The response generated by the interaction prompts produced by the decision layer after inference by the AI ​​model is filtered by the interaction layer and then output. The analysis results of the perception layer and the interaction records of the interaction layer are persisted to the data layer.

2. The teaching interaction engine system based on multimodal perception and rule reasoning according to claim 1, characterized in that, The dual-mode degradation architecture of the emotion processor is as follows: The emotion processor prioritizes loading a pre-trained Chinese sentiment analysis model to classify the input text into emotions, and outputs emotion labels and corresponding model confidence scores. When the pre-trained sentiment analysis model is unavailable or the model confidence is lower than the preset confidence threshold, the sentiment processor automatically downgrades to the keyword heuristic matching scheme. The keyword heuristic matching scheme maintains a positive keyword database and a negative keyword database, and counts the occurrence frequency of positive and negative keywords in the input text. If the number of positive keywords is greater than the number of negative keywords, it is determined to be a positive emotion; if the number of negative keywords is greater than the number of positive keywords, it is determined to be a negative emotion; if the two are equal, it is determined to be a neutral emotion.

3. The teaching interaction engine system based on multimodal perception and rule reasoning according to claim 1, characterized in that, The specific method by which the cognitive load processor estimates the cognitive load level is as follows: Calculate typing speed, where typing speed is the number of characters in the input text divided by the input time; When the typing speed is lower than a preset speed threshold, the speed score is marked as high load; otherwise, it is marked as low load. The number of long pauses whose duration exceeds a preset pause threshold is counted, and the pause score is calculated by dividing the number of long pauses by the total number of pauses. The comprehensive load score is the average of the speed score and the pause score. When the comprehensive load score is greater than a first threshold, it is determined to be a high load. When the comprehensive load score is greater than a second threshold and less than or equal to the first threshold, it is determined to be a medium load. When the comprehensive load score is less than or equal to the second threshold, it is determined to be a low load.

4. The teaching interaction engine system based on multimodal perception and rule reasoning according to claim 1, characterized in that, The rule merging method of the rule engine is as follows: The priority value for each rule is a preset non-negative integer, with higher values ​​indicating higher priority. When the rule engine is executed, it traverses all rules and collects the decision results corresponding to the rules whose condition judgment function returns true into the candidate set. The decision results in the candidate set are sorted in descending order of priority value. The highest priority decision result is the main decision, and the lower priority decision results that do not conflict with the main decision are merged into the final strategy as supplementary decisions. The rules pre-set by the rule engine include: triggering a high scaffolding, low challenge strategy when the emotion label output by the emotion processor is negative; triggering a strategy of persisting in value injection and emotional support when the cognitive load is high and the number of attempts for the current question exceeds 2; triggering a Socratic extension strategy when the answer is correct; triggering a simplified explanation strategy when the cognitive load is high; and a default Socratic guidance strategy.

5. The teaching interaction engine system based on multimodal perception and rule reasoning according to claim 1, characterized in that, The dynamic prompt word builder includes a value injection module; The value injection module maintains a value keyword mapping table, which contains the association between value categories and corresponding introductory templates; When the teaching strategy decision output by the rule engine includes value objectives, the dynamic prompt word builder automatically appends the guiding statement corresponding to the value category to the end of the user prompt word template.

6. The teaching interaction engine system based on multimodal perception and rule reasoning according to claim 1, characterized in that, The multi-level compliance checks of the security verification module are as follows: Level 1 Harmful Content Detection: Traverse the predefined harmful keyword library and detect whether the AI-generated response text contains violent, pornographic, politically sensitive, or discriminatory language keywords. If any keyword is found, it is judged as a violation. Level 2 academic misconduct detection: Traverse the predefined academic misconduct keyword library and check whether the reply text contains keywords that guide the writing of assignments or cheating. If any keyword is hit, it is determined to be a violation. Level 3 Value Alignment Check: When the teaching strategy decision output by the rule engine includes value objectives, verify whether the response text contains keywords of the corresponding value category. If not, it is determined to be a violation. When any level determines a violation, the security verification module selects a security response corresponding to the violation type from a predefined security response template library to replace the original text output.

7. The teaching interaction engine system based on multimodal perception and rule reasoning according to claim 1, characterized in that, The two-tier storage architecture for storing student profiles is as follows: The memory cache stores information using student identifiers as keys and student profile data structures as values. The student profile data structure includes learning history, proficiency scores for each knowledge point, emotional history, and cognitive load history. During a read operation, the memory cache is queried using the target student identifier. If a match is found, the student profile data structure is returned directly. If no match is found, the database is queried using the target student identifier, and the query result is written to the memory cache and then returned. During a write operation, both the memory cache and the database are updated simultaneously. If the database write operation fails, the system automatically downgrades to pure memory mode, records the downgrade log, and automatically synchronizes it after the database is restored.

8. The teaching interaction engine system based on multimodal perception and rule reasoning according to claim 1, characterized in that, The list of confusion points maintained by the dialogue state tracker includes a description text of the confusion point and a timestamp added to the confusion point; When the dialogue state tracker receives AI response text, it iterates through each confusion point description text in the confusion point list, checks whether the AI ​​response text contains keywords from the confusion point description text, and if it does, marks the corresponding confusion point as resolved and removes it from the confusion point list. The dialogue state tracker supports exporting the session state as a context dictionary, which includes the current topic, a list of remaining points of confusion, the current number of attempts, and a dialogue round count, for the rule engine to consume during conditional judgment.

9. A teaching interaction engine method based on multimodal perception and rule reasoning, applied to the teaching interaction engine system based on multimodal perception and rule reasoning as described in any one of claims 1-8, characterized in that, Includes the following steps: Step S1: Receive user input data, which includes text content, student identifier, input time, and a list of pause durations; Step S2: Query the student profile storage based on the student identifier. If no corresponding profile exists, create a new student profile. Step S3: The context enhancer calls the emotion processor to perform emotion analysis on the text content, calls the cognitive load processor to estimate the cognitive load level based on the input time and the pause duration list, and writes the emotion analysis result and the cognitive load estimation result into the enhanced metadata; Step S4: Update the dialogue state tracker, record the current dialogue round, detect and remove resolved points of confusion; Step S5: Merge the enhanced metadata with the dialogue state of the dialogue state tracker into a context dictionary, send it to the rule engine to perform rule matching and decision merging, and output the teaching strategy decision; Step S6: The dynamic prompt word builder selects a template from the prompt word template library according to the teaching strategy decision, fills in the context variables, and adds value guidance when the teaching strategy decision includes value goals, thereby generating the interactive prompt word; Step S7: Append the emotion analysis results and the cognitive load estimation results to the emotion history record and cognitive load history record of the student profile, and write them into the student profile storage; Step S8: Send the response text generated by the AI ​​model based on the interaction prompts to the security verification module to perform multi-level compliance checks; if the response text passes all checks, output the response text. If the response text fails any level of inspection, a security response template corresponding to the violation type will be selected from the security response template library for alternative output.

10. The teaching interaction engine method based on multimodal perception and rule reasoning according to claim 9, characterized in that, The specific steps for the rule engine to perform rule matching and decision merging in step S5 are as follows: Step S51: Traverse all preset rules and collect the decision results of rules whose condition judgment functions return true on the context dictionary into the candidate set; Step S52: Sort the decision results in the candidate set from high to low priority values; Step S53: Select the decision with the highest priority as the main decision; Step S54: Traverse the remaining decision results and merge the decision results that do not conflict with the main decision in terms of attributes and have a lower priority than the main decision into the final teaching strategy decision.