Education training management system and education training method

By employing multimodal perception, data fusion, a dynamic knowledge engine, and a feedback optimization system, the system addresses the limitations of existing teaching systems in terms of personalized adaptation and real-time response. It enables precise perception of learners' states and dynamic adjustment of teaching strategies, thereby improving teaching efficiency and stability.

CN120852103APending Publication Date: 2025-10-28TIANJIN HANGAN EDUCATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510786161.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing teaching systems lack a comprehensive understanding of learners' complex states, cannot accurately judge learners' actual state and immediate needs, and lack the ability to model the evolution trend of knowledge content and the dynamics of learning behavior, resulting in teaching strategies failing to achieve personalized adaptation and real-time response.

Method used

A multimodal perception module is used to collect learners' visual, speech, and interactive behavior data. Spatiotemporal alignment and feature fusion are performed through a data fusion center. A dynamic knowledge engine is used to update the domain knowledge graph in real time based on a graph neural network. Personalized teaching strategies are generated by combining deep reinforcement learning. The system parameters are dynamically adjusted through a feedback optimization system.

Benefits of technology

It enables accurate perception and personalized response to learners' multidimensional states, improves the timeliness and flexibility of teaching content, and ensures the stability and personalized adaptability of the teaching process in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852103A_ABST
    Figure CN120852103A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent education, and discloses an educational training management system and an educational training method, and the system comprises a multi-mode perception module which is used for collecting the visual, voice and interactive behavior data of a learner; the data fusion center is configured to perform space-time alignment and feature fusion on the multi-source heterogeneous data; the dynamic knowledge engine updates a domain knowledge graph in real time based on a graph neural network; the teaching decision-making core is used for generating a personalized teaching strategy by using deep reinforcement learning; the interactive execution unit is used for implementing the teaching strategy and outputting the teaching content; the feedback optimization system is used for dynamically adjusting system parameters according to the teaching effect; the multi-mode sensing module comprises an RGB-D camera array, which is used to collect the facial expression and the limb movement of the learner at a preset sampling frequency; provided is a noise reduction microphone array. According to the invention, by combining the multi-modal perception and the dynamic knowledge graph, a personalized teaching strategy can be generated in real time, and the teaching content can be dynamically adjusted according to the state of the learner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent education technology, specifically to an education and training management system and an education and training method. Background Technology

[0002] Against the backdrop of the ongoing advancement of educational informatization, intelligent teaching systems are gradually becoming important tools in education and training scenarios. Although numerous studies have attempted to combine artificial intelligence, big data, and online teaching systems to promote the personalization and intelligence of the teaching process, existing technologies still have significant shortcomings in many aspects, making it difficult to meet the comprehensive needs of personalized response, real-time adjustment, and continuous optimization in real teaching environments.

[0003] Existing teaching systems generally rely on structured behavioral data as input, lacking comprehensive means to perceive learners' complex states, especially in the acquisition and understanding of unstructured features such as body language, vocal emotions, and operational habits. This deficiency in perception directly leads to the system's inability to accurately determine the learner's actual state and immediate needs, affecting the effective generation of teaching strategies.

[0004] Meanwhile, in the process of allocating teaching resources and recommending learning paths, existing methods are mostly based on static knowledge structures or rule bases, lacking the ability to model the evolution trend of knowledge content and the dynamics of learning behavior. This makes it difficult to accurately identify the correlation and timeliness of different knowledge points, resulting in a discrepancy between recommended content and learners' actual needs, and limiting the degree of personalized adaptation.

[0005] Furthermore, existing systems often lack effective closed-loop feedback mechanisms, failing to fully utilize real-time feedback information during the learning process (such as operational errors, changes in answers, and fluctuations in attention) to refine teaching strategies. Once a teaching strategy is generated, it is often executed linearly, lacking an adaptive adjustment mechanism for process feedback. This results in a slow response to changes in learner status or unexpected situations, leading to insufficient flexibility and robustness in the teaching process. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides an education and training management system and an education and training method. It solves the problem that existing technologies, which are based on static knowledge structures or rule bases, lack the ability to dynamically model the evolution trend of knowledge content and learning behavior, making it difficult to accurately identify the correlation and timeliness of different knowledge points.

[0007] To achieve the above objectives, the present invention provides the following technical solution: an education and training management system and an education and training method, comprising:

[0008] The multimodal perception module is used to collect learners' visual, speech, and interactive behavior data;

[0009] The data fusion center is configured to perform spatiotemporal alignment and feature fusion on multi-source heterogeneous data.

[0010] A dynamic knowledge engine that updates the domain knowledge graph in real time based on graph neural networks;

[0011] The core of instructional decision-making is to generate personalized teaching strategies using deep reinforcement learning.

[0012] Interactive execution units are used to implement teaching strategies and output teaching content;

[0013] The feedback optimization system dynamically adjusts system parameters based on teaching effectiveness.

[0014] Preferably, the multimodal sensing module includes:

[0015] An RGB-D camera array captures learners' facial expressions and body movements at a preset sampling frequency;

[0016] Noise-canceling microphone array, configured with voice enhancement algorithms to eliminate ambient noise;

[0017] The interaction log collection unit records the timestamps and spatial coordinates of operation events.

[0018] Preferably, the data fusion center performs the following:

[0019] Establish a cross-modal time synchronization protocol and constrain the time deviation threshold of data from each sensor.

[0020] A gated attention mechanism is used for feature fusion, and the weight allocation of each modality feature is calculated through learnable parameters.

[0021] Preferably, the dynamic knowledge engine includes:

[0022] Knowledge graph ontology, storing domain concepts and their relationships;

[0023] The real-time data acquisition unit captures the latest information from external knowledge sources;

[0024] Graph neural network processors perform incremental updates to knowledge graphs.

[0025] Preferably, the core of the teaching decision-making process includes:

[0026] Policy networks encode system states based on attention mechanisms;

[0027] Value assessment networks predict the long-term benefits of teaching strategies;

[0028] The reward function module balances immediate feedback with long-term teaching effectiveness.

[0029] Preferably, the feedback optimization system implements:

[0030] The directional parameters of the microphone array are dynamically adjusted according to the ambient noise level.

[0031] Reassign graph node weights based on knowledge hotspot regions;

[0032] The coupling strategy updates parameters based on changes in network gradient and graph structure.

[0033] An education and training management method includes the following steps:

[0034] S1. Collect teaching scenario data through multi-source sensors;

[0035] S2. Perform time synchronization and spatial alignment processing on heterogeneous data;

[0036] S3. Construct a dynamically evolving knowledge graph;

[0037] S4. Generate personalized teaching strategies;

[0038] S5. Perform instructional interactions and collect feedback data;

[0039] S6. Optimize system operating parameters based on feedback information.

[0040] Preferably, the data processing steps include:

[0041] Convolutional neural network encoding for extracting visual features;

[0042] Calculate the Mel-frequency cepstral coefficients of the speech signal;

[0043] Cross-modal attention weight allocation is achieved through a learnable parameter matrix.

[0044] Preferably, the knowledge graph construction includes:

[0045] Identify and rank the importance of knowledge nodes;

[0046] Applying graph neural networks for relational reasoning;

[0047] Perform incremental updates based on the newly acquired data.

[0048] Preferably, the generation of the teaching strategy includes:

[0049] The fused features are jointly encoded with the graph context into a state vector;

[0050] The action probability distribution is output through the policy network;

[0051] The value of a strategy is evaluated based on cumulative rewards based on discounts.

[0052] This invention provides an education and training management system and an education and training method. It has the following beneficial effects:

[0053] 1. This invention deeply integrates multimodal perception with knowledge graph reasoning. It can perceive learners' multidimensional state characteristics at the visual, auditory, and interactive levels, and generate dynamic teaching strategies based on the knowledge graph context, thereby achieving fine-grained responses to learning behaviors. This strategy not only considers knowledge mastery but also integrates operational behavior and emotional feedback, truly embodying a learner-centered personalized teaching model.

[0054] 2. This invention constructs a real-time updated knowledge graph system through a dynamic knowledge engine and utilizes graph neural networks to model and reason about the structural relationships between knowledge points, enabling the system to possess semantic understanding and structural adaptive capabilities for teaching content. By continuously integrating external knowledge resources and learner feedback, the system can achieve the evolution and reconstruction of knowledge structures, enhancing the timeliness and relevance of teaching content.

[0055] 3. This invention integrates a virtual experimental environment, multimedia resource push, and touch operation response mechanism into the interactive execution unit. Simultaneously, it incorporates a multimodal data synchronization processing and anomaly detection module, achieving stable interaction and real-time feedback in complex scenarios. Even in the event of equipment failure, network fluctuations, or other anomalies, the system can maintain the continuity and stability of the teaching process through backup channels and rule-based strategies.

[0056] 4. This invention designs a comprehensive feedback optimization system that structurally evaluates learners' operational performance, emotional state, and knowledge acquisition during the teaching process, and uses this evaluation to adjust teaching tasks and system parameters in real time. This mechanism ensures that the teaching process can continuously adapt to changes in learners' states, constructing a closed-loop teaching system of "perception-decision-execution-feedback-optimization," thereby improving overall teaching efficiency and responsive intelligence. Attached Figure Description

[0057] Figure 1 This is a system architecture diagram of the present invention. Detailed Implementation

[0058] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0059] Please see the appendix Figure 1 This invention provides an education and training management system and an education and training method, including:

[0060] The multimodal perception module is used to collect learners' visual, speech, and interactive behavior data;

[0061] In this embodiment, the multimodal perception module achieves comprehensive data acquisition of the teaching scenario through a heterogeneous sensor network. The specific technical solution includes the following:

[0062] Visual perceptual unit

[0063] Preferably, the visual perception unit employs a multi-view RGB-D camera array, positioned at the front, left, and right sides of the teaching environment. The camera incorporates an infrared fill light device and adjusts the light intensity through an adaptive exposure control algorithm to ensure clear capture of body movement details even in low-light conditions.

[0064] During visual data processing, a keypoint detection model based on the OpenPose framework was used to extract the three-dimensional coordinates of 21 facial feature points and 17 joint points of the learner's limbs.

[0065] Speech perception unit

[0066] The voice sensing unit consists of a circular microphone array. Preferably, the array includes eight omnidirectional microphones distributed at equal angles to form 360-degree sound pickup coverage. The voice signal processing flow includes:

[0067] Noise reduction: A real-time noise reduction algorithm based on spectral subtraction is used.

[0068] Feature extraction: 39-dimensional MFCC coefficients are calculated for the denoised speech signal, including 12 cepstral coefficients, 1 log energy and its first and second order differences.

[0069] Interactive sensing unit

[0070] In this embodiment, the interactive sensing unit is composed of a capacitive touch sensor and an inertial measurement unit (IMU). The touch sensor is embedded in the surface of the experimental operating table and records the coordinates and pressure values ​​of the operating trajectory at a sampling rate of 100Hz; the IMU module is integrated into the operating tool (such as the handle of the experimental instrument) to collect triaxial acceleration and angular velocity data.

[0071] Spatiotemporal alignment mechanism

[0072] To achieve cross-modal data synchronization, this module adopts a hierarchical clock synchronization strategy:

[0073] Hardware-level synchronization: Microsecond-level clock alignment of sensor nodes is achieved through the PTP protocol;

[0074] Software-level compensation: A dynamic time warping (DTW) model is established for visual and audio frames. Preferably, the module is configured with a redundancy sensing strategy to address device failures.

[0075] When the visual sensor fails, an alternative sensing scheme based on Wi-Fi Channel State Information (CSI) is activated to infer limb movements through signal propagation characteristics;

[0076] In case of abnormal voice acquisition, the microphone of the learner's mobile device is switched as a backup signal source, and low-latency transmission is achieved through Bluetooth protocol.

[0077] Data security design

[0078] To protect privacy, this module adopts an edge computing architecture. The raw data is destroyed immediately after feature extraction is completed locally, and the feature vectors are homomorphically encrypted before transmission.

[0079] Through the above technical solutions, the multimodal perception module has achieved comprehensive, accurate, and secure collection of teaching scenario data, providing reliable input for subsequent feature fusion and teaching decisions.

[0080] The data fusion center is configured to perform spatiotemporal alignment and feature fusion on multi-source heterogeneous data.

[0081] In this embodiment, the data fusion center achieves deep integration and feature representation of heterogeneous data through a multi-level processing architecture. The specific technical solution includes the following:

[0082] Spatiotemporal alignment subsystem

[0083] Preferably, the spatiotemporal alignment subsystem employs a hierarchical synchronization strategy. At the hardware level, a precise clock protocol unifies the microsecond-level time reference of sensor nodes, while at the software level, a dynamic time warping algorithm compensates for inherent delays between devices. For the temporal mapping relationship between visual frame sequences and audio streams, a sliding window-based matching mechanism is established, and the optimal offset is determined through cross-correlation analysis to ensure temporal consistency across modal data. For spatial alignment, a multi-camera joint calibration method assisted by a calibration board is used to unify data acquired from various perspectives into a global coordinate system.

[0084] Feature extraction subsystem

[0085] The feature extraction subsystem comprises parallel multimodal processing channels. The visual channel employs a pre-trained residual neural network to extract high-level semantic features, using the activation values ​​of the last fully connected layer as the feature vector. The speech channel extracts acoustic features through Mel-frequency cepstral coefficients and further enhances key frequency band information through a self-attention mechanism. The interaction channel segments and encodes the original operation events, generating temporal features reflecting the operation mode. The output feature dimensions of each channel are normalized to form a unified scale representation.

[0086] Cross-modal fusion subsystem

[0087] This subsystem employs a gated attention mechanism to achieve dynamic feature fusion. Confidence weights for each modality feature are calculated using a learnable parameter matrix, and a temperature adjustment factor is introduced during weight allocation to avoid over-polarization. The fusion layer receives the normalized multimodal feature vectors, performs a weighted concatenation operation, and then inputs them into a bidirectional long short-term memory network to capture the temporal dependencies between cross-modal features. Preferably, the fusion network includes skip connection structures to preserve low-level details of the original features.

[0088] Quality assessment and correction mechanism

[0089] To ensure the reliability of the fusion results, this center has configured a feature quality assessment module. This module detects abnormal data input by calculating the self-consistency index of feature vectors. When the data quality of a certain modality is detected to be lower than a preset threshold, a degradation fusion strategy is initiated, reducing the fusion weight of that modality and generating a quality alarm signal. For persistent abnormal situations, a redundant data source switching mechanism is triggered, calling the backup sensor data stream.

[0090] Security and privacy protection design

[0091] During data transfer, this center implements end-to-end encrypted transmission. Feature vectors undergo homomorphic encryption before entering the fusion network to ensure the original data cannot be reverse-engineered. The fusion computation process is completed in a trusted execution environment, and key parameters employ dynamic obfuscation technology to prevent side-channel attacks. After data processing is complete, a secure erasure operation is immediately performed to eliminate residual information.

[0092] Real-time protection measures

[0093] To address the real-time requirements of teaching scenarios, this center adopts a pipelined architecture design. Each processing stage is executed asynchronously and in parallel through a circular buffer, and the computational resource allocation strategy is dynamically adjusted according to the system load. Preferably, hardware acceleration units are introduced to optimize convolution operations and matrix transformations, ensuring that the single-frame processing latency remains within an acceptable range.

[0094] Through the above technical solutions, the data fusion center has achieved efficient integration and deep correlation of multi-source heterogeneous data, providing a high signal-to-noise ratio fusion feature representation for subsequent teaching decisions, while ensuring the security and real-time performance of the data processing process.

[0095] A dynamic knowledge engine that updates the domain knowledge graph in real time based on graph neural networks;

[0096] In this embodiment, the dynamic knowledge engine supports personalized teaching decisions through a continuously updated and expanded knowledge graph. The engine is designed based on graph neural networks and natural language processing technology to achieve automatic knowledge extraction, fusion, and updating. The specific implementation details are as follows:

[0097] Construction and updating of knowledge graphs

[0098] The core component of a dynamic knowledge engine is a knowledge graph, which consists of multiple entities, attributes, and relationships. Entities in the knowledge graph represent knowledge points, teaching objectives, and learning resources; attributes define the characteristics of entities; and relationships describe the logical connections between entities. During the construction phase, the knowledge graph is described using an ontology language (such as OWL), specifically in the form of triples (entity, attribute, value). For example, an entity "function" might have an attribute "definition" with the value "In mathematics, a function is a mapping that describes the relationship between inputs and outputs."

[0099] To ensure the timeliness and breadth of the knowledge graph, the engine uses a distributed crawler cluster to fetch the latest academic literature and technical reports from multiple open resources (such as arXiv, GitHub, etc.). By using named entity recognition technology based on the BERT model, entities, relationships, and attributes in new documents are automatically extracted and added to the existing graph in the form of triples, forming incremental updates.

[0100] This incremental update uses a graph attention network-based algorithm to calculate the similarity between newly added nodes and existing nodes, thereby reasonably embedding new knowledge into the graph and avoiding redundant or erroneous data import.

[0101] Graph Embedding and Feature Representation

[0102] To facilitate efficient computation and reasoning, nodes and edges in knowledge graphs are transformed into low-dimensional vectors using graph embedding techniques. The graph embedding model constructs vector representations using the structural and attribute information of nodes, and is trained by minimizing the following objective function:

[0103]

[0104] Among them, h i and h j Let e ​​be the embedding vector of node i and i', respectively. ij Represents the edges in the graph. For edge e ij An indicator function for the existence of the graph. The optimization objective aims to maximize the vector similarity of neighboring nodes, thereby effectively preserving the structural information in the graph.

[0105] Cross-modal knowledge fusion

[0106] To enhance the expressive power of information in the knowledge graph, this embodiment introduces multimodal data fusion. Specifically, the dynamic knowledge engine performs cross-modal fusion of the knowledge graph by combining visual, speech, and interaction features acquired from the multimodal perception module. A gated attention mechanism is used to weight the features of each modality to generate a fused feature representation. The weight coefficients of each modality are dynamically adjusted by a learning mechanism to reflect the contribution of different modalities to knowledge representation.

[0107] Graph Reasoning and Intelligent Decision-Making

[0108] After embedding and fusing the graph, the dynamic knowledge engine uses graph neural networks for reasoning to obtain effective teaching strategies and decision support. The reasoning process involves multiple layers of propagation, and node representations incorporate information from surrounding nodes, thereby generating deep features that reflect global knowledge.

[0109] Based on these knowledge representations, dynamic knowledge engines can generate personalized learning suggestions tailored to the characteristics of different learners. For example, by inferring a student's mastery of a specific knowledge point, it can further recommend relevant learning resources or tasks.

[0110] Adaptive update mechanism of knowledge graph

[0111] To maintain the timeliness of the knowledge graph content, the dynamic knowledge engine employs an adaptive knowledge update mechanism. When new information enters the system, the system updates the graph through incremental learning, prioritizing knowledge points in hot areas. To this end, the dynamic knowledge engine sets up a priority queue, determining the timing of updates based on the popularity and relevance of knowledge points. Newly added knowledge is compared with existing nodes using a graph attention network algorithm to calculate similarity, thereby adjusting the representation of nodes in the knowledge graph. Through these technical solutions, the dynamic knowledge engine achieves automatic acquisition, storage, updating, and reasoning of knowledge. Combined with cross-modal fusion and graph neural network technologies, it can provide learners with personalized, real-time updated teaching support and decision recommendations.

[0112] The core of instructional decision-making is to generate personalized teaching strategies using deep reinforcement learning.

[0113] In this embodiment, the teaching decision-making core, as a core component of the system, generates personalized teaching strategies through multi-dimensional input and intelligent algorithms to meet the needs of different learners. Specific implementation methods include the following:

[0114] Input layer and data representation

[0115] The instructional decision-making core receives multimodal fusion features from the data fusion center, including visual, speech, and interactive behavior features. Each modality's features are preprocessed to form a vector with a uniform scale, represented as follows: It incorporates fused features from three perceptual channels: vision, speech, and interaction. These feature vectors contain multidimensional information such as the learner's current learning state, level of engagement, and learning behavior.

[0116] In addition to fusion features, the instructional decision core also receives a graph embedding vector from the dynamic knowledge engine. This vector reflects the knowledge structure of the current instructional content and its matching degree with the learner's knowledge background. This vector can be represented as... It contains information about the relationships between knowledge points.

[0117] Policy Network Architecture

[0118] In the core of instructional decision-making, the policy network is modeled using a Transformer-based architecture. This architecture comprises multiple encoder layers, each consisting of a self-attention mechanism and a feedforward neural network. Preferably, the network structure uses a four-layer stacked Transformer encoder, with eight attention heads per layer. This architecture enables the network to effectively capture temporal information and complex cross-modal dependencies in the input features.

[0119] Based on this, the input layer will fuse features f fusion The feature g of the knowledge graph is concatenated with the feature g to form a joint representation vector. After passing through multiple encoders, the vector yields the corresponding representation h. encoded This serves as the input for generating subsequent decisions.

[0120] Decision generation and action selection

[0121] After processing by the Transformer encoder, the network outputs a representation vector h. encoded The input data will be passed to the decoder layer, which is responsible for generating various aspects of the instructional strategy based on the learner's current needs, including the order of knowledge point explanations, the difficulty of practice questions, and recommendations for multimedia resources. Finally, the decoder outputs an action probability distribution, where the probability P(a) of each action represents the probability of performing a specific instructional action given the input features.

[0122] These instructional actions are generated using the softmax function, with the expression:

[0123]

[0124] Among them, a i W represents the i-th teaching action. a Let b be the weight matrix of the decoder layer. a As a bias term, the softmax function ensures that the sum of the probabilities of all actions is 1.

[0125] Reward Function and Strategy Optimization

[0126] To optimize the teaching decision-making process, this embodiment designs a comprehensive reward function that considers multiple dimensions, including learner accuracy, attention span, and operational proficiency. The expression for the reward function is:

[0127] R=0.6·accuracy+0.3·attention+0.1·proficiency;

[0128] Here, accuracy represents the correctness of the answers, attention represents the learner's concentration, and proficiency measures the learner's skill level in the experiment. This reward function plays a crucial role in the policy update process, guiding the model to adjust the teaching strategy based on the learner's real-time feedback.

[0129] The policy network is optimized using the proximal policy optimization (PPO) algorithm in reinforcement learning, which continuously optimizes the model parameters by maximizing the expected reward.

[0130] Dynamic adaptation and feedback mechanism

[0131] To ensure the decision-making process can flexibly adapt to changes in different learners, the instructional decision-making core also possesses dynamic adaptive capabilities. By tracking and analyzing each learner's long-term performance, the system automatically adjusts its decision-making strategies based on the learner's progress rate and learning patterns. For example, the system might adjust the difficulty of subsequent learning content based on the learner's mastery of specific knowledge points, or improve learning effectiveness by adjusting the order of practice questions. The core of the feedback mechanism lies in real-time monitoring of learner feedback signals and adjusting strategy parameters based on these signals.

[0132] Interactive execution units are used to implement teaching strategies and output teaching content;

[0133] In this embodiment, the interactive execution unit, as an important component of the education and training management system, is responsible for executing actual teaching activities based on the teaching strategies generated by the teaching decision core and for interacting with learners in real time. This unit enables the display of multimedia content, interaction with virtual experimental environments, and the delivery of personalized learning tasks. Its specific technical solutions include the following aspects:

[0134] Multimedia push system

[0135] In this embodiment, the multimedia push system in the interactive execution unit is responsible for pushing different types of teaching content to learners in real time based on the output of the teaching decision core. This content includes, but is not limited to, video tutorials, audio explanations, animated demonstrations, and text materials. All video content is streamed via the HLS (HTTP Live Streaming) protocol to ensure smooth playback under different network conditions. Audio and video content undergo adaptive bitrate adjustment based on the learner's network bandwidth to guarantee the best learning experience.

[0136] During content presentation, the video player supports interactive user control, including basic functions such as pause, fast forward, rewind, and volume adjustment, further enhancing learners' engagement.

[0137] Virtual experimental environment and interaction

[0138] The interactive execution unit also includes a virtual experimental environment where learners can conduct simulated experiments or perform actual operations. This virtual environment is built on WebGL technology and provides a highly immersive user experience through 3D graphics rendering. In the virtual experimental environment, learners can operate using a mouse or touchscreen, performing actions such as dragging, rotating, and scaling to complete tasks related to the teaching content. Each operation step is fed back to the system in real time via touch sensors and a motion capture system, thereby adjusting the learner's progress in the current task.

[0139] For example, in a chemistry experiment, learners can virtually mix chemical reagents. The system will judge the correctness of the experiment in real time based on the learner's actions, guide the next step, and provide feedback. The experimental environment supports various gesture recognition methods, accurately recording the learner's operation trajectory, pressure value, acceleration, and other data through the collaborative work of capacitive touch sensors and inertial measurement units.

[0140] Task generation and personalized push notifications

[0141] Based on the teaching strategies generated by the core teaching decision-making system, the interactive execution unit is responsible for dynamically generating personalized learning tasks and pushing them to learners. These tasks are adjusted according to the learner's current knowledge mastery, attention span, and operational proficiency. Task types include theoretical learning tasks, experimental operation tasks, and practice exercises. The system pushes corresponding task content based on the learner's learning progress and feedback data.

[0142] The task generation utilizes a deep learning-based content recommendation algorithm, enabling differentiated recommendations among learners. For example, the system recommends practice questions or experiments of appropriate difficulty levels based on a learner's mastery of a particular knowledge point, ensuring effective learning at an suitable level.

[0143] User interface and user feedback

[0144] The interactive interface design of the interactive execution unit emphasizes intuitiveness and ease of use. The interface is mainly divided into three areas: a knowledge point navigation area, a content display area, and a virtual teaching assistant area. In the knowledge point navigation area, learners can view and select their current learning progress and incomplete tasks; the content display area presents teaching resources such as videos and text; and the virtual teaching assistant area displays a real-time interactive virtual teaching assistant, providing real-time learning guidance and feedback.

[0145] Using speech recognition and natural language processing technologies, the virtual teaching assistant can answer learners' questions and provide real-time assistance. Meanwhile, the virtual teaching assistant's voice interaction is captured with high quality through an 8-channel circular microphone array, ensuring effective speech recognition even in complex environments.

[0146] During the learning process, after each task or operation step is completed, the system provides real-time feedback based on the learner's behavioral data and the correctness of their answers. Feedback includes confirmation of correct operations, suggestions for correcting errors, and suggestions for adjusting the difficulty of the task.

[0147] Operation event capture and feedback processing

[0148] The interactive execution unit captures every action and interaction event of the learner and feeds it back to the system for processing. These action events include touch operations, voice commands, and experimental operations. Each event includes a timestamp, event type (e.g., click, drag, rotation), coordinate information, and operation intensity (e.g., pressure value or acceleration). This data, as feedback on the learner's interactive behavior, is transmitted to the data fusion center for further optimization of learning strategies and task content.

[0149] In the virtual experimental environment, the capture and feedback processing of operational events are achieved through a combination of capacitive touch sensors and inertial measurement units. Information such as the trajectory and pressure values ​​of each operation are recorded in real time through high-frequency sampling, ensuring that every detail of the learner's operation is accurately recorded and fed back.

[0150] Multimodal interaction and synchronization

[0151] To ensure the interactive execution unit can efficiently coordinate multiple sensory inputs and feedbacks, this embodiment introduces a multimodal data synchronization mechanism. Different types of interactive data (such as visual data, voice data, and touch data) are synchronized through the spatiotemporal alignment mechanism of the data fusion center, ensuring the consistency of each modality's data in time and space. This mechanism achieves collaborative work across modal data through timestamp synchronization and timing alignment algorithms, thereby ensuring the smooth execution of the virtual experimental environment and multimedia content.

[0152] Through the above technical solutions, the interactive execution unit can efficiently execute personalized teaching tasks, provide dynamic display of multimedia content, support real-time interaction in virtual experimental environments, and continuously optimize the learning process through user feedback mechanisms, ensuring that learners receive continuous learning support during interaction.

[0153] The feedback optimization system dynamically adjusts system parameters based on teaching effectiveness.

[0154] In this embodiment, the feedback optimization system is used to perform real-time analysis and dynamic control of multi-dimensional feedback information collected during the teaching process, so as to achieve continuous optimization of teaching strategies and adaptive adjustment of system parameters. This system mainly works in collaboration with the multimodal perception module, the teaching decision core, and the interactive execution unit to construct a closed-loop feedback mechanism, thereby improving teaching responsiveness and personalized adaptability.

[0155] Collection and aggregation of feedback data

[0156] In this embodiment, the feedback optimization system first receives feedback data from multiple modules, including learner behavior trajectory, interaction operation log, voice expression, facial expression, answer record, operation time and completion rate.

[0157] The raw data collected by the multimodal sensing module is preprocessed by the data fusion center and then transmitted to the feedback system in a structured form. Data from different sources are uniformly mapped to the feature space required for feedback analysis to support subsequent unified evaluation and processing.

[0158] Structured analysis of feedback information

[0159] To fully utilize feedback information, the feedback optimization system has a feedback evaluation engine that categorizes received feedback data into structured classes, such as "knowledge mastery feedback," "operational proficiency feedback," and "emotional and focus feedback."

[0160] In practice, the system assigns different evaluation priorities to different types of feedback based on the specific content of the teaching process and learning tasks. By setting feedback evaluation rules, the system transforms each type of feedback information into quantitative indicators to determine the suitability and effectiveness of the current teaching strategy.

[0161] Dynamic adjustment mechanism of teaching strategies

[0162] After completing the feedback evaluation, the feedback optimization system will make necessary adjustments to the teaching strategies generated by the core teaching decision-making mechanism. If it is identified that learners are not performing well on the current task, the system will make immediate adjustments by reducing the difficulty of the content, changing the task format, or adjusting the task order.

[0163] This strategy adjustment is not a static configuration, but a dynamic optimization process driven by real-time data, ensuring that the system can quickly adapt to changes in the learner's state. For example, when the system continuously detects that the learner is distracted or interrupted in multiple tasks, it can automatically reduce the task complexity or extend the completion time.

[0164] In addition, for high-performing learners, the feedback optimization system will push their behavioral feedback to the teaching decision-making core to guide the system to generate more challenging follow-up tasks, avoiding a learning pace that is too slow or too much repetitive content.

[0165] Periodic updates of map parameters and system configuration

[0166] The feedback optimization system also includes a parameter adjustment mechanism for the knowledge engine. This mechanism is based on learners' full-cycle behavioral feedback and periodically updates the weights of nodes and the strengths of edges in the knowledge graph to reflect the current learning group's attention to and depth of mastery of different knowledge points.

[0167] For example, if a knowledge point is frequently missed in feedback from multiple learners, the system will automatically increase the priority weight of that node in the knowledge graph and prompt the teaching strategy to include reinforcement of that knowledge point. Simultaneously, the edge weights between nodes will be periodically decayed based on the actual degree of correlation between concepts to maintain the interpretability and evolution of the knowledge graph structure.

[0168] Adaptive adjustment of equipment parameters and operating status

[0169] During operation in the teaching environment, the feedback optimization system also monitors and adjusts the operating parameters of the hardware devices. For example, when a significant change in ambient light is detected, the system can automatically adjust the camera's exposure time and infrared illumination settings to maintain image acquisition quality.

[0170] Similarly, for the voice processing module, if the system detects increased ambient noise, it can automatically switch to a backup channel mode or increase the noise reduction level to avoid losing critical voice information. Under the premise of long-term stable operation of the sensing module, the system will dynamically adjust the sampling frequency and data upload interval to balance system performance and resource consumption.

[0171] Anomaly detection and emergency response mechanism

[0172] To ensure the robustness of the teaching system in complex application scenarios, this embodiment includes a comprehensive anomaly detection and handling process. When the feedback system detects missing, delayed, or abnormally sudden data from a device, it immediately issues a status alarm and activates the emergency mechanism.

[0173] When the visual channel experiences a signal interruption, the system can activate a low-resolution backup image channel to maintain data flow continuity. When the voice channel fails, the system will guide the user to upload voice input using a mobile application to achieve alternative feedback collection. If the decision module generates a strategy anomaly, the feedback optimization system will call the preset rule base to generate a fallback teaching strategy to ensure that the teaching task is not interrupted.

[0174] Closed-loop optimization and periodic learning

[0175] The feedback optimization system not only supports real-time adjustments based on real-time data, but also includes a periodic batch update mechanism. At fixed time intervals each day, the system performs batch analysis on feedback information accumulated over a period of time, combining this with the overall behavioral trends of multiple users to retrain the model and iterate the strategy.

[0176] Through this mechanism, the system can continuously evolve teaching strategies and parameter settings without affecting real-time operation, thereby improving its long-term teaching adaptability and feedback response capabilities, and achieving true closed-loop dynamic optimization.

[0177] In summary, the feedback optimization system in this embodiment possesses multi-dimensional feedback perception capabilities, dynamic strategy adjustment capabilities, anomaly tolerance capabilities, and a periodic evolution mechanism for the teaching process, ensuring the stable operation and personalized support of the teaching system under complex environments and diverse user needs.

[0178] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. An education and training management system, characterized in that, include: The multimodal perception module is used to collect learners' visual, speech, and interactive behavior data; The data fusion center is configured to perform spatiotemporal alignment and feature fusion on multi-source heterogeneous data. A dynamic knowledge engine that updates the domain knowledge graph in real time based on graph neural networks; The core of instructional decision-making is to generate personalized teaching strategies using deep reinforcement learning. Interactive execution units are used to implement teaching strategies and output teaching content; The feedback optimization system dynamically adjusts system parameters based on teaching effectiveness.

2. The education and training management system according to claim 1, characterized in that, The multimodal sensing module includes: An RGB-D camera array captures learners' facial expressions and body movements at a preset sampling frequency; Noise-canceling microphone array, configured with voice enhancement algorithms to eliminate ambient noise; The interaction log collection unit records the timestamps and spatial coordinates of operation events.

3. The education and training management system according to claim 1, characterized in that, The data fusion center performs the following: Establish a cross-modal time synchronization protocol and constrain the time deviation threshold of data from each sensor. A gated attention mechanism is used for feature fusion, and the weight allocation of each modality feature is calculated through learnable parameters.

4. The education and training management system according to claim 1, characterized in that, The dynamic knowledge engine includes: Knowledge graph ontology, storing domain concepts and their relationships; The real-time data acquisition unit captures the latest information from external knowledge sources; Graph neural network processors perform incremental updates to knowledge graphs.

5. The education and training management system according to claim 1, characterized in that, The core of the teaching decision-making process includes: Policy networks encode system states based on attention mechanisms; Value assessment networks predict the long-term benefits of teaching strategies; The reward function module balances immediate feedback with long-term teaching effectiveness.

6. The education and training management system according to claim 1, characterized in that, The feedback optimization system achieves the following: The directional parameters of the microphone array are dynamically adjusted according to the ambient noise level. Reassign graph node weights based on knowledge hotspot regions; The coupling strategy updates parameters based on changes in network gradient and graph structure.

7. An education and training management method, used in the education and training management system according to any one of claims 1-6, characterized in that, Includes the following steps: S1. Collect teaching scenario data through multi-source sensors; S2. Perform time synchronization and spatial alignment processing on heterogeneous data; S3. Construct a dynamically evolving knowledge graph; S4. Generate personalized teaching strategies; S5. Perform instructional interactions and collect feedback data; S6. Optimize system operating parameters based on feedback information.

8. The education and training management method according to claim 7, characterized in that, The data processing steps include: Convolutional neural network encoding for extracting visual features; Calculate the Mel-frequency cepstral coefficients of the speech signal; Cross-modal attention weight allocation is achieved through a learnable parameter matrix.

9. The education and training management method according to claim 7, characterized in that, The knowledge graph construction includes: Identify and rank the importance of knowledge nodes; Applying graph neural networks for relational reasoning; Perform incremental updates based on newly acquired data.

10. The education and training management method according to claim 7, characterized in that, The generation of the teaching strategy includes: The fused features are jointly encoded with the graph context into a state vector; The action probability distribution is output through the policy network; The value of a strategy is evaluated based on cumulative rewards based on discounts.