Multi-modal fusion virtual simulation experiment teaching system

The multimodal fusion virtual simulation experiment teaching system utilizes multimodal acquisition equipment and pre-trained models to generate adaptive teaching strategies, solving the problems of inaccurate capture of operational behaviors and incomplete evaluation in traditional virtual simulation experiments, thereby improving teaching effectiveness and interactivity.

CN121389787APending Publication Date: 2026-01-23CHINA UNIV OF GEOSCIENCES (BEIJING)

Patent Information

Application Number
CN202511570871.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Traditional virtual simulation experimental teaching systems cannot accurately capture learners' multimodal operational behaviors in real time, lack personalized teaching intervention, and have incomplete assessment of learning outcomes.

Method used

A virtual simulation experimental teaching system employing multimodal fusion is used to synchronously acquire gesture, voice, and tactile interaction information through multimodal acquisition devices, construct a multimodal input dataset, generate operation instruction sequences using a pre-trained multimodal fusion model, generate adaptive teaching intervention strategies by combining reinforcement learning, and evaluate the learning effect in real time.

Benefits of technology

It realizes virtual experimental simulation and intelligent teaching intervention driven by multimodal information fusion, which improves the accuracy of experimental interaction experience and learning effect evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121389787A_ABST
    Figure CN121389787A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal fusion virtual simulation experiment teaching system, and relates to the technical field of virtual simulation, and the system comprises an information collection module which is used for obtaining the interaction information of a learner; the information fusion module is used for constructing a multi-modal input data set; the instruction analysis module is used for analyzing and generating an operation instruction sequence; the teaching intervention strategy generation module is used for generating a self-adaptive teaching intervention strategy; the teaching intervention strategy execution module is used for executing a self-adaptive teaching intervention strategy; and the learning effect evaluation module is used for quantifying the learning effect evaluation score. The technical problems that in traditional virtual simulation experiment teaching, multi-modal operation behaviors of learners are difficult to accurately capture in real time, personalized teaching intervention is lacked and learning effect evaluation is incomplete can be solved, and virtual experiment simulation and intelligent teaching intervention driven by multi-modal information fusion are achieved. And the experimental interaction experience and the learning effect evaluation precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of virtual simulation, and in particular to a multi-modal fusion virtual simulation experiment teaching system. BACKGROUND

[0002] With the development of information technology and virtual reality technology, virtual simulation experiments have become an important teaching method in higher education and vocational training. Traditional virtual experiment systems usually interact through a mouse, keyboard or simple touch control, which cannot fully capture the natural operation behavior of learners, resulting in insufficient immersion in the experiment experience and difficulty in quantifying the teaching effect. In addition, existing virtual experiment systems rely on fixed processes for teaching intervention, lack personalized adaptation to the operation behavior and real-time performance of learners, and are difficult to achieve dynamic teaching guidance and effective error correction.

[0003] In terms of teaching effect evaluation, traditional systems mainly rely on experiment completion or operation results, ignoring key behavior data such as operation accuracy, task efficiency and attention distribution of learners during the experiment, and cannot fully reflect the learning effect. These problems restrict the application and promotion of virtual experiments in complex experimental skill training and high-level education. SUMMARY

[0004] The present application provides a multi-modal fusion virtual simulation experiment teaching system, which solves the technical problems of being difficult to accurately capture multi-modal operation behavior of learners in real time, lacking personalized teaching intervention, and incomplete learning effect evaluation in traditional virtual simulation experiment teaching.

[0005] The present application provides a multi-modal fusion virtual simulation experiment teaching system, which comprises: An information collection module for synchronously collecting gesture information, voice information and tactile interaction information of learners using multi-modal collection devices; an information fusion module for constructing a multi-modal input data set based on the gesture information, voice information and tactile interaction information; an instruction analysis module for inputting the multi-modal input data set into a pre-trained multi-modal fusion model, analyzing and generating an operation instruction sequence, and driving virtual experimental equipment to perform physical simulation to obtain real-time operation feedback; a teaching intervention strategy generation module for generating an adaptive teaching intervention strategy based on the real-time operation feedback and combining the behavior data of learners; a teaching intervention strategy execution module for executing the adaptive teaching intervention strategy through voice prompts and visual guidance, adjusting the teaching content and experimental process; and a learning effect evaluation module for evaluating the learning effect based on operation accuracy, task efficiency and attention distribution after the experiment is completed, and quantifying the learning effect evaluation score.

[0006] One or more technical solutions provided in the present application have at least the following technical effects or advantages: The information acquisition module is used for synchronously acquiring gesture information, speech information and tactile interaction information of a learner by using a multi-modal acquisition device; the information fusion module is used for constructing a multi-modal input data set based on the gesture information, speech information and tactile interaction information; the instruction analysis module is used for inputting the multi-modal input data set into a pre-trained multi-modal fusion model, analyzing and generating an operation instruction sequence, and driving a virtual experimental apparatus to perform physical simulation to obtain real-time operation feedback; the teaching intervention strategy generation module is used for generating an adaptive teaching intervention strategy based on the real-time operation feedback and in combination with behavior data of the learner; the teaching intervention strategy execution module is used for executing the adaptive teaching intervention strategy through speech prompting and visual guidance, and adjusting teaching content and an experimental process; and the learning effect evaluation module is used for evaluating a learning effect based on operation accuracy, task efficiency and attention distribution after the experiment is completed, and quantifying a learning effect evaluation score. The technical problems of being difficult to accurately capture multi-modal operation behaviors of a learner in real time in traditional virtual simulation experiment teaching, lacking of personalized teaching intervention, and learning effect evaluation being not comprehensive are solved, and the technical effects of realizing multi-modal information fusion driven virtual experimental simulation and intelligent teaching intervention, and improving experimental interaction experience and learning effect evaluation accuracy are achieved. BRIEF DESCRIPTION OF DRAWINGS

[0007] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0008] Figure 1 FIG. 1 is a structural schematic diagram of a multi-modal fusion virtual simulation experiment teaching system according to the present application; Figure 2 FIG. 2 is a flowchart of an instruction analysis module of a multi-modal fusion virtual simulation experiment teaching system according to the present application.

[0009] The reference signs are explained as follows: information acquisition module 11, information fusion module 12, instruction analysis module 13, teaching intervention strategy generation module 14, teaching intervention strategy execution module 15, and learning effect evaluation module 16. DETAILED DESCRIPTION

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0011] Embodiment one, as shown in FIG. 1, a multi-modal fusion virtual simulation experiment teaching system according to the present application comprises an information acquisition module, an information fusion module, an instruction analysis module, a teaching intervention strategy generation module, a teaching intervention strategy execution module and a learning effect evaluation module. Figure 1As shown, the present application provides a multi-modal fusion virtual simulation experiment teaching system, wherein the system comprises: An information collection module 11 is configured to synchronously collect gesture information, voice information and tactile interaction information of the learner by using a multi-modal collection device.

[0012] Specifically, in the information collection module 11, the multi-modal collection device is composed of a depth camera, a voice collection unit, a tactile feedback device and the like. Through the multi-modal collection device, the operation behavior information of the learner in the virtual experiment can be synchronously collected, including gesture information, voice information and tactile interaction information. The gesture information can accurately describe the operation action of the learner; the voice information can assist in understanding the operation intention of the learner; and the tactile interaction information can reflect the physical behavior characteristics of the learner interacting with the virtual equipment. By collecting these information, a basis can be provided for subsequent multi-modal fusion and operation instruction analysis, and the accuracy of adaptive intervention, learning effect evaluation and the like can be ensured.

[0013] Further, the information collection module 11 comprises: The hand action sequence of the learner is collected by using the depth camera, the gesture trajectory and posture key points are extracted as the gesture information; the voice input signal is obtained by using the voice collection unit, and the semantic analysis is performed to extract the semantic instruction and emotional characteristics as the voice information; and the tactile feedback device is used to record the tactile pressure, vibration and duration information in the process of the learner interacting with the virtual equipment as the tactile interaction information.

[0014] In a preferred embodiment, in the information collection module 11, the hand movements of the learner are continuously captured by using a depth camera or other motion capture device to form a sequence of hand movements, and each frame of image is processed by image processing and pose estimation algorithms such as convolutional neural network (CNN), regression or heat map, vector operation or rotation matrix to extract palm position, finger joint angle and gesture key points, construct a gesture key point matrix, and then analyze the trajectory based on the position information of the gesture key points in consecutive frames to obtain the hand movement trajectory and movement speed, direction and other characteristics, forming a complete gesture information dataset to reflect the learner's operation in the virtual experiment. At the same time, the speech collection unit samples the speech input of the learner during the experiment, converts the speech signal into digital audio data, and then converts the audio signal into text information through a speech recognition model such as end-to-end automatic speech recognition or pre-trained speech model, and combines natural language processing technology to analyze the semantics, extract semantic information such as operation instruction intent, request, question, etc., and in addition, emotion features in the speech can be identified by using emotion analysis methods, such as nervousness, hesitation or confidence level, to form a speech information dataset to provide auxiliary information for judging the current state of the learner. When the learner interacts with the virtual experimental equipment, the haptic feedback device such as force sensor or vibration haptic device will record the operation data of the learner, including touch pressure, force direction, vibration intensity and duration, etc. Through sampling and quantization processing of these data, a haptic interaction information dataset is formed to reflect the mechanical characteristics and operation habits of the learner in operating the virtual equipment.

[0015] The information fusion module 12 is used to construct a multi-modal input dataset based on the gesture information, speech information and haptic interaction information.

[0016] Specifically, in the information fusion module 12, after obtaining the gesture information, speech information and haptic interaction information, these information will be time-synchronized, aligning the gesture, speech and haptic on the time axis, ensuring that the learner's hand movements, speech intent and haptic interaction behavior can be corresponded at the same time point. After time synchronization, feature standardization processing and time sequence fusion are performed to obtain the final multi-modal input dataset, providing sufficient data support for intelligent analysis of learner's operation behavior and implementation of adaptive teaching intervention.

[0017] Further, the information fusion module 12 includes: The gesture information, speech information and haptic interaction information are respectively time-synchronized and feature-standardized, and the standardized results are time sequence fused based on a multi-dimensional feature alignment algorithm to output a multi-modal input dataset.

[0018] In an optional embodiment, in the information fusion module 12, in order to ensure the consistency of gesture information, speech information and tactile interaction information in time, the gesture information, speech information and tactile interaction information are interpolated or resampled to make them correspond to the corresponding operation behaviors on the same unified time axis, for example, the gesture key point data can be interpolated by a linear interpolation method to make it align with the speech frame and the tactile sampling frame, and accurate time synchronization is realized. After time synchronization, the data of the three modalities are standardized, wherein the joint angle, gesture trajectory and other features in the gesture information are mapped to a unified numerical range by Min-Max normalization; the semantic features and emotional feature vectors in the speech information are processed by mean variance standardization or Min-Max normalization; the pressure, vibration and duration parameters in the tactile interaction information are also processed by Min-Max normalization, so that the modal features have comparability in dimension, providing a basis for subsequent fusion. After that, the standardized multi-modal features are fused by a multi-dimensional feature alignment algorithm, that is, at each time step, the gesture, speech and tactile feature vectors are spliced into a single multi-modal vector according to the dimension, so as to fuse the feature sequences of the three modalities after time axis alignment into a unified multi-modal input data sequence, and form the final multi-modal input data set. This multi-modal input data set not only retains the feature information of each modality, but also accurately reflects the operation behavior and its time evolution of the learner in the experiment process, providing a sufficient data basis for subsequent multi-modal instruction analysis, virtual experiment simulation and adaptive teaching intervention.

[0019] The instruction analysis module 13 is used for inputting the multi-modal input data set into the pre-trained multi-modal fusion model, analyzing and generating an operation instruction sequence, and driving the virtual experimental equipment to perform physical simulation to obtain real-time operation feedback.

[0020] Specifically, in the instruction analysis module 13, after obtaining the multi-modal input data set, the multi-modal input data set is input into the pre-trained multi-modal fusion model, which is constructed based on the attention mechanism neural network model, can process the joint representation of gesture, speech and haptic features, and map the comprehensive operation intention of the learner to a specific operation instruction sequence. The operation instruction sequence includes specific parameters such as operation action, sequence, force and duration of virtual experimental equipment, which is used to describe the experimental steps that the learner wants to complete in the virtual experiment. Then, the generated operation instruction sequence is input into the physical simulation engine of the virtual experimental equipment. The physical simulation engine matches the equipment physical parameter library according to the operation instruction, such as mechanical parameters, heat conduction parameters, chemical reaction parameters, etc., performs numerical simulation calculation, drives the virtual equipment to present corresponding physical changes, simulates the actual influence of the learner's operation on the virtual experimental environment, and records the generated real-time operation feedback, which is used to assist subsequent adaptive teaching intervention, improve the interactivity and effectiveness of experimental teaching.

[0021] Further, as shown in Figure 2 The instruction analysis module 13 includes: Collecting multi-modal experimental operation data, the multi-modal experimental operation data taking experimental standard gesture information, experimental standard speech information and experimental standard haptic interaction information as input data, and taking experimental standard operation instruction as supervision label; training a neural network model based on attention mechanism based on the multi-modal experimental operation data, and constructing a pre-trained multi-modal fusion model; using the pre-trained multi-modal fusion model to analyze the multi-modal input data set, and generating an operation instruction sequence.

[0022] In a preferred embodiment, in the instruction parsing module 13, multi-modal experimental operation data for model training is collected, which includes experimental standard gesture information, experimental standard speech information and experimental standard haptic interaction information. Each set of experimental standard gesture information, experimental standard speech information and experimental standard haptic interaction information constitutes input data corresponding to an experimental standard operation instruction, which is used as a supervision label to train the model to identify and map the relationship between operation behavior and instructions. Subsequently, the collected multi-modal experimental operation data is used to construct a training set, and the gesture, speech and haptic features are time-aligned and standardized before being input into a deep neural network model. The deep neural network model uses a multi-modal fusion structure based on an attention mechanism, which automatically learns the importance weight of different modal features in operation instruction generation through an attention layer, achieving dynamic weighted fusion of features. During the training process, the experimental standard operation instruction is used as a supervision signal, and sequence prediction is used for optimization. By minimizing the timing difference loss function between the predicted instruction sequence and the supervision signal, the model parameters are updated so that the model can accurately generate an operation instruction sequence that conforms to the experimental procedure logic. After training, a pre-trained multi-modal fusion model can be constructed. The pre-trained multi-modal fusion model analyzes the received multi-modal input data set through the attention mechanism and time series modeling to automatically generate an operation instruction sequence, which includes detailed parameters such as operation action, sequence, force, and duration, to drive the virtual experimental equipment for physical simulation, providing reliable data support and technical foundation for subsequent virtual experimental simulation and adaptive teaching.

[0023] Further, the instruction parsing module 13 further includes: The physical simulation of the virtual experimental equipment includes at least one of mechanical simulation, heat conduction simulation and chemical reaction simulation; the physical parameter sequence of the virtual experimental equipment is determined by matching the physical parameter library of the virtual experimental equipment according to the operation instruction sequence, and the physical parameter sequence includes at least one of mechanical parameters, heat conduction parameters and chemical reaction parameters; the real-time operation feedback is generated by the physical simulation engine according to the numerical simulation of the physical parameter sequence.

[0024] In an optional embodiment, in the instruction analysis module 13, to realize that the operation behavior of the learner corresponds to the real physical response of the virtual experimental equipment, a corresponding physical parameter sequence is matched according to the operation instruction sequence, wherein the physical simulation process of the virtual experimental equipment includes at least one of mechanical simulation, heat conduction simulation and chemical reaction simulation, when the operation instruction involves behaviors such as movement, collision, extrusion and stretching of an object, a mechanical simulation unit is called; when it involves behaviors such as heating, cooling and temperature change, a heat conduction simulation unit is called; when the operation instruction contains processes such as dissolution, reaction, generation or precipitation, a chemical reaction simulation unit is called. According to the experimental content, the system can execute one of the simulations alone, or can execute multiple simulation types in parallel to realize the multi-physical field coupling of a complex experimental environment. When matching the physical parameters, the operation instruction sequence is compared with the pre-established virtual experimental equipment physical parameter library to match the corresponding parameter set. When it is mechanical simulation, mechanical parameters such as mass, elastic modulus, friction coefficient, moment of inertia and damping coefficient are extracted; when it is heat conduction simulation, thermal parameters such as thermal conductivity, specific heat capacity, density and temperature boundary condition are extracted; when it is chemical reaction simulation, chemical reaction parameters such as reaction rate constant, reaction order, product concentration change coefficient and energy change parameter are extracted. The matched parameters are summarized according to the time step to organize a physical parameter sequence, ensuring that the dynamic changes of the parameters at each stage in the simulation process can accurately reflect the physical effects of the operation behavior. Then, the physical simulation engine receives these physical parameter sequences and performs numerical calculation based on the corresponding mathematical model. For mechanical simulation, finite element method or rigid body dynamics algorithm is used to calculate the stress state, displacement, deformation and collision response; for heat conduction simulation, the temperature field distribution is solved based on the Fourier heat conduction equation; for chemical reaction simulation, the reaction rate, product generation amount and energy release are calculated based on the reaction kinetics equation and thermodynamic equilibrium model. The physical simulation engine records the simulation results at each time step in real time during the calculation process, including the state change of the experimental equipment, the reaction progress, the temperature change curve or the stress deformation, the operation error and other information, generates real-time operation feedback, which is used as basic data for teaching intervention for subsequent adaptive teaching adjustment, thereby significantly improving the immersion and interactivity of virtual experimental teaching.

[0025] Further, the inventors found in the process of implementing the embodiments of the present application that the existing virtual simulation teaching system usually faces the problems of large calculation amount and high response delay in the physical simulation process, especially when multiple physical processes (such as mechanics, heat conduction and chemical reaction) are simulated in parallel, the real-time performance of the system is difficult to guarantee, resulting in lag of operation feedback of the learner and decline of interactive experience. In view of this technical problem, the inventors add a physical simulation and real-time compromise unit to the instruction analysis module, which is used to dynamically balance between simulation accuracy and system response speed, which specifically includes the following steps: According to the complexity of the operation instruction sequence and the type of the virtual experimental equipment, a target simulation level is determined, and adaptive switching is performed between high-precision simulation and real-time approximate simulation. The calculation step and time resolution are set for physical processes such as mechanics, heat conduction, and chemical reaction, and are dynamically adjusted based on real-time performance monitoring results. When the system response delay is detected to exceed a preset threshold, a lightweight physical modeling algorithm is triggered to approximately solve part of the non-critical physical processes to ensure real-time interactive experience. At key nodes of the experiment, the high-precision simulation mode is restored, and the approximate calculation error is corrected through local recalculation and differential compensation mechanism to maintain the overall simulation accuracy.

[0026] Specifically, the compromise unit first determines the corresponding simulation level according to the complexity of the operation instruction sequence and the category of the virtual experimental equipment. When the operation involves high-complexity processes such as fine mechanics, heat conduction, or chemical reaction, the high-precision simulation mode is enabled. When the operation is a routine action or the system detects that the response delay exceeds the threshold, the real-time approximate simulation mode is switched to ensure the continuity of system interaction.

[0027] Secondly, during the simulation execution process, the system adaptively adjusts the time step and resolution of physical solving according to the current calculation load and frame rate fluctuation, and approximately calculates non-critical processes using a lightweight physical modeling algorithm. Finally, at key nodes of the experiment or stages where accurate output results are required, the system automatically restores the high-precision simulation mode, corrects the previous approximate error through local recalculation and differential compensation mechanism, and realizes the dynamic compromise between physical simulation and real-time performance. This makes the system not only ensure the scientific accuracy of the virtual experiment process, but also significantly improve the real-time interaction performance of the system, enabling students to obtain a smooth, realistic, and highly consistent feedback experimental experience.

[0028] The teaching intervention strategy generation module 14 is used to generate adaptive teaching intervention strategies based on the real-time operation feedback and the behavior data of the learners.

[0029] Specifically, in the teaching intervention strategy generation module 14, after obtaining the real-time operation feedback, the real-time operation feedback and the behavior data of the learners during the experiment are input into the decision model constructed based on reinforcement learning for dynamic evaluation of the current state, to determine whether the operation has errors, whether guidance is needed, etc. Adaptive teaching intervention strategies are generated according to the judged state, which include specific operation guidance information, voice prompt content, visual guidance scheme, etc., to ensure that learners can correct operations and optimize behaviors in virtual experiments, and improve the individualization level, interactivity, and learning effect of virtual experiment teaching.

[0030] Further, the teaching intervention strategy generation module 14 comprises: The real-time operation feedback at least includes experimental results, equipment state changes, and operation error prompts; behavior data of the learner during the experiment is collected, the behavior data at least including operation time, operation sequence, and repeated operation times; the real-time operation feedback and the behavior data are input into a decision model constructed based on reinforcement learning, the current state of the learner is dynamically judged, and a teaching intervention strategy for the current state is generated.

[0031] In a preferred embodiment, in the teaching intervention strategy generation module 14, the real-time operation feedback obtained by the physics simulation engine at least includes experimental results, equipment state changes, and operation error prompts, wherein the experimental results are whether the current operation achieves the expected experimental effect, such as whether the chemical reaction is completed or the mechanical device is correctly moved; the equipment state changes are the physical state changes of the virtual experimental equipment, such as position, deformation, temperature change, etc.; the operation error prompts are the detected deviations or errors in the operation of the learner, such as step omission, sequence error, or insufficient operation force, etc. Then, the behavior data of the learner during the experiment is collected, which at least includes operation time, operation sequence, and repeated operation times, wherein the operation time is the time consumed to complete each step of operation, reflecting the operation efficiency of the learner; the operation sequence is the sequence of the learner performing the experimental steps, which is compared with the standard experimental steps; the repeated operation times are the number of times the learner repeats the same step in the experiment, which is used to judge the proficiency or operation confusion. Subsequently, the real-time operation feedback and the behavior data of the learner are input into the decision model constructed based on reinforcement learning, which maps the operation state, experimental results, and behavior data of the learner into environmental state information, for example, when the learner makes an operation error, the step is incomplete, or the attention is scattered, etc., these behavior data are labeled as different state types for subsequent state recognition of the model. Subsequently, the decision model takes the teaching intervention strategy as the output action space, including prompting key steps, demonstrating standard operations, issuing warnings, or providing knowledge supplements, etc. different intervention methods. During the training process, the model continuously adjusts the strategy through interaction with the environment, and gives a positive reward when the selected teaching intervention can effectively improve the operation performance of the learner or shorten the task completion time; if the intervention is ineffective or leads to a decrease in learning effect, a punishment signal is given. Then, the model uses the policy iteration and value function approximation methods of reinforcement learning to continuously optimize the intervention decision, so that it can automatically select the optimal teaching intervention strategy under different learner states. After sufficient training, the decision model can judge the current state of the learner in real time during actual operation without human intervention, such as whether there is a lack of knowledge, operation error, or attention dispersion, etc., and adaptively generate the most appropriate teaching intervention measures, realize individualized and intelligent experimental teaching support, and improve the interactivity, guidance, and learning effect of virtual experimental teaching.

[0032] The teaching intervention strategy execution module 15 is used to execute the adaptive teaching intervention strategy through voice prompts and visual guidance, and adjust the teaching content and experimental process.

[0033] Specifically, in the teaching intervention strategy execution module 15, when the adaptive teaching intervention strategy is generated by the decision model, the content of the adaptive teaching intervention strategy is first processed into instructions, which is converted into an executable teaching instruction template. For strategies that require language communication, the instruction content will be converted into natural speech prompts, and the content, speed and tone of the speech prompts will be dynamically adjusted according to the current experimental stage of the learner, the error type and the attention concentration area, for example, when it is detected that the learner misses a step in the experimental operation, a voice prompt such as "Please pay attention to adding reagent" or "Please check the temperature control step" is given to help the learner correct in time. At the same time, the intervention strategy is visually presented in the form of graphics, animations or highlighted identifiers through the virtual experiment interface, for example, the equipment or area of the wrong operation is highlighted or the color is highlighted, the next action to be performed is guided by arrow, path or animation demonstration. If the learner is in a state of knowledge deficiency, a standard operation demonstration or key step explanation can be popped up in the interface to assist understanding of the experimental principle and process. Then, the difficulty or progress of the virtual experiment is dynamically modified according to the real-time performance of the learner. When it is detected that the learner completes the operation skillfully and accurately, the experiment process is automatically promoted or the experiment complexity is increased; otherwise, when it is detected that the operation error is frequent or the learning efficiency is low, the experiment pace is slowed down, more guidance prompts are provided or the task difficulty is reduced to ensure that the learner can complete the experiment task within an understandable range. Through the above-mentioned way, a multi-modal teaching intervention execution mechanism of voice and visual dual channels is realized, so that the learner can correct and consolidate knowledge under the guidance of hearing and vision, thereby effectively improving the interactivity, real-time and individualization level of virtual simulation experiment teaching.

[0034] The learning effect evaluation module 16 is used to evaluate the learning effect based on the operation accuracy, task efficiency and attention distribution after the experiment is completed, and to quantify the learning effect evaluation score.

[0035] Specifically, in the learning effect evaluation module 16, after the experiment is completed, the operation data of the learner in the entire experimental process is first extracted from the virtual experiment interaction record, including operation accuracy, task efficiency and attention distribution, wherein the operation accuracy is used to reflect the standardization and proficiency of the learner; the task efficiency is used to measure the time management and operation fluency of the learner; and the attention distribution is used to evaluate the attention concentration degree and coordination rationality of the learner in the experimental process. Then, the three indexes are normalized, and the normalized results are comprehensively calculated by weighted summation to generate a learning effect evaluation score, which is used to quantitatively reflect the operation performance of the learner, and to ensure the quality of virtual simulation experiment teaching and the accuracy of learning feedback.

[0036] Further, the learning effect evaluation module 16 includes: The operation accuracy, task efficiency and attention distribution of the learner during the whole experiment are obtained; the operation accuracy, task efficiency and attention distribution are normalized; the weight proportion is set based on the coefficient of variation method, and the normalized results are weighted and summed to obtain the learning effect evaluation score.

[0037] In a preferred embodiment, in the learning effect evaluation module 16, after the experiment is completed, the operation steps, task duration, attention area, gaze time and other data of the learner during the experiment are first extracted from the virtual experiment interaction record, and the operation accuracy, task efficiency and attention distribution are calculated based on these data. Subsequently, the operation accuracy, task efficiency and attention entropy are normalized to convert the indicators of different dimensions and value ranges to the same interval, thereby ensuring that the operation accuracy, task efficiency and attention distribution have comparability under the same evaluation system and avoiding weight deviation caused by dimensional differences. Then, the weight proportion of the three types of indicators is determined by the coefficient of variation method. The coefficient of variation reflects the volatility and discriminability of the indicator data. The greater the volatility, the higher the weight, thereby highlighting the indicators that have more significant influence on the evaluation results. For example, when the experimental task requires high operation accuracy, the weight of operation accuracy is correspondingly increased; when the experiment focuses on efficiency, the weight of task efficiency is correspondingly increased. Finally, the three indicators after normalization are weighted and summed according to the determined weight to calculate the learning effect evaluation score. This learning effect evaluation score not only reflects the comprehensive experimental performance of the learner, but also serves as a quantitative basis for subsequent personalized teaching intervention and learning path optimization, thereby providing an objective and traceable evaluation mechanism for virtual simulation experiment teaching.

[0038] Further, the learning effect evaluation module 16 includes: The correct operation step number and the total step number are subjected to ratio operation to obtain the operation accuracy; the benchmark task duration and the actual task duration are subjected to ratio operation to obtain the task efficiency; and the gaze time proportion of the attention area is subjected to entropy calculation to obtain the attention distribution.

[0039] In an optional embodiment, in the learning effect evaluation module 16, by comparing the actual operation steps of the learner with the standard experimental operation process, the ratio of the correct step number to the total step number can be obtained to calculate the operation accuracy. By recording the actual completion time of the experiment and performing ratio operation with the standard experimental preset time, the task efficiency can be obtained. The virtual experiment interface is divided into a plurality of fixed areas, such as instruments, buttons, displays and the like, and the gaze time of each attention area is obtained (which can be collected by an eye tracker). By dividing the gaze time of each attention area by the total gaze time, the gaze proportion of each attention area is obtained, and the gaze proportion is substituted into the entropy formula to calculate the attention entropy, thereby representing the attention distribution and providing reliable basic data support for subsequent comprehensive evaluation of the learning effect.

[0040] To sum up, the embodiments of the present application have at least the following technical effects: Firstly, the information collection module 11 synchronously acquires gesture information, speech information and tactile interaction information of the learner through the multi-modal acquisition device, and comprehensively records the natural behavior characteristics of the learner in the virtual experiment. Secondly, the information fusion module 12 performs feature extraction, time synchronization and standardization processing on the modal data, constructs a multi-modal input data set based on a multi-dimensional feature alignment algorithm, and realizes unified expression of data in different perception channels. Subsequently, the instruction analysis module 13 inputs the fused multi-modal input data set into the pre-trained multi-modal fusion model, generates the operation instruction sequence of the learner through model analysis, and drives the virtual experimental equipment to perform physical simulation accordingly. The simulation process includes numerical simulation of physical phenomena such as mechanics, heat conduction and chemical reaction, to generate operation feedback information in real time. Then, the teaching intervention strategy generation module 14 uses real-time operation feedback and behavior data of the learner as input, dynamically judges the current learning state of the learner based on a reinforcement learning decision model, and automatically generates a targeted adaptive teaching intervention strategy. Then, the teaching intervention strategy execution module 15 feeds back the intervention strategy to the learner in real time through voice prompts and visual guidance, dynamically adjusts the teaching content and experimental process, and realizes personalized teaching guidance and rhythm control. Finally, the learning effect evaluation module 16 quantitatively evaluates the overall performance of the learner in multiple dimensions after the experiment is completed, and comprehensively reflects the operation accuracy, proficiency and concentration of the learner. Through the cooperative operation of the above modules, the whole-process intelligent closed loop from data collection, instruction analysis, teaching intervention to effect evaluation can be realized, and the interactivity, adaptability and scientific evaluation level of virtual simulation experiment teaching can be improved.

[0041] The above description of disclosed embodiments enables one of ordinary skill in the art to make or use the application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0042] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalents, the present application also intends to include these modifications and variations.

Claims

1. A multi-modal fusion virtual simulation experiment teaching system, characterized in that, The system comprises: An information collection module for synchronously collecting gesture information, speech information and tactile interaction information of a learner by using a multi-modal collection device; An information fusion module for constructing a multi-modal input data set based on the gesture information, speech information and tactile interaction information; An instruction analysis module for inputting the multi-modal input data set into a pre-trained multi-modal fusion model, analyzing and generating an operation instruction sequence, and driving a virtual experimental apparatus to perform physical simulation to obtain real-time operation feedback; A teaching intervention strategy generation module for generating an adaptive teaching intervention strategy based on the real-time operation feedback and in combination with behavior data of the learner; A teaching intervention strategy execution module for executing the adaptive teaching intervention strategy through speech prompts and visual guidance to adjust teaching content and experimental procedures; A learning effect evaluation module for evaluating learning effect based on operation accuracy, task efficiency and attention distribution after completion of the experiment to quantitatively evaluate a learning effect evaluation score.

2. The multi-modal fused virtual simulation experiment teaching system of claim 1, wherein, The information collection module comprises: A depth camera is used to collect a hand movement sequence of the learner, and gesture trajectory and posture key points are extracted as gesture information; A speech collection unit is used to acquire a speech input signal and perform semantic analysis to extract semantic instructions and emotional features as speech information; A tactile feedback device is used to record tactile pressure, vibration and duration information during interaction between the learner and the virtual apparatus as tactile interaction information.

3. The multi-modal fused virtual simulation experiment teaching system of claim 2, wherein, The information fusion module comprises: The gesture information, speech information and tactile interaction information are respectively subjected to time synchronization and feature standardization processing; The standardized processing results are subjected to time sequence fusion based on a multi-dimensional feature alignment algorithm to output a multi-modal input data set.

4. The multi-modal fused virtual simulation experiment teaching system of claim 1, wherein, The instruction analysis module comprises: Multi-modal experimental operation data is collected, the multi-modal experimental operation data taking experimental standard gesture information, experimental standard speech information and experimental standard tactile interaction information as input data and taking experimental standard operation instructions as supervision labels; A neural network model based on an attention mechanism is trained based on the multi-modal experimental operation data to construct a pre-trained multi-modal fusion model; The pre-trained multi-modal fusion model is used to analyze the multi-modal input data set to generate an operation instruction sequence.

5. The multi-modal fused virtual simulation experiment teaching system of claim 4, wherein, The instruction analysis module further comprises: The physical simulation of the virtual experimental apparatus includes at least one of mechanical simulation, heat conduction simulation and chemical reaction simulation; A physical parameter library of the virtual experimental apparatus is matched according to the operation instruction sequence to determine a physical parameter sequence of the virtual experimental apparatus, the physical parameter sequence including at least one of mechanical parameters, heat conduction parameters and chemical reaction parameters; A physical simulation engine performs numerical simulation according to the physical parameter sequence to generate real-time operation feedback.

6. The multi-modal fused virtual simulation experiment teaching system of claim 1, wherein, The teaching intervention strategy generation module comprises: The real-time operation feedback at least includes experimental results, apparatus state changes and operation error prompts; Behavior data of the learner during the experiment is collected, the behavior data at least including operation time, operation sequence and repeated operation times; The real-time operation feedback and the behavior data are input into a decision model constructed based on reinforcement learning to dynamically determine a current state of the learner and generate a teaching intervention strategy for the current state.

7. The multi-modal fused virtual simulation experiment teaching system of claim 1, wherein, The learning effect evaluation module comprises: Operation accuracy, task efficiency and attention distribution of the learner in the whole experiment are obtained. The operation accuracy, task efficiency and attention distribution are normalized. Based on the coefficient of variation method, the weight proportion is set, and the normalized results are weighted and summed to obtain a learning effect evaluation score.

8. The multi-modal fused virtual simulation experiment teaching system of claim 7, wherein, The learning effect evaluation module comprises: The operation accuracy is obtained by ratio operation of the number of correct operation steps and the total number of steps. The task efficiency is obtained by ratio operation of the benchmark task time and the actual task time. The attention distribution is obtained by entropy calculation of the proportion of gaze time in the attention area.

Citation Information

Patent Citations

  • Multi-modal semantic fusion human-computer interaction system and method for virtual experiments

    CN111665941A

  • Intelligent simulation system and method based on multi-modal interaction and dynamic teaching strategy

    CN119624711A

  • Immersive teaching system and method based on virtual reality technology

    CN120070116A

  • Cardiopulmonary resuscitation training system based on multi-mode artificial intelligence combined with virtual reality technology

    CN120199131A

  • Online network teaching practice method based on virtual technology

    CN120317836A

Cited By

  • Learning strategy evaluation and guidance method based on experimental operation sequence pattern mining

    CN121707796A

  • Medical anthropomorphic dummy real-time feedback teaching method and system

    CN121724288A

  • Rail transit transport car operation simulation method and device

    CN122239515A