Immersive learning method based on linkage of brain-computer interface and virtual scene

Through multi-channel EEG and physiological parameter monitoring, combined with real-time data processing and adaptive adjustment of virtual scenes, the problem of insufficient real-time monitoring and feedback of learners' cognitive status in online education is solved, and personalized immersive learning experience and teaching effect improvement is achieved.

CN120469579APending Publication Date: 2025-08-12ZHEJIANG JINGHANG SHUREN CULTURE MEDIA CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510573140.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing online education system lacks real-time monitoring and feedback mechanisms for learners' internal cognitive status, and cannot make immediate adjustments based on learners' psychological and physiological status, resulting in insufficient personalized adaptation and limited learning effect and experience immersion.

Method used

Multi-channel EEG signal and physiological parameter monitoring are adopted, real-time data preprocessing and feature extraction, combined with end-to-end control logic to realize adaptive adjustment of virtual scenes, and an immersive learning method of wearable brain-computer interface and three-dimensional virtual simulation engine is built to realize closed-loop feedback of teaching content and learner status.

Benefits of technology

It realizes high-precision real-time monitoring and feedback of learners' cognitive status, dynamically optimizes the virtual learning environment, improves teaching effect and experience immersion, and provides personalized learning suggestions and immediate intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469579A_ABST
    Figure CN120469579A_ABST
Patent Text Reader

Abstract

The invention discloses an immersive learning method based on a wearable brain-computer interface and three-dimensional virtual simulation linkage. The system collects EEG, heart rate, galvanic skin, respiration and other multi-mode signals in real time, after band-pass filtering, ICA, wavelet noise reduction and feature extraction, attention indexes and cognitive loads are output within 200 ms by means of a TCN-GNN fusion model, a virtual scene is driven to highlight key knowledge points, step-by-step guidance or difficulty dynamic adjustment, and a learner keeps cardiac flow. The system further optimizes and designs interaction trajectory multi-modal fusion, generates a concentration curve, a cognitive load heat map and personalized tutoring suggestions, and pushes the concentration curve, the cognitive load heat map and the personalized tutoring suggestions to a teacher / parent end in real time through an end-edge-cloud architecture to realize objective quantification of cognitive states and precision of teaching feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of educational technology and human-computer interaction, and more particularly to an immersive learning method that deeply integrates a wearable brain-computer interface (BCI) with a three-dimensional virtual simulation scene. Based on real-time monitoring of multi-channel electroencephalogram (EEG) signals and physiological parameters, this method dynamically adjusts the virtual environment's presentation strategy through intelligent algorithms to proactively adapt to learners' focus and cognitive load, thereby providing young people with a personalized and highly interactive experience of scientific inquiry and knowledge internalization. Background Art

[0002] With the rapid development of information technology and artificial intelligence, brain-computer interface (BCI) and virtual reality (VR) / augmented reality (AR) technologies are becoming increasingly mature. Both have found widespread application in fields such as medical rehabilitation, gaming and entertainment, and military training. However, in education, research on personalized learning and cognitive intervention, particularly for adolescents, is still in its infancy. Most existing online teaching systems rely on pre-designed instructional modules, using videos, animations, and simple interactive questions to engage students' attention. However, they lack real-time monitoring and feedback mechanisms for learners' internal cognitive states. Traditional EEG signal analysis often employs offline batch processing or static threshold-based assessment methods, which can only roughly assess learners' overall focus trends and struggle to capture short-term fluctuations in focus and dynamic changes in cognitive load. VR / AR scenarios typically operate with a fixed difficulty level and pre-set interaction paths, failing to adapt instantly to learners' specific psychological and physiological states. This results in insufficient personalization, limiting learning outcomes and immersive learning experiences. In recent years, the integration of multi-channel EEG systems with portable physiological monitoring devices, such as heart rate, electrodermal conductance, and respiration, has enabled the coordinated acquisition of multi-source time-series signals. Academics have begun experimenting with the use of machine learning and deep neural networks to integrate and analyze multimodal data to identify learners' cognitive load and emotional state. However, these studies have largely remained in laboratory settings and lack effective integration with real-world teaching scenarios. On the one hand, a comprehensive, systematic solution spanning signal preprocessing, feature extraction, and virtual scene adaptation is lacking. On the other hand, there remains a gap in the practical implementation of cognitive load theory and flow theory in educational psychology, hindering the deep integration of theoretical frameworks with real-time interactive technologies. To achieve continuous tracking, analysis, and feedback on learners' multidimensional "body-mind-behavior" states during instruction and to dynamically optimize virtual learning environments based on individual differences, an immersive learning method is urgently needed that can monitor EEG and physiological parameters in real time with high precision and integrate them with virtual scenes. This approach would not only fill the gaps in existing online education in terms of immediate feedback and personalized control, but would also lay the technical foundation for future intelligent education systems driven by large-scale data. Summary of the Invention

[0003] This paper addresses the shortcomings of existing educational technologies in terms of real-time perception of learners' cognitive states and dynamic coupling with virtual scenes. It systematically proposes an immersive learning method based on the deep integration of a wearable brain-computer interface and a three-dimensional virtual simulation engine. This method uses multi-channel electroencephalogram (EEG) signals and physiological parameters such as heart rate, skin charge, and respiratory rate as input. Through real-time data preprocessing and feature extraction algorithms, it obtains the learner's attention index and cognitive load level. Then, based on an end-to-end control logic, it maps the learner's current internal cognitive state to the scene elements, interactive processes, and prompt methods in the virtual environment, adaptively adjusting them. This achieves a closed-loop feedback loop between teaching content and learner status.

[0004] In traditional online learning and simulation-based teaching, the design of teaching scenarios often relies on preset difficulty curves and fixed interaction scripts, which are unable to respond immediately to fluctuations in individual learners' concentration and changes in cognitive load during the learning process. Existing static threshold judgment methods or simple batch analysis can often only provide post-evaluation results, making it difficult to intervene immediately during the learning process, and ignore the instantaneous peaks in cognitive load caused by learners' short-term emotional changes and task switching. What's more, multimodal data fusion research is mostly limited to laboratory environments and lacks effective integration with real-world teaching scenarios, making it difficult for research results to be applied in primary and secondary schools or higher education.

[0005] To address the above technical bottlenecks, the present invention first designs a lightweight wearable brain-computer interface device that is compatible with multiple sensors. Through optimized electrode layout and low-power wireless transmission module, it ensures long-term wearing comfort and continuity of data acquisition. It also proposes a multi-stage signal preprocessing pipeline that combines bandpass filtering, independent component analysis (ICA) and wavelet transform to effectively remove motion artifacts and environmental noise, and extract key indicators including α / β wave power ratio, event-related potential (ERP) characteristics and heart rate variability from multiple perspectives in the time domain, frequency domain and time-frequency domain.

[0006] On this basis, the first part of the invention further elaborates on the mapping strategy between the learner's cognitive state and the adaptive control of the virtual scene. The system has a built-in dual decision-making model based on cognitive load theory and fluency experience theory. When the attention index is lower than the preset threshold or the cognitive load shows a continuous high trend, it automatically triggers scene dimensionality reduction or guided prompts; on the contrary, when the learner is in the optimal range of high concentration and low load, the scene complexity is maintained or moderately increased to maintain the challenge and immersion. Through this closed-loop mechanism, the present invention not only realizes the automation of the entire process from data collection to scene optimization, but also fills the gap in the field of educational technology for real-time multimodal fusion feedback and virtual interaction adaptation at the theoretical and practical levels.

[0007] At the level of the overall system architecture, the present invention innovatively integrates wearable brain-computer interface devices, edge computing nodes and three-dimensional virtual simulation engines to build an end-to-end immersive learning platform composed of three subsystems: the perception layer, the computing layer and the presentation layer. The perception layer is centered on a multi-channel EEG electrode array and heart rate, skin charge and respiration sensors built into a lightweight head-mounted device. Through optimized analog front-end circuits and high sampling rate (≥500Hz) AD conversion modules, it can complete the synchronous acquisition of original signals within 100ms. In terms of transmission, the device uses dual-mode low-power Bluetooth and Wi-Fi communication protocols for seamless switching, ensuring that the data packet loss rate is less than 0.1% in a campus teaching environment and that the single delay does not exceed 50ms to meet real-time feedback needs.

[0008] The computing layer, deployed on edge servers and a cloud-based collaborative platform within the school's local area network, undertakes key tasks such as signal preprocessing, feature extraction, multimodal fusion, and state determination. First, through a multi-stage pipeline of bandpass filtering (0.5–45Hz), independent component analysis, and a denoising wavelet threshold algorithm, the system effectively removes electromyographic artifacts and environmental noise, mapping the cleaned EEG signal to multidimensional features such as the α / β power spectrum ratio and event-related potential temporal peaks. Physiological indicators such as heart rate variability (HRV), galvanic skin response (EDA), and respiratory rate gradient are calculated. Subsequently, the invention employs a deep fusion model based on a combination of an improved temporal convolutional network (TCN) and a graph neural network (GNN). This model utilizes a cross-modal attention mechanism to weightedly fuse EEG features with physiological signals, and maps the learner's virtual environment interaction trajectory (including viewpoint movement, object manipulation, and task completion time) to graph nodes, thereby outputting the Attention Index and Cognitive Load level in real time within 200ms. The fusion model can achieve an accuracy of over 92% on offline large-scale annotated datasets, and the real-time inference latency is less than 150ms.

[0009] The presentation layer is based on the Unity / Unreal dual-engine plug-in structure, and achieves efficient rendering and interactive response by accessing the adjustment instructions issued by the computing layer. The system has a built-in parameterized scene configuration framework, which can modify the lighting intensity, key knowledge point highlighting, dynamic prompt window and level difficulty parameters online; it supports two-way communication of the WebSocket protocol to ensure millisecond-level synchronization between the rendering client and the edge server in command interaction. When the AttentionIndex is lower than 0.4 for three consecutive seconds or the CognitiveLoad exceeds 0.8, the scene immediately triggers the "dimensionality reduction visualization" mode, which helps learners quickly reconstruct knowledge links by enhancing kinematic demonstrations, simplifying physical interaction paths and introducing hierarchical guidance; when the learning state returns to the optimal range, the system can automatically restore or moderately increase the scene challenge to maximize the learner's "flow" experience.

[0010] Through the above-mentioned second technical solution, the present invention not only realizes high-precision, low-latency multi-channel signal acquisition and transmission in hardware design, but also innovatively constructs two core modules in software architecture: multimodal deep fusion and configurable virtual scene closed-loop control, providing powerful real-time, scalability and adaptability for immersive learning.

[0011] The third part of the present invention focuses on the intelligent realization of learning report generation and personalized tutoring suggestions based on multimodal fusion results. The system continuously records and stores the learner's operation logs in each virtual scene module in the background, including viewpoint switching timing, object interaction events, task completion time, and number of error corrections, etc., and synchronously maps them with the AttentionIndex and CognitiveLoad curves to construct a high-precision spatiotemporal label data set. On this basis, the learner's cognitive state and behavioral characteristics are jointly modeled using a multi-task learning framework, and the improved Transformer network with the introduction of self-attention mechanism and residual connection is used to automatically identify learning bottlenecks, such as the frequent occurrence intervals of operational errors corresponding to a certain knowledge point or the period of sudden increase in cognitive load.

[0012] Key dimensions of the model's output include a knowledge point mastery score, a concentration stability index, and an emotional regulation effectiveness assessment. The knowledge point mastery score combines the learner's correct interaction rate and the optimality of their operation path in the virtual experiment to reflect their intrinsic understanding of the core concepts. The concentration stability index measures the fluctuation of focus levels during the learning process using the variance and peak duration of the Attention Index. The emotional regulation effectiveness assessment is based on the electrodermal response curve and heart rate variability, combining the ability to recover emotionally when switching tasks or encountering difficulties.

[0013] When generating a visual report, the system uses dynamic charts and heat maps to intuitively present the above indicators: the concentration curve and cognitive load curve are superimposed on the timeline, and the interactive trajectory heat map marks the learner's high-frequency operation areas and the visual area with the longest stay time in the three-dimensional scene. At the end of the report, based on the pre-defined rule library and online learning path library, it intelligently recommends the next learning plan, including reviewing key knowledge points, optional expansion experimental scenarios, and an appropriate range of difficulty adjustment. The system uses natural language generation technology (NLG) to automatically write feedback summaries for individual learners, and exemplifies specific improvement strategies such as "It is recommended to add multi-perspective demonstrations to the molecular construction experiment module and add step-by-step guidance to the operation prompts to reduce cognitive load" or "Maintain the current difficulty in the geometric transformation scene, and introduce challenging tasks at the right time to maintain a high level of concentration", and push them in real time through the teacher / parent interface.

[0014] Through the intelligent reporting and decision-making support module in the third part, the present invention not only realizes the panoramic quantification and visualization of the learning process, but also transforms the teaching intervention from "teachers' subjective experience + fixed test paper evaluation" to "multimodal data-driven + dynamic personalized recommendation", which greatly improves the timeliness, accuracy and pertinence of teaching feedback, and provides a new technical paradigm for immersive learning for teenagers. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0017] Attachment Figure 1 It is the overall flow chart of the system.

[0018] Attachment Figure 2 This is a schematic diagram of the wearable brain-computer interface and sensor module structure.

[0019] Attachment Figure 3 It is a flow chart of multimodal fusion and real-time adjustment algorithm.

[0020] Attachment Figure 4 It is a flowchart for generating personalized learning reports. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0022] It should be noted that all directional indicators (such as up, down, left, right, front, and back) in the embodiments of the present invention are only used to explain the relative positional relationships and movement of various components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indicators will also change accordingly.

[0023] In addition, the descriptions of "first", "second", etc. in the present invention are only used to distinguish different technical features, and cannot be understood as indicating or implying their relative importance or implicitly limiting the number of the indicated technical features. Therefore, the features marked with "first" or "second" may explicitly or implicitly include at least one of the features. The technical solutions between the various embodiments can be combined with each other, but they must be based on the premise that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0024] In immersive learning systems, the real-time performance, accuracy, and scalability of the system are crucial technical indicators. Faced with the high-frequency acquisition of multi-channel EEG and physiological parameters, the dynamic rendering of large-scale virtual scenes, and the complex calculation of multimodal fusion analysis, how to achieve high-precision cognitive state judgment while ensuring low latency and seamlessly map the results to the adaptive adjustment process of the virtual environment has become the core challenge of the design of this invention. Traditional static threshold triggering or offline analysis methods can no longer meet the dual needs of real-time feedback and refined control in highly immersive learning scenarios for teenagers.

[0025] To this end, this paper proposes an immersive learning method based on the linkage of a wearable brain-computer interface and a three-dimensional virtual simulation scene. By optimizing the hardware acquisition link, building an end-edge-cloud collaborative computing architecture, designing a cross-modal deep fusion model, and combining it with a closed-loop adaptive strategy, it achieves full automation and intelligentization of the entire process, from signal acquisition, data preprocessing, multimodal fusion, dynamic scene adjustment, to personalized report generation. The technical system constructed by this invention not only takes into account system performance and user experience, but also provides a solid foundation for the future deployment of large-scale intelligent education platforms.

[0026] Next, we will describe various embodiments of the present invention in detail. These embodiments, combined with the accompanying drawings, will demonstrate technical details ranging from signal acquisition and preprocessing, multimodal fusion and state determination, virtual scene adaptive adjustment, to learning report generation and push, fully revealing the key innovations of the present invention and its application effects.

[0027] Example 1

[0028] The following is combined with Figure 1 The overall flow chart of the system shown in FIG is used to describe in detail the application example of the present invention in the virtual simulation scene of the youth chemical experiment. Figure 1In modules S101 to S107, learners first wear a wearable brain-computer interface and physiological sensing device (module S101) in a laboratory environment. The device uses a built-in multi-channel EEG electrode array and heart rate, skin conduction, and respiration sensors to collect real-time physiological and EEG signals. The data collected by the device is uploaded to the edge computing node via low-power Bluetooth or Wi-Fi, ensuring real-time upload with a network latency of no more than 50ms (associated with signal acquisition module S102).

[0029] The server performs bandpass filtering, ICA de-aliasing and wavelet noise reduction on the input raw signal according to the preset multi-stage pipeline, extracting α / β power ratio, ERP time series peak, HRV and EDA features (connected to the attached Figure 1 In a chemical reaction rate experiment, when students add reagents to a virtual titrator, there may be a sudden decrease in EEG alpha wave energy or a surge in galvanic skin response. The preprocessing module will quickly identify this transient fluctuation, ensuring signal quality and feature stability.

[0030] Next, the fusion model (attached Figure 1 The multimodal fusion module S104 in the

[104] takes the cleaned EEG features and physiological parameters, as well as the learner's interaction records in the three-dimensional chemical experiment scene - such as the test tube position movement, titration speed and dwell time - as the node input of the graph neural network, and performs weighted fusion through the temporal convolutional network and cross-modal attention mechanism to achieve real-time judgment of the attention index (AttentionIndex) and cognitive load (CognitiveLoad) (association attachment Figure 1 Module S105). In this embodiment, when AttentionIndex is continuously lower than 0.4 and CognitiveLoad exceeds 0.7, the system immediately triggers the scene dynamic adjustment strategy (see Appendix Figure 1 The scene dynamic adjustment module S106 in the titration process switches the titration process interface from the standard laboratory simulation to the "animation demonstration + step prompt" mode. By highlighting the liquid level changes in the virtual beaker and popping up a step-by-step operation guide, it helps students reconstruct the chemical reaction mechanism.

[0031] Furthermore, module S106 automatically restores the normal simulation environment and moderately increases the reaction rate comparison panel when the learner's Attention Index returns to above 0.6 and their Cognitive Load falls below 0.5, maintaining a balance between cognitive challenge and immersive experience. The entire closed-loop feedback process, from signal acquisition to scene adjustment, takes less than 250ms, ensuring continuous, seamless adaptive interaction.

[0032] After the learning phase, the system Figure 1The personalized feedback push module S107 generates a visual report of the experiment's Attention Index fluctuation curve, peak cognitive load periods, and interaction path heatmaps, and automatically pushes it to teachers and parents via the web or mobile app. The report deeply analyzes moments when students often encounter comprehension bottlenecks during the titration process and, using NLG technology, generates targeted tutoring suggestions, such as "Add an ion migration animation to the molecular reorganization step" or "Try manual titration speed control in the next experiment to improve concentration."

[0033] Through the application of this embodiment, not only is a complete closed loop achieved from multi-source physiological and EEG signal acquisition, preprocessing, deep fusion to real-time adaptation of virtual scenes, but the high-precision perception and efficient intervention of the system of the present invention on students' cognitive state are also verified in chemical experiment teaching, laying a solid foundation for improving teaching effects and learning experience.

[0034] Example 2

[0035] The following is combined with Figure 2 The diagram below shows the structure of the wearable brain-computer interface and sensor module, which further describes the hardware implementation of the head-mounted device. The device consists of six functional modules: a multi-channel EEG electrode array (S201), a heart rate sensor (S202), a galvanic skin sensor (S203), a respiratory rate sensor unit (S204), a low-power wireless communication module (S205), and a power management unit (S206). In module S201, electrodes are distributed according to the international 10-20 system in key locations such as the frontal (Fp1, Fp2), central (Cz), and parietal (Pz) lobes. There are a total of 16 dry conductive silicone electrodes, coated with an Ag / AgCl coating to improve contact stability and ensure input impedance greater than 1 GΩ and a common-mode rejection ratio (CMRR) exceeding 100 dB. The S201 signal is first preamplified by a low-noise operational amplifier (such as the TI INA333) with a programmable gain of up to ×1000 and anti-aliased by a sixth-order Butterworth low-pass filter (cutoff frequency 250Hz). It is then precisely sampled by a TIA DS1299 24-bit delta-sigma analog-to-digital converter at a channel sampling rate of 500Hz and timestamped to ensure nanosecond alignment of the data across each channel.

[0036] Module S202 utilizes a Maxim MAX30101 PPG sensor, capturing blood oxygen and heart rate signals at a raw sampling rate of 1kHz. This signal is filtered and downsampled to 200Hz by an internal DSP module, producing high-precision time-series data required for heart rate variability analysis. Module S203 utilizes two stainless steel electrodes placed on the inner temples, recording galvanic skin response (EDA) at a 10Hz rate using an ADS7124 ADC to dynamically monitor emotional fluctuations. Module S204 integrates a Bosch BMP384 MEMS differential pressure sensor into an adjustable chest strap. This sensor uses a micro-tracheal tube to sense changes in chest pressure during breathing, sampling at 100Hz to analyze respiratory rate and depth. All analog signals are time-division multiplexed before S205, where they are uploaded to the dual-mode wireless unit in real time via the DMA channel on the STM32F407 ARM Cortex-M4F microcontroller.

[0037] In module S205, the main wireless chip uses a Nordic nRF52840, enabling Bluetooth 5.2LE communication. An additional ESP32 coprocessor supports Wi-Fi 802.11b / g / n downlink. A proprietary TDMA media access protocol divides the frequency channel into fixed time slots, allocating simultaneous upload windows for up to 10 devices in a single campus environment. This keeps packet loss rates below 0.1% and round-trip latency below 50ms. The microcontroller firmware, running on a FreeRTOS kernel, processes EEG, PPG, EDA, and respiratory data in real time using priority scheduling. Software-implemented digital filtering uses 0.5–45Hz bandpass (EEG), 0.05–5Hz bandpass (EDA), 30–150Hz bandpass (PPG), and 10–40Hz high-pass (respiration). Windowed power spectra and variability indices are then calculated.

[0038] The S206 power management module houses a 3.7V, 1200mAh lithium-polymer battery and a MAX17048 fuel gauge, providing intelligent charging control and battery health monitoring. It also utilizes dynamic voltage scaling (DVS) technology to adjust the voltage of the MCU and RF chip in real time based on processing load. A micro piezoelectric energy harvester integrated into the headband structure recovers up to 50mW of energy from even the slightest head movement, extending battery life to over eight hours. The four-layer flexible printed circuit (FPC) features independent ground and power planes, and a metal shield surrounds sensitive analog signal paths to suppress electromagnetic interference and inter-channel crosstalk. The device weighs less than 350g. After testing in an environmental laboratory at 25°C and 20–80% RH, the noise floor of each sensor channel is below 0.5μV RMS, and the signal-to-noise ratio (SNR) exceeds 60dB. These features fully meet the hardware requirements for high-precision, multimodal real-time acquisition, providing a solid and reliable hardware foundation for subsequent device-edge-cloud multimodal fusion analysis.

[0039] Example 3

[0040] The following is combined with Figure 3 The multimodal fusion and real-time adjustment algorithm flow chart shown in the figure provides an in-depth description of the software algorithm implementation of the present invention. First, the system inputs the raw signals (EEG, heart rate, electrodermal conductance, and respiration) from the wearable device into the "Raw Signal Input" module, designated S301 in the figure. This module is responsible for providing a unified data format and time base for subsequent processing. Each channel signal is accompanied by a nanosecond timestamp to ensure high-precision timing alignment during the subsequent fusion stage. The signal then flows into the signal preprocessing module S302. Here, the system performs bandpass filtering to remove out-of-band noise (0.5–45Hz for EEG, 0.05–5Hz for EDA, 30–150Hz for PPG, and 10–40Hz for respiration), independent component analysis (ICA) to remove electromyographic artifacts, and multi-scale noise reduction based on wavelet thresholding to ensure clean output time series features.

[0041] The preprocessed data is fed into the S303 feature extraction module, which extracts high-dimensional feature vectors in the time, frequency, and time-frequency domains. These include, but are not limited to, alpha / beta power ratios, event-related potential (ERP) peaks, heart rate variability (HRV) statistics, and galvanic skin response peaks. Using a sliding window and overlap strategy, this module generates a new feature vector every 100ms, precisely meeting real-time requirements. The feature vectors are then aggregated into the S304 cross-modal attention fusion module, which comprises two sub-architectures: a temporal convolutional network (TCN) and a graph neural network (GNN), connected by shared attention weights. The TCN branch excels at capturing short-term patterns in each signal, while the GNN branch builds a graph structure and performs information propagation based on the learner's virtual scene interaction trajectory. The two branches are fused at the multi-head attention level, enabling temporal features and spatial interaction information to complement each other and generate a final fused representation.

[0042] The fused representation is passed to the state output module (S305). Here, the system outputs continuous values of AttentionIndex and CognitiveLoad based on a softmax classifier and normalizes them to the 0–1 range. The system then proceeds to the decision-making node (S306). If AttentionIndex falls below 0.4 or CognitiveLoad rises above 0.8, the "high load / low focus" branch is triggered; otherwise, the scene complexity is maintained or increased. In the adjustment strategy execution module (S307), the system invokes the dimensionality reduction visualization, guided prompting, and dynamic difficulty adjustment interfaces pre-installed in the virtual engine to update the 3D scene parameters with sub-millisecond instructions. The entire closed-loop processing from S301 to S307 takes less than 200ms, and after model pruning and 8-bit quantization, the memory usage is only approximately 4MB, ensuring seamless deployment on edge GPUs or high-performance embedded SoCs.

[0043] To verify the algorithm's performance, this embodiment conducted offline training and online testing on a real-world dataset of 500 adolescent participants. The model achieved an average cross-validation accuracy of 92.7% and an F1 score of 0.90 for 10,000 signal samples labeled with low, medium, and high cognitive load levels. In an online deployment environment, the response latency was within an average of 150ms, and the frame loss rate was less than 0.5%. This embodiment demonstrates that the multimodal fusion and real-time adjustment algorithm not only achieves high-precision cognitive state determination but also ensures low-latency adaptation in complex virtual scenarios, significantly improving the interactive fluency and educational effectiveness of immersive learning.

[0044] Example 4

[0045] The following is combined with Figure 4 The personalized learning report generation flow chart shown in the figure describes in detail the implementation of the present invention in the report construction and intelligent recommendation links. After the learning process is completed, the system stores all recorded multimodal time series data and interaction logs in the data storage unit of module S401, including AttentionIndex and CognitiveLoad sequences, heart rate variability, skin electrodermal response and respiratory rate curves marked with nanosecond timestamps, as well as each viewpoint switch, object interaction and task completion time of the learner in the virtual scene. The storage layer introduces a hybrid architecture that combines a distributed time series database with a graph database, which can not only efficiently execute time series queries, but also quickly retrieve interaction paths based on the graph model.

[0046] In the multi-task Transformer modeling module S402, the present invention adopts an improved structure that combines a multi-head self-attention mechanism with a residual connection, and inputs time series features and graph structures in parallel to multiple Transformer branches to process the three subtasks of cognitive state prediction, behavioral pattern mining, and emotional recovery assessment respectively. This module improves feature reuse by sharing encoder layers, and applies different task-specific heads to each task on the decoder side to generate three outputs in parallel: knowledge point mastery, concentration stability, and emotion regulation effect. During the training phase, labeled historical measured data is used, and through multi-task loss weighting and dynamic task scheduling strategies, each subtask achieves the best balance between the total loss convergence speed and the model generalization ability.

[0047] Next, in the indicator calculation module S403, the system maps the Transformer outputs into quantifiable indicators: Knowledge point mastery (S403a) combines the student's correct interaction rate and the optimality of the operation path in the virtual experiment, converted to a 0–100 score scale; Concentration stability (S403b) assesses the fluctuation of attention using the AttentionIndex variance and the duration of the peak; and Emotional Regulation Assessment (S403c) calculates the emotional recovery coefficient based on the peak-to-trough velocity of the electrodermal and heart rate variability. After all indicators are normalized, they are fed into the visualization chart synthesis module S404. This module uses a dynamic generation engine to draw superimposed curves in time series and renders a heat map on the 3D scene interaction plane. It marks the areas where students frequently operate and spend the most time, and adds markers to the curves to highlight learning bottlenecks.

[0048] Finally, in the NLG tutoring suggestion generation module S405, the system automatically writes targeted scenario improvements and subsequent learning plan suggestions based on the predefined strategy library and real-time model output, using a hybrid generation method that combines templates and deep learning. For example, when a cognitive load peak is detected in the molecular construction experiment module, the text "It is recommended to enable multi-perspective animation and refine the guidance prompts in the next operation" is generated; when concentration fluctuations are found to be drastic in geometric transformation scenario learning, the strategy of "dividing the task into two stages and adding instant feedback tests" is provided. Finally, through the S406 report push module, the complete learning report together with interactive charts is sent to the teacher / parent end in real time via WebSocket or push notification, and a simplified feedback summary is retained on the student end, realizing a closed-loop connection between teaching intervention and review planning.

[0049] Technical terms that need to be explained to help understand the present invention

[0050] (1) Brain–Computer Interface (BCI): An interactive system that converts user intent into control commands by collecting, extracting, and classifying EEG or other neural signals, without requiring peripheral nerves or muscles. The BCI device in this invention integrates dry EEG electrodes, amplification and filtering circuits, and a low-power wireless transmission unit to acquire the learner's EEG signals in real time and send data to the computing layer for cognitive status determination.

[0051] (2) Electroencephalogram (EEG): reflects the weak potential changes in the electrical activity of brain neuron groups and is usually divided into frequency bands such as δ, θ, α, β, and γ. EEG signals have high temporal resolution but low signal-to-noise ratio. They need to be preprocessed through bandpass filtering, independent component analysis (ICA), and wavelet threshold noise reduction to extract features such as the α / β wave energy ratio and event-related potential (ERP) peak, which serve as important indicators for measuring attention and cognitive load.

[0052] (3) Cognitive Load: This concept originates from cognitive psychology and refers to the inherent psychological burden an individual bears during information processing. Its level can be comprehensively assessed using multimodal parameters such as EEG characteristics (e.g., the θ / β ratio), heart rate variability (HRV), and galvanic skin response (EDA). In this invention, when the cognitive load exceeds a preset threshold, the system automatically triggers dimensionality reduction visualization of the virtual scene or guided prompts to reduce the learner's psychological pressure.

[0053] (4) Attention Index: A continuous quantitative indicator used to measure a learner's concentration at a specific moment. It is usually calculated based on the EEG α-wave and β-wave power ratio, ERP amplitude, and short-term fluctuations in heart rate variability, and is normalized and mapped to the 0–1 range. The present invention uses the Attention Index to determine whether a student is in the optimal "flow" range and dynamically adjusts the virtual scene difficulty and prompt frequency accordingly.

[0054] (5) Event-Related Potential (ERP): The temporal response of the cerebral cortex to a specific stimulus. Its components (e.g., P300, N100) reflect attention allocation and information processing. ERP has high temporal accuracy and can be used to detect learners' instantaneous responses to key knowledge points or prompts. This method uses it as an important basis for enhancing attention determination during the feature extraction stage.

[0055] (6) Heart Rate Variability (HRV): The statistical or frequency domain characteristics of short-term fluctuations in heart rate intervals (RR intervals) can reflect the interaction of the autonomic nervous system and emotional state. In this paper, HRV, together with physiological signals such as EDA and respiratory rate, constitutes a multimodal input for objectively assessing learners' cognitive load and emotional regulation ability.

[0056] (7) Electrodermal Activity (EDA): Changes in skin conductance due to sweat gland activity are often used to measure an individual's emotional arousal level. This paper records EDA signals at the temples using dual electrodes and combines them with EEG and HRV input into a fusion model to improve the accuracy of monitoring the learner's emotional state.

[0057] (8) Temporal Convolutional Network (TCN): A deep learning architecture specifically designed for processing time series signals, with causal convolution and dilated convolution mechanisms, capable of capturing long-term dependencies while maintaining low latency. The present invention uses a TCN branch in the multimodal fusion module to extract temporal patterns from preprocessed physiological and EEG features. (9) Graph Neural Network (GNN): A deep model capable of performing message passing and representation learning on graph-structured data. The present invention constructs virtual scene interaction trajectories as graph nodes and edges. The GNN branch is used to capture the learner's spatial-behavioral relationship between objects in a three-dimensional environment and performs cross-modal attention fusion with the TCN output.

[0058] (10) Natural Language Generation (NLG): A technology that automatically generates highly readable text descriptions or suggestions based on structured data output by a model. This paper combines a predefined strategy library with a deep learning generation model to implement automatic summarization of learning reports and the writing of tutoring suggestions in the report generation module.

Claims

1. An immersive learning method based on the linkage between brain-computer interface and virtual scene, characterized in that: The following steps are involved: (1) The learner wears a wearable brain-computer interface device, which is used to collect the learner's electroencephalogram (EEG) signals and physiological parameters in real time; (2) cleaning the EEG signals and physiological parameters through bandpass filtering, independent component analysis, and wavelet threshold noise reduction preprocessing, and extracting multidimensional features including α / β wave power ratio, event-related potential peak, heart rate variability, and skin electrodermal response; (3) In a three-dimensional virtual simulation scenario, the learner's attention index and cognitive load level are jointly determined based on the multidimensional features through a temporal convolutional network and a graph neural network fusion model; (4) When the attention index is lower than a preset threshold or the cognitive load is higher than a preset threshold, the scene adaptive adjustment is automatically triggered, which includes highlighting key knowledge points, popping up step-by-step guidance prompts, and dynamically reducing or delaying the scene difficulty; when the attention index and cognitive load return to the optimal range, the scene challenge is automatically restored or moderately increased; (5) performing multimodal fusion analysis on the multidimensional features and the learner's operation trajectory in the virtual scene to construct a comprehensive learning state model; (6) Generate a visual learning report based on the comprehensive learning status model, the report including a concentration curve, a cognitive load heat map and an interaction trajectory heat map, and automatically generate personalized tutoring suggestions in natural language, which are pushed in real time through the teacher / parent end.

2. The method according to claim 1, characterized in that The signal preprocessing process includes: performing 0.5–45 Hz band-pass filtering on the EEG signal, performing 0.05–5 Hz band-pass filtering on the EDA signal, performing 30–150 Hz band-pass filtering on the PPG signal, and performing 10–40 Hz high-pass filtering on the respiratory signal; using independent component analysis (ICA) to remove motion and electromyographic artifacts; and then performing multi-scale noise reduction using a wavelet threshold algorithm.

3. The method according to claim 1, characterized in that The scenario adaptive adjustment includes: when the AttentionIndex is lower than 0.4, the "dimensionality reduction visualization" mode is triggered, highlighting the target object and simplifying the interaction path; when the CognitiveLoad is higher than 0.8, a step-by-step guidance prompt pops up; when the learning state returns to a state where the AttentionIndex is higher than 0.6 and the CognitiveLoad is lower than 0.5, the normal mode is automatically restored and the difficulty parameters are appropriately increased.

4. The method according to claim 1, wherein The multimodal fusion analysis adopts a deep model that combines an improved temporal convolutional network (TCN) and a graph neural network (GNN), uses a cross-modal attention mechanism to perform weighted fusion of EEG features, physiological indicators and virtual scene interaction trajectories, and outputs AttentionIndex and CognitiveLoad within 200ms.

5. The method according to claim 1, wherein The wearable brain-computer interface device includes: a multi-channel dry EEG electrode array, a heart rate sensor (PPG), a skin charge sensor (EDA), a respiratory rate sensor unit, a low-power wireless communication module and a power management unit.

6. The method according to claim 1, characterized in that The learning report includes: a superimposed curve of the learner's AttentionIndex and CognitiveLoad, a three-dimensional scene interaction heat map, and personalized tutoring text generated by NLG.

7. An immersive learning system according to claim 1, characterized in that: include: Wearable brain-computer interface devices, edge computing nodes, cloud-based multimodal fusion servers, and virtual simulation engines work together to form an end-edge-cloud architecture consisting of the perception layer, computing layer, and presentation layer.

8. The system according to claim 7, characterized in that The wearable brain-computer interface device establishes a low-latency, low-packet-loss rate data link with the edge computing node through Bluetooth 5.2 and Wi-Fi 6E dual-mode communication.

9. An immersive learning device based on the linkage between brain-computer interface and virtual scene, characterized in that: include: EEG acquisition unit, physiological parameter monitoring unit, signal processing unit, fusion judgment unit, virtual scene presentation unit and report generation unit.

10. The device according to claim 9, characterized in that The signal processing unit has built-in bandpass filtering, ICA and wavelet threshold noise reduction algorithms, and performs timestamp synchronization and feature normalization on the collected signals before neural network inference.

Citation Information

Cited By

  • Multi-modal remote cooperative interaction system integrating artificial intelligence and virtual reality

    CN121657875A

  • Closed-loop adaptive psychological intervention method and system based on multi-modal brain-computer fusion

    CN122025028A