Multi-modal teaching resource intelligent adaptation and immersive learning interaction system
By using multimodal resource management and real-time data collection and analysis of learner cognitive perception modules, combined with an intelligent adaptation engine and immersive interactive scene generation, the problem of learner cognitive status not being addressed in existing teaching resource systems has been solved. This has enabled virtual-real integration, multi-sensory feedback, and creative interaction, forming a personalized and in-depth learning loop and improving the continuity and engagement of learning outcomes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-21
- Publication Date
- 2026-03-10
AI Technical Summary
Existing multimodal teaching resource systems fail to monitor learners' cognitive states in real time, resulting in a disconnect between resource delivery and learning needs. Immersive learning lacks integration of virtual and real worlds and multi-sensory feedback, leaving learners isolated and helpless. The learning process is disjointed, lacking creative interaction and effective feedback mechanisms, making it difficult to meet personalized and in-depth learning needs.
Through a multimodal resource management module, a learner cognition and behavior perception module, an intelligent adaptation engine, an immersive interactive scene generation and rendering module, and a learning effect closed-loop feedback module, learners' real-time data collection and analysis are realized, multimodal resources are dynamically matched, immersive interactive scenes that blend the virtual and real worlds are generated, multi-sensory feedback and creative interaction are supported, and a coherent learning trajectory is formed through cross-scene linkage.
It achieves precise resource matching driven by cognition, enhances learners' participation and focus, promotes knowledge internalization and innovation, solves the problem of isolated learning processes, forms a complete learning loop, and improves the coherence and depth of learning outcomes.
Smart Images

Figure CN121636024A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent education, in particular to a multi-modal teaching resource intelligent adaptation and immersive learning interactive system. BACKGROUND
[0002] As is known to all, with the deepening of the digital transformation of education, multi-modal teaching resources and immersive learning technologies are widely used in various educational scenarios, providing learners with more diverse knowledge acquisition approaches. However, the existing technology still has many core defects that need to be solved, which seriously affects the learning effect and experience.
[0003] Most existing systems match resources based on the static labels of learners, only considering fixed information such as grade and subject, and fail to pay attention to the real-time cognitive state of learners, such as changes in concentration, cognitive fatigue level, and specific knowledge confusion points, resulting in a serious disconnection between resource pushing and the current learning needs of learners, making it difficult to achieve true individualized instruction, and unable to adjust the presentation form, difficulty gradient, and pushing order of resources according to the real-time state of learners.
[0004] Traditional VR / AR learning scenarios are mostly independent virtual environments, lacking effective integration with real learning environments, and learners' operations in virtual scenarios are disconnected from real physical spaces, making it difficult to associate virtual learning experiences with actual application scenarios. At the same time, the interactive form is relatively single, mostly limited to passive operation responses such as clicking virtual buttons and watching preset animations, lacking multi-sensory interactive experiences, relying only on visual and auditory feedback, and failing to combine tactile, olfactory, and other sensory information to enhance immersion, resulting in limited immersion and difficulty in stimulating learners' deep participation.
[0005] Existing immersive learning is mostly based on single-person independent experience, lacking the social and collaborative scenarios commonly seen in real classrooms, and learners are isolated in the learning process, unable to deepen their understanding and application of knowledge through collaboration and communication. In addition, the three stages of pre-class preparation, in-class learning, and post-class review are disconnected, the learning trajectory is not coherent, pre-class questions cannot be naturally converted into in-class learning tasks, and in-class learning results are difficult to effectively extend to post-class review. The feedback mechanism is only for a single learning task, and it fails to form a complete closed loop of resource adaptation, interactive learning, effect evaluation, and optimization iteration, affecting the gradual deepening and consolidation of knowledge.
[0006] In the existing system, learners only passively receive knowledge as the experiencers of the scene, lack tools and platforms for autonomous creation and design of interactive content based on learned knowledge points, and cannot strengthen knowledge internalization through active output. At the same time, there is a lack of effective evaluation and feedback mechanism for the creation behavior of learners, it is difficult to find the knowledge gaps existing in the creation process, and it is also difficult to provide targeted optimization suggestions and expansion tasks, resulting in that learning stays at the level of shallow understanding, and it is difficult to improve the knowledge application and innovation ability.
[0007] These defects make it difficult for the existing technology to meet the personalized, immersive and deep learning needs, and restrict the value of multi-modal resources and immersive technology in the field of education, so there is an urgent need for an innovative learning interactive system that can solve the above problems. SUMMARY
[0008] (I) Technical problems solved In view of the deficiencies of the prior art, the present application provides a multi-modal teaching resource intelligent adaptation and immersive learning interactive system.
[0009] (II) Technical solutions To achieve the above purpose, the present application provides the following technical solutions: a multi-modal teaching resource intelligent adaptation and immersive learning interactive system, comprising: a multi-modal resource management module for integrating text, audio and video, VR components, AR components, interactive courseware type multi-modal teaching resources, and constructing a standardized metadata tag system including resource type, knowledge point association and difficulty level; A learner cognitive and behavior perception module is used to collect biological feature data, scene interaction data and learning process feedback data of the learner in real time; An intelligent adaptation engine is connected with the multi-modal resource management module and the learner cognitive and behavior perception module in bidirectional data communication, and is used to analyze the real-time cognitive state and individualized learning needs of the learner based on the collected data, and dynamically match the multi-modal teaching resources; An immersive interactive scene generation and rendering module is connected with the intelligent adaptation engine in data communication, and is used to generate a virtual-real integrated immersive interactive learning scene according to the matched multi-modal resources and learning goals, and support low-latency real-time interaction between the learner and the scene elements; A learning effect closed-loop feedback module is connected with the learner cognitive and behavior perception module, the intelligent adaptation engine and the immersive interactive scene generation and rendering module in data communication, and is used to evaluate the learning effect based on the interactive data and learning achievement data, and reversely drive the intelligent adaptation engine to dynamically optimize the resource matching strategy and scene interaction parameters.
[0010] Further, the multi-modal resource management module further includes a cross-modal semantic association sub-module, which extracts core semantic features of different modal resources through a pre-trained semantic understanding model, constructs a cross-modal semantic association graph, and realizes triggered intelligent linkage calling of text knowledge points and VR experiment resources, AR annotation resources, and audio / video explanation resources.
[0011] Further, the learner cognitive and behavior perception module includes a biological sensing unit configured with a lightweight electroencephalogram sensor and an eye tracking device, for real-time collection of learner concentration, cognitive fatigue, and knowledge confusion association point cognitive state data; and a behavior capture unit configured with a high-definition camera, a posture sensor, and a voice acquisition device, for real-time capture of learner gesture operation trajectory, body movement amplitude, voice interaction content, and interaction frequency data.
[0012] Further, the intelligent adaptation engine includes a dynamic portrait construction sub-module for fusing static basic data of learners (including but not limited to knowledge base level, learning style type, cognitive ability grade) and dynamic behavior data (including but not limited to interaction preference type, knowledge point error frequency, and scene stay duration), and constructing a real-time updated learner dynamic portrait; and a reinforcement learning adaptation sub-module for iteratively optimizing the push order, presentation form, and difficulty gradient adjustment strategy of multi-modal resources based on the learner dynamic portrait and a preset learning goal through a reinforcement learning algorithm.
[0013] Further, the immersive interactive scene generation and rendering module includes a space calculation unit that uses a simultaneous localization and mapping (SLAM) technology to real-time map physical space information of a real learning environment, and realizes accurate coordinate binding of a virtual knowledge point model and a real physical space; and a multi-sensory feedback unit for providing visual rendering feedback, auditory sound effect feedback, and tactile vibration feedback during learner interaction with a virtual scene, and the multi-sensory feedback unit can be externally connected to an olfactory module to realize scent feedback matching the scene content.
[0014] Further, the immersive interactive scene generation and rendering module further includes a creative interaction sub-module, which is composed of a set of creative tools including a 3D modeling tool, a scene logic editing tool, and a virtual character script writing tool, supports learners to design personalized interactive content based on current learning knowledge points, and real-time access the creation results to an immersive interactive learning scene.
[0015] Further, the learning effect closed-loop feedback module further includes a creation evaluation unit, which analyzes the creation content of the learner through computer vision technology and natural language processing technology, the analysis dimensions include knowledge point application accuracy, logical integrity, scene adaptability, generates targeted optimization suggestions, and simultaneously automatically generates an expandable interactive task based on the logical loopholes in the creation content.
[0016] Further, the immersive interactive scene generation and rendering module further includes a socialized collaboration unit, which is used to generate a virtual avatar that maps the facial expressions, body movements and voice tones of each learner in real time, supports multiple users to make eye contact, gesture cooperation and voice discussion in the same immersive interactive scene, and can dynamically allocate scene interaction operation permissions and resource viewing permissions according to the collaboration roles of the users.
[0017] Further, the learning effect closed-loop feedback module includes a cross-scene linkage unit, which synchronizes the virtual experiment progress data and creation achievement data in the classroom scene to the after-school mobile terminal AR scene through edge computing technology and cloud data synchronization technology, and simultaneously automatically converts the learner question data collected in the pre-class preparation stage into a targeted interactive task in the immersive scene in class, forming a coherent learning track before, during and after class.
[0018] Further, the intelligent adaptation engine and the learner cognitive and behavior perception module form a millisecond-level response closed loop: when the biological sensing unit detects that the learner is in a cognitive jam state, the intelligent adaptation engine triggers the multi-modal resource management module to generate customized explanation resources based on the current learning knowledge points in real time, and presents them through the immersive interactive scene generation and rendering module.
[0019] (Three) beneficial effects Compared with the prior art, the present application provides a multi-modal teaching resource intelligent adaptation and immersive learning interactive system, which has the following beneficial effects: The multi-modal teaching resource intelligent adaptation and immersive learning interactive system realizes cognitive-driven precise adaptation. The present application collects real-time cognitive state data such as the learner's concentration, cognitive fatigue and knowledge confusion points through biological sensing devices, constructs a real-time updated learner portrait in combination with dynamic behavior data, and dynamically optimizes resource matching strategies through the intelligent adaptation engine by using reinforcement learning algorithm. This adaptation method breaks the limitations of traditional static matching, can respond to the cognitive changes of the learner in milliseconds, and timely pushes customized multi-modal resources when the learner is confused or tired, effectively improves the matching degree of resources and learning needs, and helps the learner to efficiently absorb knowledge.
[0020] A virtual-real fusion high-immersion learning experience is constructed. With the help of SLAM technology, the application can map the physical space information of the real learning environment in real time, realize the accurate binding of virtual knowledge point models and real physical space, organically integrate virtual learning content and real environment, and avoid the problem of disconnection between virtual scene and reality. At the same time, through the multi-sensory feedback unit, visual, auditory, tactile and other multi-dimensional feedback can be provided, and according to the scene demand, olfactory feedback can also be provided, which can fully enhance the sense of reality and the sense of immersion of interaction, and significantly improve the participation and concentration of learners, solving the pain point of weak immersion of traditional virtual scene.
[0021] Deepen the creative interactive learning. The application has a rich set of creative interactive tools built-in, which supports learners to design personalized interactive content based on the current knowledge point, and changes the learners from passive scene experiencers to active scene co-builders. At the same time, the system evaluates the creation content through computer vision and natural language processing technology, generates targeted optimization suggestions, and automatically triggers extension interactive tasks based on the logic loopholes in the creation, so that the learning process becomes a combination of knowledge internalization and output, effectively strengthening the knowledge application ability and innovative thinking.
[0022] Realize socialized cooperation to improve participation. Through the socialized cooperation unit, the application generates a virtual avatar that maps the facial expressions, body movements and voice tones of each learner in real time, supports multiple users to carry out eye contact, gesture cooperation and voice discussion in the same immersive scene. The system can also dynamically allocate interactive operations and resource viewing permissions according to the learning roles, restore the collaborative scene of the real classroom, solve the problem of isolation in single immersive learning, improve the learning interest, and cultivate the cooperation ability of learners.
[0023] Ensure the coherence of cross-scene closed-loop learning. Relying on edge computing and cloud data synchronization technology, the application realizes seamless connection of learning data before, during and after class. The questions collected during pre-class preview can be automatically converted into targeted interactive tasks during class, and the experiment progress and creation achievements during class can be synchronized to the mobile AR scene after class for learners to continue to complete the extension tasks. This cross-scene linkage avoids the fragmentation of the learning process, forms a complete learning track, and helps to deepen and consolidate the knowledge points. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 It is a whole work closed-loop process schematic diagram of the application system; Figure 2 It is a multi-modal resource intelligent adaptation process schematic diagram of the application; Figure 3 It is a virtual-real fusion immersive scene generation process schematic diagram of the application; Figure 4This is a schematic diagram of the creative interactive learning process of the present invention; Figure 5 This is a schematic diagram of the cross-scenario learning closed-loop process of the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Please see Figures 1 to 5 This invention is a multimodal teaching resource intelligent adaptation and immersive learning interaction system, including: a multimodal resource management module, used to integrate multimodal teaching resources such as text, audio and video, VR components, AR components, and interactive courseware, and to construct a standardized metadata tag system that includes resource type, knowledge point association, and difficulty level; The learner cognition and behavior perception module is used to collect learners’ biometric data, scene interaction data and learning process feedback data in real time. The intelligent adaptation engine establishes bidirectional data communication connections with the multimodal resource management module and the learner cognition and behavior perception module, respectively, to analyze the learner's real-time cognitive state and personalized learning needs based on the collected data, and dynamically match and adapt multimodal teaching resources. The immersive interactive scene generation and rendering module is connected to the intelligent adaptation engine via data communication. It is used to generate a virtual-real immersive interactive learning scene based on the matched multimodal resources and learning objectives, and supports low-latency real-time interaction between learners and scene elements. The learning effect closed-loop feedback module is in data communication connection with the learner cognitive and behavior perception module, the intelligent adaptation engine, and the immersive interactive scene generation and rendering module respectively, is used for evaluating learning effect based on interactive data and learning achievement data, and reversely drives the intelligent adaptation engine to dynamically optimize resource matching strategy and scene interaction parameter. Through bidirectional data communication of the five modules of multi-modal resource management, cognitive behavior perception, intelligent adaptation engine, immersive scene generation, and closed-loop feedback, the system core workflow is constructed. The multi-modal resources are integrated and standardized tagged first, then multi-dimensional data of the learner is collected in real time, the demand is analyzed and the resources are matched by the intelligent adaptation engine, the interactive scene is generated by the immersive module, and finally the adaptation strategy is reversely optimized based on the learning data, to form a complete closed loop of collection, analysis, matching, interaction, and optimization. The overall technical framework of the system is clear, the full-link collaboration of resources, learners, and scenes is realized, the problems of rigid adaptation of traditional teaching resources and disconnection between immersive scene and learning demand are solved, a stable and expandable basic framework is provided for subsequent subdivision technology landing, and the intelligence and interactivity of the overall system are ensured.
[0027] In the scheme, the multi-modal resource management module further includes a cross-modal semantic association sub-module. The cross-modal semantic association sub-module extracts core semantic features of different modal resources through a pre-trained semantic understanding model, constructs a cross-modal semantic association graph, and realizes triggered intelligent linkage calling of text knowledge points and VR experiment resources, AR annotation resources, and audio / video explanation resources. The cross-modal semantic association sub-module is added in the multi-modal resource management module. Core semantic features of different modal resources such as text, VR / AR, and audio / video are extracted through a pre-trained semantic understanding model, an association graph is constructed, mapping relationships are established between different formats of resources based on semantic logic, linkage calling of multiple modal resources is realized for one knowledge point. The use limitation of single modal resources is broken, the cumbersome operation of manual switching between different resources by the learner is avoided, the resource calling is made more intelligent and the connection is made more smooth, the coherence of the learning process is improved, and the understanding of knowledge points is strengthened through multi-modal complementation.
[0028] In the present scheme, the learner cognitive and behavioral perception module comprises a biological sensing unit configured with a lightweight electroencephalogram sensor and an eye tracking device for real-time collection of learner concentration, cognitive fatigue, and knowledge confusion associated point position cognitive state data; and a behavior capture unit configured with a high-definition camera, a posture sensor, and a voice collection device for real-time capture of learner gesture operation trajectory, body movement amplitude, voice interaction content, and interaction frequency data. The learner cognitive and behavioral perception module is split into a biological sensing unit and a behavior capture unit, and configured with lightweight electroencephalogram sensors, eye tracking devices, and high-definition cameras, posture sensors, and other hardware, respectively, to collect cognitive state data (concentration, fatigue) and behavioral interaction data (gestures, movements, voices) in a targeted manner, and to realize all-around and non-interfering data collection of learners. Through the subdivision of hardware devices and data collection logic, the learner's internal cognitive state and external interactive behavior two-dimensional data are accurately obtained, the demand judgment deviation problem caused by traditional reliance on behavioral data is solved, and more comprehensive and accurate data support is provided for subsequent intelligent adaptation.
[0029] In the present scheme, the intelligent adaptation engine comprises a dynamic portrait construction submodule for fusing static basic data of learners (including but not limited to knowledge base level, learning style type, cognitive ability grade) and dynamic behavioral data (including but not limited to interaction preference type, knowledge point error frequency, scene stay duration), and constructing a real-time updated learner dynamic portrait; and a reinforcement learning adaptation submodule for iteratively optimizing the push order, presentation form, and difficulty gradient adjustment strategy of multi-modal resources based on the learner dynamic portrait and a preset learning goal through a reinforcement learning algorithm. The intelligent adaptation engine is built-in with a dynamic portrait construction submodule and a reinforcement learning adaptation submodule, fuses static basic data and dynamic behavioral data to construct a real-time updated learner portrait, and iteratively optimizes the resource push order, presentation form, and difficulty gradient based on the portrait and the learning goal through a reinforcement learning algorithm, to realize dynamic optimization of the adaptation strategy. The one-size-fits-all resource push mode is avoided, the adaptation strategy is upgraded from static matching to dynamic following, the real-time needs and cognitive level of learners are accurately matched, the individualization degree of resource adaptation is improved, and the learners are helped to efficiently absorb knowledge.
[0030] In the present scheme, the immersive interactive scene generation and rendering module comprises a space calculation unit that maps the physical space information of a real learning environment in real time using a simultaneous localization and mapping (SLAM) technique, and realizes accurate coordinate binding of a virtual knowledge point model and a real physical space; a multi-sensory feedback unit that provides visual rendering feedback, auditory sound effect feedback, and tactile vibration feedback in the process of interaction between a learner and a virtual scene, and can be externally connected to an olfactory module to realize smell feedback matching the scene content. The immersive interactive scene generation and rendering module comprises a space calculation unit and a multi-sensory feedback unit, realizes coordinate binding of a real environment and a virtual knowledge point through a SLAM technique, provides visual, auditory, and tactile multi-sensory feedback in the process of interaction, and can be optionally supplemented with an olfactory module to provide smell feedback, thereby constructing an immersive scene that combines virtual and real scenes and multi-sensory interaction. The boundaries between virtual and real scenes are broken, the learning scene is closer to a real environment, multi-sensory feedback enhances the sense of reality and immersion in interaction, the problems of weak immersion and single interaction in a traditional virtual scene are solved, and the participation and concentration of learners are significantly improved.
[0031] In the present scheme, the immersive interactive scene generation and rendering module further comprises a creative interactive sub-module that comprises an authoring tool set of a 3D modeling tool, a scene logic editing tool, and a virtual character script writing tool, supports learners to autonomously design personalized interactive content based on a current learning knowledge point, and realizes real-time access of the creation results to an immersive interactive learning scene. The creative interactive sub-module is added to the immersive scene module, the tool set of 3D modeling, scene logic editing, and script writing is built in, learners are supported to autonomously design interactive content based on a current knowledge point, and the creation results are real-time accessed to the immersive scene, so that learners are transformed from scene experiencers to scene co-builders. Interaction is upgraded from passive response to active creation, the learning process becomes a combination of knowledge internalization and output, the application ability of knowledge points is strengthened, the interest and depth of interaction are improved, and the creativity and exploration desire of learners are stimulated.
[0032] In the present scheme, the learning effect closed-loop feedback module further comprises a creation evaluation unit, which analyzes the creation content of the learner through computer vision technology and natural language processing technology, the analysis dimensions include knowledge point application accuracy, logical integrity, scene adaptability, generates targeted optimization suggestions, and simultaneously automatically generates an expandability interactive task based on the logical loopholes in the creation content. The learning effect closed-loop feedback module adds a creation evaluation unit, which analyzes the creation content of the learner from the dimensions of knowledge point accuracy, logical integrity, and scene adaptability through computer vision and natural language processing technology, generates optimization suggestions, and automatically generates an expandability interactive task based on the logical loopholes in the creation. Precise evaluation and targeted feedback of creative interactive results are achieved, avoiding superficiality in the creation process, while the expandability task fills in the knowledge gaps, forming a secondary learning closed loop of creation, evaluation, and completion, and further improving the learning effect.
[0033] In the present scheme, the immersive interactive scene generation and rendering module further comprises a socialized collaboration unit, which is used to generate a virtual avatar for each learner that real-time maps his facial expressions, body movements, and voice tone, supports multiple users to have eye contact, gesture collaboration, and voice discussion in the same immersive interactive scene, and can dynamically allocate scene interaction operation permissions and resource viewing permissions according to the collaboration roles of the users. The socialized collaboration unit of the immersive scene module generates a virtual avatar for each learner that real-time maps his expressions, movements, and voice, supports multiple users to have eye contact, gesture collaboration, and voice discussion in the same virtual scene, dynamically allocates interaction and viewing permissions according to the collaboration roles, and builds a virtual collaborative learning environment. Single-person immersion is upgraded to multi-person collaborative immersion, restoring the collaborative scene of a real classroom, solving the problem of isolation in traditional immersive learning, improving the interest and collaboration ability of learning through socialized interaction, and at the same time, the permission allocation ensures the orderly and efficient collaboration process.
[0034] In the present scheme, the learning effect closed-loop feedback module includes a cross-scene linkage unit, which synchronizes the virtual experiment progress data and creative achievement data in the classroom scene to the after-school mobile AR scene through edge computing technology and cloud data synchronization technology, and automatically converts the learner's question data collected during the pre-class preparation stage into targeted interactive tasks in the immersive classroom scene, forming a coherent learning track before, during and after class. The cross-scene linkage unit of the learning effect closed-loop feedback module realizes the synchronization of classroom virtual data (experiment progress, creative achievement) and after-school mobile AR scene through edge computing and cloud synchronization technology, and converts pre-class preparation questions into targeted interactive tasks in class, building a coherent learning track before, during and after class. Break the time and space limit of the learning scene, avoid the fragmentation of the learning process, make the learning content before, during and after class seamlessly connected, form a complete learning closed loop, improve the coherence and continuity of learning, and help to deepen and consolidate knowledge points.
[0035] In the present scheme, the intelligent adaptation engine and the learner's cognitive and behavior perception module form a millisecond-level response loop: when the biological sensing unit detects that the learner's cognitive jam state, the intelligent adaptation engine triggers the multi-modal resource management module to generate customized explanation resources based on the current learning knowledge points in real time, and renders and presents them in real time through the immersive interactive scene generation and rendering module. Build a millisecond-level response loop of intelligent adaptation engine and perception module, when the learner's concentration is lower than the threshold, the eye movement stays overtime, and other cognitive jam states are detected, the engine triggers the resource management module to generate customized explanation resources in real time, and renders and presents them in real time through the immersive scene. Realize the instant response of cognitive state and resource adaptation, provide targeted support at the first time when the learner is confused, avoid learning interruption or cognitive burden accumulation, significantly improve the fluency and efficiency of learning, and strengthen the real-time nature of individualized adaptation.
[0036] Reinforcement learning intelligent adaptation Value update formula, algorithm formula:
[0037] Wherein,
[0038] This formula is the core algorithm of the reinforcement learning adaptation sub-module in the intelligent adaptation engine, which is used to dynamically optimize the mapping relationship between the learner's state and the resource adaptation action, and to find the multi-modal resource adaptation strategy (resource type, presentation form, difficulty gradient) that best fits the learner's real-time cognitive state and learning goal through iterative update of the value (state-action pair value).
[0039] Core variables : State, Action, Value Function Definition: Represents the long-term cumulative reward expectation that the learner can obtain after performing the resource adaptation action when the learner is in state .
[0040] Physical Meaning: The higher the value, the more the selection of this adaptation action in this state can help the learner efficiently achieve the learning goal (such as understanding knowledge points, reducing cognitive fatigue).
[0041] Technical Association: The iterative update of the value corresponds to the dynamic optimization of resource pushing order, presentation form, and difficulty gradient, which is the core mathematical model for realizing cognitive-driven adaptation.
[0042] State Vector : Learner's real-time comprehensive state, defined as: is a high-dimensional feature vector containing three core dimensions, fully describing the current state of the learner.
[0043] : Cognitive state parameters, including concentration (0-100 points), cognitive fatigue (0-100 points), and knowledge confusion point identification (such as "circuit parallel connection" knowledge point coding), data collected in real-time through electroencephalogram sensors and eye tracking devices (sampling rate ≥ 1 times / second).
[0044] : Dynamic learner portrait feature vector, including static basic data (knowledge base level: 1-5, learning style: visual / auditory / motor) and dynamic behavior data (interaction preference: VR / text / audio / video, knowledge point error frequency: last 5 interaction error coding), with weight distribution of static 30% and dynamic 70%.
[0045] : Learning progress and goal parameters, including current knowledge point coding (such as "circuit connection" corresponding coding C003), learning goal type (understanding / application / creation), and current task completion degree (0-100%).
[0046] Adaptation Action : Resource adaptation strategy set, defined as: , represents the specific resource adaptation operation performed by the system at time , corresponding to the selection and presentation of multi-modal resources.
[0047] : Resource type selection (action space dimension 1), with optional values including text resources, VR experiment resources, AR annotation resources, audio / video explanation resources, and interactive courseware.
[0048] : Resource presentation form (action space dimension 2), optional values include 3D disassembly animation, step-by-step voice guidance, AR annotation diagram, virtual simulation experiment.
[0049] : Resource difficulty gradient (action space dimension 3), optional values include basic level (level 1), advanced level (level 2), challenge level (level 3), corresponding to "difficulty gradient adjustment strategy".
[0050] Learning rate : Update step coefficient, definition: control the amplitude of each value update, balance the weight of historical experience and new observation data.
[0051] Value range: in this system, the default setting is 0.3 (can be dynamically adjusted according to learning scene), The larger, the stronger the influence of new data on The value, the faster the adaptive strategy iteration; The smaller, The value update is more stable, avoiding strategy fluctuations caused by accidental data. Technical association: adaptive real-time update of learner profile, ensuring
[0052] The value can quickly respond to changes in learner state while maintaining strategy stability.
[0053] Immediate reward : Single-step adaptive effect feedback Definition: represents the immediate effect evaluation value collected by the learning effect closed-loop feedback module after executing the action It is the core feedback signal driving the Value update.
[0054] Decomposition formula: , where the weight (Based on a large amount of teaching data training, balancing efficiency and depth).
[0055] : Learning efficiency index (weight 0.4), including knowledge point understanding time consumption (actual time consumption / standard time consumption), interactive operation success rate (such as VR experiment operation accuracy), value range 0-1 (the higher the value, the higher the efficiency).
[0056] : Interactive depth index (weight 0.3), including creative interactive participation (such as whether to use 3D modeling tools to design content), interaction frequency (effective operation times per unit time), value range 0-1 (the higher the value, the more in-depth the interaction).
[0057] : Knowledge mastery index (weight 0.3), including instant evaluation accuracy (such as knowledge point answer score after interaction), creative content accuracy, value range 0-1 (the higher the value, the more solid the knowledge mastery).
[0058] Discount factor : Future reward weight, definition: represents the emphasis on future rewards (t+1), balances "immediate effect" and "long-term learning goal". Value setting: In this system, γ is set to 0.8 by default, meaning that the system places more emphasis on long-term cumulative rewards (such as coherent mastery of knowledge points and stable cognitive load) rather than just pursuing immediate effects of a single interaction, avoiding "short-term high efficiency but long-term cognitive burden too heavy" adaptation strategies. Technical association: support cross-scene closed-loop learning, ensure that the adaptation strategy not only adapts to the current state, but also provides support for the connection of learning trajectories before, during and after class.
[0059] : Future optimal action value, definition: represents the optimal feedback of the new state at time t+1 , all possible adaptation actions , the maximum value that can be obtained (i.e. the value of the future optimal adaptation strategy). Physical meaning: let the current value update consider the optimal feedback of the future state, ensure that the adaptation strategy has foresight (such as pushing basic level resources now, in order to efficiently receive advanced level resources later). Technical association: corresponding to millisecond-level response closed loop, when the learner's state changes from to (such as cognitive jamming removal), the system can quickly switch to the future optimal action, achieving dynamic adaptation.
[0060] Example 1: K12 junior high school physics circuit connection and troubleshooting teaching scene.
[0061] For junior high school physics series / parallel circuit knowledge points, traditional teaching has problems such as abstract and difficult to understand, practical operation risk (short circuit), lack of personalized guidance, etc. This system realizes the teaching landing of virtual and real combination + cognitive driving + creative learning.
[0062] Hardware: student side (tablet + AR glasses, light-weight EEG sensor, posture sensor, tactile feedback handle), teacher side (cloud management platform, VR monitoring terminal), real physical teaching aids (battery, wire, bulb, switch real component); Software: multi-modal resource library (circuit principle text, 3D current flow animation, short circuit risk warning video, VR fault simulation scene), creative tool set (virtual circuit component editing, logic simulation tool).
[0063] Pre-class preparation (cross-scene linkage): Students scan the textbook circuit chapter through tablet AR, and the system automatically pushes lightweight AR interactive tasks (identify circuit components, simple series animation), collects preparation doubts (such as parallel circuit current shunt principle), and synchronizes to the cloud.
[0064] In-class cognitive adaptation and virtual-real interaction: Perception collection: Students wear EEG sensors and AR glasses, and the system monitors concentration (threshold 60 points) and cognitive fatigue in real time; hand operation actions are captured through posture sensors; Scene generation: SLAM technology maps the real experiment table, accurately binds virtual current and voltage instrument panels with real circuit components (positioning error ≤5 cm), and when students hold real wires to connect, AR glasses superimpose current flow animation in real time, and when they touch the wrong connection point, the haptic handle produces vibration feedback (multi-sensory feedback unit); Cognitive response: When students repeatedly fail to connect the parallel circuit (concentration drops to 52 points, eye movement stays for more than 3 seconds), the system triggers multi-modal resources (3D disassembly animation + voice guidance: pay attention to the series relationship between branch switch and bulb) in milliseconds, and generates AR annotation pointing to the correct connection position; Creative interaction design: Students call creative toolset to design home lighting circuit (including 2 parallel branches, master switch, and fuse), and the system identifies circuit structure through computer vision, and the creation evaluation unit generates feedback from connection logic integrity and safety specification compliance: no ground protection, suggest adding ground terminal, click to view leakage protection principle, and trigger leakage fault simulation extension task; Socialized collaboration: 4 students form a group, with virtual avatars synchronizing facial expressions and hand movements, each taking on the roles of designer, operator, inspector, and recorder, and the system dynamically allocates permissions (such as operators can only perform connection operations, and inspectors can call virtual multimeter to detect voltage), and collaboratively completes complex circuit troubleshooting (system preset hidden fault: virtual connection of wires); Post-class extension (cross-scene linkage): Students scan the circuit diagram created in class through tablet AR, retrieve the unfinished leakage fault simulation task, superimpose the virtual circuit model in the real home environment, continue to optimize the design scheme, and synchronize the data to the cloud to form a complete learning track.
[0065] 85% of students receive immediate resource support when they get stuck, and the correct rate of circuit connection improves from 62% in traditional teaching to 89%; real short circuit risk is avoided, and virtual-real combination reduces the difficulty of understanding abstract concepts, with an average increase of 35% in classroom concentration; 78% of students can independently complete circuit design that meets specifications, and their knowledge application ability is significantly better than that in traditional experimental classes.
[0066] Embodiment 2: Higher education mechanical design foundation course design (multi-person collaborative scenario).
[0067] For the gear transmission mechanism design course design of mechanical engineering major, problems such as abstract drawing, long entity modeling period, low team collaboration efficiency, and difficult fault verification need to be solved. The system realizes cross-modal resource linkage + multi-person collaborative creation + virtual-real simulation verification.
[0068] Hardware: VR headset (MetaQuest3), high-definition motion capture camera, cloud GPU server (NVIDIA RTX4090), 3D printing device interface; Software: multi-modal resource library (gear design standard text, CAD drawing, VR gear meshing simulation animation, materials mechanics audio and video explanation), creative tool set (3D mechanical modeling, dynamics simulation, virtual assembly tool).
[0069] Cross-modal resource linkage: students call the gear design standard text through the system, click the modulus selection knowledge point in the text, and the system automatically links the VR gear meshing animation (cross-modal semantic association sub-module), and synchronously plays the materials mechanics audio and video explanation, realizing multi-modal understanding of text + animation + voice; Multi-person collaborative creation: 6-person team division of labor (structure design, dynamics analysis, assembly verification, optimization improvement), virtual avatar real-time mapping of limb movements and voice tone, collaboration in the same virtual scene to complete 3D gear model design: The structure design group draws gear tooth profile through 3D modeling tools, which is synchronized in real time to the view of other members; The dynamics analysis group calls the virtual simulation tool to simulate the gear transmission efficiency and generate a stress distribution heat map; Cognitive load dynamic adjustment: the system detects that the cognitive fatigue of the dynamics analysis group students exceeds 70 points through EEG sensors, automatically reduces the simulation task complexity (simplifies the secondary parameter calculation), and pushes the stress distribution simplified diagram resources. When the fatigue degree is reduced to below 50 points, the complete simulation task is restored; Virtual-real combination verification: SLAM technology maps the real 3D printing platform, students bind the virtual gear model with the printing device coordinates, preview the printing effect, and verify the gear and shaft gap (error ≤0.02mm) through the virtual assembly tool to avoid entity printing failure; Creation evaluation and optimization: the system analyzes the gear model structure through computer vision and parses the design report through natural language processing to generate evaluation results: the gear addendum circle diameter design meets the standard, but the tooth width coefficient does not consider the load distribution, suggesting adjusting to 0.8-1.2 range, and automatically generating transmission efficiency comparison virtual experiment scene with different tooth width coefficients; Cross-scene linkage: classroom design results are synchronized to mobile devices after class, and students can continue to optimize design parameters by scanning the physical printed gear with AR, superimposing virtual stress distribution animations, and finally generating a design report that meets industry standards.
[0070] The team design cycle is shortened from the traditional 4 weeks to 2 weeks, and the cross-department communication cost is reduced by 60%; virtual simulation verification improves the success rate of physical printing from 75% to 96%, and reduces design errors by 40%; students' proficiency in applying gear design standards improves from 68% to 91%, and their dynamic analysis capabilities are significantly enhanced.
[0071] Example 3: Vocational skills automobile engine maintenance training scene.
[0072] In automobile maintenance training, the real engine disassembly has high risk, high consumable cost, and difficult to reproduce fault scenarios. The system realizes virtual-real fusion fault troubleshooting + multi-sensory operation + cognitive load adaptation for efficient training.
[0073] Hardware: AR glasses (HoloLens2), haptic feedback gloves, posture sensors, olfactory modules (simulate fuel / oil smell), real engine training table; Software: multi-modal resource library (engine disassembly manual, 3D internal structure animation, fault sound library, maintenance step video), creative tool set (fault simulation design, maintenance process editing tool).
[0074] Virtual-real fusion fault labeling: SLAM technology maps the real engine training table in real time, AR glasses superimpose virtual fault labeling (such as highlighting the area corresponding to the fuel injector blockage), and play the fault sound (multi-sensory feedback) simultaneously. When touching the fault area, the haptic gloves simulate the oily and sticky touch, and the olfactory module releases a slight fuel odor; Cognitive-driven maintenance guidance: students wear EEG sensors, and the system monitors concentration and cognitive load: Initial maintenance (low cognitive load), only provide basic disassembly steps text; When disassembling to the cylinder head (concentration drops to 55 minutes), automatically push 3D disassembly animation + voice guidance: pay attention to the bolt disassembly sequence, evenly loosen the opposite corners to avoid operation errors; Creative fault design: students use the creative tool set to design engine abnormal noise fault scenarios (can set fault reasons: timing belt loose, bearing wear, valve clearance too large), and the system verifies the fault logic based on the engine working principle to generate fault diagnosis process suggestions; Multi-person collaborative maintenance: 3 students form a maintenance team, virtual avatars synchronize hand movements and voice, and take on the roles of disassembler, detector, and recorder respectively: The detector measures the fuel injector voltage with a virtual multimeter, and the data is shared in real time; The disassembler performs disassembly operations according to the test results, and the haptic feedback glove simulates the bolt tightness; Cross-scene review: After class, students scan real engine parts through mobile AR, retrieve classroom repair records and fault point annotations, and the system automatically generates weak link reinforcement tasks (such as timing belt replacement process reproduction). Fragmented review is completed on mobile devices. Effect evaluation: The learning effect closed-loop feedback module generates an evaluation report from three dimensions: maintenance efficiency, fault diagnosis accuracy, and operation standardization. For the problem of incorrect bolt disassembly sequence, customized AR practical operation videos are pushed.
[0075] Real engine consumable loss is reduced by 80%, and fault scene reproduction does not require additional costs; the fault diagnosis accuracy of new maintenance students is improved from 58% to 87%, and the maintenance time is shortened by 45%; the high temperature and high pressure risks in real maintenance are completely avoided, and the multi-sensory feedback improves the realism of practical operation.
[0076]
[0077] Although embodiments of the present application have been shown and described, it will be understood by those having ordinary skill in the art that various changes, modifications, substitutions and alterations can be made therein without departing from the principles and spirit of the application, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multimodal teaching resources intelligent adaptation and immersive learning interaction system, characterized in that, Comprise: A multi-modal resource management module for integrating text, audio and video, VR components, AR components, interactive courseware type multi-modal teaching resources, and building a standardized metadata tag system containing resource types, knowledge point association, and difficulty levels; A learner cognitive and behavior perception module for real-time collection of learner biological feature data, scene interaction data, and learning process feedback data; An intelligent adaptation engine, which is in bidirectional data communication connection with the multi-modal resource management module and the learner cognitive and behavior perception module, for analyzing the real-time cognitive state and individualized learning needs of learners based on the collected data, and dynamically matching the multi-modal teaching resources; An immersive interactive scene generation and rendering module in data communication connection with the intelligent adaptation engine, for generating a virtual-real integrated immersive interactive learning scene according to the matched multi-modal resources and learning goals, and supporting low-latency real-time interaction between learners and scene elements; A learning effect closed-loop feedback module in data communication connection with the learner cognitive and behavior perception module, the intelligent adaptation engine, and the immersive interactive scene generation and rendering module, for evaluating the learning effect based on the interaction data and learning achievement data, and driving the intelligent adaptation engine to dynamically optimize the resource matching strategy and scene interaction parameters.
2. The multi-modal teaching resource intelligent adaptation and immersive learning interaction system according to claim 1, wherein, The multi-modal resource management module further comprises a cross-modal semantic association sub-module, which extracts core semantic features of different modal resources through a pre-trained semantic understanding model, builds a cross-modal semantic association graph, and realizes trigger-type intelligent linkage calling of text knowledge points and VR experiment resources, AR annotation resources, and audio and video explanation resources.
3. The multi-modal teaching resource intelligent adaptation and immersive learning interaction system of claim 1, wherein, The learner cognitive and behavior perception module comprises a biological sensing unit configured with lightweight electroencephalogram sensors and eye tracking devices for real-time collection of learner concentration, cognitive fatigue, and knowledge confusion association point cognitive state data; and a behavior capture unit configured with high-definition cameras, posture sensors, and voice collection devices for real-time capture of learner gesture operation trajectories, body movement amplitudes, voice interaction contents, and interaction frequency data.
4. The multi-modal teaching resources intelligent adaptation and immersive learning interaction system of claim 1, wherein, The intelligent adaptation engine comprises a dynamic portrait construction sub-module for fusing learner static basic data (including but not limited to knowledge base level, learning style type, cognitive ability level) and dynamic behavior data (including but not limited to interaction preference type, knowledge point error frequency, scene stay duration), and constructing a real-time updated learner dynamic portrait; and a reinforcement learning adaptation sub-module for iteratively optimizing the push order, presentation form, and difficulty gradient adjustment strategy of multi-modal resources based on the learner dynamic portrait and preset learning goals through a reinforcement learning algorithm.
5. The multi-modal teaching resources intelligent adaptation and immersive learning interaction system according to claim 1, wherein, The immersive interactive scene generation and rendering module comprises a space calculation unit that adopts a simultaneous localization and mapping (SLAM) technology to map physical space information of a real learning environment in real time, and realizes accurate coordinate binding of a virtual knowledge point model and a real physical space; a multi-sensory feedback unit that is used to provide visual rendering feedback, auditory sound effect feedback, and tactile vibration feedback in a synchronous manner during interaction between a learner and a virtual scene, and can be externally connected to an olfactory module to realize smell feedback matched with scene content.
6. The multi-modal teaching resources intelligent adaptation and immersive learning interaction system of claim 1, wherein The immersive interactive scene generation and rendering module further comprises a creative interaction sub-module that comprises an authoring tool set of a 3D modeling tool, a scene logic editing tool, and a virtual character script writing tool, supports a learner to autonomously design personalized interactive content based on a current learning knowledge point, and realizes real-time access of authoring achievements to an immersive interactive learning scene.
7. The multi-modal teaching resource intelligent adaptation and immersive learning interaction system of claim 6, wherein, The learning effect closed-loop feedback module further comprises a creation evaluation unit that analyzes creation content of a learner by using computer vision technology and natural language processing technology, analyzes dimensions including knowledge point application accuracy, logic integrity, and scene adaptability, generates targeted optimization suggestions, and automatically generates an expandable interactive task based on a logic loophole in the creation content.
8. The multi-modal teaching resources intelligent adaptation and immersive learning interaction system of claim 1, wherein The immersive interactive scene generation and rendering module further comprises a socialized collaboration unit that is used to generate a virtual avatar that maps facial expressions, body movements, and voice tones of each learner in real time, supports multiple users to perform eye contact, gesture cooperation, and voice discussion in the same immersive interactive scene, and can dynamically allocate scene interaction operation permissions and resource viewing permissions according to collaboration roles of the users.
9. The multi-modal teaching resources intelligent adaptation and immersive learning interaction system of claim 1, wherein, The learning effect closed-loop feedback module comprises a cross-scene linkage unit that synchronizes virtual experiment progress data and creation achievement data in a classroom scene to an after-school mobile terminal AR scene by using edge computing technology and cloud data synchronization technology, automatically converts learner question data collected in a pre-class preparation stage into a targeted interactive task in an immersive scene in class, and forms a coherent learning track before class, in class, and after class.
10. The multi-modal teaching resources intelligent adaptation and immersive learning interaction system of claim 3, wherein, The intelligent adaptation engine and the learner cognitive and behavior perception module form a millisecond-level response closed loop: when the biological sensing unit detects a cognitive jam state of a learner, the intelligent adaptation engine triggers the multi-modal resource management module to generate customized explanation resources based on a current learning knowledge point in real time, and presents the resources in real time by using the immersive interactive scene generation and rendering module.