Multimodal interaction-driven remote education immersive scene construction method and system
By using a multimodal interaction-driven method to construct immersive scenarios for remote education, real-time interaction data of education participants is obtained, and a scenario resource scheduler and dynamic scenario graph are configured. This solves the problem of flexible response in the teaching process in remote education scenarios and realizes dynamic planning of teaching paths and personalized optimization of immersive experiences.
Patent Information
- Application Number
- CN202511625860.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-03
AI Technical Summary
Existing methods for constructing immersive scenarios in distance education are unable to flexibly respond to unexpected needs during the teaching process, resulting in insufficient adaptability to different situations and an inability to meet the needs of different teaching paces and learning requirements.
By acquiring multimodal interaction data of educational participants in virtual classrooms, analyzing real-time interaction intentions, configuring scene resource schedulers, constructing dynamic scene graphs, generating immersive narrative flows and interaction response rules, dynamic planning of teaching paths and flexible allocation of resources are achieved, thereby optimizing the immersive teaching experience.
It improves the adaptability of remote education scenarios, ensures real-time and personalized support for the teaching experience, avoids rigidity and lack of interaction caused by pre-set frameworks, and enhances learners' immersion and teaching effectiveness.
Smart Images

Figure CN121455342A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a multi-modal interaction-driven remote education immersive scene construction method and system, and belongs to the technical field of remote education. BACKGROUND
[0002] The remote education immersive scene is a dynamic virtual-reality fusion system serving remote teaching, mainly relying on VR / AR technology to construct, which serves the diverse teaching needs from cultural courses to skill training through high scene, real-time interaction and learning situation tracking.
[0003] At present, the commonly used remote education immersive scene construction method mainly fixes scene elements (such as virtual blackboard) through pre-written code logic, and fills scene content relying on static resources (such as fixed format video). Although this method can guarantee the stability of the scene in the initial stage through the standardized process, since its core logic completely depends on the pre-set framework, it is difficult to flexibly respond to the sudden needs in the teaching process, such as the temporary addition of group discussion links, which weakens the situational adaptability of the remote education scene to different teaching paces and different learning needs. SUMMARY
[0004] The application provides a multi-modal interaction-driven remote education immersive scene construction method and system, which can flexibly respond to the sudden needs in the teaching process and improve the situational adaptability of the remote education scene.
[0005] To achieve the above-mentioned purpose, the application provides a multi-modal interaction-driven remote education immersive scene construction method, which comprises: Obtaining multi-modal interaction data of an education participant in a virtual classroom to analyze the real-time interaction intention of the education participant, and configuring a scene resource scheduler of the virtual classroom according to the real-time interaction intention; Extracting a scene resource scheduling list and a dynamic content parameter set in the scene resource scheduler to set a teaching constraint condition of the virtual classroom, and constructing a dynamic scene graph of the virtual classroom according to the teaching constraint condition and the real-time interaction intention; Analyzing a key teaching situation corresponding to the real-time interaction intention from the dynamic scene graph, calculating a cognitive coupling coefficient of the real-time interaction intention and the key teaching situation, and generating an immersive narrative flow and an interaction response rule set corresponding to the key teaching situation based on the cognitive coupling coefficient; According to the immersive narrative flow and the interaction response rule set, a real-time rendering task queue corresponding to the dynamic scene graph is established, and an immersive teaching plot of the virtual classroom is rendered in combination with the scene resource scheduler, the immersive narrative flow and the real-time rendering task queue; Based on the dynamic scene atlas, the remote education immersive scene corresponding to the education participant is generated through the real-time rendering task queue and the immersive teaching plot.
[0006] Optionally, the constructing the dynamic scene atlas of the virtual classroom according to the teaching constraint condition and the real-time interaction intention comprises: According to the teaching constraint condition and the real-time interaction intention, the core teaching semantic feature of the virtual classroom is extracted; Based on the core teaching semantic feature, the teaching situation classification rule library of the virtual classroom is configured; From the teaching situation classification rule library, the situation code and the teaching priority corresponding to the real-time interaction intention are matched out; According to the situation code and the teaching priority, the teaching situation node of the virtual classroom is identified; The knowledge point association relationship and the activity space-time boundary in the teaching constraint condition are analyzed to label the teaching activity node of the virtual classroom; Based on the mapping rule between the situation code and the knowledge point association relationship, the logical association link between the teaching situation node and the teaching activity node is established; According to the teaching situation node, the teaching activity node and the logical association link, the dynamic scene atlas of the virtual classroom is constructed.
[0007] Optionally, the configuring the teaching situation classification rule library of the virtual classroom based on the core teaching semantic feature comprises: The feature dimension and semantic intensity of the core teaching semantic feature are analyzed, and the situation type corresponding to the core teaching semantic feature is determined; The real-time teaching state parameter of the virtual classroom is extracted, wherein the real-time teaching state parameter comprises device interaction feature, student cognitive state and teaching task promotion index; According to the feature dimension, the semantic intensity and the real-time teaching state parameter, the situation matching confidence of the core teaching semantic feature is calculated; Based on the situation matching confidence and the situation type, the target situation classification sequence corresponding to the core teaching semantic feature is mapped from the preset situation library; According to the real-time teaching state parameter, the state adaptive classification threshold of the target situation classification sequence is generated; The target situation classification sequence and the state adaptive classification threshold are fused to construct the teaching situation-state association mapping table of the virtual classroom; Based on the teaching situation-state association mapping table, the teaching situation classification rule library of the virtual classroom is configured.
[0008] Optionally, the generating, based on the cognitive coupling coefficient, an immersive narrative flow and an interaction response rule set corresponding to the key teaching situation comprises: According to the cognitive coupling coefficient, obtaining teaching target sequence data and cognitive state evolution data corresponding to the key teaching situation; Based on the teaching target sequence data, extracting narrative dynamics characteristics corresponding to the key teaching situation; According to the narrative dynamics characteristics, generating an immersive narrative flow of the key teaching situation; Based on the cognitive state evolution data, analyzing the cognitive state trajectory of the key teaching situation; Identify the key state node of the cognitive state trajectory, and determine the cognitive safety boundary corresponding to the key teaching situation through the key state node; According to the cognitive safety boundary, constructing an interaction constraint list of the key teaching situation; Combining the immersive narrative flow and the interaction constraint list, generating an interaction response rule set of the key teaching situation.
[0009] Optionally, the configuring, according to the real-time interaction intention, a scene resource scheduler of the virtual classroom comprises: Parsing the teaching situation semantics corresponding to the real-time interaction intention; Extracting the core resource demand label in the teaching situation semantics; Identifying the functional attributes and dependency relationships of various scene resources in the virtual classroom; According to the functional attributes and the dependency relationships, calculating the collaborative scheduling weight of the scene resources; Based on the core resource demand label and the collaborative scheduling weight, generating a dynamic loading priority queue of the scene resources; According to the dynamic loading priority queue, setting a batch loading scheme of the scene resources, and establishing a resource competition allocation rule of the batch loading scheme; Combining the dynamic loading priority queue, the batch loading scheme and the resource competition allocation rule, constructing a scene resource scheduler of the virtual classroom.
[0010] Optionally, the extracting a scene resource scheduling list and a dynamic content parameter set in the scene resource scheduler to set the teaching constraint condition of the virtual classroom comprises: Parsing the resource function metadata in the scene resource scheduling list and the user state vector in the dynamic content parameter set; By establishing a logical association between the resource function metadata and the user state vector, a context-resource logical topology network corresponding to the virtual classroom is constructed. Identify the key context switching nodes corresponding to the context-resource logical topology network, and define the system control parameters of the key context switching nodes; Extract the content entity boundaries, user operation permission range, and process advancement threshold from the system control parameters to define the rule instantiation strategy for the key context switching nodes; Based on the rule instantiation strategy, the teaching constraints of the virtual classroom are set.
[0011] Optionally, calculating the cognitive coupling coefficient between the real-time interactive intent and the key teaching context includes: Identify the cognitive load demand corresponding to the real-time interactive intent, and analyze the cognitive resource supply level corresponding to the key teaching scenarios; Based on the cognitive load demand and the cognitive resource supply level, determine the cognitive load fit between the real-time interactive intent and the key teaching context; The degree of structuring corresponding to the key teaching scenarios is quantitatively analyzed, and the cognitive guidance intensity coefficient corresponding to the degree of structuring is determined. Analyze the emotional motivations corresponding to the real-time interaction intentions to calculate the motivational continuity matching degree corresponding to the real-time interaction intentions; Identify the intensity of cognitive conflict corresponding to the real-time interactive intent and the urgency of knowledge assimilation corresponding to the key teaching context; Based on the intensity of cognitive conflict and the urgency of knowledge assimilation, calculate the dynamic balance margin between the real-time interactive intention and the key teaching context; By integrating the cognitive load fit, the cognitive guidance intensity coefficient, the motivation continuity matching degree, and the dynamic balance margin, the cognitive coupling coefficient between the real-time interactive intention and the key teaching context is calculated.
[0012] Optionally, establishing a real-time rendering task queue corresponding to the dynamic scene graph based on the immersive narrative flow and the set of interactive response rules includes: The key teaching event frames in the immersive narrative flow are parsed out, and the interactive hotspot areas in the set of interactive response rules are identified. Based on the key teaching event frames and the interactive hotspot areas, the rendering priority space corresponding to the dynamic scene map is divided, wherein the rendering priority space includes an instant rendering area, a high-priority streaming loading area, a regular loading area, and a background preloading area. By utilizing the plot urgency of the immersive narrative flow and the real-time interaction requirements of the interactive response rule set, resource scheduling thresholds for each region within the rendering priority space are defined. Based on the resource scheduling threshold, differentiated loading rules are set for each region within the rendering priority space; Based on the differentiated loading rules, resource loading instructions are generated for each region within the rendering priority space; Based on the rendering priority space, the differentiated loading rules, and the resource loading instructions, a real-time rendering task queue corresponding to the dynamic scene map is established.
[0013] Optionally, the step of combining the scene resource scheduler, the immersive narrative stream, and the real-time rendering task queue to render the immersive teaching plot of the virtual classroom includes: Identify the narrative evolution trigger points in the immersive narrative flow; Based on the spatiotemporal attributes of the narrative evolution trigger points, the synchronous marker points for the teaching content of the virtual classroom are determined; The collaborative rendering instructions of the real-time rendering task queue are parsed to schedule the multimodal resource package in the scene resource scheduler according to the resource loading requirements of the collaborative rendering instructions. Based on the synchronous markers of the teaching content, perform temporal alignment processing on each element in the multimodal resource package to obtain an aligned multimodal resource package; Based on the aligned multimodal resource package, an immersive teaching scenario for the virtual classroom is rendered.
[0014] To address the aforementioned problems, the present invention also provides a multimodal interaction-driven immersive scenario construction system for distance education, the system comprising: The intent recognition module is used to acquire multimodal interaction data of educational participants in the virtual classroom, so as to parse the real-time interaction intent of the educational participants and configure the scene resource scheduler of the virtual classroom according to the real-time interaction intent. The resource setting module is used to extract the scene resource scheduling list and dynamic content parameter set from the scene resource scheduler in order to set the teaching constraints of the virtual classroom, and construct the dynamic scene map of the virtual classroom according to the teaching constraints and the real-time interaction intention. The context analysis module is used to parse the key teaching context corresponding to the real-time interactive intent from the dynamic scene graph, calculate the cognitive coupling coefficient between the real-time interactive intent and the key teaching context, and generate an immersive narrative flow and interactive response rule set corresponding to the key teaching context based on the cognitive coupling coefficient. The plot rendering module is used to establish a real-time rendering task queue corresponding to the dynamic scene graph based on the immersive narrative flow and the set of interactive response rules, and to render the immersive teaching plot of the virtual classroom by combining the scene resource scheduler, the immersive narrative flow and the real-time rendering task queue. The scene generation module is used to generate a remote education immersive scene corresponding to the education participant based on the dynamic scene map, through the real-time rendering task queue and the immersive teaching plot.
[0015] Compared to the problems described in the background technology, the embodiments of the present invention, by configuring the scene resource scheduler of the virtual classroom according to the real-time interactive intent, can enhance the scene resource scheduler's intelligent perception and on-demand supply capabilities for the entire dynamic evolution process of the virtual classroom, ensuring the real-time generation of the immersive teaching experience and the matching degree of the teaching intent; furthermore, by extracting the scene resource scheduling list and dynamic content parameter set from the scene resource scheduler to set the teaching constraints of the virtual classroom, the embodiments of the present invention can realize the dynamic planning of the teaching process and the flexible allocation of teaching resources, ensuring that the generated teaching path is highly consistent with the real-time state of the classroom, effectively avoiding problems such as rigid teaching processes, lack of on-site interaction, and insufficient personalized support caused by traditional preset frameworks; the embodiments of the present invention, by constructing the dynamic scene map of the virtual classroom according to the teaching constraints and the real-time interactive intent, can clarify the dynamic coupling relationship and evolution law of teaching rule boundaries and teacher-student interaction behavior in the multimodal virtual space, and thus specifically optimize the dynamic generation logic and multimodal interaction response accuracy of the immersive scene of distance education; furthermore, the embodiments of the present invention... By calculating the cognitive coupling coefficient between the real-time interactive intent and the key teaching context, this embodiment of the invention can accurately optimize the narrative generation logic and resource scheduling strategy of the immersive scene. Based on the cognitive coupling coefficient, this embodiment generates an immersive narrative flow and a set of interactive response rules corresponding to the key teaching context. This allows for real-time optimization of core elements in multimodal interactive-driven immersive remote education scenarios, such as the rhythm design standards and content adaptation dimensions of the immersive narrative flow, as well as the feedback trigger threshold and strategy execution parameters of the interactive response rule set. This significantly improves the accuracy of adapting the individual cognitive state and learning needs of remote education learners to the immersive teaching scenario. Finally, based on the dynamic scene map, this embodiment generates a remote education immersive scene corresponding to the educational participants through the real-time rendering task queue and the immersive teaching plot. This enables the virtual teaching environment to present a personalized teaching experience with optimal efficiency and controllable cognitive load under multimodal interactive driving, while avoiding scene rigidity, lack of interaction, and context mismatch caused by relying on preset linear scripts and static resources. This greatly improves the adaptability, immersion, and teaching effectiveness of the remote education scenario. Therefore, this invention can flexibly respond to unexpected needs in the teaching process and improve the adaptability of distance education scenarios. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a method for constructing a multimodal interaction-driven immersive scenario for distance education according to an embodiment of the present invention. Figure 2 The interaction hierarchy diagram of the virtual classroom in the multimodal interaction-driven immersive scene construction method for distance education provided in an embodiment of the present invention; Figure 3 A scene structure design diagram of a virtual classroom for a multimodal interaction-driven immersive scene construction method for distance education provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the modules for implementing the multimodal interaction-driven immersive scene construction system for distance education, according to an embodiment of the present invention.
[0017] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides a method for constructing an immersive remote education scene driven by multimodal interaction. The execution subject of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a method for constructing a multimodal interaction-driven immersive remote education scene according to an embodiment of the present invention. In this embodiment, the method for constructing a multimodal interaction-driven immersive remote education scene includes: S1. Obtain multimodal interaction data of educational participants in the virtual classroom to analyze the real-time interaction intentions of the educational participants, and configure the scene resource scheduler of the virtual classroom according to the real-time interaction intentions.
[0021] This invention, through acquiring multimodal interaction data of educational participants in a virtual classroom, analyzes their real-time interaction intentions. This allows for targeted optimization of the multimodal-driven generation logic and dynamic teaching resource integration strategies for immersive remote education scenarios. The educational participants refer to individuals involved in immersive remote education activities, including but not limited to students, teachers, and other teaching support staff. The virtual classroom refers to a digital teaching space constructed using virtual reality, augmented reality, and other technologies, providing an immersive learning experience for educational participants. For example, a simulated physics laboratory virtual scene built using VR technology allows participants to conduct virtual experiments and engage in learning exchanges. The multimodal interaction data refers to the data set generated by educational participants in the virtual classroom through various interaction methods, including different modalities of interaction such as voice, gestures, facial expressions, and text input. The real-time interaction intention refers to the interaction purpose or learning behavior tendency that educational participants intend to engage in during the current teaching process in the virtual classroom, determined based on the real-time analysis of the participants' multimodal interaction data. For example, by analyzing students' voice questions and gestures pointing to virtual knowledge points in real time, it can be determined whether a student currently intends to delve deeper into the principles of a specific physics experiment.
[0022] Optionally, the real-time interactive intent of the educational participants can be analyzed using an attention mechanism fusion algorithm, such as the Cross-Attention algorithm. For example, the original feature sequences of each modality (voice feature sequence, gesture feature sequence, gaze feature sequence) are used as input, and the semantic alignment between modalities is used as the optimization objective. Collaborative attention calculation is performed on the heterogeneous interactive data to extract cross-modal association weights and context-aware vectors to determine the real-time interactive intent vector.
[0023] To understand the real-time interactive intent of educational participants, refer to Figure 2 The diagram shown is an interaction hierarchy diagram of a virtual classroom in a multimodal interaction-driven immersive scenario construction method for distance education provided by an embodiment of the present invention. By clarifying the hierarchical structure of the virtual learning scenario and the types of interactive behaviors corresponding to each level, this diagram provides a hierarchical logical framework and reference standard for analyzing the real-time interactive intentions of educational participants. Different levels correspond to different depths and types of interactive behaviors, which can help us more systematically and accurately identify and classify the multimodal interactive behaviors of educational participants in the virtual classroom (such as distinguishing between simple scene roaming, selection and manipulation of learning resources, or collaborative communication with others), and then extract key information from these interactive behaviors to better understand the real-time interactive intentions of educational participants, laying the foundation for subsequent operations such as configuring scene resource schedulers and constructing dynamic scene graphs.
[0024] Furthermore, by configuring the scene resource scheduler of the virtual classroom according to the real-time interactive intent, the embodiments of the present invention can enhance the scene resource scheduler's intelligent perception and on-demand supply capabilities for the entire dynamic evolution process of the virtual classroom, ensuring the real-time generation of the immersive teaching experience and the matching degree of the teaching intent. The scene resource scheduler refers to a core functional component that can perform real-time scheduling, allocation and optimization of various scene resources (such as dynamic teaching resources, interactive function modules, scene configuration parameters, etc.) in the virtual classroom based on the real-time interactive intent of the educational participants.
[0025] As an embodiment of the present invention, configuring the scene resource scheduler of the virtual classroom according to the real-time interaction intent includes: Analyze the semantics of the teaching context corresponding to the real-time interactive intent; Extract the core resource demand tags from the semantics of the teaching context; Identify the functional attributes and dependencies of various scenario resources in the virtual classroom; Calculate the collaborative scheduling weight of the scene resources based on the functional attributes and the dependencies; Based on the core resource requirement tags and the collaborative scheduling weights, a dynamic loading priority queue for the scene resources is generated. Based on the dynamic loading priority queue, a batch loading scheme for the scene resources is set, and a resource contention allocation rule for the batch loading scheme is established. By combining the dynamic loading priority queue, the batch loading scheme, and the resource contention allocation rules, a scene resource scheduler for the virtual classroom is constructed.
[0026] Among them, the teaching scenario semantics refers to the scenario-based semantic information directly related to the current distance education teaching activity obtained after semantic parsing of the real-time interaction intentions of educational participants, which is used to clarify the teaching links, learning objectives or operation scenarios corresponding to the interaction intentions, such as scenarios like teachers' theoretical lectures, students' hands-on experiments, and individual independent explorations; the core resource requirement label refers to the characteristic label extracted from the teaching scenario semantics and used to identify the category of scenario resources necessary to achieve the current teaching requirements. For example, for the scenario of teachers' theoretical lectures, its core resource requirement labels may include the virtual avatar of the main lecturer, the 3D courseware model, the blackboard writing trajectory generator, etc.; the scenario resources refer to the digital resource set used to construct and run virtual classrooms and support immersive distance education teaching activities, including but not limited to virtual teaching props, teaching demonstration content, interactive function modules, and scenario configuration files; the functional attribute refers to the metadata used to describe the roles and behavioral characteristics of the scenario resources in the virtual classroom, including the type of the resource (such as model / audio / script), the rendering complexity level, whether it supports dynamic interaction, and the estimated loading time-consuming, etc.; the dependency relationship refers to the associated requirements existing between different scenario resources in the virtual classroom to achieve the complete teaching function, including mandatory dependencies (i.e., the operation of one resource requires the loading of another resource as a prerequisite) and collaborative dependencies (i.e., one resource needs to be used in cooperation with other resources to enhance the teaching effect); the collaborative scheduling weight refers to the quantitative index calculated based on the functional attributes and dependency relationships of the scenario resources and used to measure the adaptation importance and collaborative necessity of a certain resource under the current teaching requirements. The higher the weight value, the stronger the supporting role of the resource in achieving the teaching requirements; the dynamic loading priority queue refers to a loading instruction sequence formed by dynamically sorting all the currently required scenario resources according to the real-time parsed teaching scenario; the batch loading scheme refers to a phased resource loading strategy that divides the to-be-loaded scenario resources into logical batches with different loading time sequences according to the dynamic loading priority queue, and sets clear loading trigger conditions and resource release strategies for each batch; the resource competition allocation rule refers to a series of predefined logics used to arbitrate and allocate these limited resources when the system computing resources (such as network bandwidth, GPU memory) are not sufficient to simultaneously meet all loading requests.
[0027] Optionally, the teaching scenario semantics corresponding to the real-time interaction intention can be parsed using the BERT semantic parsing model; the collaborative scheduling weight of the scenario resources can be calculated using the PageRank algorithm; the resource competition allocation rule of the batch loading scheme can be established using the highest response ratio first algorithm.
[0028] S2. Extract the scene resource scheduling list and dynamic content parameter set from the scene resource scheduler to set the teaching constraints of the virtual classroom. Based on the teaching constraints and the real-time interaction intention, construct the dynamic scene map of the virtual classroom.
[0029] This invention extracts the scene resource scheduling list and dynamic content parameter set from the scene resource scheduler to set the teaching constraints of the virtual classroom. This enables dynamic planning of the teaching process and flexible allocation of teaching resources, ensuring that the generated teaching path is highly consistent with the real-time classroom status. It effectively avoids problems such as rigid teaching process, lack of on-site interaction, and insufficient personalized support caused by traditional preset frameworks.
[0030] The scenario resource scheduling list refers to a structured list document generated by the scenario resource scheduler based on the real-time interactive intentions of educational participants, dynamic loading priority queues, and batch loading schemes. This list contains specific information about the scenario resources required for the current virtual classroom, loading plans, and usage specifications. The dynamic content parameter set refers to a set of variables that are dynamically updated during the operation of the virtual classroom based on the real-time status and interactive behavior of educational participants. These variables quantify the current teaching context and personalized learning status, including student learning progress values, real-time knowledge mastery scores, current interaction activity index, consumed class time, and personalized difficulty coefficients. The teaching constraints refer to a series of rule boundaries set based on the scheduling list and dynamic parameter set to ensure the orderly and effective conduct of virtual classroom teaching activities. These constraints include teaching process constraints (such as the order of experimental operation steps and the logic of knowledge point explanation), resource usage constraints (such as core resource usage permissions and non-core resource loading restrictions), and interactive behavior constraints (such as allowed interaction methods and the number of retries for operational errors).
[0031] As an embodiment of the present invention, the step of extracting the scene resource scheduling list and dynamic content parameter set from the scene resource scheduler to set the teaching constraints of the virtual classroom includes: Parse the resource function metadata in the scenario resource scheduling list and the user state vector in the dynamic content parameter set; By establishing a logical association between the resource function metadata and the user state vector, a context-resource logical topology network corresponding to the virtual classroom is constructed. Identify the key context switching nodes corresponding to the context-resource logical topology network, and define the system control parameters of the key context switching nodes; Extract the content entity boundaries, user operation permission range, and process advancement threshold from the system control parameters to define the rule instantiation strategy for the key context switching nodes; Based on the rule instantiation strategy, the teaching constraints of the virtual classroom are set.
[0032] The resource function metadata refers to a structured data set extracted from the scene resource scheduling list, used to describe the core functional characteristics of scene resources. It includes the resource's function identifier, function scope, input and output parameters, adapted teaching scene type, and functional association with other resources. The user state vector refers to multi-dimensional vector data extracted from the dynamic content parameter set, used to quantitatively describe the current learning status, interactive behavior characteristics, and demand tendencies of educational participants in the virtual classroom. It includes quantitative indicators such as learning progress dimension (e.g., knowledge point mastery), interactive behavior dimension (e.g., operation proficiency), and physiological feedback dimension (e.g., focus). The logical association refers to the correspondence between the resource function metadata and the user state vector based on teaching or interaction rules, including direct matching relationships (e.g., the matching between geometric proof needs and geometric proof tool functions in the user state vector) and conditional triggering relationships (e.g., operation error rate > 0 in the user state vector).The system includes: 5. Triggering error correction tools; 6. Enhancing collaborative relationships (e.g., the collaboration between virtual experiment tools and data calculation tools, requiring activation based on user experiment data recording status); 7. The context-resource logical topology network refers to a topology graph constructed based on the logical association between resource function metadata and user state vectors, with teaching context nodes and resource nodes as the core and association relationships as the edges; 8. Key context switching nodes refer to nodes in the context-resource logical topology network that are at turning points in the teaching process or where user state changes significantly; 9. System control parameters refer to a set of system control parameters extracted from key context switching nodes to achieve smooth switching of teaching contexts, including resource scheduling parameters (e.g., resource loading / unloading rate during switching), interaction adaptation parameters (e.g., compatibility of operation modes before and after switching), teaching rhythm parameters (e.g., transition duration during switching), and system performance parameters (e.g., peak memory usage threshold during switching); 10. Content entity boundaries refer to constraints extracted from system control parameters to define the scope of resource content and data interaction boundaries at key context switching nodes, including resource content boundaries (e.g., the range of resource types allowed before and after switching) and data sharing boundaries (e.g., different teaching stages). The system defines the user data fields that can be shared between resources and the boundaries of functional interactions (such as allowed interface call permissions between resources). The user operation permission scope refers to the constraints extracted from the system control parameters of key context switching nodes, used to define the types of operations, operation permission levels, and operation boundaries that educational participants can execute during key context switching phases in the virtual classroom (such as teaching segment switching or scene resource updates). The process advancement threshold refers to the quantitative indicators extracted from the system control parameters used to determine whether key context switching nodes meet the conditions for teaching process advancement, including user behavior thresholds (such as the percentage of operations completed in the current segment), resource status thresholds (such as the loading completion rate of required resources), and time constraint thresholds (such as the maximum dwell time in the current segment). The rule instantiation strategy refers to the specific strategy formulated based on the content entity boundaries of key context switching nodes, the user operation permission scope, and the process advancement threshold to transform abstract constraint rules into executable rules, including rule triggering conditions (such as when the process advancement threshold is met), rule execution steps (such as unloading preceding resources before loading subsequent resources), exception handling mechanisms (such as prompting strategies when the threshold is not met), and parameter dynamic adjustment schemes (such as adjusting threshold sensitivity according to user status).
[0033] Optionally, the logical association between the resource function metadata and the user state vector can be analyzed by association rule mining algorithms (such as the FP-Growth algorithm); the context-resource logical topology network corresponding to the virtual classroom can be constructed using a graph neural network (GNN) model; and the system control parameters of the key context switching nodes can be defined by dynamic priority scheduling algorithms (such as the earliest deadline first EDF algorithm).
[0034] Furthermore, by constructing a dynamic scene graph of the virtual classroom based on the teaching constraints and the real-time interactive intent, this embodiment of the invention can clarify the dynamic coupling relationship and evolution law of teaching rule boundaries and teacher-student interaction behavior in the multimodal virtual space. This allows for targeted optimization of the dynamic generation logic and multimodal interaction response accuracy of the immersive remote education scene. The dynamic scene graph refers to a structured knowledge graph constructed based on teaching constraints and real-time interactive intent in a multimodal interaction-driven immersive remote education scene, which can dynamically map the scene element composition, element association relationship, and scene evolution rules of the virtual classroom.
[0035] As an embodiment of the present invention, constructing a dynamic scene map of the virtual classroom based on the teaching constraints and the real-time interactive intent includes: Based on the teaching constraints and the real-time interaction intent, the core teaching semantic features of the virtual classroom are extracted; Based on the core teaching semantic features, configure the teaching context classification rule base for the virtual classroom; Match the context code and teaching priority corresponding to the real-time interactive intent from the teaching context classification rule base; Based on the context code and the teaching priority, identify the teaching context nodes of the virtual classroom; The relationships between knowledge points and the spatiotemporal boundaries of activities within the teaching constraints are analyzed to mark the teaching activity nodes in the virtual classroom. Based on the mapping rules between the context code and the knowledge point association, a logical association link is established between the teaching context node and the teaching activity node; Based on the teaching context nodes, the teaching activity nodes, and the logical connection links, a dynamic scene map of the virtual classroom is constructed.
[0036] The core teaching semantic features refer to a set of semantic features extracted from teaching constraints and real-time interactive intentions that can characterize the core teaching objectives and dynamic needs of the virtual classroom. These include, but are not limited to, teaching stages (such as introduction, lecturing, and practice), interaction modes (such as lecturing, question-and-answer, collaboration, and assessment), core knowledge points, and emotional atmosphere (such as exploration, competition, and collaboration). The teaching situation classification rule base refers to a set of rules constructed based on the core teaching semantic features of the virtual classroom for standardized classification of teaching situations. It includes classification dimensions (such as subject area, teaching link, and interaction mode), classification thresholds (such as determining an experimental situation if the proportion of operational behaviors in the interaction is ≥60%), and rule triggering conditions (such as when the core teaching...). The semantic features include the activation of experimental scenario classification rules when virtual operations are performed. The scenario encoding refers to the standardized string encoding of the teaching scenario corresponding to the real-time interactive intent, based on a teaching scenario classification rule base. The teaching priority refers to the priority quantification value (e.g., 0-10 points, with 10 being the highest) assigned to the matched teaching scenario based on the weight of teaching objectives in teaching constraints (e.g., core knowledge points take precedence over extended content) and the urgency of the real-time interactive intent (e.g., immediate questions take precedence over delayed review). The teaching scenario node refers to the node unit in the dynamic scene graph used to represent the current teaching scenario of the virtual classroom, containing scenario encoding, teaching priority, related subject areas, and scenario description information. The knowledge point relationships refer to the logical connections between different knowledge points in the virtual classroom, as analyzed from the teaching constraints, such as causal relationships, progressive relationships, and parallel relationships. The activity spatiotemporal boundaries refer to the constraints extracted from the teaching constraints, used to define the time and space range of teaching activities in the virtual classroom, such as a maximum duration of 15 minutes for experimental operations; experimental operations must be conducted within a designated virtual laboratory scenario. The teaching activity nodes refer to the node units in the dynamic scene graph used to represent specific teaching behaviors or tasks in the virtual classroom, including activity name, activity type (e.g., observation, operation, discussion), knowledge point association set (e.g., associated core knowledge point IDs), and spatiotemporal boundary parameters, connecting the teaching context. The mapping rules refer to a set of rules used to define the correspondence between context codes and knowledge point associations, including one-to-one mapping (e.g., a specific context code uniquely corresponds to a set of knowledge point associations), many-to-one mapping (e.g., multiple similar context codes correspond to the same core knowledge point association), and conditional mapping (e.g., when the teaching priority is > 8 points, the context code corresponds to an extended knowledge point association); the logical association link refers to a directed link established based on the mapping rules in the dynamic scenario graph, connecting teaching context nodes and teaching activity nodes, including link weight (representing association strength, such as 0.9), triggering conditions (e.g., activated when the student completes the preceding activity), and execution order (e.g., executing activity A first, then activity B).
[0037] Optionally, the context code and teaching priority corresponding to the real-time interactive intent can be matched from the teaching context classification rule base using a decision tree model; a dependency parser (such as Stanford CoreNLP) can be used to parse the knowledge point associations and activity spatiotemporal boundaries in the teaching constraints; the mapping rules between the context code and the knowledge point associations can be determined by an association rule mining algorithm (such as the Apriori algorithm).
[0038] As another embodiment of the present invention, configuring the teaching context classification rule base of the virtual classroom based on the core teaching semantic features includes: The feature dimensions and semantic strength of the core teaching semantic features are analyzed, and the context type corresponding to the core teaching semantic features is determined. Extract the real-time teaching status parameters of the virtual classroom, wherein the real-time teaching status parameters include device interaction characteristics, student cognitive status and teaching task progress indicators; Calculate the context matching confidence of the core teaching semantic features based on the feature dimensions, the semantic strength, and the real-time teaching status parameters; Based on the context matching confidence and the context type, the target context classification sequence corresponding to the core teaching semantic features is mapped from the pre-set context library; Based on the real-time teaching state parameters, generate the state adaptive classification threshold of the target context classification sequence; By integrating the target context classification sequence with the state adaptive classification threshold, a teaching context-state association mapping table for the virtual classroom is constructed. Based on the teaching context-state association mapping table, configure the teaching context classification rule base for the virtual classroom.
[0039] The feature dimension refers to the classification dimension that represents different attributes of virtual classroom teaching needs after the core teaching semantic features are structurally decomposed. This includes teaching theme dimensions (e.g., subject area, knowledge point category), teaching form dimensions (e.g., theoretical explanation, virtual experiment, group discussion), interaction method dimensions (e.g., voice commands, gesture operation, text input), and knowledge difficulty dimensions (e.g., basic cognition, application practice, extended inquiry). The semantic intensity refers to a quantitative value representing the significance or certainty level of a core teaching semantic feature on its corresponding feature dimension. The scenario type refers to the standard teaching scenario category to which the virtual classroom teaching scenario belongs, initially determined based on the feature dimension and semantic intensity of the core teaching semantic features, including… The system includes various scenario types such as lecture-based, question-and-answer interactive, group collaboration, self-directed inquiry, and skills training. The real-time teaching status parameters refer to a set of quantitative data collected in real-time during the virtual classroom operation, reflecting the equipment operation status, student learning status, and teaching progress status of distance education. These include device interaction characteristics (such as VR device response latency on the student end, multi-terminal network bandwidth stability, and device memory usage), student cognitive status (such as focus based on facial expression recognition, knowledge mastery determined by the accuracy of in-class quizzes, and operational error rate), and teaching task progress indicators (such as the overall class task completion rate, group collaboration progress deviation, and remaining time for the current segment). The scenario matching confidence level refers to the confidence level based on core teaching semantic features. The matching degree of feature dimensions, semantic strength weight, and adaptability of real-time teaching status parameters are weighted and calculated to quantify the accuracy of matching core teaching semantic features with a specific teaching context. The pre-built context library refers to a pre-constructed structured database storing templates of common teaching contexts in the field of distance education, including context identifiers (e.g., Phys-Exp-001), context types (e.g., virtual experiment contexts), feature dimension adaptation ranges (e.g., physics-kinematics-virtual operation), a list of associated resources (e.g., virtual experiment equipment models, operation guidance animations), and adaptation status thresholds (e.g., network bandwidth ≥ 1.5Mbps), etc. The target context classification sequence refers to a sequence sorted from high to low context matching confidence. An ordered list of one or more scenario types most relevant to the current core teaching semantic features; the state adaptive classification threshold refers to a quantitative threshold dynamically generated based on real-time teaching state parameters, used to determine whether each scenario in the target scenario classification sequence meets the current operating conditions, including device adaptation thresholds (such as minimum network bandwidth threshold, maximum device response latency threshold), student adaptation thresholds (such as minimum knowledge mastery threshold, minimum operation proficiency threshold), and progress adaptation thresholds (such as minimum completion rate of the current stage threshold); the teaching scenario-state association mapping table refers to a structured table constructed by integrating the target scenario classification sequence with the corresponding state adaptive classification thresholds, used to clarify the dynamic classification thresholds corresponding to each scenario type and the switching logic between them.
[0040] Optionally, the real-time teaching state parameters of the virtual classroom can be extracted using sensor data fusion algorithms (such as Kalman filters) and Hidden Markov Models. For example, by fusing multi-source asynchronous data from eye trackers, gesture sensors, and voice interaction modules using a Kalman filter, a stable student attention trajectory can be obtained. Then, a Hidden Markov Model can be used to decode the state of continuous interactive behavior sequences to identify the real-time teaching state parameters. A deep matching network (such as DSSM) can be used to map the target context classification sequence corresponding to the core teaching semantic features from a pre-set context library. The teaching context-state association mapping table of the virtual classroom can be constructed using a graph database (such as Neo4j). The state adaptive classification threshold of the target context classification sequence can be generated using a Q-learning-based reinforcement learning model.
[0041] S3. Extract the key teaching context corresponding to the real-time interactive intent from the dynamic scene graph, calculate the cognitive coupling coefficient between the real-time interactive intent and the key teaching context, and generate an immersive narrative flow and interactive response rule set corresponding to the key teaching context based on the cognitive coupling coefficient.
[0042] This invention, through parsing the key teaching contexts corresponding to the real-time interactive intentions from the dynamic scene graph, clarifies the influence mechanism of node relationships in the dynamic scene graph on the selection of teaching contexts. This allows for targeted optimization of the dynamic matching accuracy and multimodal interactive response efficiency of virtual classroom teaching contexts. The key teaching contexts refer to teaching context units parsed from the dynamic scene graph that highly match the real-time interactive intentions of educational participants and play a core supporting role in achieving the current teaching objectives. Optionally, the key teaching contexts corresponding to the real-time interactive intentions can be parsed from the dynamic scene graph using a graph neural network (GAT) based on an attention mechanism.
[0043] Furthermore, by calculating the cognitive coupling coefficient between the real-time interactive intent and the key teaching context, this embodiment of the invention can accurately optimize the narrative generation logic and resource scheduling strategy of the immersive scene. The cognitive coupling coefficient is a numerical indicator used to quantify the degree of bidirectional influence and dynamic adaptation between the real-time interactive intent and the key teaching context.
[0044] As an embodiment of the present invention, calculating the cognitive coupling coefficient between the real-time interactive intent and the key teaching context includes: Identify the cognitive load demand corresponding to the real-time interactive intent, and analyze the cognitive resource supply level corresponding to the key teaching scenarios; Based on the cognitive load demand and the cognitive resource supply level, determine the cognitive load fit between the real-time interactive intent and the key teaching context; The degree of structuring corresponding to the key teaching scenarios is quantitatively analyzed, and the cognitive guidance intensity coefficient corresponding to the degree of structuring is determined. Analyze the emotional motivations corresponding to the real-time interaction intentions to calculate the motivational continuity matching degree corresponding to the real-time interaction intentions; Identify the intensity of cognitive conflict corresponding to the real-time interactive intent and the urgency of knowledge assimilation corresponding to the key teaching context; Based on the intensity of cognitive conflict and the urgency of knowledge assimilation, calculate the dynamic balance margin between the real-time interactive intention and the key teaching context; By integrating the cognitive load fit, the cognitive guidance intensity coefficient, the motivation continuity matching degree, and the dynamic balance margin, the cognitive coupling coefficient between the real-time interactive intention and the key teaching context is calculated.
[0045] The cognitive load requirement refers to the total amount and type of cognitive resources required to complete the learning task corresponding to the real-time interactive intentions of educational participants, as parsed from those intentions. This includes information processing load (such as the amount of logical connections needed to understand complex knowledge points), operational execution load (such as the number of steps required to complete a virtual experiment), and memory storage load (such as retaining key information needed to remember experimental operating procedures). The cognitive resource supply level refers to the cognitive support capabilities inherent in the key teaching context itself, which can support learners in completing the corresponding learning tasks. This includes information presentation support (such as whether complex knowledge points are broken down into step-by-step explanations) and operational assistance support (such as whether...). Experimental procedure prompts), memory reinforcement support (such as whether key information is repeated through animation); the cognitive load fit refers to a numerical value that characterizes the matching relationship between the cognitive load demand and the cognitive resource supply level, taking a value in the (0,1) interval. For example, a complex question requiring high working memory (high demand) is matched with a situation that can provide step-by-step diagrams and formula derivations (high supply), then the cognitive load fit is high. Optionally, the cognitive load fit can be determined by a psychological effort rating scale (such as NASA-TLX) in cognitive load theory; the degree of structuring refers to the organizational logic and operational efficiency of teaching content in a quantitative key teaching situation. The quantitative indicators of the standardization of the workflow and the orderliness of information presentation include content structuring (e.g., whether knowledge points are stratified by basic to advanced levels), process structuring (e.g., whether experimental operations have a fixed sequence of steps), and information structuring (e.g., whether charts are used to replace scattered text); the cognitive guidance intensity coefficient refers to a numerical value that quantifies the constraint and guidance force of the key teaching situation on the thinking path and operational scope of educational participants; the emotional motivation refers to the emotional tendency and learning motivation that drives educational participants to initiate the interactive behavior, as analyzed from the real-time interactive intention, including interest-based motivation (e.g., asking questions out of curiosity about "the cause of circuit failure") and task-based motivation (e.g., to complete homework). The motivational continuity matching degree refers to a quantitative value that assesses the degree of fit between the emotional motivation and the expected teaching narrative flow and rhythm of the key teaching situation; the cognitive conflict intensity refers to the degree of difference between the information or question carried by the real-time interactive intention and the existing knowledge schema of the educational participant; the knowledge assimilation urgency refers to the time pressure index required by the key teaching situation to digest cognitive conflict and complete knowledge update according to its preset teaching objectives and process rhythm, and the calculation formula is: Knowledge assimilation urgency = For example, in a fast-paced in-class quiz, the knowledge assimilation urgency is high due to limited remaining time and a large number of questions, while in a self-study session, the knowledge assimilation urgency can approach 0. The dynamic balance margin refers to the value obtained by calculating the ratio or functional relationship between the cognitive conflict intensity and the knowledge assimilation urgency, representing the stable and safe boundary that the teaching system maintains between eliciting cognitive challenges and ensuring knowledge assimilation. For example, an open-ended question that evokes strong cognitive conflict (high cognitive conflict intensity) will have a high dynamic balance margin if it occurs in a discussion session with ample time (low knowledge assimilation urgency), and the system will be in a safe exploration zone; if it occurs in an exam that is about to end (high knowledge assimilation urgency), the dynamic balance margin will be extremely low, and the system will be on the verge of disharmony.
[0046] Optionally, the cognitive guidance intensity coefficient corresponding to the degree of structuring can be determined using a structured analysis framework for educational scenarios (such as the ISTD (Instructional Behavior Classification)). For example, the ISTD can be used to encode the teacher-student interaction patterns in the context, and the degree of structuring can be determined based on the continuous spectrum coordinates of "teacher control - student autonomy," which can then be mapped to the guidance intensity coefficient. Alternatively, a multimodal emotion recognition model (such as the OpenFace toolkit) and a motivational state classifier (an SVM-based motivational classification model) can be used in affective computing to calculate the motivational continuity matching degree corresponding to the real-time interactive intent. The ACE toolkit extracts action units (such as AU4 eyebrow retraction and AU12 mouth corner twitching) from facial video streams, combines them with speech features (fundamental frequency and speech rate) to identify emotional states using an SVM classifier, and uses a motivation classification model to determine motivation types based on interactive content features. Finally, it calculates the cosine similarity between the motivation type and the expected state of the teaching narrative to determine the motivation type. It can identify the cognitive conflict intensity corresponding to the real-time interactive intent through a knowledge graph semantic distance algorithm (such as TransE vectorization representation), and can identify the knowledge assimilation urgency corresponding to the key teaching contexts using teaching temporal network analysis tools (such as dynamic curriculum planners).
[0047] In another embodiment of the present invention, the cognitive coupling coefficient between the real-time interactive intent and the key teaching context can be calculated using the following formula: ; in, Represents the cognitive coupling coefficient. Represents the cognitive coupling gain coefficient. Indicates the load weighting factor. Indicates cognitive load fit. Indicates the load sensitivity index. This represents the cognitive guidance intensity coefficient. Indicates the degree of matching of motivational continuity. Indicates dynamic balance margin. Indicates the offset buffer constant. This represents the balance sensitivity index.
[0048] Specifically, the cognitive coupling gain coefficient (Q) is a preset constant used to characterize an individual learner's baseline cognitive fusion ability, with a value range of (0,1]. For example, for a student whose cognitive style is convergent and who is good at learning in structured situations, the cognitive coupling gain coefficient can be set to 0.9; while for a learner whose cognition is more divergent and who needs more guidance, the cognitive coupling gain coefficient can be set to 0.7. Optionally, the cognitive coupling gain coefficient can be assessed using the Kolb Learning Style Scale. The load weighting factor (Q) The cognitive load sensitivity index (CFI) is a coefficient used to adjust the relative contribution weight of cognitive load fit in the calculation of the overall cognitive coupling coefficient. It can be derived through multiple linear regression analysis and ranges from 0 to 1. For example, in a mathematical reasoning context that highly relies on working memory, cognitive load is a key limiting factor, and the load weight factor should be set to a high value, such as 0.8; while in an art appreciation context that focuses on emotional experience, the load weight factor can be set to a relatively low value, such as 0.3. γ refers to a power exponent used to control the sensitivity of the cognitive coupling coefficient to changes in cognitive load fit. It can be determined using the analytic hierarchy process (AHP) based on expert experience. For example, when γ > 1, it means the system is highly sensitive to cognitive load fit; a slight decrease in cognitive load fit will lead to a significant decrease in the cognitive coupling coefficient, suitable for high-precision skill training with zero tolerance for cognitive overload. When γ < 1, the system is not sensitive to load changes and can allow for some fluctuations in cognitive load fit, suitable for exploratory learning. The offset buffer constant (…) The equilibrium sensitivity index () is a preset constant greater than zero. Its main function is to ensure that the denominator of the calculation formula is not zero when the dynamic equilibrium margin approaches zero, thereby maintaining the numerical stability of the system. Based on Lyapunov stability theory, it can be used to find the boundary of the attraction domain in the system's state space that guarantees the stable operation of the cognitive system. () refers to a power exponent used to control the sensitivity of the cognitive coupling coefficient to changes in the dynamic equilibrium margin, which can be determined by policy optimization algorithms in reinforcement learning, for example, when When = 0.5, it indicates that the system has moderate sensitivity to balance margin; when When the value is greater than 0.5, the system penalizes a reduction in balance margin more severely; even if other dimensions perform well, the coupling coefficient will be significantly reduced due to insufficient margin. When the value is less than 0.5, the system is more tolerant of a certain degree of cognitive imbalance.
[0049] It should be noted that this formula introduces the load sensitivity index ( Constructing power-law terms To quantify cognitive load fit ( The nonlinear amplification effect near the critical point accurately characterizes the sharp weakening effect of cognitive resource supply and demand imbalance on overall coupling; by using the cognitive guidance strength coefficient ( ) and the degree of matching of motivational continuity ( Designed as a product term This characterizes the profound synergistic relationship between the structured constraints of instruction and the learners' intrinsic emotional drive; the absence of either will lead to a precipitous drop in the coupling effect. This is achieved by constructing a regulation term based on a saturation function. And introduce the imbalance buffer constant (D) and the balance sensitivity index (D). Together, they modeled the system's cognitive balance margin ( The dynamic response characteristics under change; D ensures the numerical stability and robustness of the system when cognitive dissonance is imminent, while This finely regulates the system's strategy selection from risk preference to risk aversion in the equilibrium state; finally, the cognitive coupling gain coefficient (Q) is used to globally calibrate the overall coupling level based on the learner's individual ability and subject characteristics, thereby realizing the quantification of dynamic coupling strength that integrates cognition, emotion and context from isolated interactive behavior assessment.
[0050] Furthermore, by generating an immersive narrative flow and interactive response rule set corresponding to the key teaching scenarios based on the cognitive coupling coefficient, this embodiment of the invention can optimize in real time the core elements of the immersive narrative flow, such as the rhythm design standards and content adaptation dimensions, as well as the feedback trigger threshold and strategy execution parameters of the interactive response rule set, in a multimodal interactive-driven immersive remote education scenario. This significantly improves the accuracy of adapting the individual cognitive state and learning needs of remote education learners to the immersive teaching scenario.
[0051] The immersive narrative flow refers to a time-series script and content sequence dynamically generated based on the cognitive coupling coefficient, used to guide the orderly evolution of virtual classroom scene elements (such as environment, roles, and events) according to specific teaching objectives and emotional tones; the interactive response rule set refers to a set of rules predefined based on the cognitive coupling coefficient, used to regulate how the virtual classroom processes and responds to various interactive operations (such as voice, gestures, and eye contact) of educational participants.
[0052] As an embodiment of the present invention, the step of generating the immersive narrative flow and interactive response rule set corresponding to the key teaching context based on the cognitive coupling coefficient includes: Based on the cognitive coupling coefficient, obtain the teaching objective sequence data and cognitive state evolution data corresponding to the key teaching scenarios; Based on the teaching objective sequence data, the narrative dynamics features corresponding to the key teaching situations are extracted; Based on the narrative dynamics characteristics, an immersive narrative flow for the key teaching scenarios is generated; Based on the cognitive state evolution data, analyze the cognitive state trajectory of the key teaching scenarios; Identify the key state nodes of the cognitive state trajectory, and determine the cognitive safety boundary corresponding to the key teaching scenario through the key state nodes; Based on the cognitive safety boundary, construct a list of interaction constraints for the key teaching scenarios; By combining the immersive narrative flow and the list of interactive constraints, a set of interactive response rules for the key teaching scenarios is generated.
[0053] The teaching objective sequence data refers to a set of teaching objectives that, based on the cognitive development logic of basic cognition, advanced understanding, application and transfer, and comprehensive innovation, are broken down into key teaching scenarios and have clear hierarchical relationships, temporal connections, and quantifiable achievement standards. The cognitive state evolution data refers to a multi-dimensional data set reflecting the dynamic changes in learners' cognitive states over time, collected in real-time by multimodal interactive devices (such as eye trackers, EEG sensors, and interactive behavior acquisition terminals) during the learning process in key teaching scenarios. This includes attention state data (such as attention concentration, duration, and frequency of attention shifts) and knowledge mastery state data (such as the accuracy rate of answering knowledge points, the distribution of error types, and the number of knowledge gaps). The data includes: area markers, thinking activity state data (such as interactive operation response time, problem thinking dwell time, and participation in higher-order thinking tasks); the narrative dynamics features refer to the core feature elements extracted from the teaching objective sequence data based on key teaching contexts, which can drive the orderly advancement of the immersive narrative flow and adapt to the learner's cognitive development rhythm, including the narrative rhythm (such as the density of knowledge points delivered per unit time), tension (such as the intensity and frequency of setting cognitive challenges), branch probability (such as the weight of providing optional exploration paths at specific nodes), and loop triggering conditions (such as the conditions for triggering the review stage when insufficient knowledge mastery is detected); the cognitive state trajectory refers to the data of discrete cognitive state evolution through interpolation or state... After fitting the state space model, a continuous and smooth cognitive state change path is formed. For example, in a three-dimensional state space, a trajectory can be described as smoothly evolving from a state point with high cognitive load, low mastery, and neutral emotion to a state point with low cognitive load, high mastery, and positive emotion. The key state nodes refer to the state points in the three-dimensional state space where a trajectory can be described as smoothly evolving from a state point with high cognitive load, low mastery, and neutral emotion to a state point with low cognitive load, high mastery, and positive emotion, including the cognitive overload threshold (e.g., the cognitive load value continuously exceeds the threshold T for 5 seconds), the insight moment (e.g., the first derivative of the knowledge mastery curve shows a significant positive peak), and the attention loss focus (e.g., the attention index rapidly increases). The cognitive safety boundary refers to the boundary range set by a threshold algorithm based on key state nodes, combined with the teaching objectives of key teaching situations and the cognitive development laws of learners, to ensure that the learner's cognitive state is always within the effective learning range. It includes an upper limit boundary and a lower limit boundary. The upper limit boundary is the cognitive challenge boundary, which avoids cognitive overload caused by excessive cognitive load (such as setting the proportion of higher-order thinking tasks ≤40% and the number of knowledge modules presented in a single session ≤5). The lower limit boundary is the cognitive protection boundary, which avoids cognitive slackness caused by excessive cognitive load (such as setting the attention concentration ≥60% and the knowledge mastery accuracy rate ≥70%).The interaction constraint list refers to a set of constraint rules set according to cognitive safety boundaries for multimodal interactive behaviors in key teaching scenarios. These rules guide learners to maintain their cognitive state within the safety boundaries. They include: 1) interaction permission constraints (e.g., restricting the ability to skip teaching steps when the cognitive state is below the lower boundary; granting access to advanced interactive tasks when the cognitive state is above the upper boundary); 2) interaction feedback constraints (e.g., triggering guiding prompts when the cognitive state is close to the lower boundary; triggering encouraging achievement feedback when the cognitive state is within the safety boundaries); and 3) interaction frequency constraints (e.g., reducing the triggering frequency of high-difficulty interactive tasks when the cognitive state is unstable; increasing the triggering frequency of exploratory interactive tasks when the state is stable).
[0054] Optionally, the immersive narrative flow of the key teaching context can be generated using a dynamic event graph model. For example, each sub-goal in the sequence of teaching objectives can be modeled as a graph node, and the teaching logic and narrative dynamics features can be modeled as directed edges and their weights. By probabilistic sampling based on the cognitive coupling coefficient, the optimal plot path sequence can be dynamically generated to form an immersive narrative flow. The cognitive safety boundary corresponding to the key teaching context can be determined based on the cognitive state space analysis method of Lyapunov stability. For example, in the cognitive state space constructed by dimensions such as cognitive load, knowledge mastery, and emotional state, an energy function can be constructed using Lyapunov stability theory. By analyzing the state points in historical teaching data that lead to cognitive collapse (such as abandoning the task), the stability domain boundary of the energy function can be calculated and defined as the cognitive safety boundary.
[0055] S4. Based on the immersive narrative flow and the set of interactive response rules, establish a real-time rendering task queue corresponding to the dynamic scene graph, and combine the scene resource scheduler, the immersive narrative flow and the real-time rendering task queue to render the immersive teaching plot of the virtual classroom.
[0056] This invention establishes a real-time rendering task queue corresponding to the dynamic scene graph based on the immersive narrative flow and the set of interactive response rules. This can accurately match the narrative progression rhythm and multimodal interactive feedback requirements of immersive remote education scenarios, and strengthen the real-time dynamic correlation between the rendering output of the dynamic scene graph and the evolution of the teaching context. The real-time rendering task queue refers to a sequence of instructions that is dynamically generated and sorted according to the plot evolution logic of the immersive narrative flow and the real-time requirements of the set of interactive response rules, and is used to instruct the graphics rendering engine to load and draw virtual scene resources in a specific order and priority.
[0057] As an embodiment of the present invention, the step of establishing a real-time rendering task queue corresponding to the dynamic scene graph based on the immersive narrative flow and the set of interactive response rules includes: The key teaching event frames in the immersive narrative flow are parsed out, and the interactive hotspot areas in the set of interactive response rules are identified. Based on the key teaching event frames and the interactive hotspot areas, the rendering priority space corresponding to the dynamic scene map is divided, wherein the rendering priority space includes an instant rendering area, a high-priority streaming loading area, a regular loading area, and a background preloading area. By utilizing the plot urgency of the immersive narrative flow and the real-time interaction requirements of the interactive response rule set, resource scheduling thresholds for each region within the rendering priority space are defined. Based on the resource scheduling threshold, differentiated loading rules are set for each region within the rendering priority space; Based on the differentiated loading rules, resource loading instructions are generated for each region within the rendering priority space; Based on the rendering priority space, the differentiated loading rules, and the resource loading instructions, a real-time rendering task queue corresponding to the dynamic scene map is established.
[0058] The key teaching event frames refer to the key time nodes and scene state combinations that have a decisive impact on the achievement of teaching objectives and the cognitive guidance process, as parsed from the immersive narrative flow. For example, in the narrative flow of cell division, key teaching event frames include moments representing key steps in the division process, such as chromatin condensation, nuclear membrane disintegration, chromosome alignment at the equatorial plate, and sister chromatid separation. The interactive hotspot areas refer to specific three-dimensional spatial ranges or sets of virtual objects in the interactive space of the virtual classroom, defined by the set of interactive response rules, that allow or require educational participants to perform specific operations (such as clicking, grabbing, or gazing). For example, in a virtual chemistry experiment, interactive hotspot areas include the grabbable area of the alcohol lamp, the clickable selection area of the reagent bottle, and the safe operation area of the experimental table. The rendering priority space refers to a logical classification framework dynamically divided according to teaching and interactive needs to manage the loading and rendering order of all virtual scene resources, including an immediate rendering area, a high-priority streaming loading area, a regular loading area, and a background preloading area. The plot urgency refers to a quantitative indicator describing the urgency and time constraints of each plot node's progression in the immersive narrative flow. For example, the urgency level of the timed quiz segment is close to 1, requiring rapid scene rendering response to maintain the tense pace; the urgency level of the knowledge review segment is close to 0.3, allowing rendering resources to be appropriately allocated to other areas; the real-time interaction requirement refers to the constraint standard on the time delay between interactive behavior and scene rendering feedback extracted from the set of interactive response rules; the resource scheduling threshold refers to the upper and lower limits of the rendering resources (such as GPU computing power, memory usage, and network bandwidth) allocated to each area within the rendering priority space based on the urgency level and the real-time interaction requirement; the differentiated loading rule refers to the specific resource management strategy formulated for the characteristics of different areas in the rendering priority space and their resource scheduling thresholds. For example, for the real-time rendering area, the rule is preemptive loading, interrupting all low-priority tasks; for the high-priority streaming loading area, the rule is to guarantee bandwidth and use a lossless compression format; for the background preloading area, the rule is to load during idle time and use a lossy compression format; the resource loading instruction refers to an executable command generated according to the differentiated loading rule, containing information such as resource identifier, target memory address, loading parameters, and completion callback function.
[0059] Optionally, the rendering priority space corresponding to the dynamic scene graph can be divided using a clustering algorithm that combines teaching time constraints and spatial location relationships. For example, the scene resources can be divided into four priority regions. By calculating the teaching criticality weight (based on narrative flow event frames) and interaction sensitivity weight (based on hotspot regions) of each resource node, K-means clustering is used to form four clusters in the two-dimensional weight space, corresponding to the real-time rendering area, high-priority area, regular loading area, and background preloading area, respectively. The differentiated loading rules for each region within the rendering priority space can be set based on a QoS level policy template library. The resource loading instructions for each region within the rendering priority space can be dynamically generated by a real-time resource scheduler.
[0060] Furthermore, by combining the scene resource scheduler, the immersive narrative stream, and the real-time rendering task queue, the embodiments of the present invention render the immersive teaching plot of the virtual classroom. This deeply couples the learner's real-time cognitive interaction feedback with the dynamic construction needs of the virtual environment. Precise resource scheduling and sequential rendering instructions can enhance the fluency of the teaching plot presentation and the accuracy of the interaction feedback, thereby improving the efficiency of teaching information transmission and the realism of context perception in the dynamic evolution of the immersive remote education scene. The immersive teaching plot refers to the smallest narrative unit and interactive experience unit that is dynamically generated and presented in the virtual classroom through the collaborative drive of the scene resource scheduler and the real-time rendering task queue, and has complete teaching semantics and emotional appeal.
[0061] As an embodiment of the present invention, the step of rendering the immersive teaching plot of the virtual classroom by combining the scene resource scheduler, the immersive narrative stream, and the real-time rendering task queue includes: Identify the narrative evolution trigger points in the immersive narrative flow; Based on the spatiotemporal attributes of the narrative evolution trigger points, the synchronous marker points for the teaching content of the virtual classroom are determined; The collaborative rendering instructions of the real-time rendering task queue are parsed to schedule the multimodal resource package in the scene resource scheduler according to the resource loading requirements of the collaborative rendering instructions. Based on the synchronous markers of the teaching content, perform temporal alignment processing on each element in the multimodal resource package to obtain an aligned multimodal resource package; Based on the aligned multimodal resource package, an immersive teaching scenario for the virtual classroom is rendered.
[0062] The narrative evolution trigger point refers to the key node identified from the immersive narrative flow of the virtual classroom that propels the teaching plot from the current stage to the next stage. For example, in the narrative flow of a free-fall experiment in physics, the narrative evolution trigger point includes the release of the ball, the ball's first contact with the ground, and the completion of data collection. The spatiotemporal attributes refer to the set of parameters describing the time and space coordinates of the narrative evolution trigger point in the virtual classroom. The time coordinates include the absolute position and relative interval of the trigger point on the teaching timeline; the space coordinates include the three-dimensional spatial position that the camera should focus on when the trigger is activated, the identification of the scene environment (such as a laboratory or starry sky), and the expected spatial state of virtual objects (such as experimental instruments). The teaching content synchronization marker point refers to the point that, based on the spatiotemporal attributes of the narrative evolution trigger point, marks the teaching content (such as video lectures) in the virtual classroom. The core function of the system is to achieve temporal coordination of multiple types of teaching content (including explanations, audio narration, text courseware, and virtual interactive elements). For example, it can align the timeline of the audio narration explaining the quadratic function formula with the timeline of the formula writing animation on the virtual blackboard. The collaborative rendering instructions refer to the operation commands parsed from the real-time rendering task queue, used to coordinate the scene resource scheduler and the virtual classroom rendering process. These operation commands include priority instructions for rendering tasks (such as prioritizing the rendering of the interactive interface of the practice session), resource coordination instructions (such as synchronously loading scene models that match the explanation content), and timing control instructions (such as completing rendering preparation within 500ms after activation of the trigger point). The resource loading requirements refer to the types, quantities, and loading standards of resources required to complete the rendering of the corresponding teaching plot in the virtual classroom, as specified by the collaborative rendering instructions. Specifically, this includes the format requirements of multimodal resources (e.g., videos must be in MP4 format, virtual models must be in FBX format), quality requirements (e.g., courseware images must have a resolution of no less than 1920×1080), loading time requirements (e.g., interactive resources must be ready to load 2 seconds before activation at the trigger point), and dependency requirements between resources (e.g., scene lighting resources must be loaded before loading the experimental scene model). The multimodal resource package refers to a collection of diverse teaching resources managed by the scene resource scheduler, used to construct immersive teaching scenarios in a virtual classroom. These resources include visual resources (e.g., virtual classroom scene models, teaching courseware images, experimental operation animations), auditory resources (e.g., teacher explanation audio, scene background sound effects, interactive feedback prompts), and interactive resources (e.g., virtual teaching aid models, interactive question-answering interfaces, operation guidance pop-ups). The aligned multimodal resource package refers to a collection of resources in the multimodal resource package that has been calibrated in terms of time axis and spatial position according to the synchronous marker points of the teaching content. For example, by aligning the time axis, the key steps of the circuit connection explanation audio and the virtual circuit connection animation are completely synchronized.
[0063] Optionally, the collaborative rendering instructions of the real-time rendering task queue can be parsed by a task dependency graph parsing algorithm; the multimodal resource packages in the scene resource scheduler can be scheduled by a resource dependency management algorithm (such as a topology sorting algorithm); and the immersive teaching scenarios of the virtual classroom can be rendered using the Unity HDRP rendering pipeline tool.
[0064] S5. Based on the dynamic scene map, the remote education immersive scene corresponding to the education participant is generated through the real-time rendering task queue and the immersive teaching plot.
[0065] This invention, through the dynamic scene map, the real-time rendering task queue, and the immersive teaching scenario, generates a remote education immersive scene corresponding to the educational participants. This enables the virtual teaching environment to present a personalized teaching experience with optimal efficiency and controllable cognitive load under the drive of multimodal interaction. At the same time, it avoids the problems of scene rigidity, lack of interaction, and context mismatch caused by relying on preset linear scripts and static resources, and greatly improves the adaptability, immersion, and teaching effectiveness of the remote education scene.
[0066] The aforementioned immersive remote education scenario refers to a virtual learning environment built for educational participants based on a structured framework of dynamic scene graphs. This environment features multimodal interaction capabilities, spatiotemporal consistency, and cognitive adaptability, achieved through resource scheduling and priority control of real-time rendering task queues, combined with the content evolution logic of immersive teaching scenarios. Optionally, the immersive remote education scenario corresponding to the educational participants can be generated using a Web-based real-time rendering framework (such as WebGL).
[0067] To clearly present the design logic and structure of the virtual classroom scenario, and to support the generation of immersive scenarios in distance education, please refer to [reference needed]. Figure 3 The diagram shows the scene structure design of a virtual classroom for a multimodal interactive-driven immersive scene construction method for distance education, as provided in an embodiment of the present invention. The spatial elements, interface elements, and technical elements in the diagram provide a basis for resource classification and association for the scene resource scheduler configuration. For example, different interface layout forms, resources corresponding to visual elements, and technical resources supporting interaction can be used to generate a scene resource scheduling list and a dynamic content parameter set, thereby setting teaching constraints. The process of "familiarizing oneself with spatial connotation → designing spatial experience → enhancing spatial perception → achieving spatial identity" and the cognitive logic of "perception → comprehension" in the diagram provide ideas for generating an immersive narrative flow and a set of interactive response rules. Based on this, a real-time rendering task queue can be better established, and combined with the scene resource scheduler, an immersive teaching plot that conforms to cognitive laws and interactive intentions can be rendered, thereby improving the immersion and effectiveness of distance education.
[0068] Compared to the problems described in the background technology, the embodiments of the present invention, by configuring the scene resource scheduler of the virtual classroom according to the real-time interactive intent, can enhance the scene resource scheduler's intelligent perception and on-demand supply capabilities for the entire dynamic evolution process of the virtual classroom, ensuring the real-time generation of the immersive teaching experience and the matching degree of the teaching intent; furthermore, by extracting the scene resource scheduling list and dynamic content parameter set from the scene resource scheduler to set the teaching constraints of the virtual classroom, the embodiments of the present invention can realize the dynamic planning of the teaching process and the flexible allocation of teaching resources, ensuring that the generated teaching path is highly consistent with the real-time state of the classroom, effectively avoiding problems such as rigid teaching processes, lack of on-site interaction, and insufficient personalized support caused by traditional preset frameworks; the embodiments of the present invention, by constructing the dynamic scene map of the virtual classroom according to the teaching constraints and the real-time interactive intent, can clarify the dynamic coupling relationship and evolution law of teaching rule boundaries and teacher-student interaction behavior in the multimodal virtual space, and thus specifically optimize the dynamic generation logic and multimodal interaction response accuracy of the immersive scene of distance education; furthermore, the embodiments of the present invention... By calculating the cognitive coupling coefficient between the real-time interactive intent and the key teaching context, this embodiment of the invention can accurately optimize the narrative generation logic and resource scheduling strategy of the immersive scene. Based on the cognitive coupling coefficient, this embodiment generates an immersive narrative flow and a set of interactive response rules corresponding to the key teaching context. This allows for real-time optimization of core elements in multimodal interactive-driven immersive remote education scenarios, such as the rhythm design standards and content adaptation dimensions of the immersive narrative flow, as well as the feedback trigger threshold and strategy execution parameters of the interactive response rule set. This significantly improves the accuracy of adapting the individual cognitive state and learning needs of remote education learners to the immersive teaching scenario. Finally, based on the dynamic scene map, this embodiment generates a remote education immersive scene corresponding to the educational participants through the real-time rendering task queue and the immersive teaching plot. This enables the virtual teaching environment to present a personalized teaching experience with optimal efficiency and controllable cognitive load under multimodal interactive driving, while avoiding scene rigidity, lack of interaction, and context mismatch caused by relying on preset linear scripts and static resources. This greatly improves the adaptability, immersion, and teaching effectiveness of the remote education scenario. Therefore, this invention can flexibly respond to unexpected needs in the teaching process and improve the adaptability of distance education scenarios.
[0069] like Figure 4 The diagram shown is a functional module diagram of a multimodal interactive-driven immersive scene construction system for distance education according to the present invention.
[0070] The multimodal interaction-driven immersive scene construction system 200 for distance education described in this invention can be installed in an electronic device. Depending on the functions implemented, the multimodal interaction-driven immersive scene construction system for distance education may include an intent recognition module 201, a resource setting module 202, a context analysis module 203, a plot rendering module 204, and a scene generation module 205. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0071] In this embodiment of the invention, the functions of each module / unit are as follows: The intent recognition module 201 is used to acquire multimodal interaction data of educational participants in a virtual classroom, so as to parse the real-time interaction intent of the educational participants and configure the scene resource scheduler of the virtual classroom according to the real-time interaction intent. The resource setting module 202 is used to extract the scene resource scheduling list and dynamic content parameter set from the scene resource scheduler in order to set the teaching constraints of the virtual classroom, and construct the dynamic scene map of the virtual classroom according to the teaching constraints and the real-time interaction intention. The context analysis module 203 is used to parse the key teaching context corresponding to the real-time interactive intent from the dynamic scene graph, calculate the cognitive coupling coefficient between the real-time interactive intent and the key teaching context, and generate an immersive narrative flow and interactive response rule set corresponding to the key teaching context based on the cognitive coupling coefficient. The plot rendering module 204 is used to establish a real-time rendering task queue corresponding to the dynamic scene graph based on the immersive narrative flow and the set of interactive response rules, and to render the immersive teaching plot of the virtual classroom by combining the scene resource scheduler, the immersive narrative flow and the real-time rendering task queue. The scene generation module 205 is used to generate a remote education immersive scene corresponding to the education participant based on the dynamic scene map, through the real-time rendering task queue and the immersive teaching plot.
[0072] In detail, the modules in the multimodal interaction-driven immersive scene construction system 200 for remote education described in this embodiment of the invention employ the same methods as described above. Figure 1 The method uses the same techniques as the multimodal interaction-driven immersive scenario construction method for distance education described above, and can produce the same technical effects, so it will not be elaborated here.
[0073] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0074] Finally, it should be noted that in the above embodiments, each embodiment can be combined with each other or independent. Deleting any one of them will not affect the technical implementation of other embodiments. The above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for constructing an immersive scenario for distance education driven by multimodal interaction, characterized in that, The method includes: Acquire multimodal interaction data of educational participants in a virtual classroom to analyze the real-time interaction intentions of the educational participants, and configure the scene resource scheduler of the virtual classroom based on the real-time interaction intentions; Extract the scene resource scheduling list and dynamic content parameter set from the scene resource scheduler to set the teaching constraints of the virtual classroom. Based on the teaching constraints and the real-time interaction intent, construct the dynamic scene graph of the virtual classroom. The key teaching context corresponding to the real-time interactive intent is parsed from the dynamic scene graph, the cognitive coupling coefficient between the real-time interactive intent and the key teaching context is calculated, and based on the cognitive coupling coefficient, an immersive narrative flow and interactive response rule set corresponding to the key teaching context are generated. Based on the immersive narrative flow and the set of interactive response rules, a real-time rendering task queue corresponding to the dynamic scene graph is established. Combining the scene resource scheduler, the immersive narrative flow, and the real-time rendering task queue, the immersive teaching plot of the virtual classroom is rendered. Based on the dynamic scene map, the remote education immersive scene corresponding to the education participant is generated through the real-time rendering task queue and the immersive teaching plot.
2. The method for constructing a multimodal interaction-driven immersive remote education scene as described in claim 1, characterized in that, The step of constructing a dynamic scene map of the virtual classroom based on the teaching constraints and the real-time interactive intent includes: Based on the teaching constraints and the real-time interaction intent, the core teaching semantic features of the virtual classroom are extracted; Based on the core teaching semantic features, configure the teaching context classification rule base for the virtual classroom; Match the context code and teaching priority corresponding to the real-time interactive intent from the teaching context classification rule base; Based on the context code and the teaching priority, identify the teaching context nodes of the virtual classroom; The relationships between knowledge points and the spatiotemporal boundaries of activities within the teaching constraints are analyzed to mark the teaching activity nodes in the virtual classroom. Based on the mapping rules between the context code and the knowledge point association, a logical association link is established between the teaching context node and the teaching activity node; Based on the teaching context nodes, the teaching activity nodes, and the logical connection links, a dynamic scene map of the virtual classroom is constructed.
3. The method for constructing a multimodal interaction-driven immersive remote education scene as described in claim 2, characterized in that, The step of configuring the teaching context classification rule base for the virtual classroom based on the core teaching semantic features includes: The feature dimensions and semantic strength of the core teaching semantic features are analyzed, and the context type corresponding to the core teaching semantic features is determined. Extract the real-time teaching status parameters of the virtual classroom, wherein the real-time teaching status parameters include device interaction characteristics, student cognitive status and teaching task progress indicators; Calculate the context matching confidence of the core teaching semantic features based on the feature dimensions, the semantic strength, and the real-time teaching status parameters; Based on the context matching confidence and the context type, the target context classification sequence corresponding to the core teaching semantic features is mapped from the pre-set context library; Based on the real-time teaching state parameters, generate the state adaptive classification threshold of the target context classification sequence; By integrating the target context classification sequence with the state adaptive classification threshold, a teaching context-state association mapping table for the virtual classroom is constructed. Based on the teaching context-state association mapping table, configure the teaching context classification rule base for the virtual classroom.
4. The method for constructing a multimodal interaction-driven immersive scenario for distance education as described in claim 1, characterized in that, The process of generating the immersive narrative flow and interactive response rule set corresponding to the key teaching scenarios based on the cognitive coupling coefficient includes: Based on the cognitive coupling coefficient, obtain the teaching objective sequence data and cognitive state evolution data corresponding to the key teaching scenarios; Based on the teaching objective sequence data, the narrative dynamics features corresponding to the key teaching situations are extracted; Based on the narrative dynamics characteristics, an immersive narrative flow for the key teaching scenarios is generated; Based on the cognitive state evolution data, analyze the cognitive state trajectory of the key teaching scenarios; Identify the key state nodes of the cognitive state trajectory, and determine the cognitive safety boundary corresponding to the key teaching scenario through the key state nodes; Based on the cognitive safety boundary, construct a list of interaction constraints for the key teaching scenarios; By combining the immersive narrative flow and the list of interactive constraints, a set of interactive response rules for the key teaching scenarios is generated.
5. The method for constructing a multimodal interaction-driven immersive scenario for distance education as described in claim 1, characterized in that, The step of configuring the scene resource scheduler of the virtual classroom according to the real-time interaction intent includes: Analyze the semantics of the teaching context corresponding to the real-time interactive intent; Extract the core resource demand tags from the semantics of the teaching context; Identify the functional attributes and dependencies of various scenario resources in the virtual classroom; Calculate the collaborative scheduling weight of the scene resources based on the functional attributes and the dependencies; Based on the core resource requirement tags and the collaborative scheduling weights, a dynamic loading priority queue for the scene resources is generated. Based on the dynamic loading priority queue, a batch loading scheme for the scene resources is set, and a resource contention allocation rule for the batch loading scheme is established. By combining the dynamic loading priority queue, the batch loading scheme, and the resource contention allocation rules, a scene resource scheduler for the virtual classroom is constructed.
6. The method for constructing a multimodal interaction-driven immersive scenario for distance education as described in claim 1, characterized in that, The step of extracting the scene resource scheduling list and dynamic content parameter set from the scene resource scheduler to set the teaching constraints of the virtual classroom includes: Parse the resource function metadata in the scenario resource scheduling list and the user state vector in the dynamic content parameter set; By establishing a logical association between the resource function metadata and the user state vector, a context-resource logical topology network corresponding to the virtual classroom is constructed. Identify the key context switching nodes corresponding to the context-resource logical topology network, and define the system control parameters of the key context switching nodes; Extract the content entity boundaries, user operation permission range, and process advancement threshold from the system control parameters to define the rule instantiation strategy for the key context switching nodes; Based on the rule instantiation strategy, the teaching constraints of the virtual classroom are set.
7. The method for constructing a multimodal interaction-driven immersive scenario for distance education as described in claim 1, characterized in that, The calculation of the cognitive coupling coefficient between the real-time interactive intent and the key teaching context includes: Identify the cognitive load demand corresponding to the real-time interactive intent, and analyze the cognitive resource supply level corresponding to the key teaching scenarios; Based on the cognitive load demand and the cognitive resource supply level, determine the cognitive load fit between the real-time interactive intent and the key teaching context; The degree of structuring corresponding to the key teaching scenarios is quantitatively analyzed, and the cognitive guidance intensity coefficient corresponding to the degree of structuring is determined. Analyze the emotional motivations corresponding to the real-time interaction intentions to calculate the motivational continuity matching degree corresponding to the real-time interaction intentions; Identify the intensity of cognitive conflict corresponding to the real-time interactive intent and the urgency of knowledge assimilation corresponding to the key teaching context; Based on the intensity of cognitive conflict and the urgency of knowledge assimilation, calculate the dynamic balance margin between the real-time interactive intention and the key teaching context; By integrating the cognitive load fit, the cognitive guidance intensity coefficient, the motivation continuity matching degree, and the dynamic balance margin, the cognitive coupling coefficient between the real-time interactive intention and the key teaching context is calculated.
8. The method for constructing a multimodal interaction-driven immersive scenario for distance education as described in claim 1, characterized in that, The step of establishing a real-time rendering task queue corresponding to the dynamic scene graph based on the immersive narrative flow and the set of interactive response rules includes: The key teaching event frames in the immersive narrative flow are parsed out, and the interactive hotspot areas in the set of interactive response rules are identified. Based on the key teaching event frames and the interactive hotspot areas, the rendering priority space corresponding to the dynamic scene map is divided, wherein the rendering priority space includes an instant rendering area, a high-priority streaming loading area, a regular loading area, and a background preloading area. By utilizing the plot urgency of the immersive narrative flow and the real-time interaction requirements of the interactive response rule set, resource scheduling thresholds for each region within the rendering priority space are defined. Based on the resource scheduling threshold, differentiated loading rules are set for each region within the rendering priority space; Based on the differentiated loading rules, resource loading instructions are generated for each region within the rendering priority space; Based on the rendering priority space, the differentiated loading rules, and the resource loading instructions, a real-time rendering task queue corresponding to the dynamic scene map is established.
9. The method for constructing a multimodal interaction-driven immersive scenario for distance education as described in claim 1, characterized in that, The process of rendering the immersive teaching scenario of the virtual classroom by combining the scene resource scheduler, the immersive narrative stream, and the real-time rendering task queue includes: Identify the narrative evolution trigger points in the immersive narrative flow; Based on the spatiotemporal attributes of the narrative evolution trigger points, the synchronous marker points for the teaching content of the virtual classroom are determined; The collaborative rendering instructions of the real-time rendering task queue are parsed to schedule the multimodal resource package in the scene resource scheduler according to the resource loading requirements of the collaborative rendering instructions. Based on the synchronous markers of the teaching content, perform temporal alignment processing on each element in the multimodal resource package to obtain an aligned multimodal resource package; Based on the aligned multimodal resource package, an immersive teaching scenario for the virtual classroom is rendered.
10. A multimodal interaction-driven immersive scenario construction system for distance education, characterized in that, The system includes: The intent recognition module is used to acquire multimodal interaction data of educational participants in the virtual classroom, so as to parse the real-time interaction intent of the educational participants and configure the scene resource scheduler of the virtual classroom according to the real-time interaction intent. The resource setting module is used to extract the scene resource scheduling list and dynamic content parameter set from the scene resource scheduler in order to set the teaching constraints of the virtual classroom, and construct the dynamic scene map of the virtual classroom according to the teaching constraints and the real-time interaction intention. The context analysis module is used to parse the key teaching context corresponding to the real-time interactive intent from the dynamic scene graph, calculate the cognitive coupling coefficient between the real-time interactive intent and the key teaching context, and generate an immersive narrative flow and interactive response rule set corresponding to the key teaching context based on the cognitive coupling coefficient. The plot rendering module is used to establish a real-time rendering task queue corresponding to the dynamic scene graph based on the immersive narrative flow and the set of interactive response rules, and to render the immersive teaching plot of the virtual classroom by combining the scene resource scheduler, the immersive narrative flow and the real-time rendering task queue. The scene generation module is used to generate a remote education immersive scene corresponding to the education participant based on the dynamic scene map, through the real-time rendering task queue and the immersive teaching plot.
Citation Information
Cited By
Classroom teaching behavior analysis method and system based on multi-modal data fusion
CN121724812A