Online teaching interaction method based on multi-modal knowledge graph, medium and equipment
By constructing a multimodal knowledge graph and dynamically adjusting teaching strategies, the problem of insufficient fragmentation of knowledge and interactive intelligence in virtual simulation teaching is solved, and personalized and intelligent online teaching is realized.
Patent Information
- Application Number
- CN202510764531.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing virtual simulation teaching system has shortcomings in the systematic construction of knowledge systems and the intelligence of teaching interactions, and it is difficult to achieve organic integration and personalized teaching of multimodal teaching resources.
The online teaching method based on multimodal knowledge graph is adopted to construct a multimodal knowledge graph through cross-modal semantic alignment processing, and a three-dimensional teaching situation sub-graph is generated based on user interaction behavior data, and the teaching strategy is dynamically adjusted through the cognitive state tracking matrix to generate a personalized learning navigation map.
It improves the teaching effect of virtual simulation teaching, realizes systematic integration and personalized teaching of knowledge elements, and enhances the intelligence and adaptability of teaching.
Smart Images

Figure CN120339011A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ideological and political education, and particularly to an online teaching interaction method, medium and device based on a multi-modal knowledge graph. Background Art
[0002] In the field of ideological and political education informatization, a teaching mode combining video lectures, courseware display and online testing is mainly adopted. With the development of education informatization, virtual simulation technology has become an important means to improve teaching effects, and various online platforms carry out teaching activities through three-dimensional scene simulation, interactive operations and other means. Currently, the mainstream virtual teaching systems mainly rely on preset programs to implement basic teaching functions. Although they can provide a certain immersive experience, there are obvious deficiencies in the systematic construction of knowledge systems and the intelligence of teaching interactions. The teaching elements in virtual scenes often lack deep semantic associations and are difficult to construct a systematic knowledge system. Especially in dealing with multi-modal teaching resources and realizing personalized teaching, the existing virtual simulation systems are difficult to achieve the organic integration of knowledge elements and generally show problems of insufficient adaptability. Summary of the Invention
[0003] In view of the above problems, the present invention provides an online teaching interaction method, medium and device based on a multi-modal knowledge graph, which solves the problems of knowledge fragmentation and insufficient interaction intelligence in virtual simulation teaching.
[0004] To achieve the above object, in a first aspect, the present application provides an online teaching interaction method based on a multi-modal knowledge graph, including: Collecting original teaching information, performing cross-modal semantic alignment processing on the original teaching information to obtain structured teaching information; Constructing a multi-modal knowledge graph, and deeply fusing text concept entities, video key frame feature vectors and speech text transcripts through a graph neural network to obtain cross-modal semantic association edges, where the cross-modal semantic association edges include concept-visual feature association edges, concept-speech segment association edges and cross-modal similarity association edges; Receiving user interaction behavior data, where the user interaction behavior data includes speech question content, virtual teaching operation sequences and interface click trajectories, and converting the user interaction behavior data into a knowledge graph query vector; Extracting a three-dimensional teaching scenario sub-graph from the multi-modal knowledge graph according to the knowledge graph query vector, activating the forward reasoning path of the three-dimensional teaching scenario sub-graph to generate progressive teaching content when a standard teaching operation is detected, and triggering the counterfactual reasoning mechanism of the three-dimensional teaching scenario sub-graph to generate comparative teaching content when an incorrect teaching operation is detected; Record the interactive behavior data of users in the three-dimensional teaching scenario sub-graph in real time, denoted as scenario interaction data, match the mapping relationship between the scenario interaction data and the multi-modal knowledge graph, and construct a cognitive state tracking matrix; Dynamically adjust the teaching strategy according to the cognitive state tracking matrix, strengthen the cross-modal display combination of concept nodes with low mastery, and extend the high-order knowledge associated with high-proficiency operation sequences; In addition, aggregate the cognitive state tracking matrix according to the group behavior analysis model, adjust the teaching path recommendation weights in the multi-modal knowledge graph, and generate a personalized learning navigation graph, which includes the optimal learning path planning and multi-modal resource recommendation scheme.
[0005] Furthermore, the original teaching information includes textbook text information, teaching video information, and classroom voice information; Collect the original teaching information, perform cross-modal semantic alignment processing on the original teaching information to obtain structured teaching information, including: Obtain the textbook text information, extract teaching concept entities from it through named entity recognition, and generate concept feature vectors by performing word embedding encoding on each teaching concept entity; In addition, obtain the teaching video information, perform key frame sampling on the video stream to obtain a video key frame sequence, and the video key frame sequence includes multiple first video key frames; In addition, obtain the classroom voice information, perform segmentation processing on the voice stream to obtain a voice segment set, and the voice segment set includes multiple first voice segments; Construct a cross-modal semantic alignment model, including: Calculate the first attention weights between the concept feature vectors and the visual feature vectors of all video key frames, and calculate the second attention weights between the concept feature vectors and the acoustic feature vectors of all voice segments; According to the first attention weights and the second attention weights, screen out the video key frames and voice segments strongly associated with the teaching concept entities, denoted as the second video key frames and the second voice segments; Generate structured teaching information according to the cross-modal semantic alignment model. The structured teaching information includes teaching concept entities and concept feature vectors, second video key frames and first attention weights, second voice segments and second attention weights, visual association relationships between teaching concept entities and second video key frames, and voice association relationships between teaching concept entities and second voice segments.
[0006] Furthermore, construct a multi-modal knowledge graph, and deeply fuse text concept entities, video key frame feature vectors, and voice text transcriptions through a graph neural network to obtain cross-modal semantic association edges, including: Construct an initial node set of the multi-modal knowledge graph, and the initial node set includes: Taking the teaching concept entity as the concept node, the node feature is the concept feature vector; Taking the second video key frame as the visual node, the node feature is the visual feature vector; Taking the second voice segment as the voice node, the node feature is the voice feature vector; Establish cross-modal semantic association edges, including: According to the first attention weight value of the concept node, establish a concept-visual association edge with the visual node, and the weight value of the concept-visual association edge is the first attention weight after normalization; According to the second attention weight value of the concept node, establish a concept-voice association edge with the voice node, and the weight value of the concept-voice association edge is the second attention weight after normalization; For the visual node and the voice node with a temporal correspondence relationship, establish a visual-voice temporal edge, and the weight value of the visual-voice temporal edge is calculated according to the time overlap degree between the second video key frame and the second voice segment; Perform deep fusion of multi-modal features through a graph neural network, including: Input the initial node set into the graph attention network; Perform feature propagation along the concept-visual association edge and the concept-voice association edge, and calculate the cross-modal similarity score between nodes through the multi-head attention mechanism; Update the cross-modal feature representation of the initial node set according to the cross-modal similarity score, and generate knowledge graph nodes with fused features; Obtain the final multi-modal knowledge graph.
[0007] Furthermore, convert the user interaction behavior data into a knowledge graph query vector, including: Convert the voice question content into query text through a speech recognition engine; Use named entity recognition to extract the core concept entity in the query text; Map the core concept entity to the concept node in the multi-modal knowledge graph, and use the feature vector of the successfully mapped concept node as the text query base vector; And, parse the operation object and operation action in the virtual teaching operation sequence to obtain the parsing result; Locate the corresponding teaching process node in the multi-modal knowledge graph according to the parsing result, and associate the feature vector of the teaching process node as the operation query base vector; And, identify the feature vector of the knowledge graph node associated with the click area of the interface click trajectory, denoted as the click feature vector, and generate a heat weight according to the click frequency; Generate a click query base vector according to the heat weight and the click feature vector; Weightedly fuse the text query base vector, the operation query base vector, and the click query base vector to obtain a fused feature vector, and input the fused feature vector into the query encoder to generate the final knowledge graph query vector.
[0008] Furthermore, extract a three-dimensional teaching scenario sub-graph from the multi-modal knowledge graph according to the knowledge graph query vector, including: Perform node matching in the multi-modal knowledge graph based on the knowledge graph query vector to locate the initial query node; Perform graph traversal along the concept-visual association edge and the concept-audio association edge to extract the associated visual query node and audio query node; Establish a spatio-temporal association relationship according to the visual-audio temporal edge to form a three-dimensional teaching scenario sub-graph; When a standard teaching operation is detected, activate the forward inference path of the three-dimensional teaching scenario sub-graph to generate progressive teaching content, including: Start from the initial query node and perform reasoning along the standard teaching path in the knowledge graph; Combine the associated concept query nodes, visual query nodes, and audio query nodes in the teaching logic order; Generate a progressive teaching sequence including basic concept explanations, example demonstrations, and audio guidance; When an incorrect teaching operation is detected, trigger the counterfactual reasoning mechanism of the three-dimensional teaching scenario sub-graph to generate comparative teaching content, including: Identify the abnormal node in the knowledge graph corresponding to the incorrect operation; Disable the standard teaching path edge directly associated with the abnormal node; Activate the counterfactual reasoning edge to generate teaching content including error analysis, correct demonstration, and comparative explanation.
[0009] Furthermore, the scenario interaction data includes the access frequency and residence duration of the concept query node, the viewing completeness of the visual query node, the repeated playback times of the audio query node, and the selection ratio of the standard teaching path and the counterfactual reasoning path; Match the mapping relationship between the scenario interaction data and the multi-modal knowledge graph to construct a cognitive state tracking matrix, including: Establish an initial cognitive state tracking matrix. The rows of the initial cognitive state tracking matrix represent the concept nodes in the multi-modal knowledge graph, and the columns include the proficiency index, the operation proficiency index, the error pattern index, and the attention distribution index; Map the access records of the concept query nodes to the concept nodes, calculate the average test correct rate and the access duration score of the concept nodes to obtain the proficiency index; Map the viewing records of the visual query nodes to the visual nodes, and count the operation success rate and the average time consumption of the teaching process nodes to obtain the operation proficiency index; Mapping the teaching path selection record to the standard teaching path or the counterfactual reasoning path, recording the triggering times of the counterfactual reasoning path and the subsequent operation improvement, and obtaining the error mode index; Map the playback records of the voice query node to the voice node, and obtain the attention distribution index according to the frequency of interaction events among the voice node, concept node, and visual node within a unit time. The proficiency index, the operation proficiency index, the error pattern index, and the attention distribution index are updated to the columns of the initial cognitive state tracking matrix to obtain an updated cognitive state tracking matrix.
[0010] Furthermore, the teaching strategy is dynamically adjusted according to the cognitive state tracking matrix to strengthen the cross-modal presentation combination of low-proficiency concept nodes and extend the high-level knowledge associated with high-proficiency operation sequences, including: Extracting a set of low-mastery concept nodes whose proficiency index is lower than a first preset threshold from the cognitive state tracking matrix; Perform a cross-modal reinforcement step for each low-mastery concept node in the set of low-mastery concept nodes: Increase the display duration and frequency of visual nodes associated with low-mastery concept nodes; Improve the playback priority of speech nodes associated with low-mastery concept nodes; Produce intensive instructional packages that include textual explanations, visual demonstrations, and audio guidance; Extracting a set of high-proficiency concept nodes whose operation proficiency index is higher than a second preset threshold from the cognitive state tracking matrix; Perform the high-order knowledge extension step for each high-proficiency concept node in the high-proficiency concept node set: Extracting high-order visual nodes along concept-visual association edges; Extracting higher-order phonetic nodes along concept-phonetic association edges; Generate a high-level teaching package that includes advanced knowledge explanations, in-depth case analysis, and extended exercises; According to the error mode indicators in the cognitive state tracking matrix, the triggering conditions of the counterfactual reasoning path are adjusted, including: Add early warning nodes for high-frequency error modes; Optimize the comparison display method between incorrect operations and correct demonstrations; According to the attention distribution indicators in the cognitive state tracking matrix, the teaching rhythm is dynamically adjusted, including: Simplify teaching content for low attention areas; Add interactive links to high-attention areas; Update the existing teaching strategy configuration to obtain an updated teaching configuration and execute it. The updated teaching configuration includes an enhanced teaching plan for low - mastery concept nodes, an advanced extension plan for high - proficiency concept nodes, optimized parameters for counterfactual reasoning paths, and adjustment parameters for teaching rhythm.
[0011] Furthermore, aggregate the cognitive state tracking matrices according to the group behavior analysis model, adjust the teaching path recommendation weights in the multi - modal knowledge graph, and generate a personalized learning navigation graph, including: Obtain the cognitive state tracking matrices of multiple users, denoted as the group cognitive state tracking matrix. The group cognitive state tracking matrix includes group concept nodes. Calculate the average mastery index of each group concept node and statistically analyze the high - frequency error patterns and group - associated nodes; Identify the hot areas of group attention distribution after analyzing the common features of the optimal learning paths; Calculate the average mastery index, high - frequency error patterns, and characteristics of the optimal learning paths of the group cognitive state tracking matrix, and adjust the teaching path recommendation weights in the multi - modal knowledge graph in combination with the hot areas of group attention distribution, including: Reduce the basic teaching weight for concept nodes with high group mastery, and increase the counterfactual reasoning weight for group high - frequency error nodes; Increase the recommendation priority for the associated edges in the group optimal path, and optimize the resource display weight for the hot areas of group attention; Extract the sub - graph structure related to the current user from the multi - modal knowledge graph according to the adjusted teaching path recommendation weights; Recalculate the path scores of the teaching paths, mark the key concept nodes on the recommended learning paths, and mark the error - prone node areas that need to be focused on; Generate a learning navigation graph including the main learning path and alternative paths with personalized recommendations, recommended concept nodes, recommended visual nodes, recommended voice node combinations, recommended path weight scores, estimated learning durations, and prompt marks for key reinforcement areas.
[0012] In a second aspect, the present invention also provides a computer - readable storage medium storing computer program instructions, which, when executed by a processor, implement the method described in the first aspect.
[0013] In a third aspect, the present invention also provides an electronic device including a memory and a processor. The memory is used to store one or more computer program instructions, and the one or more computer program instructions are executed by the processor to implement the method described in the first aspect.
[0014] Different from the prior art, the above - mentioned technical solution has the following beneficial effects: The above technical solution provides an online teaching interaction method, medium and device based on a multimodal knowledge graph. The method first collects original teaching information and performs cross-modal semantic alignment processing to obtain structured teaching information; then constructs a multimodal knowledge graph including text concept entities, video key frame feature vectors, and speech text transcripts, forming concept-visual feature association edges, concept-speech segment association edges, and cross-modal similarity association edges; converts user interaction behavior data into a knowledge graph query vector, extracts a three-dimensional teaching scenario subgraph from the multimodal knowledge graph, and generates progressive or comparative teaching content according to the user operation type; finally, dynamically adjusts the teaching strategy by constructing a cognitive state tracking matrix and generates a personalized learning navigation graph. The above technical solution solves the problems of fragmented knowledge and insufficient interaction intelligence in virtual simulation teaching, and improves the teaching effect of ideological and political education.
[0015] The above description of the invention content is only an overview of the technical solution of this application. In order to enable those of ordinary skill in the art to more clearly understand the technical solution of this application, and then can be implemented according to the content recorded in the description and the drawings, and in order to make the above objects, other objects, features and advantages of this application more easily understood, the following is described in conjunction with the specific implementation manners and drawings of this application. Brief Description of the Drawings
[0016] The drawings are only used to show the principles, implementation methods, applications, features and effects of the specific implementation manners of the present invention and other related contents, and should not be considered as a limitation to this application.
[0017] In the accompanying drawings of the specification: Figure 1 It is a method step diagram of steps S101 to S106 of the online teaching interaction method described in the specific implementation manner; Figure 2 It is a method step diagram of steps S201 to S203 of the online teaching interaction method described in the specific implementation manner; Figure 3 It is a method step diagram of steps S301 to S306 of the online teaching interaction method described in the specific implementation manner; Figure 4 It is a method step diagram of steps S401 to S405 of the online teaching interaction method described in the specific implementation manner; Figure 5 It is a structural schematic diagram of the electronic device described in the specific implementation manner.
[0018] The reference numerals involved in the above-mentioned drawings are explained as follows: 1. Electronic device; 11. Memory; 12. Processor. Detailed implementation manners
[0019] To describe in detail the possible application scenarios, technical principles, specific implementable solutions, achievable purposes and effects of the present application, the following will be described in detail with reference to the specific examples listed and in conjunction with the accompanying drawings. The embodiments described herein are only used to more clearly illustrate the technical solutions of the present application, so they are only examples and cannot be used to limit the protection scope of the present application.
[0020] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The term "embodiment" appearing in various positions in the specification does not necessarily refer to the same embodiment, nor does it particularly limit its independence or relevance to other embodiments. In principle, in the present application, as long as there is no technical contradiction or conflict, the technical features mentioned in each embodiment can be combined in any way to form corresponding implementable technical solutions.
[0021] Unless otherwise defined, the meanings of the technical terms used herein are the same as those commonly understood by those skilled in the technical field to which the present application belongs; the use of the relevant terms herein is only for describing specific embodiments and is not intended to limit the present application.
[0022] In the description of the present application, the phrase "and / or" is an expression used to describe the logical relationship between objects, indicating that there can be three relationships. For example, A and / or B means: there is A, there is B, and there is both A and B at the same time. In addition, the character " / " herein generally represents an "or" logical relationship between the associated objects before and after.
[0023] In the present application, terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual quantity, primary or secondary, or order relationship between these entities or operations.
[0024] Without more limitations, in the present application, the open expressions such as "including", "comprising", "having" or other similar expressions used in the statement are intended to cover non-exclusive inclusion. These expressions do not exclude that there may be other elements in the process, method or product including the said elements, so that the process, method or product including a series of elements may not only include those defined elements, but also include other elements not explicitly listed, or also include elements inherent in this process, method or product.
[0025] Similar to the understanding in the "Examination Guidelines", in this application, expressions such as "greater than", "less than", and "exceeding" are understood to exclude the base number; expressions such as "above", "below", and "within" are understood to include the base number. In addition, in the description of the embodiments of this application, the meaning of "multiple" is two or more (including two), and similar expressions related to "many" are also understood in this way, such as "multiple groups", "multiple times", etc., unless otherwise specifically defined.
[0026] Please refer to Figure 1 , in the first aspect, this embodiment provides an online teaching interaction method based on a multimodal knowledge graph, including: S101. Collect the original teaching information, perform cross-modal semantic alignment processing on the original teaching information to obtain structured teaching information; S102. Construct a multimodal knowledge graph, and deeply fuse the text concept entities, video key frame feature vectors, and speech text transcripts through a graph neural network to obtain cross-modal semantic association edges, where the cross-modal semantic association edges include concept-visual feature association edges, concept-speech segment association edges, and cross-modal similarity association edges; S103. Receive user interaction behavior data, where the user interaction behavior data includes speech question content, virtual teaching operation sequences, and interface click trajectories, and convert the user interaction behavior data into a knowledge graph query vector; S104. Extract a three-dimensional teaching scenario sub-graph from the multimodal knowledge graph according to the knowledge graph query vector. When a standard teaching operation is detected, activate the forward reasoning path of the three-dimensional teaching scenario sub-graph to generate progressive teaching content. When an incorrect teaching operation is detected, trigger the counterfactual reasoning mechanism of the three-dimensional teaching scenario sub-graph to generate comparative teaching content; S105. Real-time record the interaction behavior data of the user in the three-dimensional teaching scenario sub-graph, denoted as scenario interaction data, match the mapping relationship between the scenario interaction data and the multimodal knowledge graph, and construct a cognitive state tracking matrix; S106. Dynamically adjust the teaching strategy according to the cognitive state tracking matrix, strengthen the cross-modal display combination of low-mastery concept nodes, and extend the high-proficiency operation sequence-related high-order knowledge; And, aggregate the cognitive state tracking matrix according to the group behavior analysis model, adjust the teaching path recommendation weights in the multimodal knowledge graph, and generate a personalized learning navigation graph, where the learning navigation graph includes an optimal learning path plan and a multimodal resource recommendation scheme.
[0027] In step S101, the original teaching information includes text textbooks, teaching videos, experimental operation videos, and classroom voice explanations; preferably, cross-modal semantic alignment processing uses a Transformer-based multi-modal encoder to map heterogeneous teaching resources to a unified semantic representation space to ensure the consistency of different modal information at the semantic level; the structured teaching information includes text concept entities, video key-frame feature vectors, and voice text transcription content. This step realizes the standardized processing of heterogeneous data through a cross-modal feature extraction network. Preferably, text concept entities are extracted using named entity recognition technology, video key-frame feature vectors are extracted through a convolutional neural network, and voice text transcription content is obtained through conversion by a speech recognition model.
[0028] In step S102, the constructed multi-modal knowledge graph uses a graph attention network to achieve feature fusion. The concept-visual feature association edge contains spatial attention weights, and the concept-speech segment association edge contains temporal alignment information. The spatial attention weights are calculated through a visual-text alignment network, reflecting the correspondence between text concepts and video regions; the temporal alignment information is established through a dynamic time warping algorithm to ensure the time synchronization of speech segments and text concepts. The cross-modal similarity association edge is constructed through a contrastive learning framework to quantify the semantic consistency between different modal features.
[0029] In step S103, the user interaction behavior data is converted into a knowledge graph query vector through a multi-modal encoder. Preferably, the voice question content is encoded using a pre-trained language model, the virtual teaching operation sequence extracts features through a spatio-temporal convolutional network, and the interface click trajectory is converted into a spatial distribution vector of knowledge nodes. The fusion query vector of the three interaction features of voice question content, virtual teaching operation sequence, and interface click trajectory provides a retrieval basis for subsequent situation sub-graph extraction.
[0030] In step S104, the three-dimensional teaching situation sub-graph includes concept nodes, association edge weights, and cross-modal resource links. The sub-graph extraction process preferably uses an attention sampling algorithm based on the query vector to dynamically construct a knowledge sub-network related to the current teaching situation. Preferably, the counterfactual reasoning mechanism of the three-dimensional teaching situation sub-graph generates a contrastive situation that violates teaching rules through the conditional masking of knowledge graph edges. The counterfactual reasoning mechanism is realized through a trainable edge mask matrix, which can selectively mask specific semantic relationships to construct teaching counterexamples. Preferably, the counterfactual reasoning mechanism realizes path redirection through a graph convolutional network, specifically including: constructing a trainable edge mask matrix, when an abnormal node is detected, calculating the path blocking probability based on the graph attention mechanism, and dynamically masking the standard teaching path edges; at the same time, constructing a contrastive teaching segment generator through a pre-trained adversarial generation network, and performing adversarial fusion on the wrong operation feature vector and the correct demonstration feature vector to generate a multi-modal contrast case containing difference annotations.
[0031] In step S105, the cognitive state tracking matrix includes a proficiency rating, an operation proficiency index, and an error pattern classification. The matrix element values are jointly determined by the access depth of the knowledge nodes and the activation frequency of the associated edges. Preferably, the access depth is calculated based on the position of the node in the interaction path, and the activation frequency is the count of the triggering times of the associated edges within a unit time. Further, the proficiency rating is calculated based on the residence time of the knowledge nodes and the access path depth; the operation proficiency index is evaluated based on the completion time of the operation sequence and the matching degree with the standard process; and the error pattern classification is achieved by comparing the difference patterns between the user operations and the standard operations.
[0032] In step S106, the teaching strategy adjustment module implements differentiated teaching according to the cognitive state tracking matrix. For the concept nodes with low proficiency, it enhances the visual prominence and voice explanations in their cross-modal display combinations; for the high-proficiency operation sequences, it extends the reasoning paths associated with the higher-order knowledge nodes. Preferably, the generation of the personalized learning navigation map uses a collaborative filtering algorithm based on swarm intelligence to optimize the recommended path by integrating individual cognitive characteristics and group learning patterns.
[0033] This embodiment realizes the deep fusion of text concept entities, video key-frame feature vectors, and speech text transcripts through the cross-modal semantic alignment processing of structured teaching information. The constructed multi-modal knowledge graph establishes a cross-modal semantic association network through concept-visual feature associated edges, concept-speech segment associated edges, and cross-modal similarity associated edges. By converting the speech question content, virtual teaching operation sequences, and interface click trajectories into knowledge graph query vectors, it realizes the accurate mapping of user interaction behavior data and the knowledge graph. Through the forward reasoning path and counterfactual reasoning mechanism of the three-dimensional teaching scenario sub-graph, the system can dynamically generate progressive teaching content and comparative teaching content. The construction of the cognitive state tracking matrix is based on the mapping relationship between the scenario interaction data and the multi-modal knowledge graph, accurately reflecting the user's proficiency rating, operation proficiency index, and error pattern classification. The differentiated teaching strategy adjustment implemented according to the cognitive state tracking matrix effectively strengthens the cross-modal display combinations of the low-proficiency concept nodes and extends the higher-order knowledge associated with the high-proficiency operation sequences. The learning navigation map generated by combining the group behavior analysis model realizes the personalized adaptation of the teaching process through the optimal learning path planning and multi-modal resource recommendation scheme. This embodiment forms a complete closed-loop from teaching resource organization, learning state assessment to teaching strategy optimization through the construction and application of the multi-modal knowledge graph, significantly improving the intelligent level and teaching effect of the online teaching system.
[0034] In some embodiments, the original teaching information includes textbook text information, teaching video information, and classroom voice information; Collect the original teaching information, perform cross-modal semantic alignment processing on the original teaching information to obtain structured teaching information, including: Obtain the textbook text information, extract teaching concept entities from it through named entity recognition, and generate concept feature vectors by performing word embedding encoding on each teaching concept entity; In addition, obtain the teaching video information, perform key frame sampling on the video stream to obtain a video key frame sequence, and the video key frame sequence includes multiple first video key frames; In addition, obtain the classroom voice information, perform segmentation processing on the voice stream to obtain a voice segment set, and the voice segment set includes multiple first voice segments; Build a cross-modal semantic alignment model, including: Calculate the first attention weights between the concept feature vectors and the visual feature vectors of all video key frames, and calculate the second attention weights between the concept feature vectors and the acoustic feature vectors of all voice segments; Select the video key frames and voice segments strongly associated with the teaching concept entities according to the first attention weights and the second attention weights, and record them as the second video key frames and the second voice segments; Generate structured teaching information according to the cross-modal semantic alignment model. The structured teaching information includes teaching concept entities and concept feature vectors, the second video key frames and the first attention weights, the second voice segments and the second attention weights, the visual association relationship between the teaching concept entities and the second video key frames, and the voice association relationship between the teaching concept entities and the second voice segments.
[0035] In this embodiment, the textbook text information extracts teaching concept entities through named entity recognition technology. Preferably, each entity generates concept feature vectors through word embedding encoding by a pre-trained language model, and the concept feature vectors represent the distribution characteristics of teaching concepts in the semantic space.
[0036] During the processing of the teaching video information, performing key frame sampling on the video stream to obtain a video key frame sequence, including: Extract visual feature vectors through a pre-trained convolutional neural network; Detect and recognize the text annotation information in the picture; Establish the temporal correspondence relationship between the video key frames and the teaching concept entities.
[0037] When processing the classroom voice information, performing segmentation processing on the voice stream to obtain a voice segment set, including: Convert it into voice text content through a voice recognition engine; Extract the acoustic feature vectors of the voice segments; Establish the semantic correspondence relationship between the voice segments and the teaching concept entities.
[0038] The construction process of the cross-modal semantic alignment model realizes the association of multi-modal features through the attention mechanism. The first attention weight calculates the correlation between the conceptual feature vector and the visual feature vector of the video key frame, reflecting the matching degree between the teaching concept and the visual content; the second attention weight measures the association strength between the conceptual feature vector and the acoustic feature vector of the speech segment, representing the semantic consistency between the speech content and the teaching concept. The selected second video key frame and the second speech segment are both strongly associated with the teaching concept entity, ensuring the accuracy of cross-modal alignment.
[0039] The generation process of structured teaching information establishes a systematic association between the teaching concept entity and multi-modal resources. The visual association relationship quantifies the matching degree between the teaching concept and the video key frame through the first attention weight; the speech association relationship reflects the corresponding strength between the teaching concept and the speech segment through the second attention weight. The visual association relationship and the speech association relationship together constitute the semantic network foundation of cross-modal teaching knowledge.
[0040] In this embodiment, through systematic cross-modal alignment processing, the originally discrete teaching resources are transformed into a structured information system with clear semantic associations. When processing a certain teaching concept, the system automatically associates the key frames of relevant scenes in the video and the speech segments explaining the concept, forming a multi-dimensional combination of teaching resources. It not only retains the richness of the original teaching resources, but also establishes a systematic association between knowledge elements through semantic alignment, providing a standardized data basis for the subsequent construction of a multi-modal knowledge graph. This embodiment realizes the transformation process from the original teaching resources to a structured knowledge system, ensuring a high degree of consistency at the semantic level of different modal teaching elements.
[0041] Please refer to Figure 2 , in some embodiments, a multi-modal knowledge graph is constructed. Through a graph neural network, the text concept entity, the video key frame feature vector, and the speech text transcription content are deeply fused to obtain cross-modal semantic association edges, including: S201. Construct an initial node set of the multi-modal knowledge graph. The initial node set includes: Taking the teaching concept entity as the concept node, and the node feature as the conceptual feature vector; Taking the second video key frame as the visual node, and the node feature as the visual feature vector; Taking the second speech segment as the speech node, and the node feature as the speech feature vector; S202. Establish cross-modal semantic association edges, including: According to the first attention weight value of the concept node, establish a concept-visual association edge with the visual node, and the weight value of the concept-visual association edge is the first attention weight after normalization processing; According to the second attention weight value of the concept node, establish a concept-speech association edge with the speech node, and the weight value of the concept-speech association edge is the second attention weight after normalization; For visual nodes and speech nodes with a temporal correspondence relationship, establish a visual-speech temporal edge, and the weight value of the visual-speech temporal edge is calculated based on the time overlap degree between the second video key frame and the second speech segment; S203. Perform deep multi-modal feature fusion through a graph neural network, including: Input the initial node set into the graph attention network; Perform feature propagation along the concept-visual association edge and the concept-speech association edge, and calculate the cross-modal similarity score between nodes through the multi-head attention mechanism; Update the cross-modal feature representation of the initial node set according to the cross-modal similarity score to generate knowledge graph nodes with fused features; Obtain the final multi-modal knowledge graph.
[0042] In step S201, constructing the initial node set of the multi-modal knowledge graph means converting each element in the structured teaching information into graph structure data. Among them, the concept node is based on the teaching concept entity, and its concept feature vector is obtained by encoding through a pre-trained language model, representing the semantic features of the teaching concept; the visual node corresponds to the second video key frame, and its visual feature vector is extracted through a convolutional neural network, containing the visual semantic information of the video content; the speech node comes from the second speech segment, and the speech feature vector is extracted through an acoustic model, reflecting the acoustic features and semantic content of the speech content. The above initial nodes provide basic data units for subsequent cross-modal feature fusion.
[0043] In step S202, establishing the cross-modal semantic association edge is to construct the topological structure of the multi-modal knowledge graph through attention weights and temporal relationships. Among them, the weight value of the concept-visual association edge is determined by the first attention weight after normalization, reflecting the semantic correlation degree between the teaching concept and the video content; the weight value of the concept-speech association edge comes from the second attention weight after normalization, representing the matching strength between the teaching concept and the speech explanation. The establishment of the visual-speech temporal edge is based on the principle of multimedia synchronization, and its weight value is obtained by calculating the time overlap degree between the second video key frame and the second speech segment to ensure the consistency of the audiovisual content in the time dimension.
[0044] In step S203, the deep multi-modal feature fusion through the graph neural network refers to the realization of cross-modal knowledge representation learning by leveraging the characteristics of the graph attention network. Among them, the feature propagation process is carried out along the concept-visual association edges and concept-audio association edges, and the cross-modal similarity scores between nodes are calculated through the multi-head attention mechanism. Preferably, the cross-modal similarity scores comprehensively consider semantic relevance and modal complementarity; the feature update of the knowledge graph nodes adopts a gating mechanism to dynamically adjust the fusion ratio of each modal feature. Finally, the generated fusion features not only retain the original modal characteristics but also contain cross-modal interaction information.
[0045] The finally obtained multi-modal knowledge graph includes concept nodes, visual nodes, and audio nodes with fusion features, concept-visual association edges and their weight values, concept-audio association edges and their weight values, and visual-audio temporal edges and their weight values, presenting the cross-modal association basis for users.
[0046] This embodiment realizes the multi-modal structured representation and deep fusion of teaching knowledge. When processing a certain concept node, the system fuses its features with the key frames of the relevant theoretical explanation video (visual nodes) and voice commentary segments (audio nodes) through the graph attention network, generating a unified knowledge representation that includes text concepts, visual examples, and voice explanations. It not only establishes explicit associations between teaching elements but also realizes complementary enhancement of knowledge representation through deep fusion at the feature level, providing a rich semantic basis for subsequent contextual teaching interactions and enabling the system to understand and present teaching content from different perspectives.
[0047] Please refer to Figure 3 , in some embodiments, converting user interaction behavior data into a knowledge graph query vector includes: S301. Converting the voice question content into query text through a voice recognition engine; S302. Using named entity recognition to extract the core concept entities in the query text; S303. Mapping the core concept entities to the concept nodes in the multi-modal knowledge graph, and using the feature vectors of the successfully mapped concept nodes as the text query base vectors; And, parsing the operation objects and operation actions in the virtual teaching operation sequence to obtain the parsing result; S304. Locating the corresponding teaching process nodes in the multi-modal knowledge graph according to the parsing result, and associating the feature vectors of the teaching process nodes as the operation query base vectors; And, identifying the feature vectors of the knowledge graph nodes associated with the click areas of the interface click trajectories, denoted as click feature vectors, and generating a heat weight according to the click frequency; S305. Generating a click query base vector according to the heat weight and the click feature vectors; S306. Perform weighted fusion on the text query base vector, operation query base vector, and click query base vector to obtain a fused feature vector, and input the fused feature vector into the query encoder to generate the final knowledge graph query vector.
[0048] In step S301, the voice question content is converted into a query text through a speech recognition engine, and the user's voice interaction input is converted into a processable text form. The speech recognition engine preferably adopts an end-to-end deep learning model, which can adapt to the professional terms and diverse pronunciations in the teaching scenario. The converted query text retains the semantic integrity of the original voice question and provides an input basis for subsequent concept entity extraction.
[0049] In step S302, named entity recognition is used to extract the core concept entities in the query text, and a pre-trained language model is used to identify the key concepts related to teaching. The core concept entities correspond to the concept nodes in the multi-modal knowledge graph, ensuring that the user's question can be accurately mapped to the knowledge system. The named entity recognition model is optimized for the teaching field and can accurately identify the professional terms and core concepts in ideological and political education.
[0050] In step S303, the core concept entities are mapped to the concept nodes in the multi-modal knowledge graph through semantic similarity calculation. The feature vectors of the successfully mapped concept nodes are used as the text query base vectors, which contain the deep semantic features of teaching concepts. At the same time, the operation object and operation action in the virtual teaching operation sequence are parsed through an action semantic analysis model, and the parsing result reflects the user's teaching behavior intention in the virtual environment.
[0051] In step S304, locating the corresponding teaching process node in the multi-modal knowledge graph according to the parsing result is achieved through behavior-knowledge association rules. The feature vectors of the associated teaching process nodes are used as the operation query base vectors, representing the knowledge content corresponding to the teaching operation. The feature vectors of the knowledge graph nodes associated with the click area of the click trajectory on the recognition interface are completed through a spatial mapping algorithm, and the heat weight generated by the click frequency reflects the user's interest preference.
[0052] In step S305, generating the click query base vector according to the heat weight and click feature vector is achieved through a weighted aggregation algorithm. The heat weight is normalized to ensure the balanced contribution of different interaction behaviors. The click query base vector comprehensively reflects the user's explicit interest orientation.
[0053] In step S306, the weighted fusion of the text query base vector, the operation query base vector, and the click query base vector is achieved through an attention mechanism. The fused feature vector is input into the query encoder to generate the final knowledge graph query vector. Among them, the weight of the text query base vector is the confidence score of the voice question; the weight of the operation query base vector is the normality score of the operation step; the weight of the click query base vector is the normalized popularity weight. The query encoder adopts a multi-layer perceptron structure and can learn the deep associations of different interaction features.
[0054] This embodiment realizes the intelligent conversion from multi-modal interaction behavior to knowledge graph query. When the user asks a question by voice and clicks on the relevant case area, the system converts the voice content into a text query base vector, combines it with the click query base vector generated by the click area, and finally forms a knowledge graph query vector that comprehensively reflects the user's intention. This embodiment fully considers the semantic contribution degrees of different interaction behaviors, ensuring that the query results not only meet the explicit needs of users but also can explore potential learning interests. By building an intelligent bridge from user behavior to knowledge retrieval, this embodiment provides an accurate query basis for personalized teaching.
[0055] Please refer to Figure 4 , in some embodiments, extracting a three-dimensional teaching scenario sub-graph from the multi-modal knowledge graph according to the knowledge graph query vector includes: S401. Perform node matching in the multi-modal knowledge graph based on the knowledge graph query vector to locate the initial query node; S402. Perform graph traversal along the concept-visual association edge and the concept-voice association edge to extract the associated visual query node and voice query node; S403. Establish a spatio-temporal association relationship according to the visual-voice temporal edge to form a three-dimensional teaching scenario sub-graph; S404. When a standard teaching operation is detected, activate the forward reasoning path of the three-dimensional teaching scenario sub-graph to generate progressive teaching content, including: Start from the initial query node and perform reasoning along the standard teaching path in the knowledge graph; Combine the associated concept query nodes, visual query nodes, and voice query nodes in the teaching logic order; Generate a progressive teaching sequence including basic concept explanations, example demonstrations, and voice guidance; S405. When an incorrect teaching operation is detected, trigger the counterfactual reasoning mechanism of the three-dimensional teaching scenario sub-graph to generate comparative teaching content, including: Identify the abnormal node in the knowledge graph corresponding to the incorrect operation; Disable the standard teaching path edge directly associated with the abnormal node; Activate the counterfactual reasoning edge to generate teaching content that includes error analysis, correct demonstrations, and comparative explanations.
[0056] In step S401, node matching is performed in the multi-modal knowledge graph based on the knowledge graph query vector. The initial query node that is most relevant to the query intent is located by calculating the vector similarity. The initial query node serves as the starting point for extracting the situational subgraph, and its matching accuracy directly affects the accuracy of the subsequent teaching content. Preferably, the approximate nearest neighbor search algorithm is used in the node matching process to ensure the reliability of the query results while guaranteeing real-time performance.
[0057] In step S402, preferably, graph traversal along the concept-visual association edge and the concept-audio association edge is achieved through the breadth-first search algorithm. The extracted visual query nodes may include visual examples related to the teaching concept, and the audio query nodes provide corresponding audio explanations. During the traversal process, path filtering is performed according to the association edge weight values to ensure that the extracted nodes have sufficient semantic relevance.
[0058] In step S403, preferably, establishing the spatio-temporal association relationship based on the visual-audio temporal edge is completed through the time alignment algorithm. The spatio-temporal association relationship ensures that the video demonstration is synchronized with the audio explanation, forming a three-dimensional teaching situational subgraph with spatio-temporal consistency. The establishment of the spatio-temporal association relationship takes into account the logical coherence and time continuity requirements of the teaching content.
[0059] In step S404, preferably, activating the forward reasoning path of the three-dimensional teaching situational subgraph to generate progressive teaching content is achieved through the teaching logic engine. The standard teaching path is predefined by domain experts and reflects the progressive relationship of knowledge points. Optionally, the generation of the progressive teaching sequence follows the teaching mode of "concept explanation - example demonstration - practice guidance" to ensure the systematicness and coherence of the learning process.
[0060] In step S405, the abnormal nodes in the knowledge graph identify the knowledge blind spots involved in the wrong operations. Disabling the standard teaching path edges can prevent the spread of wrong demonstrations. Optionally, a graph anomaly detection model is used to identify the abnormal nodes corresponding to the wrong operations, and the abnormal paths are verified by combining temporal analysis. The activation of the counterfactual reasoning edge is based on the teaching rule base, and the generated comparative teaching content highlights the differences between wrong and correct operations, strengthening the cognitive correction effect.
[0061] In this embodiment, the initial query node is accurately located through the query vector of the knowledge graph, and the associated nodes are extracted along the concept-vision association edge and the concept-speech association edge, and a three-dimensional teaching scenario subgraph with spatio-temporal consistency is constructed based on the vision-speech temporal edge. When a standard teaching operation is detected, the system generates a progressive teaching sequence including concept explanation, example demonstration, and voice guidance along the preset teaching path; when an incorrect operation is recognized, contrastive teaching content is generated through the counterfactual reasoning mechanism. This embodiment realizes the intelligent organization and dynamic generation of teaching content, ensuring the systematicness of standard teaching and the pertinence of error correction. By precisely matching the user's query intention, maintaining the spatio-temporal association of multimodal teaching resources, and supporting positive and negative teaching reasoning, this embodiment significantly improves the accuracy and adaptability of online teaching, not only ensuring the systematic integrity of knowledge transfer but also effectively intervening in learning misunderstandings to form a closed-loop optimized intelligent teaching mechanism.
[0062] In some embodiments, the situational interaction data includes the access frequency and residence duration of the concept query node, the viewing completeness of the visual query node, the repeated playback times of the voice query node, and the selection ratio of the standard teaching path and the counterfactual reasoning path; Match the mapping relationship between the situational interaction data and the multimodal knowledge graph to construct a cognitive state tracking matrix, including: Establish an initial cognitive state tracking matrix, where the rows of the initial cognitive state tracking matrix represent the concept nodes in the multimodal knowledge graph, and the columns include the proficiency index, operation proficiency index, error pattern index, and attention distribution index; Map the access records of the concept query nodes to the concept nodes, calculate the average test correct rate and access duration score of the concept nodes to obtain the proficiency index; Map the viewing records of the visual query nodes to the visual nodes, and count the operation success rate and average time consumption of the teaching process nodes to obtain the operation proficiency index; Map the teaching path selection records to the standard teaching path or the counterfactual reasoning path, and record the trigger times of the counterfactual reasoning path and the subsequent operation improvement degree to obtain the error pattern index; Map the playback records of the voice query nodes to the voice nodes, and obtain the attention distribution index according to the interaction event frequency of the voice nodes, concept nodes, and visual nodes within a unit time; Update the proficiency index, operation proficiency index, error pattern index, and attention distribution index to the columns of the initial cognitive state tracking matrix to obtain the updated cognitive state tracking matrix.
[0063] In this embodiment, the situational interaction data refers to the behavior records generated during the interaction between the user and the three-dimensional teaching situational sub-graph, including key indicators such as the access frequency and residence duration of concept query nodes, the viewing completeness of visual query nodes, and the repeated playback times of voice query nodes. The situational interaction data is collected by multi-modal sensors and reflects the user's cognitive processing process of different teaching elements. The selection ratio of the standard teaching path and the counterfactual reasoning path characterizes the user's learning strategy preference and is an important basis for evaluating teaching effectiveness.
[0064] The construction process of the cognitive state tracking matrix quantifies the user's learning state through multi-dimensional indicators. The cognitive state tracking matrix includes proficiency scores, operation proficiency indicators, and error mode classifications, and the matrix element values are jointly determined by the access depth of knowledge nodes and the activation frequency of associated edges. The row and column structures of the initial cognitive state tracking matrix ensure that various evaluation indicators can be systematically organized. Among them, the proficiency is based on the test accuracy rate and access depth of concept query nodes; the operation proficiency is based on the operation accuracy and completion speed of teaching process nodes; the error mode is based on the trigger frequency and correction effect of the counterfactual reasoning path; the attention distribution is based on the interaction duration of each query node and the hot zone click data. Through structured design, a multi-angle characterization of the cognitive state is achieved.
[0065] Preferably, the mapping statistical method is adopted in the index calculation process to realize data conversion. The average test accuracy rate and access duration score of concept nodes are calculated by the weighted algorithm to reflect the knowledge mastery degree; the operation success rate statistics of teaching process nodes consider the normalization processing of task complexity; the trigger times record of the counterfactual reasoning path combines the time decay factor to highlight the influence of recent behaviors; the attention distribution index is obtained through the spatio-temporal clustering analysis of interaction events to identify the user's focus of attention.
[0066] This embodiment realizes the precise evaluation of the learning process through systematic cognitive state modeling. For example, when a student shows a high access frequency but a low test accuracy rate at a certain concept node, the system determines that the student is superficially familiar but lacks in-depth understanding; when a certain counterfactual reasoning path is frequently triggered, it is recognized that there is a cognitive misunderstanding in this method. It not only captures explicit behavior characteristics but also can dig out deep cognitive states, providing a reliable basis for the adjustment of personalized teaching strategies. This embodiment constructs a complete analysis chain from interaction behaviors to cognitive evaluation to support the intelligent decision-making of the teaching system.
[0067] In some embodiments, the teaching strategy is dynamically adjusted according to the cognitive state tracking matrix, strengthening the cross-modal display combination of concept nodes with low mastery, and extending the high-order knowledge associated with high-proficiency operation sequences, including: Extracting a set of low-mastery concept nodes whose proficiency indicators in the cognitive state tracking matrix are lower than the first preset threshold; Perform a cross-modal reinforcement step for each low-mastery concept node in the set of low-mastery concept nodes: Increase the display duration and frequency of visual nodes associated with low-mastery concept nodes; Improve the playback priority of speech nodes associated with low-mastery concept nodes; Produce intensive instructional packages that include textual explanations, visual demonstrations, and audio guidance; Extracting a set of high-proficiency concept nodes whose operation proficiency index is higher than a second preset threshold from the cognitive state tracking matrix; Perform the high-order knowledge extension step for each high-proficiency concept node in the high-proficiency concept node set: Extracting high-order visual nodes along concept-visual association edges; Extracting higher-order phonetic nodes along concept-phonetic association edges; Generate a high-level teaching package that includes advanced knowledge explanations, in-depth case analysis, and extended exercises; According to the error mode indicators in the cognitive state tracking matrix, the triggering conditions of the counterfactual reasoning path are adjusted, including: Add early warning nodes for high-frequency error modes; Optimize the comparison display method between incorrect operations and correct demonstrations; According to the attention distribution indicators in the cognitive state tracking matrix, the teaching rhythm is dynamically adjusted, including: Simplify teaching content for low attention areas; Add interactive links to high-attention areas; The existing teaching strategy configuration is updated to obtain and execute an updated teaching configuration, which includes a reinforcement teaching plan for low-mastery concept nodes, a high-order extension plan for high-proficiency concept nodes, optimization parameters for counterfactual reasoning paths, and adjustment parameters for teaching rhythm.
[0068] In this embodiment, the dynamic adjustment of the teaching strategy is based on the in-depth analysis of the cognitive state tracking matrix. The low-mastery concept node set refers to the knowledge nodes whose proficiency index is lower than the first preset threshold, and the first preset threshold is determined by analyzing the mastery distribution of typical learning difficulties in historical teaching data. The cross-modal reinforcement step constructs a multi-dimensional reinforced teaching combination by enhancing the display time of visual nodes and improving the playback priority of voice nodes, so as to improve students' understanding and memory of low-mastery knowledge points in a targeted manner.
[0069] The high - order knowledge extension step is implemented for high - proficiency concept nodes whose operation proficiency index is higher than the second preset threshold, and the second preset threshold is set according to the skill mastery standard required by the teaching syllabus. The realization of high - order knowledge extension relies on the hierarchical topology characteristics of the knowledge graph, and multi - hop reasoning is carried out along the concept association edges to discover knowledge expansion paths that conform to the law of cognitive development. Specifically, high - order visual nodes and high - order speech nodes are obtained by expanding along the association edges of the knowledge graph, including deeper teaching content and more complex application cases. During the extension process, the core concept remains unchanged, only the depth of argumentation and the breadth of application are increased, ensuring the coherence of the knowledge system, and strengthening the teaching combination and high - order teaching together constitute a differentiated teaching plan to achieve individualized teaching.
[0070] The optimization of the counterfactual reasoning path focuses on improving the immediacy and pertinence of error correction. The added position of the warning prompt node is determined according to the position of the knowledge graph where the error occurs. Further, the implantation position of the warning prompt node is preferably the nearest common ancestor node in the knowledge graph to ensure the relevance between the prompt content and the root cause of the error. Preferably, the optimization of the comparison display method adopts a dual - channel presentation technology to maintain the integrity of the standard teaching content while highlighting the key differences.
[0071] The dynamic adjustment of the teaching rhythm is based on the attention distribution index, that is, the dynamic adjustment of the teaching rhythm is based on the trend analysis of the attention distribution pattern and is realized through the coordinated adjustment of content density and interaction frequency for progressive optimization. Specifically, a content simplification strategy is adopted for low - attention areas. Preferably, the concept decomposition and example - focusing strategies are adopted in low - attention areas; for high - attention areas, more interactive designs are added. Preferably, cognitive challenges and transfer training are implemented in high - attention areas. Further, the above - mentioned method of dynamically adjusting the teaching rhythm can be realized through the content presentation engine.
[0072] In this embodiment, the dynamic optimization of teaching strategies is driven by the cognitive state tracking matrix, realizing precise individualized teaching. Cross - modal reinforcement is implemented for low - mastery concept nodes, and a multi - dimensional teaching combination is constructed by increasing the display duration of associated visual nodes, raising the playback priority of speech nodes, etc.; for high - proficiency concept nodes, high - order knowledge is extended, and in - depth teaching content is extracted along the association edges of the knowledge graph; at the same time, the counterfactual reasoning path and the teaching rhythm are optimized. Through the differential adjustment mechanism based on the cognitive state, not only the solid mastery of basic knowledge is ensured, but also the continuous development of high - order capabilities is supported. This embodiment continuously optimizes the counterfactual reasoning path and the teaching rhythm, forming a progressive teaching closed - loop from assessment to low - mastery concepts, then to high - order knowledge extension, and further optimizing incorrect operations, ensuring that the teaching strategy always maintains the best match with the learner's cognitive state.
[0073] In some embodiments, the cognitive state tracking matrix is aggregated according to the group behavior analysis model, and the teaching path recommendation weight in the multimodal knowledge graph is adjusted to generate a personalized learning navigation graph, including: Obtain the cognitive state tracking matrix of multiple users, recorded as the group cognitive state tracking matrix, the group cognitive state tracking matrix includes group concept nodes, calculate the average mastery index of each group concept node and count the high-frequency error patterns and group-related nodes; Analyze the common features of the optimal learning path and identify the hot spots of group attention distribution; Calculate the average mastery index, high-frequency error pattern, and optimal learning path characteristics of the group cognitive state tracking matrix, and adjust the teaching path recommendation weight in the multimodal knowledge graph based on the group attention distribution hotspots, including: The weight of basic teaching is reduced for concept nodes with high mastery of the group, and the weight of counterfactual reasoning is increased for nodes with high frequency of group errors; Improve the recommendation priority of the associated edges in the group's optimal path, and optimize the resource display weight for the group's attention hotspot areas; According to the adjusted teaching path recommendation weight, the subgraph structure related to the current user is extracted from the multimodal knowledge graph; After recalculating the path score of the teaching path, mark the key concept nodes on the recommended learning path and mark the error-prone node areas that need to be focused on; Generate a learning navigation map that includes personalized recommended main learning paths and alternative paths, recommended concept nodes, recommended visual nodes, recommended voice node combinations, recommended path weight scores and estimated learning time, and key reinforcement area prompt marks.
[0074] In this embodiment, the group cognitive state tracking matrix is a group learning feature database formed by aggregating the individual cognitive state tracking matrices of multiple users, and extracts the learning rules at the group level through statistical analysis methods, among which the average mastery index of the group concept node reflects the general difficulty of the knowledge point, the high-frequency error pattern reveals the common cognitive misunderstanding, and the group association node represents the internal connection between the knowledge points. The above group feature data provides an objective basis for the optimization of the teaching path.
[0075] The analysis of the common features of the optimal learning path is achieved by mining the data of the cognitive state tracking matrix of users with high learning effectiveness. Preferably, the common access sequences of high-effectiveness learners in the knowledge graph are extracted through data mining techniques. These sequences usually show an optimized combination pattern of specific concept nodes, visual nodes, and speech nodes. Combining with path evaluation indicators (such as the learning efficiency improvement rate, knowledge transfer effect, etc.), the node distribution characteristics of high-quality paths with universality are screened out. Combining with the interaction heat map data of all users, the area modules in the knowledge graph that continuously attract high attention are located, which are the hot areas of group attention distribution. The finally identified hot areas of group attention distribution usually correspond to the core concept groups or typical application scenarios in the knowledge system, providing an important basis for the subsequent optimization of the allocation of teaching resources.
[0076] The adjustment process of the teaching path recommendation weight comprehensively considers multiple factors such as group mastery, error patterns, and attention distribution, ensuring that the recommendation strategy not only conforms to the general learning rules but also can be optimized for common problems.
[0077] The adjustment process of the teaching path recommendation weight adopts a comprehensive analysis mechanism of multi-dimensional group learning characteristics. Specifically, the knowledge difficulty gradient is identified by calculating the group average mastery index, and the basic teaching content of high-mastery nodes is appropriately streamlined; the cognitive weak links are determined by analyzing the high-frequency error patterns, and the display intensity of the counterfactual reasoning path is enhanced accordingly. At the same time, based on the associated edge characteristics of the optimal learning path, the recommendation priority of the core knowledge link is improved, and the spatial layout weight of teaching resources is optimized according to the group attention hot area. By dynamically marking the key concept nodes and error-prone areas, a visual learning guidance plan with clear primary and secondary and prominent key points is formed.
[0078] The generation process of the personalized learning navigation graph adopts a hierarchical recommendation strategy. The main learning path is generated based on the adjusted recommendation weight, covering the core knowledge points preferentially; the alternative paths provide alternative solutions for different learning styles. The combination of recommended concept nodes, recommended visual nodes, and recommended speech nodes follows the multi-modal matching principle to ensure the coordinated presentation of various teaching resources. The recommended path weight score is calculated by comprehensively considering knowledge importance, learning difficulty, and user suitability, and the estimated learning duration is estimated according to the average time consumption in historical learning data. By adjusting the teaching path recommendation weight in the multi-modal knowledge graph in the above way, not only the topological structure of the knowledge graph is retained, but also the group learning behavior characteristics are incorporated, making the finally generated learning navigation graph able to intelligently balance the integrity of the knowledge system and the personalized needs of the learning path.
[0079] This embodiment realizes the organic combination of group wisdom and personalized learning. For example, when most students make frequent mistakes at a certain node, the system automatically increases the recommended weight of relevant counterfactual reasoning paths; for a certain concept node that the group has mastered well, the display intensity of basic teaching resources is appropriately reduced. By dynamically optimizing the recommended weights of teaching paths and marking key concept nodes and error-prone areas, the generated learning navigation map can not only reflect collective learning experience but also meet individual differentiated needs, effectively improving learning efficiency and quality.
[0080] In a second aspect, this embodiment also provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the method described in the first aspect.
[0081] The computer program involved in this embodiment can be stored in a computer-readable storage medium of a computer device. The computer-readable storage medium of the computer device includes but is not limited to magnetic disks, magnetic tapes, magnetic cards, floppy disks, flash memories, optical discs, optical cards, read-only memories (ROMs), random access memories (RAMs), erasable programmable ROMs (EPROMs), and electrically erasable programmable ROMs (EEPROMs), etc., and also includes other biological, physical, or chemical structures that can achieve the same or equivalent functions as the above-listed storage media, such as units with information storage capabilities like DNA, RNA, proteins, etc. In a specific embodiment, the storage medium involved can be one of the above medium types or a combination of the above medium types. In different embodiments, the computer program involved in the embodiment can be stored centrally in a single medium or distributedly in multiple media. The memory containing the computer-readable storage medium of the computer device can be a non-volatile memory or a random access memory. These computer-readable storage media of the computer device can be built into the device or can be an external device or a part of an external device connected to the device involved in the embodiment. In some embodiments, the memory with the computer-readable storage medium of the computer device is deployed locally; in other embodiments, a scheme of deploying the memory away from the processor can also be adopted, such as a network-attached memory accessed via an RF circuit or an external port and a communication network, where the communication network can be the Internet, one or more internal networks, local area networks (LANs), wide area wireless networks (WLANs), storage area networks (SANs), etc., or an appropriate combination thereof, as long as the computer device can access the memory. In addition, the computer program involved in the embodiment can be stored in plaintext / ciphertext form or can be designed as training data and be integrally recombined and implicitly stored in the parameter states of a deep neural network or other machine learning models through model training.
[0082] Please refer to Figure 5, in a third aspect, the present embodiment further provides an electronic device 1, including a memory 11 and a processor 12. The memory 11 is used to store one or more computer program instructions. Among them, the one or more computer program instructions are executed by the processor 12 to implement the method described in the first aspect.
[0083] The processor described in this embodiment can be implemented by hardware, firmware, software, or a combination thereof. It can use circuits, a single or multiple application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), central processing units (CPUs), controllers, microcontrollers, microprocessors, etc. It also includes other physical, biological, or chemical structures that can achieve functions similar to or equivalent to the above-listed processors, such as biological neurons, quantum computing units, DNA computing units, etc., so that the processor can execute some steps, all steps, or any combination of the steps mentioned in the computer programs or methods involved in the various embodiments of the present application.
[0084] Different from the prior art, the above technical solution has the following beneficial effects: By constructing a multi-modal knowledge graph to achieve the deep integration and intelligent organization of teaching resources, dynamically generating progressive and comparative teaching content based on the three-dimensional teaching scenario sub-graph, accurately evaluating the learning effect using the cognitive state tracking matrix and optimizing the teaching strategy, and finally forming a closed-loop adaptive intelligent teaching system. The method ensures the consistency of heterogeneous teaching resources through cross-modal semantic alignment processing, realizes the multi-modal fusion of knowledge representation with the help of graph neural networks, generates accurate knowledge graph query vectors according to user interaction behavior data, and optimizes the personalized learning navigation graph based on group behavior analysis, enabling the system to implement differentiated teaching for different cognitive states, strengthening the cross-modal display of low-mastery concept nodes, extending high-order knowledge for high-proficiency operation sequences, and continuously optimizing the counterfactual reasoning path and teaching rhythm to achieve a complete teaching closed-loop from knowledge presentation to learning assessment and then to strategy adjustment. The above technical solution effectively improves the accuracy, adaptability, and personalization level of online teaching, not only ensuring the systematic integrity of the knowledge system but also effectively intervening in learning misunderstandings, providing a reliable technical support for intelligent education.
[0085] Finally, it should be noted that although the above embodiments have been described in the text and drawings of the specification of this application, the patent protection scope of this application cannot be limited thereby. Any technical solutions obtained by equivalent structure or equivalent process substitution or modification based on the substantial concept of this application and using the content recorded in the text and drawings of the specification of this application, as well as directly or indirectly implementing the technical solutions of the above embodiments in other related technical fields, etc., are all included in the patent protection scope of this application.
Claims
1. An online teaching interaction method based on a multi-modal knowledge graph, characterized in that Including: Collecting original teaching information, performing cross-modal semantic alignment processing on the original teaching information to obtain structured teaching information; Constructing a multi-modal knowledge graph, and deeply fusing text concept entities, video key frame feature vectors, and speech text transcripts through a graph neural network to obtain cross-modal semantic association edges. The cross-modal semantic association edges include concept-visual feature association edges, concept-speech segment association edges, and cross-modal similarity association edges; Receiving user interaction behavior data, where the user interaction behavior data includes speech question content, virtual teaching operation sequences, and interface click trajectories, and converting the user interaction behavior data into a knowledge graph query vector; Extracting a three-dimensional teaching scenario sub-graph from the multi-modal knowledge graph according to the knowledge graph query vector. When a standard teaching operation is detected, activating the forward reasoning path of the three-dimensional teaching scenario sub-graph to generate progressive teaching content. When an incorrect teaching operation is detected, triggering the counterfactual reasoning mechanism of the three-dimensional teaching scenario sub-graph to generate comparative teaching content; Real-time recording the interaction behavior data of the user in the three-dimensional teaching scenario sub-graph, denoted as scenario interaction data, matching the mapping relationship between the scenario interaction data and the multi-modal knowledge graph, and constructing a cognitive state tracking matrix; Dynamically adjusting the teaching strategy according to the cognitive state tracking matrix, strengthening the cross-modal display combination of concept nodes with low mastery, and extending the high-order knowledge associated with high-proficiency operation sequences; And, aggregating the cognitive state tracking matrix according to the group behavior analysis model, adjusting the teaching path recommendation weights in the multi-modal knowledge graph, and generating a personalized learning navigation graph, where the learning navigation graph includes an optimal learning path planning and a multi-modal resource recommendation scheme.
2. The online teaching interaction method based on the multimodal knowledge graph according to claim 1, characterized in that The original teaching information includes textbook text information, teaching video information, and classroom voice information; Collecting original teaching information, performing cross-modal semantic alignment processing on the original teaching information to obtain structured teaching information, including: Obtaining textbook text information, extracting teaching concept entities from it through named entity recognition, and performing word embedding encoding on each teaching concept entity to generate concept feature vectors; And, obtaining teaching video information, performing key frame sampling on the video stream to obtain a video key frame sequence, where the video key frame sequence includes a plurality of first video key frames; And, obtaining classroom voice information, performing segmentation processing on the voice stream to obtain a voice segment set, where the voice segment set includes a plurality of first voice segments; Constructing a cross-modal semantic alignment model, including: Calculating the first attention weights between the concept feature vectors and all video key frame visual feature vectors, and calculating the second attention weights between the concept feature vectors and all voice segment acoustic feature vectors; Selecting, according to the first attention weights and the second attention weights, the video key frames and voice segments strongly associated with the teaching concept entities, denoted as second video key frames and second voice segments; Generate structured teaching information according to the cross-modal semantic alignment model, where the structured teaching information includes teaching concept entities and concept feature vectors, second video key frames and first attention weights, second speech segments and second attention weights, visual association relationships between teaching concept entities and second video key frames, and speech association relationships between teaching concept entities and second speech segments.
3. The online teaching interaction method based on the multi-modal knowledge graph according to claim 2, characterized in that, Construct a multi-modal knowledge graph, and deeply fuse text concept entities, video key frame feature vectors, and speech text transcripts through a graph neural network to obtain cross-modal semantic association edges, including: Construct an initial node set of the multi-modal knowledge graph, where the initial node set includes: Use teaching concept entities as concept nodes, and node features as concept feature vectors; Use second video key frames as visual nodes, and node features as visual feature vectors; Use second speech segments as speech nodes, and node features as speech feature vectors; Establish cross-modal semantic association edges, including: According to the first attention weight value of the concept node, establish a concept-visual association edge with the visual node, and the weight value of the concept-visual association edge is the first attention weight after normalization; According to the second attention weight value of the concept node, establish a concept-speech association edge with the speech node, and the weight value of the concept-speech association edge is the second attention weight after normalization; For visual nodes and speech nodes with a temporal correspondence relationship, establish a visual-speech temporal edge, and the weight value of the visual-speech temporal edge is calculated according to the time overlap degree between the second video key frame and the second speech segment; Perform deep fusion of multi-modal features through a graph neural network, including: Input the initial node set into the graph attention network; Perform feature propagation along the concept-visual association edge and the concept-speech association edge, and calculate the cross-modal similarity score between nodes through the multi-head attention mechanism; Update the cross-modal feature representation of the initial node set according to the cross-modal similarity score, and generate knowledge graph nodes with fused features; Obtain the final multi-modal knowledge graph.
4. The online teaching interaction method based on a multi-modal knowledge graph according to claim 1, wherein Convert user interaction behavior data into a knowledge graph query vector, including: Convert the speech question content into query text through a speech recognition engine; Use named entity recognition to extract core concept entities in the query text; Map the core concept entities to the concept nodes in the multi-modal knowledge graph, and use the feature vectors of the successfully mapped concept nodes as text query basis vectors; Moreover, parse the operation objects and operation actions in the virtual teaching operation sequence to obtain the parsing result; Locate the corresponding teaching process node in the multi-modal knowledge graph according to the parsing result, and associate the feature vector of the teaching process node as an operation query basis vector; Moreover, identify the feature vectors of the knowledge graph nodes associated with the click area of the interface click trajectory, denote them as click feature vectors, and generate a heat weight according to the click frequency; Generate a click query basis vector according to the heat weight and the click feature vector; The text query base vector, operation query base vector, and click query base vector are weighted and fused to obtain a fused feature vector, and the fused feature vector is input into a query encoder to generate a final knowledge graph query vector.
5. The online teaching interaction method based on a multimodal knowledge graph according to claim 1, wherein Extract a three-dimensional teaching scenario sub-graph from the multi-modal knowledge graph according to the knowledge graph query vector, including: Based on the knowledge graph query vector, perform node matching in the multi-modal knowledge graph to locate the initial query node; Perform graph traversal along the concept-visual association edge and concept-audio association edge to extract the associated visual query node and audio query node; Establish a spatio-temporal association relationship according to the visual-audio time series edge to form a three-dimensional teaching scenario sub-graph; When a standard teaching operation is detected, activate the forward reasoning path of the three-dimensional teaching scenario sub-graph to generate progressive teaching content, including: Start from the initial query node and perform reasoning along the standard teaching path in the knowledge graph; Combine the associated concept query nodes, visual query nodes, and audio query nodes in the order of teaching logic; Generate a progressive teaching sequence including basic concept explanations, example demonstrations, and audio guidance; When an incorrect teaching operation is detected, trigger the counterfactual reasoning mechanism of the three-dimensional teaching scenario sub-graph to generate comparative teaching content, including: Identify the abnormal node in the knowledge graph corresponding to the incorrect operation; Disable the standard teaching path edge directly associated with the abnormal node; Activate the counterfactual reasoning edge to generate teaching content including error analysis, correct demonstration, and comparative explanation.
6. The online teaching interaction method based on a multimodal knowledge graph according to claim 5, characterized in that, The scenario interaction data includes the access frequency and stay duration of the concept query node, the viewing completeness of the visual query node, the repeated playback times of the audio query node, and the selection ratio of the standard teaching path and the counterfactual reasoning path; Match the mapping relationship between the scenario interaction data and the multi-modal knowledge graph to construct a cognitive state tracking matrix, including: Establish an initial cognitive state tracking matrix, where the rows of the initial cognitive state tracking matrix represent the concept nodes in the multi-modal knowledge graph, and the columns include the proficiency index, operation proficiency index, error pattern index, and attention distribution index; Map the access records of the concept query node to the concept node, calculate the average test correct rate and access duration score of the concept node to obtain the proficiency index; Map the viewing records of the visual query node to the visual node, and count the operation success rate and average time consumption of the teaching process node to obtain the operation proficiency index; Map the teaching path selection record to the standard teaching path or the counterfactual reasoning path, and record the triggering times of the counterfactual reasoning path and the subsequent operation improvement degree to obtain the error pattern index; Map the playback records of the audio query node to the audio node, and obtain the attention distribution index according to the interaction event frequency of the audio node, concept node, and visual node within a unit time; Update the proficiency index, operation proficiency index, error pattern index, and attention distribution index to the columns of the initial cognitive state tracking matrix to obtain the updated cognitive state tracking matrix.
7. The online teaching interaction method based on a multi-modal knowledge graph according to claim 6, wherein Dynamically adjust teaching strategies based on the cognitive state tracking matrix, strengthen the cross-modal presentation combination of low-proficiency concept nodes, and extend the high-level knowledge associated with high-proficiency operation sequences, including: Extracting a set of low-proficiency concept nodes whose proficiency index is lower than a first preset threshold from the cognitive state tracking matrix; Performing a cross-modal reinforcement step on each low-mastery concept node in the set of low-mastery concept nodes: Increase the display duration and frequency of visual nodes associated with low-mastery concept nodes; Improve the playback priority of speech nodes associated with low-mastery concept nodes; Produce intensive instructional packages that include textual explanations, visual demonstrations, and audio guidance; Extracting a set of high-proficiency concept nodes whose operation proficiency index is higher than a second preset threshold from the cognitive state tracking matrix; Performing a high-level knowledge extension step on each high-proficiency concept node in the high-proficiency concept node set: Extracting high-order visual nodes along concept-visual association edges; Extracting higher-order phonetic nodes along concept-phonetic association edges; Generate a high-level teaching package that includes advanced knowledge explanations, in-depth case analysis, and extended exercises; According to the error mode indicator in the cognitive state tracking matrix, the triggering condition of the counterfactual reasoning path is adjusted, including: Add early warning nodes for high-frequency error modes; Optimize the comparison display method between incorrect operations and correct demonstrations; According to the attention distribution index in the cognitive state tracking matrix, the teaching rhythm is dynamically adjusted, including: Simplify teaching content for low attention areas; Add interactive links to high-attention areas; The existing teaching strategy configuration is updated to obtain and execute an updated teaching configuration, wherein the updated teaching configuration includes a reinforcement teaching plan for low-mastery concept nodes, a high-order extension plan for high-proficiency concept nodes, optimization parameters for counterfactual reasoning paths, and adjustment parameters for teaching rhythm.
8. The online teaching interaction method based on a multi-modal knowledge graph according to claim 1, wherein According to the group behavior analysis model, the cognitive state tracking matrix is aggregated, the teaching path recommendation weight in the multimodal knowledge graph is adjusted, and a personalized learning navigation graph is generated, including: Obtaining a cognitive state tracking matrix of multiple users, recorded as a group cognitive state tracking matrix, wherein the group cognitive state tracking matrix includes group concept nodes, calculating an average mastery index of each group concept node and counting high-frequency error patterns and group-related nodes; Analyze the common features of the optimal learning path and identify the hot spots of group attention distribution; Calculate the average mastery index, high-frequency error pattern, and optimal learning path characteristics of the group cognitive state tracking matrix, and adjust the teaching path recommendation weight in the multimodal knowledge graph based on the group attention distribution hotspots, including: The weight of basic teaching is reduced for concept nodes with high mastery of the group, and the weight of counterfactual reasoning is increased for nodes with high frequency of group errors; Improve the recommendation priority of the associated edges in the group's optimal path, and optimize the resource display weight for the group's attention hotspot areas; According to the adjusted teaching path recommendation weight, the subgraph structure related to the current user is extracted from the multimodal knowledge graph; After recalculating the path score of the teaching path, mark the key concept nodes on the recommended learning path and mark the error-prone node areas that need to be focused on; Generate the learning navigation graph including the main learning path and alternative paths with personalized recommendations, recommended concept nodes, recommended visual nodes, combinations of recommended voice nodes, recommended path weight scores, estimated learning durations, and key reinforcement area prompt markers.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.
10. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Multi-modal knowledge graph establishment method and application
CN117131933A
Multi-modal knowledge graph representation learning method based on cross-modal semantic alignment
CN117435744A
Safety propaganda and education training knowledge graph and data management method and system based on AI
CN118585658A
Knowledge association learning method and system based on knowledge graph and virtual reality
CN119166830A
Question and answer method and device based on large language model and nonvolatile storage medium
CN119168068A
Cited By
Intelligent display terminal interaction method and system applying AI model
CN120523335A
Knowledge graph-based low-altitude economic domain question and answer teaching interaction method and system
CN120561254A
Cross-modal data query method
CN120632128A
Dynamic classroom optimization method based on real-time space-time semantic graph tracking
CN120688701A
Teacher AI accomplishment multi-mode diagnosis and accomplishment path generation system
CN120823081A