Visual intelligent interactive education model system

By using a cross-modal transformer network and a self-supervised comparative learning optimization module, multimodal data analysis collaborative between the edge and cloud was achieved. This solved the problems of global semantic modeling, missing modality processing, and insufficient real-time performance in existing technologies, and improved the accuracy and real-time performance of classroom interactive analysis.

CN121997250APending Publication Date: 2026-05-08CHANGCHUN INST OF ELECTRONIC TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGCHUN INST OF ELECTRONIC TECH
Filing Date
2025-12-25
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing multimodal classroom data analysis methods have shortcomings in global semantic modeling, missing modality handling, real-time performance and scalability. They are difficult to achieve efficient cross-modal fusion, lack real-time interactive analysis capabilities, and their performance degrades when faced with complex and dynamic teaching scenarios.

Method used

It employs a cross-modal transformer network and a multi-head cross-attention mechanism for data fusion, combined with a self-supervised comparative learning optimization module, to achieve collaborative operation between the edge and the cloud. It has cross-modal bidirectional fusion and self-supervised alignment capabilities, and supports real-time interactive analysis.

Benefits of technology

It improves the accuracy and real-time performance of classroom emotion recognition and interaction prediction, enhances the robustness and adaptability of the system, enables feature completion when modalities are missing, and reduces network latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997250A_ABST
    Figure CN121997250A_ABST
Patent Text Reader

Abstract

The invention discloses a visual intelligent interactive education model system, which realizes real-time acquisition, fusion and intelligent analysis of classroom multi-modal data in combination with edge computing and cloud collaboration. According to the system, a multi-modal data acquisition and processing environment is cooperatively constructed at an edge end and a cloud end, texts, voices, videos, images, three-dimensional postures, interaction events and the like in a classroom are subjected to synchronous preprocessing and input into a cross-modal transformer network, bidirectional fusion between modals is realized under a multi-head cross attention mechanism, and a multi-modal data acquisition and processing system is established. Generating holographic representation of the unified semantic space; a self-supervised contrast learning strategy is adopted to complete semantic alignment and discrimination enhancement of multi-modal features; based on holographic representation, emotion recognition, behavior analysis and interaction intention prediction can be carried out, and mode missing completion is supported; the edge end deploys a lightweight model to realize low-delay processing, and the cloud end performs high-precision reasoning and optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of interactive education model technology, and specifically relates to a visual intelligent interactive education model system. Background Technology

[0002] Currently, multimodal education data analysis technology has been widely used in scenarios such as smart classrooms, online education, and blended learning. In traditional methods, a common approach is to model classroom multimodal data based on graph convolutional neural networks.

[0003] Graph convolutional neural networks utilize graph structures to capture local correlations between data from different modalities and achieve feature fusion to some extent; however, this type of method has the following shortcomings:

[0004] The graph convolutional neural network (GNN) has limited capabilities in handling time series and long-distance dependencies, making it difficult to fully capture global semantic relationships across modalities and time periods in the classroom. When data in a certain modality is missing or of poor quality, the overall performance of the GNN is prone to significant decline, lacking effective modality completion and inference mechanisms. Real-time performance is limited: the computational overhead of GNNs is large, resulting in delays in real-time processing and feedback of large-scale classroom data, affecting the interactive teaching experience. In complex scenarios where video, audio, text, and 3D pose data are simultaneously accessed, traditional GNN solutions lack flexibility in model structure adjustment and feature fusion strategy optimization.

[0005] As shown in the existing multimodal classroom behavior analysis system, "as per application number: CN201711469436.5 A multimodal student classroom behavior analysis system and method", this patent can analyze student classroom behavior through video and audio feature extraction. However, it adopts the method of extracting features independently for each modality and then performing statistical analysis. The integration degree is low, and the adaptability to modality missing and real-time requirements is limited. It has no advantage when deep cross-modal interaction is required or when dealing with complex dynamic teaching scenarios.

[0006] Therefore, there is an urgent need for a new type of visual intelligent education model system that can achieve efficient cross-modal fusion, has the ability to adaptively complete missing modalities, and supports real-time interactive analysis, so as to improve the quality of classroom interaction and the accuracy of analysis. Summary of the Invention

[0007] (a) Technical problems to be solved

[0008] This invention aims to overcome the shortcomings of existing multimodal classroom data analysis methods based on graph convolutional neural networks in terms of global semantic modeling, missing modality processing, real-time performance, and scalability. It provides a visual intelligent interactive education model system that can run collaboratively at the edge and in the cloud, and has cross-modal bidirectional fusion and self-supervised alignment capabilities, thereby improving the accuracy and real-time performance of classroom emotion recognition, behavior analysis, and interaction prediction.

[0009] (II) Technical Solution

[0010] The present invention is achieved through the following technical solution: The present invention proposes a visual intelligent interactive education model system, characterized in that: a multimodal data acquisition and preprocessing module is used to synchronously acquire multimodal data in classroom teaching scenarios and to perform noise reduction processing on the data;

[0011] The lightweight edge processing module, deployed on the local computing device in the classroom, is used to perform real-time feature extraction and compression encoding on preprocessed multimodal data, reducing network transmission latency and bandwidth consumption.

[0012] The cross-modal feature fusion and representation generation module, including a cross-modal transformer network and a multi-head cross-attention mechanism, is used to perform bidirectional information interaction and fusion of features from different modalities to generate holographic representation vectors in a unified semantic space.

[0013] The self-supervised contrastive learning optimization module is used to maximize the similarity between different modal representations of the same event and minimize the similarity between different event representations by constructing multimodal positive and negative sample pairs. This optimizes the parameters of the feature encoder and fusion network, achieving semantic alignment and discriminative enhancement.

[0014] The classroom intelligent analysis module is used to perform classroom emotion recognition, behavior analysis, and interactive intent prediction based on holographic representation vectors, and generate corresponding analysis results;

[0015] The visualization and teaching management interface module is used to display classroom analysis results, emotional change trends, and interaction predictions in a structured and visualized manner, and to interact with the cloud-based teaching management system to realize teaching playback, behavior statistics, and teaching optimization.

[0016] The cloud-edge collaboration and data synchronization module is used to achieve bidirectional synchronization of multimodal data, model parameters and analysis results between the edge and the cloud. It includes clock synchronization, timestamp marking and buffer scheduling mechanisms to ensure the real-time performance and consistency of cloud-edge collaboration.

[0017] (III) Beneficial Effects

[0018] By employing a cross-modal transformer network and a multi-head cross-attention mechanism, this invention can simultaneously capture long-distance dependencies across time and modality in classroom data, achieving more accurate global information fusion than graph convolutional neural networks. Furthermore, by introducing a modality prediction and generation mechanism, even if data in a particular modality is lost or its quality degrades, the system can still infer and complete features through a fusion decoder, significantly improving robustness. A lightweight model is deployed at the edge for real-time feature extraction and preliminary inference, while the cloud handles full model optimization and complex analysis, ensuring both low-latency interaction and analytical accuracy. Attached Figure Description

[0019] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0020] Figure 1 This is a system module block diagram of the present invention.

[0021] Figure 2 This is a system flowchart of the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Example 1

[0024] Please refer to the following: A classroom interaction analysis system based on cross-modal transformers and self-supervised comparative learning. Figure 1 The present invention provides a technical solution: the system includes an edge acquisition and lightweight processing module, a cloud cross-modal fusion and optimization module, an emotion and interaction prediction module, and a visualization feedback module.

[0025] Step 1: Multimodal Data Acquisition and Synchronous Processing. In a classroom environment, camera equipment, microphone arrays, posture sensors, and interaction detection devices can be deployed at different locations within the classroom. The camera equipment is used to acquire video streams of participants' facial expressions and body movements. The microphone array is used to acquire speech signals and reduce environmental noise. The posture sensors can acquire two-dimensional or three-dimensional skeleton information of participants. The interaction detection device can collect event signals such as clicking, writing, and page turning. The acquired modal data are aligned using a timestamp server, with a time error preferably not exceeding ±5ms to ensure synchronization between modalities. Video data can be used for keyframe extraction and skeleton recognition, speech data can be used for frame segmentation and feature extraction, text data can be used for word segmentation and grammatical parsing, posture data can be normalized, and interactive events can be converted into structured event sequences.

[0026] Step 2: Input the preprocessed multimodal data into the cross-modal transformer network, realize bidirectional information fusion between modes under the multi-head cross-attention mechanism, and generate holographic representation vectors in a unified semantic space;

[0027] The cross-mode transformer network includes:

[0028] Multimodal encoder: converts each mode (Text, audio, video, gestures, and interaction events) are respectively mapped to feature sequences:

[0029]

[0030] in This represents the number of time steps for this mode. For feature dimensions.

[0031] Multi-head cross-attention mechanism: any two modalities and The correlation weights between them are:

[0032]

[0033] in , , , For learnable parameters, This is the scaling factor.

[0034] Fusion decoder: concatenates and maps the fused representation vectors from different modalities to a unified semantic space.

[0035]

[0036] in This represents vector concatenation, FFN represents a feedforward neural network, and the output is a holographic representation vector. Compared to graph convolutional neural networks: graph convolutional neural networks require the construction of an adjacency matrix in multimodal fusion. And through the formula

[0037]

[0038] Traditional methods rely on fixed adjacency structures for propagation, making it difficult to dynamically adapt to different modal relationships. The cross-modal transformer of this invention does not require an explicit adjacency graph and learns inter-modal dependencies adaptively through attention weights, resulting in stronger adaptability.

[0039] Among them, the transformer network can also perform emotion recognition and interaction intent prediction, specifically manifested as follows:

[0040] Based on holographic representation vector The emotion recognition module outputs a probability vector:

[0041]

[0042] The interactive intent prediction module uses the time-series modeling capabilities of transformers to predict the future. Interaction probability per second:

[0043]

[0044] in This is the Sigmoid activation function.

[0045] Step 3: Employ a self-supervised contrastive learning strategy to achieve semantic alignment and discrimination enhancement of multimodal features by maximizing the similarity of different modal representations of the same teaching event and minimizing the similarity of different event representations.

[0046] During the training phase, the system utilizes a self-supervised contrastive learning strategy to construct multimodal "positive sample pairs" and "negative sample pairs," and optimizes the parameters of the multimodal encoder and cross-modal fusion module. This makes the features of different modalities of the same event closer in semantic space, while the features of different events are further apart. This method does not require additional labeled data, reduces manual costs, and significantly improves the generalization performance of the model in unseen classroom scenarios.

[0047] Construct positive sample pairs Different modal features from the same teaching event, negative sample pairs From different events. Optimized using the InfoNCE loss function:

[0048]

[0049] in: Cosine similarity; This is the temperature coefficient (typically taken as 0.07). For batch size.

[0050] The optimization process simultaneously updates the parameters of the multimodal encoder and the cross-modal fusion module, ensuring that the same semantic events are similar in a unified semantic space, while different semantic events are far apart.

[0051] Step 4: The system transforms the holographic representation vector into structured visual data, including emotion change curves, interactive event timelines, and participation heatmaps. Teachers can review the lessons after class to analyze student participation and emotion trends, providing a basis for optimizing teaching strategies.

[0052] Example 2

[0053] This system also has the ability to predict missing modes and tolerate network anomalies;

[0054] Step 1: Missing Mode Detection

[0055] When a certain modality of data is detected in continuous during the inference phase... Missing within a time slice, or its confidence level being lower than a preset threshold. When this occurs, the system triggers missing mode compensation.

[0056] Step 2: Feature-level compensation

[0057] In another experimental environment, this embodiment further optimizes the system in Example 1 by adding a self-supervised contrastive learning strategy and a modality missing completion mechanism. During the training phase, the system extracts multimodal data fragments from different classrooms, forming positive sample pairs from different modalities originating from the same time slice or the same teaching event, and forming negative sample pairs from data from different events. The goal of training is to enable the model to shorten the distance between positive sample pairs and widen the distance between negative sample pairs in a unified semantic space, thereby achieving stronger semantic alignment capabilities. The core contrastive loss function is:

[0058]

[0059] in:

[0060] The feature vector representing a positive sample pair

[0061] The feature vector of any sample

[0062] Represents a similarity function (such as cosine similarity).

[0063] Temperature coefficient

[0064] Indicates the number of samples in the batch

[0065] During the inference phase, if a certain modality of data (such as video stream) is missing, the system will predict the feature vector of that modality based on the obtained modality data through the fusion decoder, and may even directly generate the missing video or audio segments, thereby ensuring the integrity and stability of the holographic representation.

[0066] Experimental results show that even with network latency or partial sensor failure, the system can still maintain an interaction prediction accuracy of approximately 92% and an emotion recognition accuracy of over 90%, effectively improving the robustness and continuity of classroom interaction analysis.

[0067] In a smart classroom at a university, the visual intelligent interactive education model system of this invention was used to analyze the classroom interaction in a "Advanced Mathematics" course. The system is equipped with multiple cameras, microphone arrays, and posture sensors to collect multimodal data such as video, voice, and posture in real time. If a convolutional neural network is used to fuse and analyze this data, the video needs to be input frame by frame into the convolutional neural network for local feature extraction; the voice data needs to be converted into spectrograms and input into another convolutional neural network; during the fusion stage, only weighted averaging or simple splicing of local features can be performed.

[0068] There are two problems with doing this:

[0069] Convolutional neural networks have a limited receptive field, capturing short-term local features, making it difficult to accurately model long-term classroom interaction patterns (such as the trend of student attention changes from the start of class to 40 minutes later).

[0070] Video and speech features are extracted independently after processing by convolutional neural networks, and are only simply combined during the fusion stage. This makes it impossible to establish deep cross-modal associations in the early stages of feature extraction.

[0071] Advantages of Transformer Networks

[0072] In the system of this invention, video frame sequences, speech feature sequences, and pose data sequences are first mapped into a unified feature space, and then jointly modeled through a multimodal transformer network. The transformer network adopts a self-attention mechanism, which can directly calculate the correlation between different time points in the global scope of the sequence. For example, there may be a causal relationship between a student's laughter at the 5th minute of class and a teacher's interactive question at the 20th minute. The transformer network can directly model this cross-time-period correlation, which is difficult for convolutional neural networks to capture. In the multi-head attention mechanism of transformer networks, features of a particular frame of a video modality can directly "attention" to related speech or gesture features, rather than just focusing on data of the same modality. This means that, for example, the synchronicity between changes in a student's facial expressions and changes in tone of voice can be detected and utilized earlier, resulting in higher accuracy in emotion recognition than convolutional neural network fusion methods. In a network latency simulation experiment, the system lost 15 seconds of video data. Since the transformer network has learned the long-term dependencies between different modalities during the training phase (such as the correspondence between speech intonation and facial expression), it can infer the missing video features using speech and posture features, thereby ensuring the continuity of the sentiment analysis curve. In contrast, the convolutional neural network method lacks the ability to model globally across modalities, resulting in significantly distorted recovery results.

[0073] Test Scenario Accuracy of Convolutional Neural Network Fusion Model Accuracy of transformer network fusion model Normal data input 85% 92% Video modality missing (15 seconds) 68% 89% Speech modality missing (15 seconds) 70% 88% Cross-time period sentiment trend prediction (30 minutes) 61% 87%

[0074] Experimental results show that the transformer network used in this invention significantly outperforms traditional convolutional neural network methods in global context modeling, cross-modal dependency learning, and missing data recovery. The above descriptions are merely preferred embodiments of this invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A visual intelligent interactive education model system, characterized in that: Includes the following modules: The multimodal data acquisition and preprocessing module is used to simultaneously acquire multimodal data in classroom teaching scenarios and perform noise reduction processing on the data; The lightweight edge processing module, deployed on the local computing device in the classroom, is used to perform real-time feature extraction and compression encoding on preprocessed multimodal data, reducing network transmission latency and bandwidth consumption. The cross-modal feature fusion and representation generation module, including a cross-modal transformer network and a multi-head cross-attention mechanism, is used to perform bidirectional information interaction and fusion of features from different modalities to generate holographic representation vectors in a unified semantic space. The self-supervised contrastive learning optimization module is used to maximize the similarity between different modal representations of the same event and minimize the similarity between different event representations by constructing multimodal positive and negative sample pairs. This optimizes the parameters of the feature encoder and fusion network, achieving semantic alignment and discriminative enhancement. The classroom intelligent analysis module is used to perform classroom emotion recognition, behavior analysis, and interactive intent prediction based on holographic representation vectors, and generate corresponding analysis results; The visualization and teaching management interface module is used to display classroom analysis results, emotional change trends, and interaction predictions in a structured and visualized manner, and to interact with the cloud-based teaching management system to realize teaching playback, behavior statistics, and teaching optimization. The cloud-edge collaboration and data synchronization module is used to achieve bidirectional synchronization of multimodal data, model parameters and analysis results between the edge and the cloud. It includes clock synchronization, timestamp marking and buffer scheduling mechanisms to ensure the real-time performance and consistency of cloud-edge collaboration.

2. The visual intelligent interactive education model system according to claim 1, characterized in that: The cross-modal transformer network includes a multimodal encoder, a multi-head cross-attention module, and a fusion decoder. The multimodal encoder is used to independently extract the feature representations of each modality, the multi-head cross-attention module is used to calculate the correlation weights between modalities and achieve bidirectional fusion, and the fusion decoder is used to generate a globally unified semantic representation vector.

3. The visual intelligent interactive education model system according to claim 1, characterized in that: The self-supervised contrastive learning optimization module is based on the contrastive loss function and simultaneously optimizes the parameters of each modal feature encoder and the cross-modal fusion module to enhance the semantic consistency and discriminative ability of multimodal representations.

4. The visual intelligent interactive education model system according to claim 1, characterized in that: The classroom intelligent analysis module is used to perform comprehensive calculations on classroom emotional states based on the holographic representation vector. The calculations include extracting emotion-related information from different modal features and fusing them to generate an overall emotional state label for the classroom.

5. The visual intelligent interactive education model system according to claim 1, characterized in that: The cloud-edge collaboration and data synchronization module has a clock synchronization error of less than 1 millisecond and is equipped with a multimodal data buffer queue and a time alignment scheduler for timing alignment and missing data completion of different modal data streams.

6. The visual intelligent interactive education model system according to claim 1, characterized in that: When data for a particular modality is missing during the inference phase, the fusion decoder can predict the feature vector of that modality or directly generate the content of that modality to ensure the integrity of the multimodal representation.

7. The visual intelligent interactive education model system according to claim 1, characterized in that: The lightweight edge processing module and the cloud-based cross-modal transformer network adopt a parameter hierarchical update mechanism. When network conditions are abnormal, the edge can continuously output analysis results based on cached data and local inference models until the network recovers.

Citation Information

Patent Citations

  • A multi-mode student class behavior analysis system and method

    CN108090857A