Classroom behavior and emotion collaborative analysis system based on multi-modal AI
Through multimodal AI technology, combined with multiple data collection devices and edge computing, high-precision real-time analysis of classroom behavior and emotions is achieved, solving the problems of single-modal limitations and static analysis, providing low-latency teaching intervention strategies, and improving classroom interaction efficiency and data privacy protection.
Patent Information
- Application Number
- CN202510597896.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-26
AI Technical Summary
The existing collaborative analysis system of classroom behavior and emotions has single-modal limitations, static analysis defects and inefficient data fusion problems, and cannot achieve high-precision real-time analysis of behavior and emotions.
It adopts multimodal AI technology, through dynamic weight adjustment, cross-modal graph neural network and edge computing, combined with infrared cameras, directional microphone arrays, wearable devices and electronic whiteboards to synchronously collect multiple data, perform timestamp synchronization, wavelet denoising and feature extraction, and use improved BiGRU network, LSTM network and BERT model to extract features, and construct multimodal association map through graph neural network to achieve real-time feedback.
It achieves multi-dimensional real-time monitoring of students' classroom behavior and emotions, enhances the refinement and accuracy of analysis, provides low-latency teaching intervention strategies, ensures data privacy and security, and improves the immediacy and effectiveness of classroom interactions.
Smart Images

Figure CN120708116A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the intersection of artificial intelligence and educational technology, and specifically to a classroom behavior and emotion collaborative analysis system based on multimodal AI. Background Art
[0002] Classroom behavior and emotion collaborative analysis involves observing and analyzing students' behavior and emotional changes in the classroom to understand their learning and psychological state, thereby optimizing teaching strategies and improving the classroom atmosphere. This analysis method is particularly important in primary and secondary education because it helps teachers better understand students' needs, improve teaching effectiveness, and promote students' all-round development.
[0003] The existing classroom behavior and emotion collaborative analysis system has the following defects:
[0004] Single-modality limitations: Traditional systems (such as CN113569805A video-based behavioral analysis) rely on a single data source and cannot capture the correlation between behavior and emotions.
[0005] Static analysis flaws: Existing systems mostly rely on offline or periodic analysis, lacking the ability to model the dynamic interaction between behavior and emotions over time, leading to delayed feedback.
[0006] Inefficient data fusion: The multimodal feature fusion method does not implement dynamic cross-modal weight allocation and has poor environmental adaptability.
[0007] To this end, a classroom behavior and emotion collaborative analysis system based on multimodal AI is proposed. Summary of the Invention
[0008] The present invention aims to solve the problems raised in the background technology and provides a classroom behavior and emotion collaborative analysis system based on multimodal AI. Through dynamic weight adjustment, cross-modal graph neural network and edge computing optimization, it solves the limitations of single modality and the inefficiency of data fusion, and realizes high-precision real-time analysis of classroom behavior and emotion.
[0009] The specific technical solutions are as follows:
[0010] A multimodal AI-based classroom behavior and emotion collaborative analysis system, including:
[0011] Multimodal data acquisition module, used to simultaneously collect visual, voice, physiological signals and text data through infrared cameras, directional microphone arrays, wearable devices and electronic whiteboards;
[0012] A preprocessing module, performing timestamp synchronization, wavelet denoising and normalization processing on the data;
[0013] Feature extraction module, including:
[0014] The visual feature extraction unit uses an improved BiGRU network to process the facial micro-expression sequences generated by the optical flow method;
[0015] Speech feature extraction unit, which extracts intonation and speech rate features from Mel spectrum based on LSTM network;
[0016] Text feature extraction unit, which generates semantic sentiment vectors through the BERT model;
[0017] Physiological signal feature extraction unit, which uses wavelet transform to extract HRV frequency domain features and GSR time domain slope;
[0018] Collaborative analysis modules, including:
[0019] Dynamic weight adjustment unit, dynamically assigning weights to each modality based on the ambient noise level;
[0020] Cross-modal fusion network, using graph neural network GNN to construct multimodal association graph;
[0021] The real-time feedback module generates behavioral intervention strategies based on edge computing nodes, with a response time of ≤200ms.
[0022] In the above-mentioned multimodal AI-based classroom behavior and emotion collaborative analysis system, the dynamic weight adjustment unit calculates the speech modality weight using the following formula:
[0023]
[0024] Wherein, N is the real-time noise decibel value, N0=60dB, and k=0.1 is the adjustment coefficient.
[0025] In the above-mentioned multimodal AI-based classroom behavior and emotion collaborative analysis system, in the graph neural network of the cross-modal fusion network:
[0026] Nodes include visual, speech, text, and physiological signal feature vectors;
[0027] The edge weight is calculated through the multi-head attention mechanism, the formula is:
[0028]
[0029] Among them, w ij represents the attention weight of the i-th element to the j-th element; Q i is the query vector; K j is the key vector (keyvector); is the transpose of the key vector; d is the feature dimension; It is a scaling factor used to stabilize the gradient; the Softmax function normalizes the dot product result.
[0030] In the above-mentioned multimodal AI-based classroom behavior and emotion collaborative analysis system, in the physiological signal feature extraction unit:
[0031] HRV frequency domain characteristics include the ratio of low-frequency power LF to high-frequency power HF;
[0032] The GSR time domain slope is the average rate of change of the skin conductance rising segment, and the calculation formula is:
[0033]
[0034] Where ΔGSR represents the change in skin conductance, and Δt represents the corresponding time change.
[0035] In the aforementioned multimodal AI-based classroom behavior and emotion collaborative analysis system, the intervention strategy generation of the real-time feedback module includes:
[0036] When a student's distracted behavior lasts for more than 3 minutes, the wearable device will vibrate to remind them;
[0037] When the proportion of negative emotions in the group is greater than 40%, suggestions for adjusting classroom activities will be sent to the teacher.
[0038] In the above-mentioned multimodal AI-based classroom behavior and emotion collaborative analysis system, the improved BiGRU network introduces an attention pooling operation in the output layer, and the calculation formula is:
[0039]
[0040] Among them, a t is the attention weight at the tth time step, h t is the BiGRU hidden state.
[0041] The aforementioned multimodal AI-based classroom behavior and emotion collaborative analysis system also includes a privacy protection module:
[0042] Using the federated learning framework, physiological feature data is processed only on the local device;
[0043] Differential privacy technology is used to add Gaussian noise to the uploaded classroom analysis results.
[0044] The classroom behavior-emotion collaborative analysis method of the multimodal AI-based classroom behavior and emotion collaborative analysis system includes the following steps:
[0045] Step 1: Align multimodal data streams through timestamp synchronization algorithm;
[0046] Step 2: Use wavelet transform to extract the time-frequency domain features of physiological signals;
[0047] Step 3: Generate behavior-emotion association graph based on GNN cross-modal fusion network;
[0048] Step 4: Output teaching optimization strategies in real time through edge computing nodes.
[0049] In the above-mentioned multimodal AI-based classroom behavior and emotion collaborative analysis system, in step three:
[0050] The node feature update formula of the association graph is:
[0051]
[0052] Where σ is the ReLU activation function; W (k) is the trainable weight matrix; i (k+1) represents the feature representation of node i at the k+1 layer; N(i) represents the set of neighbor nodes of node i; A ij Represents the adjacency relationship between node i and node j; represents the feature representation of node j at the kth layer.
[0053] In the above-mentioned multimodal AI-based classroom behavior and emotion collaborative analysis system, the implementation of the timestamp synchronization algorithm in step 1 includes:
[0054] Assign an independent clock source to each type of modal data stream and calibrate the clock deviation of each device through the NTP protocol;
[0055] Dynamic time warping (DTW) is used to align non-uniformly sampled multimodal data streams, with an alignment error of ≤10ms.
[0056] Add unified timing labels to the synchronized data streams and generate a time-modality joint index table.
[0057] The present invention has the following beneficial effects:
[0058] 1. Multi-dimensional behavior-emotion analysis: By integrating multimodal data such as vision, voice, text, and physiological signals, the system achieves comprehensive monitoring of students' classroom behavior (such as concentration and interaction frequency) and emotional state (such as positive and negative), overcoming the limitations of traditional unimodal analysis.
[0059] 2. Dynamic environmental adaptability: A dynamic weight adjustment mechanism is used to adaptively allocate modal weights based on real-time environmental parameters (such as noise level) to improve the analysis robustness in complex classroom scenarios.
[0060] 3. Cross-modal collaborative modeling: Cross-modal fusion technology based on graph neural networks (GNN) captures the dynamic correlation between behavior and emotions, enhancing the refinement and accuracy of joint analysis.
[0061] 4. Low-latency real-time feedback: Combined with edge computing technology, it enables the rapid generation and push of teaching intervention strategies, significantly improving the immediacy and effectiveness of classroom interactions.
[0062] 5. Privacy and security protection: Through federated learning and differential privacy technologies, we ensure the local processing and desensitized transmission of sensitive biometric data to meet education data compliance requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 A schematic diagram of the architecture of a multimodal AI-based classroom behavior and emotion collaborative analysis system provided by an embodiment of the present invention;
[0064] Figure 2 A graph showing the relationship between speech modality weight and noise level for the multimodal AI-based classroom behavior and emotion collaborative analysis system provided in an embodiment of the present invention;
[0065] Figure 3 A radar chart showing the advantages of the multimodal AI-based classroom behavior and emotion collaborative analysis system provided by an embodiment of the present invention;
[0066] Figure 4 A behavior intervention strategy triggering frequency diagram for the multimodal AI-based classroom behavior and emotion collaborative analysis system provided by an embodiment of the present invention;
[0067] Figure 5 A curve chart showing the improvement in the technical effects of the multimodal AI-based classroom behavior and emotion collaborative analysis system provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0068] The technical solution of the present invention will be further described below with reference to the accompanying drawings and through specific implementation methods.
[0069] Among them, the drawings are only used for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting this patent; in order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0070] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "inside", "outside" and the like indicate an orientation or position relationship based on the orientation or position relationship shown in the drawings, it is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting this patent. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0071] In the description of the present invention, unless otherwise expressly specified or limited, when the term "connection" or the like appears to indicate a connection relationship between components, such term should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be internal communication between two components or an interaction between two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood in specific circumstances.
[0072] The classroom behavior and emotion collaborative analysis system based on multimodal AI provided in this embodiment is as follows: Figure 1-Figure 5 As shown, it includes: multimodal data acquisition module, preprocessing module, feature extraction module, collaborative analysis module and real-time feedback module, among which:
[0073] The multimodal data acquisition module is used to synchronously collect visual, voice, physiological signals and text data through infrared cameras, directional microphone arrays, wearable devices and electronic whiteboards;
[0074] The preprocessing module is used to perform timestamp synchronization, wavelet denoising and normalization processing on the data;
[0075] Feature extraction module, including visual feature extraction unit, speech feature extraction unit, text feature extraction unit and physiological signal feature extraction unit;
[0076] Among them, the visual feature extraction unit uses an improved BiGRU network to process the facial micro-expression sequence generated by the optical flow method;
[0077] Among them, the speech feature extraction unit extracts the intonation and speech speed features of the Mel spectrum based on the LSTM network;
[0078] Among them, the text feature extraction unit generates semantic sentiment vectors through the BERT model;
[0079] Among them, the physiological signal feature extraction unit uses wavelet transform to extract HRV frequency domain features and GSR time domain slope;
[0080] Collaborative analysis module, including a dynamic weight adjustment unit and a cross-modal fusion network, where:
[0081] The dynamic weight adjustment unit is used to dynamically allocate the weight of each mode according to the environmental noise level;
[0082] The cross-modal fusion network is used to construct a multimodal association graph using graph neural network (GNN);
[0083] The real-time feedback module is used to generate behavioral intervention strategies based on edge computing nodes, with a response time of ≤200ms.
[0084] The multimodal AI-based classroom behavior and emotion collaborative analysis system that adopts the above solution realizes multi-dimensional real-time monitoring of classroom behavior and emotions by integrating multimodal data acquisition and processing modules; at the same time, it combines dynamic weight adjustment and graph neural network fusion technology to improve the comprehensiveness and accuracy of behavior-emotion correlation analysis; at the same time, the real-time feedback module based on edge computing ensures low-latency output of teaching intervention strategies and enhances the practicality of the system.
[0085] The dynamic weight adjustment unit calculates the speech modality weight using the following formula:
[0086]
[0087] Wherein, N is the real-time noise decibel value, N0=60dB, and k=0.1 is the adjustment coefficient.
[0088] By adopting the above scheme, the speech weight adjustment formula that is adaptive to environmental noise can effectively suppress noise interference and improve the reliability of speech modality in complex classroom scenarios; the dynamic weight allocation mechanism can enhance the system's robustness to environmental changes and avoid misjudgment caused by fixed weights.
[0089] Among them, in the graph neural network of the cross-modal fusion network:
[0090] Nodes include visual, speech, text, and physiological signal feature vectors;
[0091] The edge weight is calculated through the multi-head attention mechanism, the formula is:
[0092]
[0093] Among them, w ij represents the attention weight of the i-th element to the j-th element; Q i is the query vector; K j is the key vector (keyvector); is the transpose of the key vector; d is the feature dimension; It is a scaling factor used to stabilize the gradient; the Softmax function normalizes the dot product result.
[0094] By adopting the above scheme, the graph neural network based on the multi-head attention mechanism significantly enhances the interactive modeling ability of cross-modal features; through the dynamic calculation of edge weights, it captures the deep correlation between behavior and emotions, and improves the refinement of joint analysis.
[0095] Among them, in the physiological signal feature extraction unit:
[0096] HRV frequency domain characteristics include the ratio of low-frequency power LF to high-frequency power HF;
[0097] The GSR time domain slope is the average rate of change of the skin conductance rising segment, and the calculation formula is:
[0098]
[0099] Where ΔGSR represents the change in skin conductance, and Δt represents the corresponding time change.
[0100] The above scheme adopts the combined extraction method of HRV frequency domain features and GSR time domain slope to enhance the characterization ability of physiological signals for emotional states; the calculation formula of quantified physiological indicators enhances the specificity of emotion classification and reduces the impact of environmental factors on physiological data.
[0101] The intervention strategy generation of the real-time feedback module includes:
[0102] When a student's distracted behavior lasts for more than 3 minutes, the wearable device will vibrate to remind them;
[0103] When the proportion of negative emotions in the group is greater than 40%, suggestions for adjusting classroom activities will be sent to the teacher.
[0104] By adopting the above solution, precise teaching intervention can be achieved based on the trigger mechanism of the duration of distracting behavior and the proportion of group emotions; through vibration reminders of wearable devices and strategy push on the teacher side, the efficiency of classroom interaction and student concentration can be improved.
[0105] Among them, the improved BiGRU network introduces the attention pooling operation in the output layer, and the calculation formula is:
[0106]
[0107] Among them, a t is the attention weight at the tth time step, h t is the BiGRU hidden state.
[0108] Using the above scheme, the improved BiGRU network is combined with the attention pooling operation to optimize the feature extraction of facial micro-expression sequences; enhance the visual modality's ability to capture subtle behaviors (such as distraction and interaction) and reduce misidentification caused by motion blur.
[0109] Among them, it also includes the privacy protection module:
[0110] Using the federated learning framework, physiological feature data is processed only on the local device;
[0111] Differential privacy technology is used to add Gaussian noise to the uploaded classroom analysis results.
[0112] By adopting the above solution, through the combination of federated learning and differential privacy technology, it is possible to ensure the local processing and desensitized transmission of biometric data; while ensuring the accuracy of analysis, it can meet the legal requirements for education data privacy protection and reduce the risk of data leakage.
[0113] The classroom behavior-emotion collaborative analysis method of the multimodal AI-based classroom behavior and emotion collaborative analysis system includes the following steps:
[0114] Step 1: Align multimodal data streams through timestamp synchronization algorithm;
[0115] Step 2: Use wavelet transform to extract the time-frequency domain features of physiological signals;
[0116] Step 3: Generate behavior-emotion association graph based on GNN cross-modal fusion network;
[0117] Step 4: Output teaching optimization strategies in real time through edge computing nodes.
[0118] By adopting the above solution, the standardized process of time series synchronization, feature extraction and GNN fusion is used to realize the process operation of behavior-emotion analysis; the real-time feedback mechanism based on edge computing greatly shortens the strategy generation delay and significantly improves the instant response capability of teaching scenarios.
[0119] Among them, in step three:
[0120] The node feature update formula of the association graph is:
[0121]
[0122] Where σ is the ReLU activation function; W (k) is the trainable weight matrix; i (k+1) represents the feature representation of node i at the k+1 layer; N(i) represents the set of neighbor nodes of node i; A ij Represents the adjacency relationship between node i and node j; represents the feature representation of node j at the kth layer.
[0123] By adopting the above scheme, the hierarchical interactive learning ability of multimodal data can be enhanced through the design of the graph neural network node feature update formula; through the aggregation and nonlinear activation of neighbor node features, the dynamic modeling effect of the association graph can be significantly improved.
[0124] The implementation of the timestamp synchronization algorithm in step 1 includes:
[0125] Assign an independent clock source to each type of modal data stream and calibrate the clock deviation of each device through the NTP protocol;
[0126] Dynamic time warping (DTW) is used to align non-uniformly sampled multimodal data streams, with an alignment error of ≤10ms.
[0127] Add unified timing labels to the synchronized data streams and generate a time-modality joint index table.
[0128] The above scheme can significantly improve the time alignment accuracy of multimodal data through the dual calibration mechanism of the NTP protocol and the DTW algorithm. By constructing a joint index table of time series labels, it provides structured input for cross-modal analysis and reduces the data matching error rate.
[0129] The connection relationship between the modules in the multimodal AI-based classroom behavior and emotion collaborative analysis system provided in this embodiment is as follows:
[0130] Multimodal data acquisition module → preprocessing module → feature extraction module → collaborative analysis module → real-time feedback module;
[0131] Hardware and algorithm collaboration:
[0132] After the infrared camera, microphone array and other hardware collect the raw data, the pre-processing module performs time stamp synchronization, denoising and normalization processing to form a standardized multimodal data stream.
[0133] The time-aligned data output by the preprocessing module are respectively input into the four sub-units of the feature extraction module (vision, speech, text, and physiological signals). Each sub-unit extracts high-dimensional features through a dedicated algorithm (such as BiGRU, LSTM, BERT, and wavelet transform).
[0134] Cross-modal interaction design:
[0135] The output vector of the feature extraction module is input into the collaborative analysis module. The dynamic weight adjustment unit dynamically allocates the weight of each modality according to the environmental noise level (such as the speech weight formula) to generate a weighted feature matrix.
[0136] The cross-modal fusion network (GNN) receives the weighted feature matrix, models the inter-modal correlation through graph structure (such as the temporal matching of visual action features and physiological signals), and outputs a joint behavior-emotion feature vector.
[0137] Real-time feedback loop:
[0138] The joint feature vector of the collaborative analysis module is input into the real-time feedback module. The edge computing node generates an intervention strategy based on preset rules (such as the distraction behavior duration threshold) and pushes it through the teacher-side interface or wearable devices.
[0139] The privacy protection module intervenes during the data stream transmission process to ensure that the biometric data is processed locally (federated learning) and only the desensitized results are uploaded to the cloud storage.
[0140] Key module interaction details
[0141] (1) Connection between preprocessing module and feature extraction module
[0142] Timestamp synchronization mechanism:
[0143] The preprocessing module calibrates the clocks of each hardware device through the NTP protocol, and uses the DTW algorithm to align non-uniformly sampled data (such as low-frequency samples of physiological signals and high-frequency frames of videos) to generate a time-modality joint index table.
[0144] This index table serves as the input basis for the feature extraction module, ensuring that the features of different modalities are strictly aligned in the time dimension and avoiding timing misalignment during cross-modal analysis.
[0145] (2) Internal interaction of collaborative analysis modules
[0146] Dynamic weight adjustment and GNN collaboration:
[0147] The dynamic weight adjustment unit calculates the weight of each modality (weight 2 formula) according to the real-time environmental parameters (such as noise decibel value) and generates a weighted feature matrix to input into the GNN.
[0148] GNN dynamically updates node features based on the attention mechanism (weight 3 formula), aggregates cross-modal information (such as the correlation between visual distraction behavior and negative emotions in speech) through multi-layer graph convolution, and finally generates a behavior-emotion association map.
[0149] (3) Closed-loop control of real-time feedback module
[0150] Edge computing and policy generation:
[0151] The edge computing node receives the association graph output by the GNN and generates personalized strategies (such as vibration reminders) based on historical data (such as individual student behavior patterns).
[0152] The feedback strategy is pushed to the terminal device through a low-latency channel (response time ≤ 200ms), and at the same time triggers the privacy protection module to anonymize sensitive data.
[0153] 3. Physical connection design between hardware and algorithm
[0154] Data collection end:
[0155] The infrared camera is connected to the pre-processing module via the USB3.0 interface to transmit the video stream;
[0156] Wearable devices (bracelets) upload physiological signal data in real time via Bluetooth 5.0 protocol;
[0157] The microphone array inputs voice data through the audio interface, and the electronic whiteboard transmits classroom text records via Wi-Fi.
[0158] Edge computing nodes:
[0159] Deployed on a local server, it receives pre-processed multimodal data via Gigabit Ethernet and runs lightweight models (such as pruned GNN) to generate real-time feedback.
[0160] Interaction between cloud and terminal:
[0161] The desensitization analysis results are uploaded to cloud storage (blockchain evidence) via the HTTPS protocol, and the teacher interface receives real-time policy updates via WebSocket.
[0162] In summary, the multimodal AI-based classroom behavior and emotion collaborative analysis system provided in this embodiment has the following advantages:
[0163] 1. Multi-dimensional behavior-emotion analysis: By integrating multimodal data such as vision, voice, text, and physiological signals, the system achieves comprehensive monitoring of students' classroom behavior (such as concentration and interaction frequency) and emotional state (such as positive and negative), overcoming the limitations of traditional unimodal analysis.
[0164] 2. Dynamic environmental adaptability: A dynamic weight adjustment mechanism is used to adaptively allocate modal weights based on real-time environmental parameters (such as noise level) to improve the analysis robustness in complex classroom scenarios.
[0165] 3. Cross-modal collaborative modeling: Cross-modal fusion technology based on graph neural networks (GNN) captures the dynamic correlation between behavior and emotions, enhancing the refinement and accuracy of joint analysis.
[0166] 4. Low-latency real-time feedback: Combined with edge computing technology, it enables the rapid generation and push of teaching intervention strategies, significantly improving the immediacy and effectiveness of classroom interactions.
[0167] 5. Privacy and security protection: Through federated learning and differential privacy technologies, we ensure the local processing and desensitized transmission of sensitive biometric data to meet education data compliance requirements.
[0168] The workflow of the multimodal AI-based classroom behavior and emotion collaborative analysis system is as follows:
[0169] S1: Data collection:
[0170] Synchronous multimodal data acquisition: Real-time acquisition of multi-source classroom data through infrared cameras (vision), directional microphone arrays (voice), wearable devices (physiological signals) and electronic whiteboards (text).
[0171] S2: Preprocessing:
[0172] Timing alignment and denoising: Use the NTP protocol to calibrate device clock deviations and the Dynamic Time Warping (DTW) algorithm to align non-uniformly sampled data streams, keeping errors to a very low level.
[0173] Data normalization: Perform wavelet denoising and normalization to generate high-quality multimodal input data.
[0174] S3: Feature extraction:
[0175] Visual features: An improved BiGRU network combined with optical flow method is used to extract dynamic features of facial micro-expression sequences;
[0176] Speech features: LSTM network analyzes Mel spectrum to capture changes in intonation and speaking speed;
[0177] Text features: The BERT model extracts semantic sentiment vectors and analyzes the contextual associations of classroom questions and answers;
[0178] Physiological signal characteristics: Wavelet transform is used to extract HRV frequency domain characteristics and GSR time domain slope to quantify emotion-related physiological indicators.
[0179] S4: Collaborative analysis:
[0180] Dynamic weight allocation: adjusts the speech modality weight according to the ambient noise level to suppress interference;
[0181] Cross-modal fusion: GNN is used to model multimodal association graphs, and the attention mechanism is used to dynamically calculate edge weights to generate behavior-emotion joint features.
[0182] S5: Real-time feedback:
[0183] Strategy generation: Based on the GNN output, trigger distraction behavior reminders (such as wearable device vibration) or group emotion intervention (such as activity adjustment suggestions on the teacher side);
[0184] Edge computing optimization: Calculations are completed at local nodes to ensure extremely short response times and meet real-time requirements.
[0185] S6: Privacy protection:
[0186] Federated learning: Physiological signal data is processed only on the local device, preventing the transmission of raw data.
[0187] Differential privacy: Add noise to uploaded analysis results to prevent sensitive information leakage.
[0188] The above are only preferred embodiments of the present invention and do not limit the implementation mode and protection scope of the present invention. For those skilled in the art, it should be aware that all solutions obtained by equivalent substitutions and obvious changes made using the description and illustrations of the present invention should be included in the protection scope of the present invention.
Claims
1. A multimodal AI-based classroom behavior and emotion collaborative analysis system, characterized by: include: Multimodal data acquisition module, used to simultaneously collect visual, voice, physiological signals and text data through infrared cameras, directional microphone arrays, wearable devices and electronic whiteboards; A preprocessing module, performing timestamp synchronization, wavelet denoising and normalization processing on the data; Feature extraction module, including: The visual feature extraction unit uses an improved BiGRU network to process the facial micro-expression sequences generated by the optical flow method; Speech feature extraction unit, which extracts intonation and speech rate features from Mel spectrum based on LSTM network; Text feature extraction unit, which generates semantic sentiment vectors through the BERT model; Physiological signal feature extraction unit, which uses wavelet transform to extract HRV frequency domain features and GSR time domain slope; Collaborative analysis modules, including: Dynamic weight adjustment unit, dynamically assigning weights to each modality based on the ambient noise level; Cross-modal fusion network, using graph neural network GNN to construct multimodal association graph; Real-time feedback module generates behavioral intervention strategies based on edge computing nodes.
2. The multimodal AI-based classroom behavior and emotion collaborative analysis system according to claim 1 is characterized in that: The dynamic weight adjustment unit calculates the speech modality weight using the following formula: Wherein, N is the real-time noise decibel value, N0=60dB, and k=0.1 is the adjustment coefficient.
3. The multimodal AI-based classroom behavior and emotion collaborative analysis system according to claim 1 is characterized in that: In the graph neural network of the cross-modal fusion network: Nodes include visual, speech, text, and physiological signal feature vectors; The edge weight is calculated through the multi-head attention mechanism, the formula is: Among them, w ij represents the attention weight of the i-th element to the j-th element; Q i is the query vector; K j is the key vector; is the transpose of the key vector; d is the feature dimension; It is a scaling factor used to stabilize the gradient; the Softmax function normalizes the dot product result.
4. The multimodal AI-based classroom behavior and emotion collaborative analysis system according to claim 1 is characterized in that: In the physiological signal feature extraction unit: HRV frequency domain characteristics include the ratio of low-frequency power LF to high-frequency power HF; The GSR time domain slope is the average rate of change of the skin conductance rising segment, and the calculation formula is: Where ΔGSR represents the change in skin conductance, and Δt represents the corresponding time change.
5. The multimodal AI-based classroom behavior and emotion collaborative analysis system according to claim 1 is characterized in that: The intervention strategy generation of the real-time feedback module includes: When a student's distracted behavior lasts for more than 3 minutes, the wearable device will vibrate to remind them; When the proportion of negative emotions in the group is greater than 40%, suggestions for adjusting classroom activities will be sent to the teacher.
6. The multimodal AI-based classroom behavior and emotion collaborative analysis system according to claim 1 is characterized in that: The improved BiGRU network introduces an attention pooling operation in the output layer, and the calculation formula is: Among them, a t is the attention weight at the tth time step, h t is the BiGRU hidden state.
7. The multimodal AI-based classroom behavior and emotion collaborative analysis system according to claim 1 is characterized in that: Also includes privacy protection module: Using the federated learning framework, physiological feature data is processed only on the local device; Differential privacy technology is used to add Gaussian noise to the uploaded classroom analysis results.
8. The multimodal AI-based classroom behavior and emotion collaborative analysis system according to claim 1 is characterized in that: The classroom behavior-emotion collaborative analysis method based on the multimodal AI classroom behavior and emotion collaborative analysis system includes the following steps: Step 1: Align multimodal data streams through timestamp synchronization algorithm; Step 2: Use wavelet transform to extract the time-frequency domain features of physiological signals; Step 3: Generate behavior-emotion association graph based on GNN cross-modal fusion network; Step 4: Output teaching optimization strategies in real time through edge computing nodes.
9. The multimodal AI-based classroom behavior and emotion collaborative analysis system according to claim 8 is characterized in that: In the step three: The node feature update formula of the association graph is: Where σ is the ReLU activation function; W (k) is the trainable weight matrix; i (k+1) represents the feature representation of node i at the k+1 layer; N(i) represents the set of neighbor nodes of node i; A ij Represents the adjacency relationship between node i and node j; represents the feature representation of node j at the kth layer.
10. The multimodal AI-based classroom behavior and emotion collaborative analysis system according to claim 8, characterized in that: The implementation of the timestamp synchronization algorithm in step 1 includes: Assign an independent clock source to each type of modal data stream and calibrate the clock deviation of each device through the NTP protocol; Dynamic time warping (DTW) is used to align non-uniformly sampled multimodal data streams, with an alignment error of ≤10ms. Add unified timing labels to the synchronized data streams and generate a time-modality joint index table.
Citation Information
Patent Citations
Emotion assessment method and device based on wearable electrocardio monitoring
CN110353704A
Intelligent violent behavior detection method and device based on multi-modal feature fusion
CN114882409A
Distributed false news detection system based on federal learning
CN116775865A
Financial big data optimization storage method
CN118363961A
Voice-based emotion recognition method, system and device and storage medium
CN118398033A