Classroom attention detection method and system based on multi-modal data fusion
Through multimodal data fusion and edge computing, a lightweight classroom attention detection system is built, which solves the problems of accuracy, privacy and hardware costs in the existing technology, and achieves efficient and accurate classroom attention monitoring.
Patent Information
- Application Number
- CN202510588952.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-19
AI Technical Summary
The existing classroom attention detection technology has problems such as insufficient accuracy, risk of privacy leakage, high hardware costs, and poor scenario adaptability, making it difficult to meet the high real-time needs of teaching management in real-time monitoring scenarios for multiple students.
Multimodal data fusion technology is adopted to build a hybrid evaluation model through non-invasive acquisition and lightweight processing of visual, voiceprints, and physiological signals, combined with dynamic weighting algorithms and edge computing, to reduce the computational complexity and achieve millisecond response, support edge device deployment, and adopt privacy protection measures to avoid data leakage.
It significantly improves the accuracy and universality of attention detection, reduces hardware costs, reduces misjudgment rates and privacy risks, and optimizes the adaptability and resource allocation of teaching scenarios.
Smart Images

Figure CN120508974A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence education technology, and specifically relates to a classroom attention detection method and system based on multimodal data fusion. Background Art
[0002] Current classroom student attention monitoring technologies face multiple challenges. Monitoring methods based on a single data source (such as cameras tracking eye movements or wearable wristbands collecting heart rate data) struggle to accurately distinguish between learning behavior and distraction. A student looking down could indicate note-taking or mobile phone use, leading to false positives in simple visual algorithms exceeding 30%. Physiological signals are easily affected by environmental interference, leading to false alarms. To improve accuracy, a common approach is to use high-precision sensors to continuously collect data, but this leads to privacy and energy consumption challenges. For example, continuous high-definition video recording requires dedicated storage servers, and unmasked facial data poses a risk of misuse. Furthermore, high-power devices impose excessive energy consumption. A more significant limitation is that existing technologies lack dynamic adaptability to teaching scenarios and cannot autonomously identify classroom phases (e.g., lectures, study sessions, and breaks). This can lead to false triggering of alerts during non-teaching periods, seriously disrupting instructional flow. Furthermore, most systems rely heavily on manual intervention: teachers must pre-load class schedules, manually annotate student behavior data to train models, and even wear specialized equipment. This invasive design not only increases the user burden but also easily fosters resistance from both teachers and students, hindering the widespread adoption of this technology. The above defects collectively point to the systematic deficiencies of existing solutions in terms of accuracy, privacy compliance, scenario intelligence and human-machine friendliness.
[0003] Compared to existing technologies, patent publication CN115631074A proposes a visual analysis-based online classroom monitoring solution. This solution uses cameras to collect data on students' eye movements, facial expressions, and body movements, and combines this with an algorithmic model to assess their attention and classroom status. However, this solution has the following limitations: First, it relies on high-precision cameras to continuously capture biometric data, which is susceptible to interference from factors such as ambient light and occlusion, leading to data fluctuation errors. Second, the algorithms used to calculate the eye-corneal ratio and analyze the duty cycle are complex and require significant computing resources, making them difficult to process in real time on low-profile terminals. Furthermore, this solution only models classroom behavior data and lacks comprehensive consideration of multidimensional factors such as students' home environment and independent learning habits. This results in a one-sided state attribution analysis and makes it difficult to accurately identify the root causes of fluctuations in learning outcomes.
[0004] In response to the above-mentioned deficiencies, the present invention has carried out multi-dimensional optimization innovation: First, it adopts non-invasive multi-source data fusion technology, and under the premise of protecting privacy, constructs an attention model through indirect indicators such as device usage logs and interactive response time, thereby reducing dependence on high-sensitivity visual data; second, it designs a lightweight dynamic weight algorithm, which reduces the computational complexity by 40% while introducing an adaptive noise filtering mechanism to improve the robustness of analysis under different terminal environments; third, it innovatively integrates multi-dimensional features such as vision, voiceprint, and physiology to construct a hybrid evaluation model, which significantly improves the accuracy of state analysis and provides cross-scenario data support for the generation of personalized teaching strategies. Through the above-mentioned technological breakthroughs, this solution significantly improves the universality of the system and the pertinence of educational intervention while ensuring the accuracy of evaluation.
[0005] The comparative patent (publication number CN118230259A) discloses a practical teaching management system based on Internet of Things technology, which realizes student action recognition through a dual-path convolution model combined with a relational feature fusion module, and uses a squeeze excitation block and a mobile visual transformer for attention detection.
[0006] However, it still has the following technical defects:
[0007] 1. Insufficient model computational efficiency: The dual-path convolution module needs to run a three-dimensional convolutional network and a fast regional convolutional network simultaneously, resulting in high computational complexity. Especially in the real-time monitoring scenario of multiple students, the system delay is significant, making it difficult to meet the high real-time requirements of teaching management.
[0008] 2. Limited adaptability to complex scenes: Although the relational feature fusion block models the interaction between students and the environment, it relies on linear layers and pooling operations. It is difficult to effectively capture the dynamic relationship between multiple objects in crowded and occluded scenes, and is prone to misjudgment due to the loss of local features.
[0009] 3. Feature redundancy and hardware dependence: Although the cascade structure of the squeeze excitation block and the mobile vision transformer integrates global and local features, the redundant features between channels are not fully filtered. In addition, the model has a large number of parameters and requires high-performance GPU devices, resulting in high deployment costs.
[0010] 4. Privacy and classroom disruption issues: Using cameras to directly monitor student behavior and using light-emitting diodes to remind distracted students poses a risk of privacy leakage, and the interference of light-emitting diodes can easily distract other students and affect the continuity of teaching.
[0011] In response to the above-mentioned shortcomings, this application proposes the following innovative improvements:
[0012] 1. Lightweight dynamic inference architecture: A single-path spatiotemporal fusion network with adaptive pruning replaces dual-path convolution. A dynamic channel selection mechanism reduces computational complexity by 70%, achieving millisecond-level response while ensuring action recognition accuracy, meeting real-time management needs.
[0013] 2. Multimodal relationship reasoning mechanism: Preprocessing and feature extraction of visual, voiceprint, and physiological multimodal data, including: extracting key points of the upper body skeleton, calculating the eye aspect ratio (EAR) and Euler angles of the head; determining the teacher's dominant voice period through voiceprint matching, which is used to dynamically determine the teaching time; performing Butterworth filtering to denoise heart rate variability, and extracting time domain (SDNN) and frequency domain (LF / HF) features; and automatically calculating weights after multimodal data extraction is completed.
[0014] 3. Feature optimization and hardware adaptation: Design a channel-space dual-dimensional compression module to filter redundant information through differentiable feature masks, reducing the number of model parameters by 50%, supporting edge device deployment, and reducing hardware costs.
[0015] 4. Privacy protection and non-intrusive feedback: Vibration prompts on student wristbands replace LED light reminders to avoid classroom disruptions. At the same time, a data encryption module is added to ensure student privacy compliance.
[0016] Current classroom attention detection technology is limited by systemic flaws such as a single data dimension, insufficient scenario adaptability, and an imbalance between privacy and energy consumption. Monitoring solutions that rely on a single modality (such as visual or physiological signals) have difficulty distinguishing between learning behavior and distraction, with misjudgment rates generally exceeding 30%. Furthermore, high-precision sensors continuously collect data, leading to significant privacy risks and equipment energy consumption issues. Existing technologies lack dynamic scene perception capabilities and are unable to autonomously identify teaching phases. The false trigger rate is high during non-teaching periods, seriously disrupting teaching continuity. Furthermore, most systems rely on manually labeled data, pre-set class schedules, and intrusive hardware (such as retrofitting desks with light-emitting diodes), disrupting normal classroom order and placing a heavy burden on teachers and students, hindering the implementation of the technology.
[0017] In the comparative patent (publication number CN118230259A), although the methods based on dual-path convolution or complex visual analysis attempt to improve accuracy, they suffer from poor real-time performance and soaring costs due to computational redundancy (such as the dual modules of three-dimensional convolution + regional convolution), occlusion sensitivity (local feature loss rate exceeds 37%), and high-end hardware dependence (GPU server required). In addition, direct biometric collection and physical reminder design (such as LED lights) further exacerbate the risks of privacy leakage and classroom disruption. According to actual measurements, in a classroom scenario with 50 people, the average processing delay of the dual-path convolution module of the comparative solution is 1.5 seconds, and the GPU memory occupies 8GB. However, the delay of the solution of the present invention is only 0.8 seconds, and the memory occupies 3.2GB. Summary of the Invention
[0018] This invention provides a low-cost, low-privacy-risk, and accurate classroom attention anomaly detection solution that optimizes detection efficiency and resource allocation through multimodal data fusion and dynamic identification of non-teaching time. The core innovations of this invention are:
[0019] (1) Constructing a multimodal spatiotemporal alignment framework: By mapping the three-dimensional space of visual gaze vectors, heart rate variability features, and voiceprint attention orientation, a multimodal attention representation model is established to address the problem that a single signal source is easily affected by environmental interference;
[0020] (2) Design a dynamic weight allocation mechanism: Based on the context-aware technology of teaching scene classification, the confidence weights of each modality are adjusted in real time, which improves the accuracy of attention assessment to 92.3%;
[0021] (3) While performing edge computing, cloud data collaboration is achieved: a model sharding deployment strategy is adopted to migrate more than 50% of the computing load to the edge nodes, reducing the average system response time to 800ms (only 53% of the traditional solution); and a communication interface is reserved on the edge computing device, so that a cloud server can be set up at any time to carry out data collaboration in each classroom.
[0022] (4) Develop a gradient-preserving privacy protection algorithm: Implement differential privacy processing in the feature extraction stage to ensure that the information entropy retention rate of sensitive biometric features exceeds 85%, while meeting the ISO / IEC 24760-3L3 security standard.
[0023] The present invention discloses a classroom attention detection method and system based on multimodal data fusion, comprising:
[0024] In a first aspect, the present invention provides a multimodal fusion classroom attention intelligent monitoring method, comprising:
[0025] Through the multispectral vision unit deployed in the classroom environment, students' facial micro-expression image sequences are collected in real time, and eye gaze vector features are extracted; through the voiceprint perception module of the ring array, voice signals in the teaching scene are synchronously collected, and the voiceprint attention pointing parameters are calculated; heart rate variability characteristics and limb movement data are continuously obtained through wearable biosensors; the above multimodal features are input into the dynamic weight allocation module for spatiotemporal alignment and confidence fusion, and the comprehensive attention evaluation coefficient is output.
[0026] Preferably, the multispectral vision unit adopts a combination of an OV2740 CMOS sensor (1 / 3 inch target surface) and an f / 2.0 fisheye lens, and downsamples the original resolution to 640×480 through a bilinear interpolation algorithm. The specific calculation formula is:
[0027]
[0028] The weight coefficient matrix w(i,j) is generated according to the Gaussian distribution of σ=0.8, and the processing delay is ≤15ms.
[0029] Preferably, the voiceprint sensing module includes four microphones deployed in a ring (with a spacing of 15 cm ± 5%), adopts a generalized sidelobe cancellation algorithm for noise suppression, and the sound source localization is achieved by the time delay summation formula:
[0030]
[0031] Where d = 0.15 m is the array radius, c = 343 m / s is the speed of sound, and the error of the measured azimuth angle estimation is ≤ 3°.
[0032] In a second aspect, the present invention provides a dynamic weight allocation mechanism, including:
[0033] Based on the TextCNN model, the scene classification of the blackboard OCR text is performed to generate the teaching stage probability distribution P=(p lecture ,p discussion ,p self-study ); activate the corresponding weight calculation matrix according to the scene label, strengthen the visual modality weight in the explanation mode, and improve the voiceprint modality contribution in the discussion mode; use the improved DS evidence theory to perform decision-level fusion of multimodal confidence.
[0034] Preferably, the method for constructing the TextCNN model includes:
[0035] 1. Collect historical blackboard text images to build a training set, and generate text sequences through OCR conversion;
[0036] 2. Use the Word2Vec model to generate 300-dimensional word vectors;
[0037] 3. Build parallel convolution layers with kernel sizes of 3, 5, and 7, each with 128 kernels;
[0038] 4. Optimize model parameters through cross-validation until the classification accuracy is ≥ 93%.
[0039] Preferably, the specific formula for calculating the dynamic weight in the explanation mode is:
[0040]
[0041] where θ t is the real-time head yaw angle, μ θ It is the mean of a 10-second sliding window, N=30 (corresponding to a 30fps sampling rate).
[0042] In a third aspect, the present invention provides a method for training a hybrid detection model, comprising:
[0043] A cross-modal encoder consisting of a Transformer encoder and TCN is constructed, and the attention mechanism is implemented at the bottleneck layer to fuse features. A phased training strategy is adopted, first fixing the visual encoder parameters to train the physiological signal branch, and then jointly fine-tuning all parameters. Distributed model updates are achieved through a federated learning framework.
[0044] Preferably, the initialization method of the Transformer encoder includes:
[0045] 1. Set the patch size to 16×16, the number of heads h=8, and the feedforward layer dimension to 2048;
[0046] 2. Based on the ImageNet pre-trained weights, fine-tune using the classroom gaze dataset;
[0047] 3. Compress the ResNet-50 model to <4.5M parameters through knowledge distillation.
[0048] Preferably, the implementation steps of the federated learning are:
[0049] 1. Edge nodes calculate local gradients
[0050] 2. Apply gradient Gaussian noise;
[0051] 3. Cloud-based security aggregation:
[0052]
[0053] Wherein the temperature coefficient τ=0.1, σ is the sigmoid function.
[0054] In a fourth aspect, the present invention provides an edge computing deployment solution, including:
[0055] The TensorRT acceleration engine was deployed on the Jetson Xavier NX hardware platform, and FP16 quantization of the model was implemented. A dynamic degradation mechanism was established to disable the HRV frequency domain feature calculation branch when the GPU utilization rate exceeded 85%. The visual sampling rate was reduced from 30fps to 15fps when the temperature exceeded 75°C.
[0056] Preferably, the implementation method of the FP16 quantization includes:
[0057] 1) Statistical model weight distribution histogram to determine the dynamic range;
[0058] 2) Implement symmetric quantization on convolutional layer weights:
[0059]
[0060] 3) Insert a calibration layer to compensate for the accuracy loss, ensuring that the classification accuracy drops by <1.2%.
[0061] The beneficial effects of the present invention are: proposing a multimodal lightweight fusion architecture, constructing an attention model through non-invasive multi-source data, combining adaptive pruning algorithm and edge computing optimization, achieving millisecond-level response while reducing the amount of computation; innovatively introducing graph convolution relational reasoning and differential privacy encryption technology, breaking through the bottleneck of occlusion scene modeling (misjudgment rate reduced by 42%) and ensuring data anonymization processing; designing a scene self-perception module, dynamically identifying the teaching stage and adjusting the detection strategy, reducing the invalid alarm rate to below 12%, and finally forming a classroom attention detection method and system based on multimodal data fusion with coordinated optimization of accuracy (accuracy +25%), privacy compliance (data desensitization ≥98%) and deployment economy (hardware cost reduced by 60%). BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 : The flow chart of the present invention shows the six key steps of the system operation;
[0063] Figure 2 : Schematic diagram of the multimodal data preprocessing and feature extraction process of the present invention, which includes four key modules: multi-source data spatiotemporal alignment, visual feature processing, voiceprint feature extraction, and physiological signal processing;
[0064] Figure 3 : A schematic diagram of the multimodal fusion deep learning model flow of the present invention, which includes three key modules: visual spatiotemporal encoder construction, cross-modal feature fusion, and dynamic weight allocation;
[0065] Figure 4 : The dynamic weight allocation algorithm flow chart of the present invention includes five submodules: feature splicing, gate value calculation, dynamic weight generation, scene adaptation enhancement, and real-time confidence correction;
[0066] Figure 5 : The system architecture diagram of the present invention shows the connection relationship between the data acquisition module, edge computing module, and dynamic feedback module;
[0067] Figure 6 : The algorithm flow diagram of Example 1 shows the execution process of the method for detecting abnormal attention of classroom students according to the present invention;
[0068] Figure 7 : The system architecture diagram of Example 2 shows the architecture of each hardware in the system. DETAILED DESCRIPTION
[0069] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0070] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0071] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.
[0072] Example 1: Classroom attention detection method based on multimodal fusion
[0073] like Figure 6 As shown, the present invention provides a multimodal classroom attention anomaly detection method, and the specific implementation steps are as follows:
[0074] S100: Synchronous acquisition of multi-source data
[0075] The sensor network deployed in the teaching space collects three types of data in real time:
[0076] (1) Visual data collection:
[0077] Two sets of wide-angle cameras (Sony IMX586, 120° field of view) are installed at the front and back diagonals of the classroom to continuously capture students' upper body postures at a resolution of 2560×1440 and a frame rate of 30fps.
[0078] In the exemplary implementation, each camera is equipped with an infrared fill light module (wavelength 850nm, illumination 0.1lux) to ensure effective imaging in low-light environments.
[0079] (2) Physiological signal collection:
[0080] Students wore a custom smart bracelet (integrated with a MAX30102 PPG sensor and an LSM6DS3 inertial unit) that sampled heart rate variability at 1 Hz and three-dimensional acceleration data at 50 Hz;
[0081] In specific implementation, the bracelet establishes a connection with the edge computing node through Bluetooth 5.0, and the transmission delay is controlled within 80ms.
[0082] (3) Audio data collection:
[0083] A ring microphone array (ReSpeaker 6, 6 MEMS microphones) was used to capture the teacher’s speech signal at a sampling rate of 16kHz;
[0084] Implementation details: The beamforming algorithm focuses on the teacher's voice source, improving the signal-to-noise ratio to over 15dB.
[0085] S110: Implementation of space-time synchronization mechanism
[0086] The specific steps include:
[0087] (1) Hardware-level synchronization:
[0088] (1) Xilinx Artix-7 FPGA is used to generate a 10MHz reference clock signal and distribute it to each sensor through the LVDS interface
[0089] (2) Establish a unified timeline: visual data is marked every 33ms, audio is synchronized every 512 sampling points (32ms), and physiological signals are aligned every 1000ms
[0090] (2) Software compensation mechanism, using computer language to implement the following algorithm:
[0091] The first step is to calculate the maximum transmission delay t max :
[0092] t max =max(t camera ,t audio ,t biosensor )
[0093] t camera is the timestamp of the visual sensor data, t audio is the timestamp of the sound sensor data; t biosensor The timestamp of the smart bracelet data.
[0094] The second step is to calculate and apply linear interpolation compensation:
[0095]
[0096] Where: α=0.78 (experience compensation coefficient). k is the original timestamp (k∈{c,a,b}), t max is the maximum transmission delay, t′ k The synchronization timestamp after compensation.
[0097] The maximum delay is calculated using a three-dimensional extreme value function, and the synchronization process performs linear interpolation compensation on each data channel.
[0098] In actual tests, the multimodal data synchronization error is controlled within ±12.3ms (3σ confidence interval).
[0099] S120: Multimodal Feature Extraction
[0100] (3) Visual feature extraction:
[0101] The improved OpenPose algorithm is used to detect the students' eye gaze vector features and calculate the blink frequency and gaze direction;
[0102] Key parameter settings:
[0103] 1) The eye aspect ratio threshold is set to 0.22;
[0104] 2) The horizontal gaze calibration range interval is set to (15°, 30°).
[0105] Implementation results: Single-frame image processing takes 42ms (Jetson Xavier NX platform)
[0106] (IV) Physiological signal processing:
[0107] 1) Heart rate variability analysis:
[0108] Extract the LF / HF band power ratio as an attention indicator (threshold setting: focused state > 1.2, distracted state < 0.8)
[0109] 2) Motion feature detection:
[0110] Sliding window RMS algorithm (window length 2s) is used to detect abnormal body movement frequency
[0111] (5) Audio feature processing:
[0112] Implement the voiceprint matching process:
[0113] 1) Extract the teacher's speech MFCC features (26-dimensional coefficients);
[0114] 2) Use the CUDA-accelerated DTW algorithm for real-time matching (matching a 5-second speech clip takes 76ms).
[0115] S130: Dynamic Weight Fusion Modeling
[0116] Specific implementation includes:
[0117] 1. Build a hierarchical Transformer model:
[0118] Network layer Parameter configuration Output dimension Spatial coding layer Number of heads = 4, d_model = 256 256 Temporal Convolutional Layer Kernel size = 5, dilation factor = 2 128
[0119] 2. Dynamic weight adjustment implementation:
[0120]
[0121] W v ,W a ,W p ∈[0,1] represents the normalized weights of visual, audio, and physiological signals, respectively, satisfying W v +W a +W p =1.
[0122] phase: teaching phase classification variable (discrete enumeration type), Phase∈{lecture, discussion, self-study, break, dismissal}.
[0123] Implementation effect: The F1-score of the model validation set reached 0.891 (±0.023).
[0124] S140: Anomaly Detection and Feedback
[0125] (1) Implementation of the three-level early warning mechanism:
[0126] Based on the detection results, the system adopts a three-level early warning mechanism for early warning. The early warning measures of each level of the early warning mechanism are as follows:
[0127]
[0128] (2) Privacy protection implementation:
[0129] Visual data desensitization: retain the eye area (100×100 pixels), and apply σ=3.0 Gaussian blur to the rest of the area;
[0130] Federated learning framework: 10 edge nodes aggregate the model every 24 hours, with differential privacy parameter ε = 0.5.
[0131] Example 2: Classroom Attention Detection System
[0132] like Figure 7 As shown, this system includes the following modules:
[0133] (1) Data acquisition module (data acquisition layer):
[0134] The data acquisition module consists of the following three units:
[0135] 1) Visual sensor unit: Contains two sets of wide-angle cameras and infrared fill light components;
[0136] 2) Audio frequency sensor unit: ring microphone array and FPGA synchronization controller;
[0137] 3) Physiological sensing unit (wearable bracelet): The smart bracelet has built-in PPG (photoplethysmography) and acceleration sensor.
[0138] The visual sensor unit and audio sensor unit can be connected to the router in the edge computing module through network cable and Wi-Fi modes; the physiological sensor unit (wearable bracelet) can be connected to the router in the edge computing module through Wi-Fi and Bluetooth modes.
[0139] (2) Edge computing module (data processing layer):
[0140] The edge computing module includes edge computing devices and routers. The edge computing devices are equipped with NVIDIA Jetson Xavier NX and run a fusion model accelerated by TensorRT. Resource monitoring shows that during the teaching period, the average CPU usage is 38.2%, and the peak memory usage is 1.8GB.
[0141] The router must support Bluetooth 5.2 (BLE) and WiFi 6 (802.11ax) dual-mode transmission. Huawei AX3pro, Xiaomi Router AX9000 and other higher-performance products are acceptable.
[0142] (3) Interaction module (data feedback layer):
[0143] The interactive module is mainly divided into two parts:
[0144] 1) Teacher terminal display: 10.1-inch touch screen real-time rendering of attention heat map;
[0145] 2) Student wearable smart bracelet: After receiving the reminder data transmitted by the edge computing module, it vibrates to remind;
[0146] In the data interaction module, the data update frequency is 1Hz and the positioning accuracy is ±0.5m.
[0147] The teacher's terminal can be connected to the router via network cable or Wi-Fi; the student's wearable smart bracelet can be connected to the router via Wi-Fi or Bluetooth.
[0148] After 72 hours of continuous testing in a real classroom environment, the test results are as follows:
[0149] 1. Detection performance comparison:
[0150] index This system Traditional single-modal solution Accuracy 93.1% 67.2% False alarm rate 2.3% 31.7% Response delay 760ms 1420ms
[0151] 2. Resource consumption test:
[0152] Network bandwidth usage: peak during teaching hours: 18.7 Mbps, and during idle hours: 1.3 Mbps.
[0153] Power consumption: The average power consumption of edge nodes is 11.3W, meeting energy-saving requirements.
[0154] The above implementations enable those skilled in the art to fully implement the present invention, and adjustments to the relevant parameters within a ±20% range remain within the scope of protection of the present invention. The core innovation of the present invention lies in significantly improving the accuracy and practicality of classroom attention detection through multimodal data fusion and a dynamic weighting mechanism.
Claims
1. A classroom attention detection method based on multimodal data fusion, characterized in that: The following steps are involved: Step S1: Collecting multimodal student data in the classroom scene through visual sensors, audio sensors, and a wearable wristband equipped with a PPG (photoplethysmography) sensor and an accelerometer, including facial expressions, eye movements, head posture, teacher voice signals, and heart rate variability; Synchronously collect teacher teaching behavior data, including voice waveforms, limb skeleton key points, and blackboard content. Teacher teaching behavior data is obtained synchronously only through audio sensors and visual sensors, without the need for additional equipment; Step S2: Preprocessing and feature extraction of the multimodal data, including: extracting upper body skeletal key points based on a lightweight OpenPose model, calculating the eye aspect ratio (EAR) and head Euler angles; determining the teacher's dominant voice period through voiceprint matching using MFCC feature extraction and dynamic time warping (DTW) algorithm, for dynamic determination of teaching time; performing Butterworth filtering to denoise heart rate variability, and extracting time domain (SDNN) and frequency domain (LF / HF) features; the lightweight OpenPose model is a lightweight OpenPose model based on the MobileNetV3 backbone network, with an input resolution of 368×368 pixels and output of 23 upper body key point coordinates; Step S3: Construct a multimodal fusion model consisting of a Transformer encoder and a bidirectional GRU network. The following operations are performed: Use the Transformer encoder to perform spatiotemporal modeling of visual features; Use the bidirectional GRU network to fuse physiological and behavioral temporal features; Dynamically adjust the contribution weight of each modality based on the scene label output by the detector during the classroom phase; Step S4: Based on the fused feature vector, an autoencoder-support vector machine (AE-SVM) hybrid model is used to detect attention anomalies. When the reconstruction error exceeds a dynamic threshold calculated based on sliding window statistics, an early warning is triggered. Step S5: Privacy protection measures are implemented. For visual data, only the eye and hand regions (ROIs) are retained, and the background area is Gaussian blurred. After extracting MFCC features from audio data, the original speech waveform is immediately deleted. The model parameters are updated using a federated learning framework, and local data transmission is prohibited. Step S6: Feedback the anomaly detection results to the teacher terminal and the academic affairs terminal in real time to generate a visual heat map and teaching strategy recommendations.
2. The method according to claim 1, characterized in that The multimodal data preprocessing and feature extraction in step S2 are performed according to the following rules: Step S201: Spatiotemporal alignment of multi-source data: A hardware-triggered synchronization mechanism is used to generate a 10 MHz synchronous clock signal through the FPGA to align the time domain data streams of the visual sensor (30 fps), audio sensor (16 kHz), and wearable wristband (1 Hz), and establish a unified timestamp system. The synchronous clock signal meets the following conditions under the classroom environment conditions of -20-50°C and 40-70% RH: Δt sync =max(|t vid -t aud |,|t aud -t bio |,|t vid -t bio |)≤15ms Step S202: Visual feature processing: A lightweight OpenPose model based on the MobileNetV3 backbone network extracts 23 skeletal key points of the upper body and calculates the eye aspect ratio (EAR): (where p1-p6 are the coordinates of the key points of the eyes) and calculate the Euler angles of the head: Step S203: Voiceprint feature extraction: Extract MFCC coefficients using a Mel filter bank (40 triangular filters, frequency range 80-8000 Hz), retaining the first 13-dimensional static features + 13-dimensional Δ dynamic features; use an improved dynamic time warping (DTW) algorithm to calculate the matching degree between the student's voice segment and the teacher's voiceprint template: The slope constraint is: |ij|≤0.2·max(i,j); Step S204: Physiological signal processing: Apply fifth-order Butterworth bandpass filtering (cutoff frequency 0.04-0.4 Hz) to the heart rate variability (HRV) and calculate the SDNN: Calculate time domain features:
3. The method according to claim 2, wherein: The lightweight OpenPose model based on the MobileNetV3 backbone network is compressed using channel pruning technology to meet the following requirements: input resolution: 368×368 pixels, model parameter count <4.5M, and inference speed ≥25FPS (NVIDIA Jetson Nano platform); The improved DTW algorithm is implemented through CUDA acceleration, and its computational efficiency on the GTX 4090 card is: the matching time of 10-second voice clips is less than 80ms, and the voiceprint false recognition rate (FAR) is ≤1.2% (under the condition of SNR>10dB). This computational efficiency is obtained through benchmark testing in the CUDA12.4 environment, with a test sample size of 1000 groups of 10-second voice clips. The Butterworth filter adopts a zero-phase filter design, and the specific implementation steps include: a) Forward filtering: b) Reverse filtering: c) Output: y[n] = y2[N-n+1] where b k is the forward coefficient, a k is the reverse coefficient.
4. The method according to claim 1, wherein The multimodal fusion deep learning model in step S3 operates according to the following rules: Step S301: Visual spatiotemporal encoder construction: A layered Transformer architecture is used to process visual features, including: Spatial encoding layer: 4-head self-attention mechanism processes the spatial relationship of skeleton key points: Temporal coding layer: Sliding window TCN (temporal convolutional network) extracts head posture temporal features, convolution kernel width w = 5, expansion factor d = 2 i . Step S302: Cross-modal feature fusion: a) Physiological-behavioral temporal fusion: A bidirectional GRU network is used to process HRV and head motion features, with a hidden layer dimension of h gru =128, gate calculation formula: z t =σ(W z [h t-1 ,x t ]) r t =σ(W r [h t-1 ,x t ]) b) Voiceprint Attention Weighting: The cross-attention mechanism is used to associate the student’s voice features with the teacher’s voiceprint template. (s is cosine similarity) Step S303: Dynamic weight allocation module, based on the scene label C∈{lecture, discussion, self-study, break, dismissal} output by the classroom stage detector, this module uses a dynamic weight allocation algorithm to perform dynamic weight allocation.
5. The method according to claim 4, characterized in that The dynamic weight allocation algorithm generates modal weights through a gating network. The specific execution steps are as follows: Step S3031: Feature splicing, visual feature f visual , audio features f audio , physiological characteristics physio and the embedding vector e of the scene label c To splice: Step S3032: Calculation of gate value: G=ReLU(W g ·f concat +b g ) Where W g ∈R^(3×d) is the trainable parameter matrix, and d is the sum of the feature dimensions. g is the trainable bias vector. Step S3033: Dynamic weight generation: γ = 1.2 is the temperature coefficient, which is used to adjust the steepness of the weight distribution. Step S3034: scene adaptation enhancement: The MLP is a two-layer fully connected network, and α = 0.3 controls the intensity of scene knowledge injection. Step S3035: Real-time confidence correction: η=0.6 is the dynamic adjustment factor, SNR m is the signal-to-noise ratio of each mode.
6. The method according to claim 4, wherein: The layered Transformer architecture is trained using a curriculum learning strategy: The first stage: fix the spatial coding layer and only train the temporal coding layer (lr = 1e-4); Second stage: Joint fine-tuning of all parameters (lr = 5e-6); Loss function: L = 0.6L cls +0.4L recon ; The bidirectional GRU network introduces a residual connection mechanism to meet the following requirements: And the hidden layer dimension h gru Satisfy the constraints: h gru ≥2dim(f visual ). The classroom stage detector runs the CRNN text recognition network through the edge computing device, performs real-time detection and recognition of the blackboard area, generates a text sequence and inputs it into the TextCNN model for classification, including an embedding layer (300-dimensional GloVe word vector), parallel convolution kernels (size 3 / 5 / 7), and a classification head output scene probability distribution.
7. The method according to claim 1, wherein: The dynamic weight allocation module in step S3 operates according to the following rules: during the teacher's explanation phase: visual weight 0.6, audio weight 0.3, physiological weight 0.1; during the group discussion phase: visual weight 0.4, audio weight 0.5, physiological weight 0.1; The threshold adjustment formula is: Among them, Confidence i is the real-time confidence score of the i-th modal data. The specific calculation method is: The dynamic threshold setting method in step S4 includes: calculating the reconstruction error mean μ and standard deviation σ based on historical normal data, dynamic threshold = μ + 3σ; setting the threshold according to the class type: theoretical class threshold = μ + 2σ, experimental class threshold = μ + 3σ; updating μ and σ every 10 minutes to achieve adaptive adjustment of the threshold.
8. The method according to claim 1, characterized in that The visual sensor in step S1 is a wide-angle camera, and its data acquisition strategy includes: during the teaching period: 1080P resolution, 30fps frame rate; during the non-teaching period: switching to 480P resolution, 5fps frame rate, and performing real-time mosaic processing on the video stream.
9. A classroom attention detection system based on multimodal data fusion that implements the method of any one of claims 1 to 8, characterized in that: The system consists of three modules: data acquisition, edge computing, and dynamic feedback. The data acquisition module integrates a wide-angle camera, an omnidirectional microphone array, and a low-power wristband. It transmits data via Bluetooth and WiFi, supporting dual-mode transmission via Bluetooth 5.2 (BLE) and WiFi 6 (802.11ax), with a dynamically adjustable transmission bandwidth range of 20-160MHz. The edge computing module deploys a lightweight deep learning model (MobileNetV3+GRU) and supports TensorRT acceleration, with an end-to-end latency of ≤800ms. The dynamic feedback module includes a visual dashboard on the teacher side (real-time heat map, abnormal student location) and a non-invasive reminder device on the student side (wearable wristband vibration). The model optimization strategy adopted by the edge computing module is: using knowledge distillation technology to compress the ResNet-50 teacher model into the MobileNetV3 student model, reducing the number of parameters by 82%; based on the dynamic inference mechanism of channel pruning, the secondary branch calculation is automatically unloaded according to the GPU load. The working modes of the data acquisition module include a teaching time mode and a non-teaching time mode. The characteristics of the two modes are: the teaching time mode starts full-function data acquisition, and the transmission bandwidth occupies ≤20Mbps; the non-teaching time mode only collects key sensor data, and the transmission bandwidth occupies ≤2Mbps. The abnormal warning strategy of the dynamic feedback module includes the following warning classification method and warning action: when the first-level warning (mild distraction) occurs, the student's wristband vibrates at a low frequency (1 time / minute); when the second-level warning (abnormal attention) occurs, the wristband vibrates at a high frequency (3 times / minute) and a red dot mark is superimposed on the teacher's terminal; When the third-level warning (group attention loss) occurs, the teacher's terminal will pop up teaching strategy suggestions (such as re-teaching of knowledge points and interactive Q&A). The system also includes a scene adaptation module, which is used to predict the teaching stage based on the LSTM-TCN hybrid model, making the detection accuracy ≥ 94%; it includes dynamic adjustment of sensor sampling frequency and model computing resource allocation, reducing system power consumption by 60%.
Citation Information
Patent Citations
Information-based network science and education method, system and equipment
CN115631074A
Practical teaching management system based on Internet of Things technology
CN118230259A
Cited By
Classroom abnormity early warning method, system and device based on AI and storage medium
CN120612214A
Classroom multi-modal data processing method and system based on deep learning
CN120850228A
Classroom state recognition method and system based on multi-modal pre-screening and visual fine tuning
CN121259701A
Control method and artificial intelligence experiment system
CN121354556A
Optical character recognition system and method
CN121459365A