Student classroom behavior detection method based on Transform and improved VGG16 model
Through the improved VGG16 network and spatiotemporal Transformer module, combined with the dynamic threshold mechanism, the accuracy and real-time problems of classroom behavior detection in the existing technology are solved, and high-precision and robust classroom behavior detection are achieved.
Patent Information
- Application Number
- CN202510344950.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing student classroom behavior detection methods have shortcomings in spatial feature extraction and timing modeling, and it is difficult to achieve high accuracy and real-time performance in complex classroom environments, and the error detection rate is high.
The improved VGG16 network is used in combination with the spatiotemporal Transformer module to improve spatial feature extraction capabilities through channel attention mechanism and multi-scale feature pyramid, and optimize behavior classification with dynamic threshold mechanism.
It significantly improves the accuracy and robustness of classroom behavior detection, reduces the missed detection rate and false detection rate, and meets the needs of real-time classroom monitoring.
Smart Images

Figure CN120236239A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of student classroom behavior detection, and particularly relates to a student classroom behavior detection method based on a Transformer and an improved VGG16 model, which is applicable to the real-time detection and analysis of student behaviors in an educational scenario. Background Art
[0002] With the rapid development of artificial intelligence, the automatic detection technology of student classroom behaviors has become the key to improving teaching quality and student management efficiency. At present, many researchers have applied deep learning to student classroom behavior detection. Existing classroom behavior detection methods mostly adopt time series models such as traditional convolutional neural networks (CNNs) or long short-term memory networks (LSTMs). Although these methods can capture the behavior characteristics of students to a certain extent, there are still significant technical bottlenecks. First, traditional CNNs have deficiencies in spatial feature extraction. Their convolution operations mainly rely on local receptive fields and do not fully consider the correlation between channels, resulting in poor robustness to complex scenarios such as illumination changes and occlusions. Second, although LSTM has certain advantages in time series modeling, it is difficult to effectively capture long-distance time-dependent relationships, and due to the high computational complexity of the recursive structure, it is difficult to meet the real-time requirements. In addition, existing methods have poor adaptability in complex classroom environments (such as multi-view, dynamic background, multi-person interaction, etc.) and a high false detection rate, making it difficult to meet the actual application requirements. Summary of the Invention
[0003] In view of the above problems, the present invention proposes a student classroom behavior detection method based on spatio-temporal Transformer and improved VGG16. By combining the improved VGG16 network and the spatio-temporal Transformer module, while improving the detection accuracy of small targets, the missed detection rate in high-occlusion scenarios is reduced, providing an efficient solution for real-time classroom behavior monitoring.
[0004] The method steps are as follows: S1. Collect video data through a classroom monitoring camera, and perform preprocessing on video frames, including denoising, illumination equalization, and key frame extraction; S2. Construct an improved VGG16 network as a spatial feature extraction module, and optimize the original VGG16 structure by introducing a channel attention mechanism and dilated convolution; S3. Design a spatio-temporal Transformer module to extract the time series features of consecutive video frames through a multi-head self-attention mechanism; S4. Fuse spatial features and spatio-temporal features, and construct a multi-scale feature pyramid for behavior classification; S5. Adopt a dynamic threshold mechanism to output the behavior detection results, and classify four types of classroom behaviors: "focused", "distracted", "interactive", and "out of seat".
[0005] Preferably, the specific steps of S1 include the following: S1-1, Video data collection and standardization: Collect the original video data through dual cameras (resolution 1920×1080, frame rate 25fps) deployed in the front and back of the classroom, uniformly convert the video to the MP4 format and store it in the local server to ensure data format compatibility and integrity; S1-2, Video frame preprocessing: Use the non-local means filtering (NLM) algorithm to denoise the video frames. The formula is as follows: In the formula, Ω is the neighborhood window, h is the smoothing parameter (default value is 10); C(x) is the normalization coefficient; Subsequently, use the CLAHE algorithm (block size 8×8, contrast limit 2.0) for illumination equalization to solve the problem of uneven illumination; S1-3, Key frame optimization and annotation: Calculate the motion energy of adjacent frames based on the frame difference method. The calculation formula is as follows: Select the key frames with energy higher than the threshold (τ = 15%·E max ), perform data augmentation on them (random rotation ±10°, horizontal flipping, brightness adjustment ±20%), and finally manually annotate them into four types of classroom behaviors (concentrated, distracted, interactive, leaving the seat), and divide the training set, validation set, and test set according to the ratio of 7:2:1.
[0006] Preferably, the specific steps of the improved VGG16 network of S2 include the following: S2-1, Replace the third convolutional block of the original VGG16 with a dilated convolutional layer, and set the dilation rate to 2. The resolution calculation of the output feature map of the dilated convolution is: In the formula, I is the input size, p is the padding number, d = 2 is the dilation rate, k is the convolutional kernel size, and s is the stride; S2-2, Insert a channel attention module after the fourth convolutional block, generate channel weights through global average pooling, and the weight calculation is: w c = σ(W1δ(W2GAP(F c ))), where Fc is the feature map of the c-th channel, GAP is the global average pooling, W1 and W2 are learnable parameters, δ is the ReLU activation function, and σ is the Sigmoid activation function; S2-3, Modify the fully connected layer to a 1×1 convolutional layer, and the output feature map size is: H out = H in , W out = W in , to retain the spatial information and improve the calculation efficiency.
[0007] Preferably, the spatio-temporal Transformer module of S3 includes the following: S3-1. Flatten the feature maps of consecutive T frames into a spatio-temporal sequence for input, with the input dimension being T×H×W×C; S3-2. Inject spatio-temporal position information through position encoding, and the encoding formula is: where pos is the position index, and d model = 512 is the feature dimension; S3-3. Design alternating spatio-temporal attention heads, and the self-attention calculation is: where Q, K, and V are the query, key, and value matrices respectively, and d k = 64 is the key vector dimension.
[0008] Preferably, the multi-scale feature fusion of S4 specifically includes: S4-1. Fuse features of different levels through a Feature Pyramid Network (FPN), and the feature fusion formula is as follows: P i = Conv(C i ) + Upsample(P i +1), where C i is the feature of the i-th layer, and P i is the fused feature; S4-2. Use a Bidirectional Feature Pyramid Network (BiFPN) to enhance the feature transfer efficiency, and the weight calculation is: where α i is the learnable parameter of the i-th layer, and N = 3 is the number of feature layers; S4-3. Adopt an adaptive feature fusion mechanism to dynamically adjust the weights of features at each level to enhance the robustness of the model to different scale changes.
[0009] Preferably, the dynamic threshold mechanism of S5 specifically includes: S5-1. Dynamically adjust the classification threshold according to the complexity of the classroom environment, and the threshold update formula is as follows: where θ base = 0.5 is the base threshold, λ = 0.3 is the adjustment coefficient, and NoiseLevel is the environmental noise level; S5-2. Adopt a sliding window mechanism to smooth the behavior detection results of consecutive frames, and the average confidence in the window is: where W = 5 is the window size, and s i is the confidence score of the i-th frame; S5-3. Introduce a confidence scoring mechanism to filter out detection results with low confidence. The scoring formula is as follows: s = max(Softmax(f fusion )) where f fusion is the fused feature vector. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 This is a flowchart of the classroom behavior detection method based on spatio-temporal Transformer and improved VGG16 according to the present invention.
[0011] Figure 2 This is a structural diagram of the improved VGG16 network according to the present invention.
[0012] Figure 3 This is a schematic diagram of the spatio-temporal Transformer module according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] The technical solutions of the present invention will be described in detail below with reference to the accompanying drawings, so that those skilled in the art can implement them according to the description in the specification.
[0014] The first stage: data collection and preprocessing. In this stage, the present invention collects classroom video data by deploying a dual-camera system (a panoramic camera and a close-up camera). The panoramic camera is installed at a height of 3.5 meters on the back wall of the classroom, with a depression angle of 15°, covering a range of 8 meters × 6 meters; the close-up camera is installed on both sides of the podium, with a height of 2.2 meters and a horizontal viewing angle of 60°, supporting automatic tracking of the teacher's position. The two cameras achieve clock synchronization through the IEEE 1588 Precision Time Protocol (PTP), with a time error ≤ 1 ms. The video stream is transmitted to the server through the RTSP protocol and stored as an MP4 file encoded in H.264. To extract key frames, an improved ViBe algorithm is used for motion saliency detection. In the initialization stage, the background is modeled by a Gaussian mixture model (GMM), and the background model is updated every 10 frames, with a learning rate set to 0.05. The background probability density function is:
[0014]
[0015] The motion energy is evaluated by calculating the sum of absolute differences (SAD) between consecutive frames:
[0016]
[0017] The dynamic threshold is set to the average motion energy of the first 30 seconds of the video plus 3 times the standard deviation When the inter-frame difference exceeds the threshold, it is marked as a key frame, and five frames before and after are intercepted as a segment. To eliminate the influence of uneven illumination, the CLAHE algorithm is used for illumination equalization. The block size is 8×8, the contrast limit is 2.0, and the histogram clipping threshold is calculated as:
[0018]
[0019] At the same time, non-local means denoising is used to denoise the video frames. The search window is 21×21, the similar block is 7×7, and the filtering parameter h = 10.
[0020] The second stage: construction and training of the improved VGG16 network. In this stage, based on the key frame segments extracted in the first stage, the present invention constructs an improved VGG16 network as the spatial feature extraction module. First, the third convolutional block of the original VGG16 is replaced with a dilated convolutional layer with a dilation rate d = 2, and the output feature map size is calculated by the formula to ensure that the resolution remains unchanged:
[0021] In the formula, the input size I = 112, the padding p = 2, the convolutional kernel size k = 3, the stride s = 1, and the calculated output size O = 112. A channel attention module is inserted after the fourth convolutional block. The channel descriptor is generated through global average pooling, and the channel weight is generated after dimensionality reduction and dimensionality increase through a fully connected layer. The weight is calculated as:
[0022] w c = σ(W1δ(W2GAP(F c ))), where W1 ∈ R 512×32 、W2 ∈ R 32×512 , δ is the ReLU activation function, and σ is the Sigmoid function. During the training process, the parameters of the first two convolutional blocks of VGG16 are frozen, and only conv3 and subsequent layers are fine-tuned. The dataset is jointly trained using the self-built classroom dataset (200 hours of video, 120,000 annotated segments) and Kinetics-400. Data augmentation includes random horizontal flipping, rotation, scaling, as well as temporal frame extraction and time warping. The optimizer uses AdamW, the initial learning rate is 0.001, the weight decay is 0.01, and the learning rate scheduling uses the cosine annealing strategy, with the minimum learning rate being 0.0001. The formula is:
[0023]
[0024] The focal loss is used as the loss function, with parameters α = 0.25 and γ = 2, defined as:
[0025] FL(p t ) = -a(1 - p t )γ log(p t )
[0026] The third stage: implementation of the spatio-temporal Transformer module. In this stage, the present invention designs a spatio-temporal Transformer module to extract the temporal features of consecutive video frames. First, the 16 key-frame features (14×14×512 per frame) extracted in the second stage are flattened into a spatio-temporal sequence input with a dimension of 16×196×512. Position encoding is generated for each spatial position through sine-cosine encoding, and the encoding formula is:
[0027]
[0028] And combined with learnable temporal position encoding, spatio-temporal position information is injected into the feature vector. The spatio-temporal Transformer adopts a hierarchical structure, alternately calculating self-attention in the spatial and temporal dimensions. The spatial attention heads focus on the relationships between 14×14 spatial positions within a single frame, and the temporal attention heads capture the temporal dependencies across 16 frames at the same spatial position. In the self-attention calculation, the query matrix Q, the key matrix K, and the value matrix V are generated through linear projection, with a dimension of 64 for each head. The attention output is fused through concatenation and a linear layer, and the calculation formula is:
[0029] In the formula, d k = 64. After residual connection, it is normalized by LayerNorm, and finally the spatio-temporal fusion features are output.
[0030] The fourth stage: multi-scale feature fusion. In this stage, the present invention realizes multi-scale feature fusion through a bidirectional feature pyramid (BiFPN). First, the features of the conv3, conv4, and conv5 layers of the improved VGG16 (with resolutions of 28×28, 14×14, and 7×7 respectively) are input into the BiFPN. In the top-down path, the high-level features are upsampled by a factor of 2 through bilinear interpolation and added to the mid-level features, and the mid-level features are upsampled and added to the low-level features; in the bottom-up path, the low-level features are downsampled by a factor of 2 through max pooling and added to the mid-level features, and the mid-level features are downsampled and added to the high-level features. Learnable parameters α1, α2, α3 are introduced at each fusion node to calculate the weights of the features at each level through fast normalization:
[0030]
[0031] The weighted fusion formula is:
[0032]
[0033] Finally, the fused features are output through a 1×1 convolution, retaining the spatial information and improving the computational efficiency.
[0034] Phase 5: Dynamic Threshold Classification and Deployment. In this phase, the present invention introduces a dynamic threshold mechanism to enhance the robustness of behavior detection. First, the proportion of the foreground motion area is extracted through background modeling to evaluate the environmental noise level:
[0035]
[0036] The classification threshold is dynamically adjusted according to the noise level. The basic thresholds are set as focus (θ base = 0.7), distraction (0.6), interaction (0.65), leaving the seat (0.75), the adjustment coefficient λ = 0.2, and the threshold update formula is:
[0037]
[0038] The sliding window mechanism is adopted to smooth the detection results of consecutive frames. The window size is 5 frames, and the average confidence within the window is calculated:
[0039]
[0040] To meet the edge deployment requirements, the teacher model (improved VGG16 + spatio-temporal Transformer) is compressed into a student model (MobileNetV3-Small) through knowledge distillation. The distillation loss function is the weighted sum of the focal loss and the mean square error:
[0041] L KD = 0.7·L CE + 0.3·MSE(f teacher , f student )
[0042] The final model is quantized and converted to INT8 precision through TensorRT. The model size is compressed to 2.8MB, the processing speed is increased to 62FPS (Jetson AGX Xavier), and the processing delay ≤ 200ms.
[0043] To verify the effectiveness of the classroom behavior detection method based on spatio-temporal Transformer and improved VGG16, the experiment was conducted on a self-built classroom behavior dataset. The dataset was collected from real classroom surveillance videos of multiple universities. The video resolution was 1920×1080, and the frame rate was 30 FPS. The videos were converted into a sequence of consecutive frames by uniform sampling. Each sequence contained 10 frames of images with a time span of 1 second to capture the dynamic behaviors of students. The final dataset contained 5000 video sequences, covering 10 common classroom behaviors: playing with mobile phones, lowering the head, sleeping, raising the hand, standing, turning the head, yawning, whispering to each other, reading textbooks, and listening attentively. All data were divided into a training set, a validation set, and a test set in a ratio of 7:2:1, and the LabelStudio tool was used to annotate the student positions and behavior categories in each frame of the image.
[0044] The model training was divided into two stages: in the first stage, the improved VGG16 was used to extract spatial features, and in the second stage, the spatio-temporal Transformer was used to fuse temporal information. The improved VGG16 introduced a channel attention module (SEBlock) and a multi-scale feature pyramid (FPN) on the basis of the original network. Specifically, SEBlocks were inserted after the fourth and fifth convolutional blocks of VGG16 respectively to enhance key features by adaptively adjusting channel weights; the FPN structure fused high-resolution features from shallow layers with semantic features from deep layers to improve the detection ability of small targets. The spatio-temporal Transformer consisted of an encoder-decoder architecture. The encoder layer modeled the spatio-temporal dependencies between frames through the multi-head self-attention mechanism, and the decoder layer output the behavior classification results. During training, the Adam optimizer was used with an initial learning rate of 0.001, a batch size of 8, and 200 iterations. The input sequence size was uniformly adjusted to 224×224×3.
[0045] The training results showed that the loss function of the model converged rapidly during the training process. The classification losses of the training set and the validation set tended to be stable after 50 iterations, and the final loss values were 0.12 and 0.15 respectively. In terms of performance metrics, the mean average precision (mAP@0.5) of the model on the test set reached 94.7%, the precision was 96.2%, and the recall was 93.5%. Compared with single-frame image detection, the introduction of the spatio-temporal Transformer significantly improved the recognition rate of dynamic behaviors (such as raising the hand and whispering to each other). Among them, the AP value of the "whispering to each other" category increased from 82.1% to 91.3%.
[0046] In the comparative experiment, this method was compared horizontally with the original VGG16, ResNet50, and YOLOv5s. All models were evaluated on the same test set, and the hardware environment was NVIDIA RTX 3090GPU. The experimental results show that this method performs best in accuracy and robustness (as shown in Table 1). The original VGG16 lacks temporal modeling capabilities, and its mAP is only 86.4%; ResNet50 reaches 89.7% with a deeper network structure, but it still lags behind this method. Although YOLOv5s has an advantage in detection speed (FPS=105), its mAP (90.1%) and dynamic behavior detection capabilities are obviously insufficient. Under the condition of FPS of 48, this method achieves a balance between accuracy and speed, meeting the real-time detection needs of the classroom.
[0047] Table 1 Comparative experimental results of different algorithms
[0048] In the output stage of behavior detection results, this method uses a dynamic threshold mechanism to further divide the 10 specific behaviors detected into four categories of classroom behavior states: "focus", "distraction", "interaction" and "leaving seat". Specifically, "focus" includes normal listening and reading textbooks; "distraction" includes playing with mobile phones, lowering heads, sleeping and yawning; "interaction" includes raising hands, whispering and turning heads; "leaving seat" includes standing and leaving seats. The dynamic threshold mechanism adaptively adjusts the classification threshold according to the student's behavior confidence score and temporal continuity in each frame image to ensure the accuracy and robustness of behavior classification. For example, for the "distraction" behavior, if "lowering heads" or "playing with mobile phones" is detected in multiple consecutive frames and the confidence score is higher than 0.8, it is judged as a "distraction" state; for the "interaction" behavior, if "raising hands" or "whispering" is detected and the duration exceeds 2 seconds, it is judged as an "interaction" state.
[0049] Case analysis further verifies the advantages of this method. Under lighting changes and occlusion scenarios, the original VGG16 has a high miss detection rate (about 15%) for the "head-down" behavior of students in the back row, while this method reduces the miss detection rate to 4% through multi-scale feature fusion and attention mechanism. For fast actions (such as "raising hands"), the spatiotemporal Transformer effectively captures the continuity of arm movements, and the false detection rate is reduced by 32% compared with single-frame detection. In addition, in the back row area of the classroom containing dense small targets, the detection box coverage and classification accuracy of this method are better than the comparison model, and the minimum detectable target size is 6×6 pixels, which is more capable of capturing details than YOLOv5s (8×8 pixels). Experimental results show that this method takes into account both high precision and robustness in complex classroom scenarios, providing reliable technical support for student behavior analysis.
[0050] By improving the combination of VGG16 and spatio-temporal Transformer, this method significantly enhances the accuracy and robustness of classroom behavior detection. The introduction of a dynamic threshold mechanism further optimizes the accuracy of behavior classification, especially performing excellently in dealing with complex scenarios (such as lighting changes, occlusions, and small object detection). Experimental results show that this method achieves a good balance between detection speed (FPS = 48) and accuracy (mAP@0.5 = 94.7%), can meet the requirements of real-time classroom monitoring, provide timely student behavior feedback for teachers, and contribute to the improvement of teaching quality.
Claims
1. A classroom behavior detection method based on spatiotemporal Transformer and improved VGG16, characterized in that: include: Collect multiple classroom surveillance videos of students and convert them into continuous frame sequences, each sequence containing 10 frames of images with a time span of 1 second to capture the dynamic behavior of students; The positions and behaviors of the students are marked on the continuous frame sequence to construct an initial sample set containing 10 types of specific behaviors, including playing with mobile phones, lowering heads, sleeping, raising hands, standing, turning heads, yawning, whispering, reading textbooks, and normal listening; An improved VGG16 network is used to extract spatial features. The improved VGG16 network introduces a channel attention module (SE Block) after the fourth and fifth convolution blocks, respectively, enhances the key feature expression by adaptively adjusting the channel weights, and combines a multi-scale feature pyramid (FPN) structure to fuse shallow high-resolution features with deep semantic features to improve small target detection capabilities. The spatiotemporal dependency between frames is modeled using a spatiotemporal Transformer, which consists of an encoder-decoder architecture. The encoder captures the spatiotemporal information between frames through a multi-head self-attention mechanism, and the decoder outputs the behavior classification result. A dynamic threshold mechanism is designed to classify the 10 specific behaviors detected into four classroom behavior states: "focus", "distraction", "interaction" and "leaving seat". The dynamic threshold mechanism adaptively adjusts the classification threshold based on the confidence score and temporal continuity. The specific rules are: if "head-down" is detected for 3 consecutive frames and the confidence is ≥0.85, it is judged as "distraction"; if the "hand-raising" action lasts ≥2 seconds, it is judged as "interaction"; The initial sample set is used to train a model composed of an improved VGG16 and a spatiotemporal Transformer to obtain a student classroom behavior detection model; Collect real-time monitoring videos of students in the classroom and convert them into continuous frame sequences; The student classroom behavior detection model is used to perform target detection on the processed continuous frame sequence, identify student behaviors in the real-time classroom monitoring video, and output behavior classification results.
2. The classroom behavior detection method based on spatiotemporal Transformer and improved VGG16 according to claim 1 is characterized in that: The channel attention module (SE Block) includes: A global average pooling layer to aggregate the spatial information of feature maps; Fully connected layer, used to compress the spatial dimension of the input feature map; Sigmoid activation function, used to generate channel attention weights; The channel attention weight is multiplied by the input feature map to obtain a weighted feature map.
3. The classroom behavior detection method based on spatiotemporal Transformer and improved VGG16 according to claim 2 is characterized in that: The multi-scale feature pyramid (FPN) structure includes: The first convolutional layer is used to extract shallow high-resolution features; The second convolutional layer is used to extract deep semantic features; The feature fusion module is used to fuse shallow high-resolution features with deep semantic features to generate multi-scale feature maps.
4. The classroom behavior detection method based on spatiotemporal Transformer and improved VGG16 according to claim 3 is characterized in that: The encoder of the spatiotemporal Transformer includes: a multi-head self-attention mechanism for modeling spatiotemporal dependencies between frames; Feedforward neural network for further feature extraction; The decoder comprises: Fully connected layer, used to output behavior classification results; Softmax layer, used to generate the final behavior category probability distribution.
5. The classroom behavior detection method based on spatiotemporal Transformer and improved VGG16 according to claim 4 is characterized in that: The dynamic threshold mechanism includes: A confidence score calculation module, used to calculate the confidence score of the behavior in each frame of the image; Temporal continuity analysis module, used to analyze the duration and continuity of behavior; The adaptive classification threshold adjustment module is used to dynamically adjust the classification threshold according to the confidence score and temporal continuity.
6. The classroom behavior detection method based on spatiotemporal Transformer and improved VGG16 according to claim 5 is characterized in that: The size of the input continuous frame sequence is 224×224×3.
Citation Information
Patent Citations
End-to-end student behavior detection method and system based on large convolution kernel
CN119068548A