Fine-grained teaching behavior identification method for smart classroom
Through the combination of improved YOLOv8 model and fine-grained feature extraction, multi-scale feature fusion and small object detection layer, the accuracy and real-time problems of object detection algorithm in smart classrooms in small objects and complex backgrounds are solved, and more efficient and accurate classroom behavior detection is achieved.
Patent Information
- Application Number
- CN202411981932.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing object detection algorithms are difficult to achieve high accuracy and real-time performance in smart classroom teaching environments, especially when dealing with small goals and complex backgrounds, missed detection and missed detection problems are prone to occur.
A fine-grained teaching behavior recognition method for smart classrooms is proposed. Based on the improved YOLOv8 model, combined with the fine-grained feature extraction module, multi-scale feature fusion module and small object detection layer, the detection accuracy and real-timeness of the model in complex scenarios are improved.
While maintaining efficient detection speed, the accuracy of classroom behavior detection is significantly improved, especially in multi-objective detection and small-objective recognition tasks, reducing missed detection and missed detection problems.
Smart Images

Figure CN119942640A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection and relates to a fine-grained teaching behavior recognition method for smart classrooms. Background Art
[0002] In the smart classroom teaching environment, real-time monitoring of student behavior is crucial to establishing a good teacher-student relationship. However, due to the complex classroom environment, large number of students, and large classroom depth, current behavior detection algorithms are insufficient in accuracy and real-time performance. Existing classroom behavior detection methods mainly rely on computer vision and machine learning technologies to identify students' actions by analyzing video streams or image data. Common detection algorithms include posture estimation and key point detection, action recognition based on convolution and time series models, and ensemble learning and multi-model fusion.
[0003] Although these methods often perform well in laboratory environments or certain specific scenarios, they are difficult to directly apply to teaching behavior recognition in smart classrooms in actual teaching environments due to problems such as multi-target detection, frequent action changes, and fine-grained feature retention. With the advancement of artificial intelligence technology, it has become a trend to use convolutional neural networks (CNNs) and time series models to detect student behavior. These methods have shown great potential in improving detection accuracy and adapting to complex environments. At present, target detection methods are mainly divided into two categories, namely single-stage detection methods and multi-stage detection methods.
[0004] The single-stage detection method eliminates the step of generating candidate regions by directly extracting features from the image and predicting the location and category of the target. Another innovative single-stage method is DETR (Detection Transformer) proposed by Nicolas Carion et al. of Facebook AI Research, which uses the Transformer architecture to achieve end-to-end target detection without the need for additional region proposals or post-processing steps. However, these methods are less effective in detecting small targets, especially in scenes with complex backgrounds or dense objects. They are prone to accuracy degradation due to unclear object features, small target size, or overlapping adjacent objects, leading to problems such as missed detection and false detection, making it difficult to meet the accuracy requirements of classroom detection.
[0005] The multi-stage detection method usually includes two stages: candidate region generation and subsequent fine detection. Although this method has improved accuracy compared to the single-stage method, it has the problems of slow detection speed and high complexity, and it is difficult to meet the real-time requirements of classroom action detection. In classroom behavior analysis, target detection faces many challenges, especially when dealing with small targets and complex backgrounds. Traditional target detection algorithms often perform poorly in these situations, resulting in significant missed detection and false detection problems.
[0006] In order to solve these problems, the YOLO series, as an efficient target detection algorithm, has gradually attracted the attention of researchers, especially YOLOv8, which has achieved a good balance between detection accuracy and speed. As an important node of the YOLO series target detection algorithm, YOLOv8 significantly improves accuracy and speed while maintaining efficient detection. It performs well in processing targets of various scales, and at the same time has strong real-time performance, which is suitable for application scenarios with high requirements for speed. However, YOLOv8 still has some shortcomings in practical applications, especially when processing small targets and complex backgrounds, the accuracy is often not as expected. In the classroom teaching scenario faced by the present invention, it is specifically reflected in the fact that the detection ability of small targets in the back row is still defective, which is manifested as frequent missed detection and false detection, which ultimately leads to poor detection of small targets, and the problems of missed detection and false detection are more prominent. Summary of the invention
[0007] The purpose of this invention is to propose a fine-grained teaching behavior recognition method for smart classrooms, which can better adapt to complex scenarios in the classroom environment while maintaining high efficiency and provide more accurate behavior detection results.
[0008] In order to achieve the above object, the present invention adopts the following technical scheme:
[0009] A fine-grained teaching behavior recognition method for smart classrooms includes the following steps:
[0010] Step 1. Obtain classroom teaching video data stream and construct a data set for training the following model;
[0011] Step 2. Build a teaching behavior detection model based on the improved YOLOv8 model. The model includes a feature extraction layer, a feature fusion layer, and a detection head, where there are four detection heads.
[0012] First, the video frame is input into the feature extraction layer for feature extraction. The processing process is as follows:
[0013] In the feature extraction layer, the video frame first passes through a convolution module, followed by multiple sets of convolution and fine-grained feature extraction modules for full feature extraction, and finally undergoes feature enhancement through fast pyramid pooling to complete feature extraction;
[0014] Then, the feature fusion layer combines multiple multi-scale feature fusion modules, multiple upsampling modules, multiple connection modules, and multiple convolution modules to perform feature fusion at different depths.
[0015] Finally, the features at different depths are passed through four detection heads to obtain the detection result information;
[0016] Step 3. Train the teaching behavior detection model built based on the data set in step 1, and then use the trained model to detect the input video images in the real classroom scene to obtain the student behavior detection results.
[0017] In addition, based on the above-mentioned fine-grained teaching behavior recognition method for smart classrooms, the present invention also proposes a computer device, which includes a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it is used to implement the above-mentioned fine-grained teaching behavior recognition method for smart classrooms.
[0018] In addition, based on the above-mentioned fine-grained teaching behavior recognition method for smart classrooms, the present invention also proposes a computer-readable storage medium on which a program is stored.
[0019] When the program is executed by a processor, it is used to implement a fine-grained teaching behavior recognition method for smart classrooms.
[0020] The present invention has the following advantages:
[0021] As described above, the present invention relates to a fine-grained teaching behavior recognition method for smart classrooms. The method is based on YOLOv8 and proposes a lightweight teaching behavior detection model that integrates contextual fine-grained features, namely the EDU-YOLO model, to improve the detection accuracy and real-time performance of the model in complex classroom scenarios. In order to enhance the detection capability of the model in the specific complex application scenario of smart classrooms, the present invention designs a fine-grained feature extraction module, which combines deep convolution with EMA attention mechanism to effectively extract fine-grained features while optimizing the utilization efficiency of computing resources. In addition, in view of the problem of large scale span of detection targets caused by the depth of the classroom in the classroom environment, the present invention designs a multi-scale feature fusion module, which effectively improves the model's detection capability for actions of different scales, especially in dense scenes. The method of the present invention is superior to existing detection algorithms in terms of accuracy and real-time performance, especially in multi-target detection and small target recognition tasks. The present invention improves the detection accuracy while maintaining the detection speed by designing a fine-grained feature extraction module, a multi-scale feature fusion module, and a fine-grained classroom teaching behavior detection module specifically for small targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 1 is an overall processing flow chart of a fine-grained teaching behavior recognition method for a smart classroom in an embodiment of the present invention;
[0023] Figure 2 It is a structural diagram of the teaching behavior detection model EDU-YOLO built in an embodiment of the present invention;
[0024] Figure 3 It is a structural diagram of a fine-grained feature extraction FGE module in an embodiment of the present invention;
[0025] Figure 4 It is a structural diagram of the Bottleneck module in an embodiment of the present invention;
[0026] Figure 5 It is a structural diagram of a multi-scale feature fusion MSE module in an embodiment of the present invention;
[0027] Figure 6 Schematic diagram of a fine-grained classroom teaching behavior detection module (detection head) in an embodiment of the present invention;
[0028] Figure 7 This is a data distribution diagram in an embodiment of the present invention;
[0029] Figure 8 This is a comparison chart of the experimental results of the model in the embodiment of the present invention on POCO;
[0030] Fig. 9 This is a comparison diagram of the confusion matrices of EDU-YOLO and YOLOv8 on UOCO in an embodiment of the present invention;
[0031] Fig.10 This is a comparative display diagram of the detection results of EDU-YOLO on the data set in an embodiment of the present invention.
[0032] Fig.11 This is a comparison diagram showing the detection results of EDU-YOLO in real scenes in an embodiment of the present invention. DETAILED DESCRIPTION
[0033] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0034] With the in-depth integration of educational informatization and artificial intelligence technology, real-time monitoring and evaluation of students' classroom behavior has become an important direction for the development of intelligent education. However, due to the spacious space and large depth of university classrooms, the detection accuracy of traditional target detection algorithms has decreased, especially small targets are easily ignored, and large targets cause redundant calculation problems.
[0035] like Figure 1 As shown, in response to the technical challenges faced by student classroom behavior detection in a smart teaching environment, the present invention proposes a fine-grained teaching behavior recognition method for smart classrooms. In the method, an improved target detection model EDU-YOLO is proposed. A fine-grained feature extraction module FGE, a multi-scale feature fusion module MSE and a small target detection layer are introduced into the model, which improves the detection accuracy and robustness of the model in cross-scale and fine-grained environments.
[0036] First, the FGE module combines a deep convolutional network with an efficient multi-scale attention module EMA to effectively extract fine-grained detail features of student behavior and enhance the ability to recognize diverse actions. Then, the multi-scale feature fusion technology improves the fusion performance of features of different scales, especially in dense classroom scenes, and improves the accurate recognition of fine-grained target behaviors. Finally, the small target detection layer optimizes the ability to capture long-distance and small actions, effectively improving the overall detection accuracy.
[0037] Compared with traditional algorithms such as FasterR-CNN and SSD, this invention adopts multi-level feature fusion in terms of ideas, fully utilizes feature information of different scales, enhances the detection ability of complex situations and small targets, and structurally introduces fine-grained feature extraction module FGE and multi-scale feature fusion module MSE, which improves the network's receptive field and feature expression ability while saving computing resources, so that the model can still accurately and quickly detect targets in complex backgrounds. In addition, the introduction of the small target detection layer significantly reduces the missed detection rate and improves the detection performance in dense scenes.
[0038] like Figure 1 As shown, the fine-grained teaching behavior recognition method for smart classrooms in this embodiment includes the following steps:
[0039] Step 1. Obtain the classroom teaching video data stream and split it into frames, and use annotation tools (such as Make Sense) to build a dataset for training the following model. Each picture in the dataset is carefully annotated to ensure the accuracy of the action category.
[0040] The labels of the images include ten categories, namely, raisehead, bowhead, turnhead, phone, play phone, book, lookbook, sleep, raise hand, bend. The entire dataset is divided into three parts: training set, test set and validation set.
[0041] Step 2. Build the teaching behavior detection model EDU-YOLO based on the improved YOLOv8 model. Its overall structure is as follows: Figure 1 As shown, the model consists of four parts: input layer, feature extraction layer, detection layer and output layer.
[0042] The EDU-YOLO model significantly improves detection precision and accuracy by designing a more efficient feature extraction and fusion mechanism, especially for small targets and subtle movements in classroom behavior detection.
[0043] Specifically, in classroom behavior detection, due to the large depth of the classroom, in the image captured by the camera, the students at the front of the screen (i.e., close to the camera) occupy a larger pixel area on the image, forming a larger target object; on the contrary, the students at the back of the screen (i.e., far from the camera) only occupy a smaller pixel area on the image due to perspective deformation and increased distance, becoming a smaller target object. This significant difference in target scale caused by differences in spatial position greatly increases the difficulty and complexity of the target detection algorithm, and ultimately results in the model being unable to effectively detect students at the back of the screen.
[0044] By introducing the FGE module, MSE module and small target layer of the present invention, EDU-YOLO can better adapt to complex scenarios in the classroom environment while maintaining high efficiency and provide more accurate behavior detection results.
[0045] The network structure of the EDU-YOLO model is as follows Figure 2 As shown. In order to solve the problem of low accuracy in classroom behavior detection, the present invention designs an FGE module based on real-time requirements. It effectively extracts cross-scale behavior features by combining the cross-stage partial fusion module C2f with the efficient multi-scale attention module EMA, and then enhances the ability to integrate and utilize key behavior features of different scales through the MSE module, thereby improving the detection performance of the model in complex classroom scenarios. In addition, in order to solve the problem that small-scale targets cannot be effectively detected, a fine-grained classroom behavior detection module is specially designed to optimize the recognition ability of distant targets and tiny target behaviors, effectively improve the overall detection accuracy, and meet the actual needs of classroom behavior detection.
[0046] like Figure 2 As shown, the EDU-YOLO model built in this embodiment includes a feature extraction layer, a feature fusion layer, and a detection head, where there are four detection heads. The overall processing flow of the EDU-YOLO model is as follows:
[0047] First, the input video frame is input to the feature extraction layer for feature extraction. In the feature extraction layer, the video frame image first passes through a convolution module, followed by multiple sets of convolution and fine-grained feature extraction modules for full feature extraction, and finally passes through fast pyramid pooling for feature enhancement to complete feature extraction. Then, the feature fusion layer combines multiple multi-scale feature fusion modules with upsampling modules and multiple multi-scale feature fusion modules with convolution modules to perform feature fusion at different depths. Finally, features at different depths pass through four detection heads to obtain detection result information.
[0048] The feature extraction layer includes a convolution module, four groups of convolution and fine-grained feature extraction modules, and a fast pyramid pooling SPPF module. The four groups of convolution and fine-grained feature extraction modules are defined as the first, second, third, and fourth groups of convolution and fine-grained feature extraction modules, respectively, and each group contains a convolution module and a fine-grained feature extraction module FGE.
[0049] The video frame first passes through a convolution module for feature extraction, and then passes through the first, second, third, and fourth groups of convolutions and fine-grained feature extraction modules FGE for feature extraction, and finally passes through the fast pyramid pooling SPPF module for feature enhancement to complete the feature extraction.
[0050] There are seven multi-scale feature fusion modules MSE, which are defined as the first, second, third, fourth, fifth, sixth, and seventh multi-scale feature fusion modules; there are three upsampling modules, which are defined as the first, second, and third upsampling modules.
[0051] There are three convolution modules, which are defined as the first, second and third convolution modules; there are seven connection modules, which are defined as the first, second, third, fourth, fifth, sixth and seventh connection modules respectively.
[0052] The output features of the SPPF module are first connected with the output features of the fourth group of convolution and fine-grained feature extraction FGE modules through a first connection module, and the output features of the first connection module enter the first upsampling module for upsampling.
[0053] The output features of the first upsampling module are connected to the output features of the FGE module in the third group of convolution and fine-grained feature extraction modules through the second connection module, and the output features of the second connection module enter the second upsampling module for upsampling.
[0054] The output features of the second upsampling module are connected with the output features of the FGE module in the second group of convolution and fine-grained feature extraction modules through the third connection module, and the output features of the third connection module enter the third upsampling module for upsampling.
[0055] The output features of the third upsampling module pass through the fourth connection module and are connected with the output features of the FGE module in the first group of convolution and fine-grained feature extraction modules. The output features of the fourth connection module pass through the first multi-scale feature fusion module and the first convolution module in sequence.
[0056] The output features of the first convolution module enter the fifth connection module; the output features of the third connection module enter the fifth connection module through a second multi-scale feature fusion module and are connected with the output features of the first convolution module.
[0057] The output features of the fifth connection module pass through the third multi-scale feature fusion module and the second convolution module in sequence.
[0058] The output features of the second convolution module enter the sixth connection module; the output features of the second connection module enter the sixth connection module through a fourth multi-scale feature fusion module and are connected with the output features of the second convolution module.
[0059] The output features of the sixth connection module pass through the fifth multi-scale feature fusion module and the third convolution module in sequence.
[0060] The output features of the third convolution module enter the seventh connection module; the output features of the first connection module enter the seventh connection module through a sixth multi-scale feature fusion module and are connected with the output features of the third convolution module.
[0061] The output features of the seventh connection module enter the seventh multi-scale feature fusion module for processing.
[0062] The first, third, fifth and seventh multi-scale feature fusion modules are respectively connected to a detection head.
[0063] EMA preserves the information of each channel and improves the pixel-level feature representation without adding significant computational burden. The module divides the channel dimension of the input feature map into multiple sub-feature groups and evenly distributes spatial semantic features within each group. The EMA module specifically designs two parallel branches: one uses 1×1 convolution to encode channel attention, and the other uses 3×3 convolution to capture multi-scale feature representations. The outputs of these two branches are further aggregated through a cross-dimensional interaction method to enhance the interaction and information fusion between features. Experiments on widely used benchmark datasets such as CIFAR100, MSCOCO, and VisDrone2019 show that the efficient multi-scale attention module EMA exhibits excellent performance in both image classification and object detection tasks. EMA not only significantly improves the accuracy of the model while maintaining computational efficiency, but also adapts well to network architectures of different depths.
[0064] Based on the design concept of EMA, the present invention designs a fine-grained feature extraction module FGE (Fine-grained feature extraction module), which combines the lightweight advantage of the C2f module with the refined feature extraction capability of the EMA attention mechanism to enhance the feature expression capability of the network; the C2f module has a strong feature reuse capability, which improves the generalization ability of the model while maintaining computational efficiency; and the multi-scale feature capture capability of EMA enables it to more accurately identify diversified behaviors, especially for small target detection in complex scenes.
[0065] like Figure 3 As shown in the figure, the processing flow of the fine-grained feature extraction FGE module is as follows:
[0066] The fine-grained feature extraction module FGE first compresses the input features through the convolution Conv operation, and uses the split function Split to divide the features into two branches. Such a branch design helps to enhance the nonlinear ability and expression ability of the network, thereby improving the network's modeling ability for complex data. One branch is not processed, and the other branch contains multiple Bottleneck or Bneck modules, which are used to deeply extract the input features to enhance the overall feature extraction ability of the network. In the figure, n represents the number of repetitions of the Bneck module. These Bneck modules transmit feature information through parallel paths, and splice the outputs of the two branches together through the connection operation. Then, part of the features of the obtained feature group are fused with the original features after the EMA attention mechanism to obtain a new feature group. Finally, the features are fused and output with the convolution Conv structure through feature splicing Concat. This structure integrates an efficient multi-scale attention mechanism to dynamically adjust the weights between channels, saving computing resources while highlighting key features. The efficient multi-scale attention module EMA helps the network focus on important channels, thereby enhancing the detection ability of the model in complex scenarios. In addition, the module retains gradient information through residual connections to ensure that important features can be effectively transmitted in deep networks. This design ensures that rich gradient flow information can still be obtained after multiple cross-layer connections, enhancing the ability to capture small objects and fine-grained features.
[0067] like Figure 4 The structure of each Bottleneck module is shown. The features of the input Bottleneck module are transmitted through parallel paths. The main branch path passes through two convolution modules. The output features of the main branch path are residually connected with the original input features through the connection module. The specific calculation process of the fine-grained feature extraction FGE module is as follows:
[0068] First, the features of the input fine-grained feature extraction FGE module are processed, and the input features are denoted as X∈RC×H×W; where C, H and W represent the number of channels, height and width, respectively.
[0069] First, the input features are processed through a convolution operation to obtain the features Figure X , the convolutional features Figure X Then it is divided into two parts through the channel splitting operation Split, namely:
[0070] X′=Conv1(X);
[0071] X1,X2=Split(X′).
[0072] Conv1 is a standard convolutional layer for preliminary feature extraction. X1 and X2 represent the two branches of the input feature X after segmentation, which are used for subsequent processing by different paths. Then, the segmented X2 is sent to multiple Bottleneck layers to further extract deep features to obtain X2′. After the Bneck layer extracts features, X1 and X′2 are concatenated together through the Concat operation to generate a new feature map Y1, Y2, …, Y m , the formula is as follows:
[0073] X′2=Bottleneck n (X2);
[0074] Y1,Y2,…,Y m =Concat(X1,X′2);
[0075] Wherein, n is the number of repetitions of the Bneck module, and in this embodiment, n is 2.
[0076] Then, feature Y m Input the efficient multi-scale attention module EMA and add the result to the original feature group. The efficient multi-scale attention module EMA is used to dynamically weight the features between channels and highlight the channel features that are critical to the detection task. Finally, all the features of the obtained feature group are concat-connected to obtain a processed feature map, which is then passed through a convolutional network to obtain the final feature map Y. After such a series of processing, the output and input sizes of the module remain consistent, that is:
[0077] Y′=Concat(Y1,Y2,…,EMA(Y m ));
[0078] Y=Conv2(Y′).
[0079] Among them, EMA is the attention mechanism function, which is responsible for assigning different weights to channels and emphasizing important features.
[0080] The multi-scale feature fusion module MSE is used to enhance and fuse features of different scales. The input features are fused and enhanced by combining a multi-layer perceptron, a random dropout strategy, and an efficient multi-scale attention mechanism. This design makes the connection between features closer and enhances the network's ability to fuse and utilize detailed features of different scale behaviors, especially in complex environments with multiple targets, which can improve the accuracy of target detection.
[0081] like Figure 5 The structure of the multi-scale feature fusion module MSE is shown. The processing flow of the MSE module is as follows:
[0082] First, the input features are adjusted to meet the model's size calculation requirements through the dynamic adjustment module AD, and then enter the partial convolutional network P_Conv for feature compression processing to generate a compressed feature map. Then, these features are linearly transformed through a multi-layer perceptron MLP containing two convolutional layers to further extract features. Subsequently, the path random dropout DropPath strategy is introduced to enhance the robustness of the model and reduce overfitting. The efficient multi-scale attention module EMA in the module is used to further highlight important features by dynamically adjusting the correlation between channels. The partial convolutional network in this module uses a specified block method for spatial mixing to achieve adaptive learning of local features. With the help of residual connections, the output of the data after passing through the MLP module and the efficient multi-scale attention module EMA will be fused with the input features to ensure efficient information transmission. In summary, this structure effectively improves the model's ability to focus on important features while enhancing the model's generalization ability through the combination of residual connections, efficient multi-scale attention mechanisms, path random dropout strategies, and multi-layer perceptrons. The specific calculation process of the multi-scale feature fusion module MSE is as follows:
[0083]
[0084] If the input channels do not match the expected output channels, a 1×1 convolution is used to adjust them.
[0085] X min =P_Conv(X adj );
[0086] Y mlp =Conv(ReLU(Conv(X min )));
[0087] Y attn =EMA(DropPath(Y mlp ));
[0088] Y final =X adj +Y attn .
[0089] Where X is the input feature map, X adj It is the feature map after channel adjustment. When the number of input channels inc does not match the number of target channels dim, a 1×1 convolution layer is used to adjust the number of channels. minis the spatially mixed feature map obtained by processing through a custom P_Conv layer. ReLU is the activation function. DropPath is a function that discards part of the path in the feature map with a certain probability (set to 0.1) to prevent overfitting. Efficient Multiscale Attention (EMA) performs the attention mechanism operation to enhance the important parts of the features and suppress the unimportant parts. Y mlp , Y attn , Y final They are multi-layer perceptron MLP, efficient multi-scale attention module EMA and fusion features.
[0090] In addition, in order to meet the challenge of small target detection in smart teaching scenarios, the present invention designs a fine-grained classroom behavior detection module, namely the small target layer. This layer combines the FGE module and the MSE module proposed in this paper to efficiently extract and fuse fine-grained features, effectively improving the model's recognition accuracy for small targets. In complex backgrounds and dense scenes, small targets are often ignored because their features are not obvious. This layer optimizes the network receptive field while enhancing the model's ability to detect the behavior of distant students, thereby reducing missed detections. Its representation diagram is shown in the figure below. Figure 6 shown.
[0091] The fine-grained classroom behavior detection module is particularly critical in the network structure. Figure 2 , through a layer of fine-grained feature extraction module, a multi-scale feature fusion module and a small target detection head to fully extract, fuse and detect the fine-grained features of small-scale targets. This design enables the feature map to effectively retain more fine-grained detail information after passing through the fine-grained feature extraction module, and at the same time, the fine-grained features are enhanced and fused by the multi-scale feature fusion module, thereby enhancing the detectability of small targets. Subsequently, the fused fine-grained features are detected by the small target detection head and the detection results are obtained. In a dense classroom scene, small targets at a distance are usually difficult to detect due to feature fuzziness, and the small target detection layer significantly reduces the omission of small targets by extracting fine-grained features and enhancing and fusing multi-scale features, thereby improving the overall performance of the model in the classroom behavior detection task. The EDU-YOLO model constructed by the present invention not only improves the accuracy of classroom behavior detection, but also provides teachers with a more effective teaching management tool. This tool can help teachers understand students' classroom performance in a timely manner and formulate more targeted teaching strategies and feedback mechanisms based on data analysis results.
[0092] Step 3. Based on the data set in step 1, the lightweight teaching behavior detection model that integrates contextual fine-grained features is trained, and then the trained lightweight teaching behavior detection model is used to detect the input video images in the real classroom scene to obtain the classroom behavior detection results in the form of pictures with confidence.
[0093] In addition, in order to verify the effectiveness of the method proposed in the present invention, the method of the present invention is also compared with several mainstream target detection algorithms to comprehensively evaluate the performance of the method of the present invention on different indicators. Through the analysis of the results of the comparative experiment and the ablation experiment, the effectiveness of the model and the effect of the improvement measures on the improvement of the model detection accuracy will be further verified.
[0094] To ensure the fairness of the experiment, all uses of the present invention were completed on a Linux system equipped with a 13th Gen Intel (R) Core (TM) i9-13900K CPU and dual NVIDIA GeForce RTX4090 GPUs.
[0095] The parameters are set as follows: intersection over union (iou) is 0.7, initial learning rate (lr0) is 0.01, learning rate decay rate (lrf) is 0.01, optimizer weight decay (weightdecay) is 0.0005, batch size (batchsize) is 16, training rounds (epochs) is 500, image input size (imgsz) is 640. The original size of the input image is 1920*1080.
[0096] The evaluation indicators include mean Average Precision (mAP), precision (P), recall (R), F1 score (F1Score) and intersection over union (IoU).
[0097] Precision refers to the proportion of all targets predicted to be positive samples that are actually positive samples. The calculation formula is:
[0098]
[0099] TP (TruePositive) stands for true positive, which means the number of targets correctly detected by the model, that is, the detected target box and the true box meet the IoU threshold requirement (usually 0.5).
[0100] FP (False Positive) stands for false positive, which indicates the number of targets detected by the model incorrectly, that is, the IoU between the detected target box and the real box is less than the set threshold, or there is no corresponding target in the real scene.
[0101] The recall rate indicates the proportion of all positive samples that are correctly detected. The calculation formula is:
[0102]
[0103] Among them, FN (False Negative) stands for false negative, which means the number of real targets that the model failed to detect, that is, the target actually exists but the model fails to recognize it.
[0104] mAP is the most commonly used evaluation indicator in target detection. It is the weighted sum of the average precision of various types of targets and can be expressed as a formula. mAP is usually divided into mAP 0。5 and mAP 0。5:0。95 :
[0105]
[0106] Where AP(j) represents the AP of the jth category. In particular, mAP 0。5 Refers to the average precision when the IoU (Intersection over Union) is 0.5, with special attention to the performance of the model when the predicted bounding box has a high overlap with the true bounding box. mAP 0。5:0。95 It is a more stringent evaluation standard, which means the average precision calculated when IoU ranges from 0.5 to 0.95 with a step size of 0.05, which can better reflect the comprehensive performance of the model in detecting the positioning accuracy of the frame.
[0107] The F1 score is the harmonic mean of precision and recall, and is used to balance the evaluation criteria of precision and recall.
[0108]
[0109] The intersection-over-union ratio is an indicator commonly used in target detection to measure the degree of overlap between the predicted box and the actual labeled box. The calculation formula is:
[0110]
[0111] Where Aoverlap represents the area of the overlap between the predicted box and the true box. Aunion represents the area of the union of the predicted box and the true box, which is the total area of the two boxes minus the area of the overlap.
[0112] The comparative experiment is as follows: The POCO dataset collects the real postures of college students in the classroom under different environments, classrooms, subjects, number of people and camera angles. It contains very complex teaching scenes, which can provide strong support for the study of students' status in the classroom. It contains ten categories in total, namely, raise head, bow head, turn head, phone, play phone, book, lookbook, sleep, raise hand, bend. The entire dataset is divided into three parts: training set, test set and validation set.
[0113] The training set contains 1146 RGB images and 83566 labels, the test set contains 378 RGB images and 26909 labels, and the validation set contains 379 RGB images and 27485 labels.
[0114] Figure 7 (a) shows the data distribution in the POCO dataset. In order to verify the effectiveness of the method, Faster R-CNN, Single Stage Detector (SSD), Residual Network (Res2Net), Transformer-based Object Detection Algorithm (DETR), YOLOv8n, YOLOv10n, YOLO11n and RE-DETR are selected as comparison models and verified on the POCO dataset. The experimental results are shown in Table 1.
[0115] Table 1 POCO comparison test results
[0116]
[0117] It can be seen from the experimental results that the method of the present invention shows significant improvement in various indicators. Specifically, the EDU-YOLO model in the present invention has significant improvements in precision, recall, mAP, and 0。5 and mAP 0。5:0。95 The performance in all aspects is better than the existing YOLO series models. First, in terms of precision, EDU-YOLO reached 0.856, which is significantly higher than other models, especially compared with YOLOv8 and YOLOv10, which increased by 2.3% and 5.0% respectively. And it also increased by 1.7% compared with the most advanced YOLO11 model. This improvement shows that the method of the present invention performs well in reducing false detections. Secondly, in terms of recall rate, EDU-YOLO achieved a high score of 0.797, which exceeded other models, especially compared with YOLOv8's 0.721 and YOLOv10's 0.730, the improvement is significant. The above shows that the method of the present invention has a low missed detection rate when detecting targets and has a strong target detection capability. In the most important mAP 0。5 and mAP 0。5:0。95 On the other hand, EDU-YOLO achieves 0.849 and 0.603 respectively, especially in mAP. 0。5:0。95 It is significantly improved compared with 0.537 of YOLOv8 and 0.540 of YOLO11, indicating that the proposed method has good generalization ability under different IoU thresholds.
[0118] In general, the method of the present invention effectively improves the detection accuracy and recall rate by introducing the FGE module, MSE and small target detection layer, and further enhances the performance of the model in complex scenarios.
[0119] Figure 8 The figure shows the changing trends of the evaluation indicators of the comparison model and the EDU-YOLO model of the present invention during the 500 epoch training process. Figure 8 (a) and 8(b), in comparison, EDU-YOLO has better accuracy and mAP. 0。5 It performs better in this aspect, and its performance tends to be stable after 100 epochs.
[0120] In order to further verify the performance of the EDU-YOLO model proposed in the present invention, a new data set UOCO was produced, which contains 300 pictures collected in a real classroom environment. The data set covers six specific action categories, namely: using a mobile phone (playphone), turning head (turnhead), sleeping (sleep), standing (stand), bowing head (lowhead) and raising head (raisehead). The selection of these action categories is intended to cover the common behavior patterns of students in the classroom, and provide rich annotation data to support the training and evaluation of the model. Each picture in the data set of the present invention is carefully annotated to ensure the accuracy of the action category. The construction process of the data set follows strict standards to ensure that each action has good performance in different environments and angles, reflecting the diversity of real classrooms. This diversity not only improves the adaptability of the model to complex scenes, but also helps to evaluate the generalization performance of the model in different teaching situations. Figure 7 (b) shows the label distribution in the UOCO dataset. The comparative experimental results of different models on UOCO are shown in Table 2.
[0121] Table 2 UOCO comparative test results
[0122]
[0123] According to the experimental results, EDU-YOLO significantly outperforms the YOLO series comparison models YOLOv8n, YOLOv10n and YOLO11n in all evaluation indicators. First, in terms of precision, EDU-YOLO reached 0.632, which is significantly higher than YOLOv8's 0.554, YOLOv10's 0.543 and YOLO11's 0.528. This improvement shows that EDU-YOLO excels in reducing false detections and can identify targets more accurately. Secondly, in terms of recall, EDU-YOLO also leads, achieving a high score of 0.562, showing stronger target detection capabilities compared to other models (YOLOv8's 0.528, YOLOv10's 0.486 and YOLO11's 0.511). The above shows that the EDU-YOLO model can more effectively capture the various behaviors of students in complex classroom scenarios. In mAP 0。5 and mAP 0。5:0。95 In terms of indicators, EDU-YOLO reached 0.552 and 0.256 respectively, both higher than the performance of the comparison model. This shows that EDU-YOLO has good generalization ability at different IoU thresholds and can adapt to various detection tasks. At the same time, the present invention introduces the confusion matrix of YOLOv8 and EDU-YOLO, such as Fig. 9 (a) and 9(b) are shown in order to more intuitively compare the classification performance of the two on the UOCO dataset. Fig. 9 It can be seen that EDU-YOLO is better at detecting small movements.
[0124] In addition, in order to verify the impact of each improved module of the present invention on the model performance, an ablation experiment was conducted on the basis of the baseline model, gradually removing the fine-grained feature extraction (FGE) module, the multi-scale feature fusion (MSE) module and the small object detection layer to observe the changes in model performance. Table 3 shows the model performance in terms of Precision, Recall, and mAP under different module combinations. 0。5 and mAP 0。5:0。95 Performance on other indicators.
[0125] Table 3 Ablation experiment results
[0126]
[0127] From the results in Table 3, we can see that each module in the present invention contributes to improving the model performance to varying degrees:
[0128] (1) Fine-grained feature extraction module FGE: When the FGE module is removed, the model's precision is 0.846 and the recall rate is 0.792, indicating that both the false positive rate and the missed positive rate have increased. This is because the module acts on the feature extraction process. After removing the module, the model does not fully extract features, resulting in a lack of fine-grained features in the feature fusion and detection process, which in turn leads to a decline in model performance.
[0129] (2) Multi-scale feature fusion module MSE: When the MSE module is removed, the accuracy of the model is 0.845, and the other three indicators are also reduced. This is because the features extracted in the feature extraction process are not effectively utilized, which leads to the inability to effectively detect targets containing fine-grained features and targets with complex features, which ultimately manifests as reduced accuracy. This shows that the multi-scale feature fusion (MSE) module plays a positive role in the feature fusion utilization process.
[0130] (3) Fine-grained classroom behavior detection module (i.e., small target layer): When the small target detection layer is removed, all indicators of the model are significantly reduced compared to the complete model, indicating that the model's detection effect on fine-grained small targets has been greatly reduced. This proves that this module effectively improves the model's ability to detect small targets, especially in complex backgrounds. The detection accuracy has been significantly improved.
[0131] In summary, it can be seen from the experimental results that the FGE, MSE module and fine-grained classroom behavior detection module (small target layer) proposed in the present invention have played a positive role in improving the performance of the model, significantly improving the detection accuracy, recall rate and generalization ability of the model, especially in complex backgrounds and fine-grained target detection scenarios.
[0132] In addition, in order to further intuitively verify the detection effect of the EDU-YOLO model, the present invention demonstrates the detection effect in the data set and real scene and demonstrates the classroom teaching behavior analysis system. The detection results are presented in three parts: image detection results in the data set, image detection effect in the real classroom scene, as shown in the following figure. Fig.10 As shown. Fig.10 In the figure, from left to right, they represent teaching scenes in different classrooms. The first picture shows the detection results of students' real classroom behaviors collected in the POCO dataset. It can be seen that YOLOv8 missed the small distant targets in this scene, and the detection of student behaviors close to the camera was relatively fuzzy, while EDU-YOLO accurately detected the behaviors of all students. The second picture shows EDU-YOLO's accurate detection of the behaviors of multiple students in complex scenes. Compared with other models, EDU-YOLO significantly reduces false detections and missed detections. Fig.11In the figure, the detection effect of a real classroom scene is shown. The upper left, upper right, lower left, and lower right correspond to the detection results of YOLOv8, YOLOv10, YOLO11, and EDU-YOLO respectively. YOLOv8 and YOLOv10 perform poorly in small target detection and have some missed detections; while EDU-YOLO can accurately capture the behavior of students in the distance, and can better distinguish the different behaviors of multiple students, especially in dense scenes.
[0133] The effectiveness and practicality of the method of the present invention are verified through experiments on the POCO dataset and the self-made UOCO dataset as well as applications in real classroom scenarios. The experimental results show that EDU-YOLO is more effective in precision, recall, and mAP. 0。5 and mAP 0。5:0。95 It is superior to the existing YOLO series models in key indicators, especially in multi-target detection and small target recognition tasks, and significantly improves the detection accuracy of large and small targets in complex classroom environments. Compared with the most advanced YOLO11 model, it has increased by 1.7%, 7.1%, 6.7% and 6.3% respectively. The EDU-YOLO algorithm significantly improves the accuracy and real-time performance of student behavior detection by introducing innovative feature extraction and fusion mechanisms.
[0134] Example 2
[0135] This embodiment 2 describes a computer device, which includes a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, the steps of the fine-grained teaching behavior recognition method for smart classrooms in the above embodiment 1 are implemented.
[0136] In this embodiment, the computer device is any device or apparatus with data processing capability, which will not be described in detail here.
[0137] Example 3
[0138] This embodiment 3 describes a computer-readable storage medium on which a program is stored. When the program is executed by a processor, it is used to implement the steps of the fine-grained teaching behavior recognition method for smart classrooms in the above-mentioned embodiment 1.
[0139] The computer-readable storage medium can be an internal storage unit of any device or apparatus with data processing capabilities, such as a hard disk or memory, or an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card, etc. equipped on the device.
[0140] Of course, the above description is only a preferred embodiment of the present invention, and the present invention is not limited to the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any technician familiar with the field under the guidance of this specification fall within the essential scope of this specification and should be protected by the present invention.
Claims
1. A fine-grained teaching behavior recognition method for smart classrooms, characterized by: The steps include: Step 1. Obtain classroom teaching video data stream and construct a data set for training the following model; Step 2. Build a teaching behavior detection model based on the improved YOLOv8 model. The teaching behavior detection model includes a feature extraction layer, a feature fusion layer, and detection heads of different scales, where there are four detection heads; First, the video frame is input into the feature extraction module for feature extraction. The process is as follows: In the feature extraction layer, the video frame image first passes through a convolution module, then passes through multiple groups of convolution and fine-grained feature extraction modules for full feature extraction, and finally passes through fast pyramid pooling for feature enhancement to complete feature extraction; Then, the feature fusion layer combines multiple multi-scale feature fusion modules, multiple upsampling modules, multiple connection modules, and multiple convolution modules to perform feature fusion at different depths. Finally, the features at different depths are passed through four detection heads to obtain the detection result information; Step 3. Train the teaching behavior detection model constructed based on the data set in step 1, and then use the trained model to detect the input video images in the real classroom scene to obtain the teaching behavior detection results.
2. According to the fine-grained teaching behavior recognition method for smart classrooms according to claim 1, it is characterized in that: The feature extraction layer includes a convolution module, four groups of convolution and fine-grained feature extraction modules and a fast pyramid pooling module; the four groups of convolution and fine-grained feature extraction modules are defined as the first, second, third and fourth groups of convolution and fine-grained feature extraction modules, respectively, and each group includes a convolution module and a fine-grained feature extraction module; The video frame first passes through the convolution module to extract features, and then passes through the first, second, third, and fourth groups of convolution and fine-grained feature extraction modules to extract features, and finally passes through the fast pyramid pooling module for feature enhancement to complete feature extraction.
3. The fine-grained teaching behavior recognition method for smart classroom according to claim 2 is characterized in that: There are seven multi-scale feature fusion modules, which are defined as the first, second, third, fourth, fifth, sixth and seventh multi-scale feature fusion modules; There are three upsampling modules, which are defined as the first, second, and third upsampling modules; There are three convolution modules, which are defined as the first, second and third convolution modules; there are seven connection modules, which are defined as the first, second, third, fourth, fifth, sixth and seventh connection modules respectively; The output features of the fast pyramid pooling module are first connected with the output features of the fine-grained feature extraction module in the fourth group of convolution and fine-grained feature extraction modules through a first connection module; The output features of the first connection module enter the first upsampling module for upsampling; The output features of the first upsampling module are connected with the output features of the fine-grained feature extraction module in the third group of convolution and fine-grained feature extraction modules through a second connection module; The output features of the second connection module enter the second upsampling module for upsampling; The output features of the second upsampling module are connected with the output features of the fine-grained feature extraction module in the second group of convolution and fine-grained feature extraction modules through a third connection module; The output features of the third connection module enter the third upsampling module for upsampling; The output features of the third upsampling module are connected with the output features of the fine-grained feature extraction module in the first group of convolution and fine-grained feature extraction modules through a fourth connection module; The output features of the fourth connection module are sequentially passed through the first multi-scale feature fusion module and the first convolution module; The output features of the first convolution module enter the fifth connection module; the output features of the third connection module enter the fifth connection module through a second multi-scale feature fusion module and are connected with the output features of the first convolution module; The output features of the fifth connection module are sequentially passed through the third multi-scale feature fusion module and the second convolution module; The output features of the second convolution module enter the sixth connection module; the output features of the second connection module enter the sixth connection module through a fourth multi-scale feature fusion module and are connected with the output features of the second convolution module; The output features of the sixth connection module are sequentially passed through the fifth multi-scale feature fusion module and the third convolution module; The output features of the third convolution module enter the seventh connection module; the output features of the first connection module enter the seventh connection module through a sixth multi-scale feature fusion module and are connected with the output features of the third convolution module; The output features of the seventh connection module enter the seventh multi-scale feature fusion module for processing; The first, third, fifth, and seventh multi-scale feature fusion modules are respectively connected to one of the detection heads.
4. The fine-grained teaching behavior recognition method for smart classroom according to claim 1 is characterized in that: The fine-grained feature extraction module combines the lightweight advantages of the C2f module with the refined feature extraction capability of the attention mechanism in the efficient multi-scale attention module EMA to enhance the feature expression capability of the network.
5. According to claim 4, the fine-grained teaching behavior recognition method for smart classroom is characterized in that: The processing flow of the fine-grained feature extraction module is as follows: The fine-grained feature extraction module first compresses the input features through convolution operations and divides the features into two branches using a split function, which helps to enhance the nonlinear ability and expression ability of the network; One of the branches contains multiple Bottleneck modules, namely Bneck modules, which are used to perform deep extraction of input features to enhance the overall feature extraction capability of the network. The Bneck modules transmit feature information through parallel paths. The other branch is not processed, and the outputs of the two branches are spliced together through the connection operation; Next, some features of the obtained feature group are fused with the original features after passing through the EMA attention mechanism to obtain a new feature group, and finally the features are fused with the convolution structure through feature splicing and output.
6. The fine-grained teaching behavior recognition method for smart classroom according to claim 5 is characterized in that: The specific calculation process of the fine-grained feature extraction module is as follows: First, the features of the input fine-grained feature extraction module are processed, and the input features are recorded as X∈RC×H×W; where C, H and W represent the number of channels, height and width respectively; First, the input features are processed through a convolution operation to obtain a feature map X. The convolution-processed feature map X is then divided into two parts through the channel splitting operation Split, namely: X′=Conv1(X); X1,X2=Split(X′); Conv1 is a standard convolutional layer for preliminary feature extraction. X1 and X2 represent the two branches of the input feature X after segmentation, which are used for subsequent processing by different paths. Then, the segmented X2 is sent to multiple Bottleneck layers to further extract deep features to obtain X2′. After the Bneck layer extracts features, X1 and X′2 are concatenated together through the Concat operation to generate a new feature map Y1, Y2, …, Y m , the formula is as follows: X2′=Bottleneck n (X2); Y1,Y2,…,Y m =Concat(X1,X′2); Where n is the number of repetitions of the Bneck module; Then, feature Y m Input the efficient multi-scale attention module EMA, which is used to dynamically weight the features between channels and highlight the channel features that are critical to the detection task; The result EMA(Y m ) is added to the original feature group; Finally, all the features of the obtained feature group are concat-connected to obtain a processed feature map, which is then passed through a convolutional network to obtain the final feature map Y. The output and input sizes of the module are consistent, that is: Y′=Concat(Y1,Y2,…,EMA(Y m )); Y=Conv2(Y′).
7. The fine-grained teaching behavior recognition method for smart classroom according to claim 1 is characterized in that: The multi-scale feature fusion module includes a dynamic adjustment module AD, a partial convolutional network P_Conv, a multi-layer perceptron MLP and an efficient multi-scale attention module EMA; the processing flow of the multi-scale feature fusion module is as follows: First, the input feature is adjusted to meet the model's size requirements through the dynamic adjustment module AD, and then enters the partial convolution network P_Conv for feature compression processing to generate a compressed feature map; Then, the features are linearly transformed by a multi-layer perceptron (MLP) containing two convolutional layers to further extract features; Subsequently, the DropPath strategy is used to enhance the robustness of the model and reduce overfitting. The efficient multi-scale attention module EMA is used to highlight important features by dynamically adjusting the correlation between channels; With the help of residual connection, the output features after passing through the MLP module and the efficient multi-scale attention module EMA are fused with the input features of the multi-scale feature fusion module to ensure efficient transmission of information.
8. The fine-grained teaching behavior recognition method for smart classroom according to claim 7 is characterized in that: The specific calculation process of the multi-scale feature fusion module is as follows: Where X is the input feature map, X adj It is the feature map after channel adjustment. When the number of input channels inc does not match the number of target channels dim, the number of channels is adjusted through a 1×1 convolution layer; If the input channel does not match the desired output channel, a 1×1 convolution is used for adjustment. The process formula is expressed as follows: X min =PartialConv(X adj ); Y mlp =Conv(ReLU(Conv(X min ))); AND attn =EMA(DropPath(Y mlp )); AND final =X adj +Y attn ; Among them, X min It is the feature map output by the partial convolution network P_Conv, which is processed by the defined PartialConv layer. adj After processing, ReLU is the activation function, DropPath is the drop function, and EMA is the attention mechanism operation; Y mlp , Y attn , Y final They are multi-layer perceptron MLP, efficient multi-scale attention module EMA and fusion features.
9. A computer device comprising a memory and one or more processors; an executable code is stored in the memory; characterized in that: When the processor executes the executable code, it is used to implement the fine-grained teaching behavior recognition method for smart classrooms as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a program stored thereon; characterized in that: When the program is executed by a processor, it is used to implement the fine-grained teaching behavior recognition method for smart classrooms as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Classroom learning behavior identification method based on improved YOLOv8
CN117671781A
Intelligent classroom student classroom behavior identification method based on improved YOLOv8
CN118247730A
Product small target defect detection method based on YoloV8
CN118628885A
Target detection method based on fine-grained features
CN118887378A
Cited By
Lightweight identification method for cavity diseases in road
CN121074632A
Virtual teaching scene-oriented refined behavior identification method and system
CN121861730A
A refined behavior recognition method and system for virtual teaching scenarios
CN121861730B