Student classroom behavior analysis method based on improved YOLOv8

By combining OpenPose and the improved YOLOv8 model, students' classroom behavior analysis methods are constructed, and the problem of low efficiency in classroom attendance and listening status evaluation in the existing technology is solved, accurate identification and teaching optimization of students' behavior are achieved, and teaching quality and learning effect are improved.

CN120299079APending Publication Date: 2025-07-11SHIHEZI UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311373976.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-23
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing classroom attendance and listening status evaluation methods are inefficient, difficult to fully reflect students' learning quality, and it is impossible to adjust teaching strategies in a timely manner.

Method used

Combining OpenPose and the improved YOLOv8 model, through pose estimation and object detection technology, students' classroom behavior analysis methods are constructed, including data preprocessing, bone key point image extraction and improved YOLOv8 recognition model training to achieve accurate identification and classification of student behavior.

Benefits of technology

Provide more accurate student classroom behavior analysis tools to help teachers identify abnormal behaviors, optimize teaching strategies, and improve teaching quality and learning effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299079A_ABST
    Figure CN120299079A_ABST
Patent Text Reader

Abstract

The invention discloses an improved YOLOv8-based student classroom behavior analysis method. The method comprises the steps of obtaining an initial data set based on a monitoring video; preprocessing the initial data set to obtain a student classroom behavior data set; acquiring skeleton key point images corresponding to different classroom behaviors based on the student classroom behavior data set; constructing an improved YOLOv8 recognition model, and training the improved YOLOv8 recognition model based on the skeleton key point image to obtain a trained recognition model; and obtaining a student classroom image, and obtaining a student classroom behavior classification result based on the student classroom image and the trained identification model. According to the method, classroom behaviors of students are classified through the improved YOLOv8 model, and a more accurate behavior analysis tool can be provided for teachers and education researchers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of student classroom behavior recognition, and particularly to a method for analyzing student classroom behavior based on improved YOLOv8. Background Art

[0002] In the context of big data, educational informatization needs to further explore new teaching models based on information technology. Although offline teaching remains the core of education, the learning efficiency of students in the classroom and the quality of classroom teaching have a direct impact on their knowledge acquisition. However, in actual teaching activities, the number of students in class at the same time is huge, and it is difficult for teachers to accurately grasp the listening situation of each student. Traditional classroom attendance methods and methods for evaluating students' listening states mainly rely on means such as teachers' classroom observations, manual observation of surveillance videos in the classroom, and questionnaires. This method has a certain lag, low efficiency, and cannot comprehensively reflect the learning quality of students in the classroom. At the same time, it is also difficult to collect and analyze data, and it is impossible to obtain students' classroom behaviors in a timely manner and adjust and optimize teaching strategies based on students' classroom behaviors.

[0003] Therefore, in the current educational context, it is very important and necessary to explore new teaching models based on information technology, which can promote the intelligent, personalized, and high-quality development of education and provide better learning environments and teaching experiences for students.

[0004] In this context, new teaching models based on information technology are urgently needed to be explored. Summary of the Invention

[0005] In view of the existing technical problems, the present invention proposes a method for analyzing student classroom behavior based on OpenPose and improved YOLOv8. This method combines pose estimation and object detection technologies to analyze the poses and actions of students, and can accurately identify the behaviors and performances of students in the classroom.

[0006] A method for analyzing student classroom behavior based on improved YOLOv8 provided by the present invention includes:

[0007] Obtaining an initial data set based on a surveillance video;

[0008] Preprocessing the initial data set to obtain a student classroom behavior data set;

[0009] Obtaining skeletal key point images corresponding to different classroom behaviors based on the student classroom behavior data set;

[0010] Constructing an improved YOLOv8 recognition model, and training the improved YOLOv8 recognition model based on the skeletal key point images to obtain a trained recognition model;

[0011] Obtain the student classroom image, and obtain the student classroom behavior classification result based on the student classroom image and the trained recognition model.

[0012] Optionally, the process of obtaining the initial dataset based on the monitoring video includes:

[0013] Perform frame extraction on the monitoring video to obtain a picture dataset;

[0014] Perform object annotation on the picture dataset based on the YOLOv8 object detection model to obtain the initial dataset.

[0015] Optionally, the process of preprocessing the initial dataset to obtain the student classroom behavior dataset includes:

[0016] Perform data cleaning on the initial dataset to obtain the first dataset;

[0017] Perform data cropping on the first dataset to obtain the second dataset;

[0018] Perform data augmentation on the second dataset to obtain the third dataset;

[0019] Perform feature partitioning on the third dataset to obtain the student classroom behavior dataset.

[0020] Optionally, the methods of data cleaning include the CNN model training method and the manual observation method.

[0021] Optionally, use the openpose pose estimation algorithm to perform feature recognition on the student classroom behavior dataset to obtain the skeletal key point images corresponding to different classroom behaviors.

[0022] Optionally, the classroom behaviors include: listening attentively, reading a book, writing, looking around, raising a hand, and standing.

[0023] Optionally, the improved YOLOv8 recognition model includes introducing the CBAM attention mechanism and deformable convolution in the Backbone network layer of the YOLOv8 model.

[0024] Optionally, the CBAM attention mechanism performs comprehensive attention weighting on image features based on channel attention and spatial attention;

[0025] Among them, the calculation formula for performing comprehensive attention weighting is:

[0026]

[0027] In the formula, M s (F) represents the spatial attention module, and f 7×7 represents the convolution operation with a convolution kernel of 7×7. represents the global average pooling feature, represents the max pooling feature.

[0028] Optionally, the process of training the improved YOLOv8 recognition model based on the skeletal key point images further includes evaluating the classification results based on an evaluation model;

[0029] wherein, the evaluation metrics of the evaluation model include accuracy, recall rate, and F1 value.

[0030] The present invention has the following technical effects:

[0031] By using the improved YOLOv8 model for classifying students' classroom behaviors, a more accurate behavior analysis tool can be provided for teachers and educational researchers. This helps to evaluate students' performance and behaviors in the classroom, identify abnormal behaviors, and provide more targeted guidance and adjustment for teaching. This will promote the improvement of teaching quality, optimize the learning environment, and enhance students' learning effects and participation. Description of the Drawings

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 is a flowchart of the method for analyzing students' classroom behaviors based on the improved YOLOv8 in the embodiments of the present invention;

[0034] Figure 2 is the Backbone model diagram of the improved YOLOv8 in the embodiments of the present invention. Detailed Embodiments

[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0036] The present invention includes the following steps: Step 1: Collect multiple classroom behavior monitoring videos of students, and perform data preprocessing and data partitioning on these videos. This includes operations such as data annotation, data augmentation, and data cropping, and a student classroom behavior dataset is obtained through partitioning. Step 2: Use the OpenPose pose estimation algorithm to extract the pose information of students, such as the positions and movements of the arms, neck, and head. Step 3: Construct an improved YOLOv8 model. The accuracy of the model is improved mainly by introducing the CBAM attention mechanism and deformable convolution (DCN) in the Backbone network layer to detect and classify students' behaviors. Step 4: Use the dataset samples to train the improved YOLOv8 model, thereby obtaining a student classroom behavior detection model and classification results. The method of the present invention is of great significance, can accurately evaluate the performance and behaviors of students in the classroom, identify abnormal behaviors, and assist teachers in classroom teaching to improve teaching quality. The following is a specific description through specific embodiments.

[0037] Embodiment 1

[0038] This embodiment provides a method for analyzing students' classroom behaviors based on the improved YOLOv8, including the following steps:

[0039] Obtain an initial dataset based on the monitoring video, and the specific process includes:

[0040] Obtain students' classroom teaching videos;

[0041] Perform frame extraction on the video data, and convert the video into a picture dataset at a certain time interval;

[0042] Perform target annotation on the student targets in the picture dataset. Detect all student targets using the YOLOv8 object detection model, and crop and save each target to obtain the initial dataset.

[0043] Perform preprocessing on the initial dataset to obtain a student classroom behavior dataset, and the specific process includes:

[0044] Preprocess the initial dataset, including data cleaning, data cropping, and data augmentation to obtain the student classroom behavior dataset. Data cleaning of the initial dataset yields the first dataset, specifically: In the data cleaning stage, by using the method of training with a CNN model and manual observation, those image data that are severely occluded or of poor quality are removed. This can reduce the impact of noise and invalid information and improve the accuracy and reliability of subsequent analysis. Data cropping of the first dataset yields the second dataset, specifically: In the cropping stage, while keeping the image ratio unchanged, with the center point of each image as the reference, the scaled image is uniformly cropped to a size of 224×224 pixels. The purpose of this is to ensure the consistency of the input data, enabling the model to better learn and understand the features and information in the images. Data augmentation of the second dataset yields the third dataset, specifically: Augmentation involves operations such as flipping the dataset to increase the diversity and richness of data samples, which can generate more training samples and improve the generalization ability and robustness of the model.

[0045] Feature partitioning of the third dataset yields the student classroom behavior dataset, specifically: The preprocessed dataset is divided into six types of behaviors, namely listening attentively in class, reading a book, writing, looking around, raising a hand, and standing, through feature partitioning to obtain the student classroom behavior dataset.

[0046] Obtain the skeletal key point images corresponding to different classroom behaviors based on the student classroom behavior dataset, specifically including:

[0047] Input the student classroom behavior dataset into the openpose pose estimation algorithm to obtain the skeletal key point pictures of the students, including: nose, neck, right shoulder, right elbow, right wrist, left shoulder, left elbow, left wrist, right hip, right eye, left eye, right ear, left ear. Classification is performed through different human skeletal diagrams.

[0048] Construct an improved YOLOv8 recognition model, and train the improved YOLOv8 recognition model based on the skeletal key point images to obtain a trained recognition model, specifically including:

[0049] Construct an improved YOLOv8 model. To improve the accuracy of the model, the CBAM attention mechanism and deformable convolution (DCN) are introduced into the Backbone network layer of YOLOv8.

[0050] The specific implementation is as follows: Among them, the CBAM attention mechanism multiplies the channel attention and spatial attention:

[0051]

[0052] Among them, M s (F) is the spatial attention module, f7×7 Represents a convolution operation with a 7×7 convolution kernel, represents global average pooling features, represents max pooling features.

[0053] The CBAM attention mechanism can perform comprehensive attention weighting on image features, enabling the model to pay more attention to important channel features and spatial locations and suppressing unimportant features. This helps improve the model's perception and extraction ability of key features, thereby enhancing the model's accuracy and robustness.

[0054] Deformable Convolution (DCN) introduces a deformable convolution kernel. Deformable convolution automatically adjusts the sampling method of the convolution kernel at different positions by learning the parameters of the offset, thereby achieving flexible modeling of the object shape in the image. In the improved YOLOv8 model, introducing deformable convolution can enable the model to more flexibly adapt to different postures and action changes of students in the classroom and improve the accurate detection and classification ability of students' behaviors.

[0055] Use the dataset samples to train the improved YOLOv8 model. Optimize the model parameters through an iterative training process to enable it to accurately identify and classify different behaviors of students. During the training process, the CBAM attention mechanism and the introduced deformable convolution are added. These improvement measures help improve the performance and accuracy of the model. During the training process, the model learns the images and labels in the dataset samples and gradually adjusts the weights and biases of the model to minimize the difference between the prediction results and the actual labels. Through continuous iterative training, the model gradually improves its understanding and discrimination ability of students' behaviors, thereby achieving the purpose of more accurately classifying students' classroom behaviors.

[0056] Obtain students' classroom images, and based on the students' classroom images and the trained recognition model, obtain the classification results of students' classroom behaviors, specifically including:

[0057] After training is completed, the obtained model can be applied to the classification task of students' classroom behaviors. The model can input an image of students' classroom behaviors and, through the inference process of the model, accurately judge the behavior category of the student, such as listening attentively, reading a book, writing, looking around, raising a hand, and standing. The optimized model has higher accuracy and robustness and can better adapt to different scenarios and changes in students' behaviors.

[0058] The present invention can achieve automated analysis and evaluation of students' behaviors in the classroom by introducing advanced computer vision and machine learning technologies, namely OpenPose and the improved YOLOv8. By collecting and processing and analyzing students' classroom behavior monitoring videos, more accurate and real-time information about students' behaviors can be provided, providing more precise classroom guidance and personalized teaching for teachers.

[0059] This education informatization model based on big data can not only improve teaching efficiency and quality, but also conveniently collect and analyze a large amount of education data, providing a scientific basis for education decision-making and improvement. At the same time, students can also obtain personalized learning feedback and guidance through these technologies to enhance learning effects.

[0060] Example Two

[0061] As Figure 1 shown, this embodiment provides a method for analyzing students' classroom behaviors based on the improved YOLOv8, including the following steps:

[0062] Obtain an initial data set based on the surveillance video. The specific process includes: To meet the establishment of the students' classroom behavior data set, the students' data set is obtained from the classroom surveillance video of a primary school in Anhui Province. The video resolution is 2560×1440, and each video has 60 minutes of video data. Perform frame extraction on the surveillance video to obtain a picture data set. Specifically, it includes: converting the video stream data into several frames of images, and according to the classroom behaviors of each student in the images, using a data annotation tool to annotate the images, including using a rectangular box to frame the position of each student and indicating the name of the student's behavior, and dividing the data set into a training set and a test set.

[0063] Perform target annotation on the picture data set based on the YOLOv8 object detection model to obtain an initial data set. Specifically, it includes: annotating the images through the data annotation tool LabelImg, and accurately annotating the position of each student target in the pictures with bounding boxes. Detect each student target by using the YOLOv8 object detection model, and crop and save each target to obtain the initial data set.

[0064] Preprocess the initial data set to obtain the students' classroom behavior data set. The specific process includes: Perform data cleaning on the initial data set to obtain the first data set. Specifically, it includes: preprocessing the data set, first removing those picture data that are severely occluded or of poor quality through CNN model training and the manual observation method. Perform data cropping on the first data set to obtain the second data set. Specifically, it includes: cropping the remaining data, and uniformly cropping the scaled images to a size of 224×224 pixels. Perform data augmentation on the second data set to obtain the third data set. Specifically, it includes: performing operations such as flipping on the second data set to increase the diversity and richness of the data samples. Perform feature division on the third data set to obtain the students' classroom behavior data set. Specifically, it includes: dividing the third data set into six types of behaviors: listening attentively in class, reading books, writing, looking around, raising hands, and standing, to obtain the students' classroom behavior data set.

[0065] Based on the student classroom behavior dataset, obtain the skeletal key point images corresponding to different classroom behaviors, specifically including:

[0066] On the dataset established above, use the OpenPose human pose estimation algorithm to obtain the skeletal key point images of each category.

[0067] Construct an improved YOLOv8 recognition model, and train the improved YOLOv8 recognition model based on the skeletal key point images to obtain a trained recognition model, specifically including:

[0068] According to the previous scheme, add an attention mechanism and deformable convolution to the YOLOv8 model. The student classroom behavior classification model obtained is the improved YOLOv8 student classroom behavior classification model shown in Table 1 below. Among them, each row represents a layer of the backbone network, including the module type and parameters. P1 / 2 indicates that the size of the output feature map is half of the size of the input image.

[0069] Table 1

[0070] module args Output size Conv [64,3,2] P1 / 2 Conv [128,3,2] P2 / 4 C2f [128,True] P2 / 4 Conv [256,3,2] P3 / 8 C2f_DCN [256,True] P3 / 8 Conv [512,3,2] P4 / 16 C2f_DCN [512,True] P4 / 16 Conv [1024,3,2] P5 / 32 C2f_DCN [1024,True] P5 / 32 CBAMBlock [16,7] P5 / 32

[0071] Using the student classroom behavior dataset, different classification algorithms can be trained, and the classification training effects of the datasets with and without using the OpenPose algorithm can be compared, including:

[0072] The classification training effect of the dataset with or without using the OpenPose algorithm on the improved YOLOv8 algorithm;

[0073] The comparison of the classification model effects of the dataset using the OpenPose algorithm on YOLOv8 and the improved YOLOv8 respectively. The experimental results are shown in Table 2.

[0074] Table 2

[0075] Accuracy Recall F1 score YOLOv8 0.872 0.793 0.816 YOLOv8+openpose 0.889 0.847 0.866 Optimized YOLOv8 0.880 0.823 0.838 Optimized YOLOv8+openpose 0.902 0.863 0.879

[0076] As Figure 2 shown, the present invention makes full use of the advantages of introducing the CBAM attention mechanism and deformable convolution (DCN) in the Backbone network layer of YOLOv8. The CBAM attention mechanism improves the model's attention and discrimination ability for the key areas of student behaviors by allocating the weights of channel features and the weights of spatial features of the feature map, suppressing the weights of invalid features and increasing the weights of useful features, thereby improving the overall accuracy of object detection. The deformable convolution (DCN) enhances the model's accurate detection ability for student behaviors by introducing deformable convolution kernels.

[0077] The improved YOLOv8 model uses the student classroom behavior dataset during the training process. By iteratively training and optimizing the model parameters, it can accurately identify and classify different behaviors of students. Compared with the dataset using the OpenPose algorithm, the model can also achieve good classification results on the dataset without using the OpenPose algorithm, indicating that the improved model has a certain generalization ability.

[0078] In this embodiment, the dataset is divided into a training set and a test set in a ratio of 4:1. The training set is used to train the improved YOLOv8 model, and the test set is used for testing to obtain a classification model for student classroom behavior detection. By evaluating the performance of the model on the test set, including indicators such as accuracy, recall rate, and F1 value, the overall performance and generalization ability of the model are evaluated.

[0079] Obtain student classroom images, and based on the student classroom images and the trained recognition model, obtain the classification results of student classroom behaviors.

[0080] In addition, the experimental results show that on the dataset using the OpenPose algorithm, the improved YOLOv8 model shows better performance in student behavior classification compared to the original YOLOv8 model. This indicates that by introducing the CBAM attention mechanism and deformable convolution (DCN), the model can more accurately distinguish the behavior characteristics of students and improve the accuracy and reliability of classification.

[0081] Therefore, the improved YOLOv8 model demonstrates good performance and application prospects in the task of student classroom behavior classification. It can effectively assist teachers in evaluating students' performance and behaviors in the classroom, identifying abnormal behaviors, and providing data support and decision-making basis, thereby improving teaching quality and learning effects.

[0082] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for analyzing students' classroom behaviors based on improved YOLOv8, characterized in that, Including: Obtaining an initial data set based on the surveillance video; Preprocessing the initial data set to obtain a student classroom behavior data set; Obtaining skeletal key point images corresponding to different classroom behaviors based on the student classroom behavior data set; Constructing an improved YOLOv8 recognition model, training the improved YOLOv8 recognition model based on the skeletal key point images to obtain a trained recognition model; Obtaining a student classroom image, and obtaining a student classroom behavior classification result based on the student classroom image and the trained recognition model.

2. The method for analyzing students' classroom behaviors based on the improved YOLOv8 according to claim 1, wherein The process of obtaining an initial data set based on the surveillance video includes: Performing frame extraction on the surveillance video to obtain a picture data set; Performing target annotation on the picture data set based on the YOLOv8 object detection model to obtain an initial data set.

3. The method for analyzing students' classroom behaviors based on the improved YOLOv8 according to claim 1, characterized in that The process of preprocessing the initial data set to obtain a student classroom behavior data set includes: Performing data cleaning on the initial data set to obtain a first data set; Performing data cropping on the first data set to obtain a second data set; Performing data augmentation on the second data set to obtain a third data set; Performing feature partitioning on the third data set to obtain a student classroom behavior data set.

4. The method for analyzing students' classroom behaviors based on the improved YOLOv8 according to claim 3, wherein, The data cleaning method includes the CNN model training method and the manual observation method.

5. The method for analyzing students' classroom behaviors based on the improved YOLOv8 according to claim 1, wherein, Using the openpose pose estimation algorithm to perform feature recognition on the student classroom behavior data set to obtain skeletal key point images corresponding to different classroom behaviors.

6. The method for analyzing students' classroom behaviors based on the improved YOLOv8 according to claim 5, wherein, The classroom behaviors include: listening attentively in class, reading books, writing, looking around, raising hands, and standing.

7. The method for analyzing students' classroom behaviors based on the improved YOLOv8 according to claim 1, wherein The improved YOLOv8 recognition model includes introducing the CBAM attention mechanism and deformable convolution in the Backbone network layer of the YOLOv8 model.

8. The method for analyzing students' classroom behaviors based on the improved YOLOv8 according to claim 7, wherein, The CBAM attention mechanism performs comprehensive attention weighting on image features based on channel attention and spatial attention; Among them, the calculation formula for performing comprehensive attention weighting is: where M s (F) represents the spatial attention module, and f 7×7 represents the convolution operation with a 7×7 convolution kernel, represents the global average pooling feature, represents the max pooling feature.

9. The method for analyzing students' classroom behaviors based on the improved YOLOv8 according to claim 1, characterized in that The process of training the improved YOLOv8 recognition model based on the skeletal key point images further includes evaluating the classification result based on an evaluation model; Among them, the evaluation metrics of the evaluation model include accuracy, recall rate, and F1 value.

Citation Information

Cited By

  • Construction method of recognition model based on YOLOv8 and security recognition system

    CN120976532A