Student classroom behavior classification method based on improved lightweight network

By collecting video and physiological information data in the classroom and using improved lightweight network models for data fusion, in-depth, comprehensive and personalized analysis of students' classroom behavior is achieved, and the problem of difficulty in identifying and monitoring students' individual differences in the existing technology is solved, which improves the pertinence and effectiveness of education and teaching.

CN119964051APending Publication Date: 2025-05-09CHANGCHUN UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510026594.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

It is difficult for existing technology to achieve in-depth, comprehensive and personalized analysis of students' classroom behaviors, and it is impossible to effectively identify and monitor students' individual differences, making it difficult to implement personalized education.

Method used

The student classroom behavior classification method is adopted based on an improved lightweight network. By installing a camera in the classroom to collect video data and equipped with a smart bracelet to collect physiological information data. Combining MobileNetV1 as the baseline model, the image data and bracelet data are fused by the feature dimension alignment method to achieve accurate classification of classroom behavior.

Benefits of technology

Real-time and accurate detection of students' classroom behaviors is achieved, and timely feedback on students' learning status can help teachers take personalized intervention measures to improve students' learning quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964051A_ABST
    Figure CN119964051A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of education data mining, and particularly relates to a student classroom behavior classification method based on an improved lightweight network, which comprises the following steps of: 1, placing a camera with a built-in SD (Secure Digital) card at the center of a classroom blackboard, acquiring a student classroom video as a video data source, and intercepting the video into a video frame image; before class, a student wears a customized smart bracelet on the wrist, and dynamic physiological information data of the student during video recording is collected as a physiological information data source. By fusing multi-source data, individual difference requirements of students are met, and the analysis method is more comprehensive and scientific; on the basis of a traditional MobileNet lightweight algorithm, another data mode is fused, complementation of feature levels is achieved, and the defect of single-mode data is effectively overcome. Meanwhile, compared with a MobileNet baseline model, the performance is more excellent, and the practicability and effectiveness in actual classroom application are manifested.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of educational data mining, and in particular relates to a student classroom behavior classification method based on an improved lightweight network. Background Art

[0002] With the rapid development of modern technologies such as artificial intelligence, the Internet of Things, and big data and their application in all walks of life, the teaching process is gradually moving towards intelligence, precision, and personalization. The classroom is an important place for teachers to teach and students to acquire knowledge. Learners' input status, participation methods, and learning attitudes in the classroom are increasingly becoming the starting point and driving force for the occurrence and development of learning. Using information technology to analyze the learning status of students in the classroom can not only play a role in regulating and guiding students' behavior in the classroom, but also quantitatively reflect the degree of classroom activity from multiple dimensions, so that teachers can efficiently and intuitively grasp the situation of students' learning behavior input, and provide data support for subsequent optimization of teaching design and implementation of teaching intervention.

[0003] In traditional classrooms, teachers often use homework, questions and observation to roughly understand students' learning status and learning needs and carry out teaching for the whole class. While taking into account the teaching content, it is difficult for teachers to pay attention to the performance of each student, such as frequent bowing, whispering, sleepiness and other behaviors that are not focused, which in turn affects the students' learning status. At the same time, students' learning interests, knowledge levels, learning abilities, learning psychology and emotional states are diverse and complex. Teachers often lack sufficient energy to pay attention to the individual characteristics and differences of all students, and it is difficult to take timely intervention measures that match them, and personalized education is difficult to truly implement. In addition, most of the existing studies on students' classroom behavior using intelligent means are based on cameras or sensors, and some are assisted by other equipment. There are limitations such as complex and expensive equipment, low recognition due to camera angle problems, inconvenient classroom power supply, and inability to solve individual differences among students.

[0004] In view of this, in the context of the continuous pursuit of precision and personalized development in the field of education, in order to achieve a more in-depth, comprehensive and individual analysis of students' classroom behavior, an innovative classroom behavior classification method is urgently needed to achieve accurate identification and all-round monitoring of classroom behavior, thereby providing more targeted and effective support for educational and teaching activities. Summary of the invention

[0005] 1. Technical issues to be resolved

[0006] In view of the deficiencies of the prior art, the present invention provides a method for classifying student classroom behaviors based on an improved lightweight network, which solves the problems raised in the above-mentioned background technology.

[0007] (II) Technical solution

[0008] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0009] A method for classifying student classroom behaviors based on an improved lightweight network includes the following steps:

[0010] Step 1: Place a camera with a built-in SD card in the center of the classroom blackboard, collect students' classroom videos as the video data source, and capture the videos into video frame images; before class, let students wear customized smart bracelets on their wrists, and collect students' dynamic physiological information data during video recording as the physiological information data source;

[0011] Step 2: Divide the video clips by frames, capture the single person image in each frame, and use different Arabic numeral labels to represent different classroom behaviors with higher frequency; according to the timeline, arrange the collected disordered wristband physiological data in the order of the time when the data was received, crop the frames containing the wristband data timestamp, and then crop the single person image of the wristband wearer that matches the IMEI number of each wristband, and format and name the images containing the wristband physiological information, and then classify them according to the classroom behavior categories set above to form a multi-source fusion data set;

[0012] Step 3: Read all images in the data set in order by category, convert the images into RGB format, unify the image size, parse the formatting and naming of the images, and perform data preprocessing on the unified size images, parsed physiological information data (image formatting and naming), and label (class folder name) data respectively;

[0013] Step 4: Using MobileNetV1 as the baseline model, a feature dimension alignment method is used to fuse the image data and the bracelet. The residual is pulled out after the first Dw convolution, and the data is fused before the last Pw convolution. The bracelet data is used as the residual branch data and finally fused into the main image data. Finally, the final classification result is output through the fully connected layer.

[0014] Step 5: Evaluate the effect of classroom behavior classification through accuracy, precision, recall rate and F1 value to achieve comprehensive and accurate detection of students' classroom behavior.

[0015] Furthermore, the smart bracelet customized in step 1 has built-in different sensors for heart rate, body temperature, and blood pressure. During the video recording process, the background synchronously obtains the physiological information data of each student in real time, and uploads it to the cloud platform through the NBIOT network. The cloud platform is configured with the URL address of the computer terminal and the corresponding receiving port. The API interface is used to send the received physiological information data to the server for reception, parsing and saving in a csv file. Considering the class duration and analysis effect, the data sending frequency is set to once per minute.

[0016] Furthermore, in step 2, OpenCV is used to segment the video clip into 1s / frame; among the classroom behaviors appearing in the video, listening, lowering the head, talking and sleeping are the students' classroom behaviors that appear more frequently, so the numbers 1-4 are used to represent these four types of actions, 1 represents raising the head to listen, 2 represents lowering the head, 3 represents talking, and 4 represents sleeping; the picture is formatted and named to include the wristband physiological information, indicating the wristband physiological data with the same timestamp as the cropped subject image, and each physiological data is connected with &.

[0017] Furthermore, in step 3, OpenCV is used to unify the image size and adjust all of them to 224×224 size; then the formatting and naming of the picture are parsed, and the images after the unified size, the parsed physiological information data (image formatting and naming) and the label (class folder name) data are converted into NumPy array format in turn, and saved as NumPy array files respectively, and stored in a list; then the three lists are exported into three files in npy format: image_data.npy, labels.npy, and numeric_data.npy, which store the image, bracelet and label data respectively; then the above three npy files are read, and pytorch is used to define the data structure of the multi-source fusion data set, and the image and bracelet data are aligned and encapsulated together for easy model training; finally, the class of the data set is defined and the data set object is instantiated, and train_DataLoader and val_DataLoader are generated according to the ratio of training set: validation set = 8:2 for model experiments.

[0018] Furthermore, the method for aligning the feature dimensions in step 4 is to first change the number of feature channels of the five physiological data collected by the bracelet from 5 to 16 through a linear layer and an activation function layer (ReLu), and then reshape the bracelet data into a tensor consistent with the image feature channel through a reshape operation, and then let the bracelet data pass through a layer of deep separable convolution to fully learn the feature distribution of the feature space image, and add the bracelet data to the network as a feature of the image.

[0019] Furthermore, the specific formulas for accuracy, precision, recall and F1 value in step 5 are as follows:

[0020]

[0021]

[0022] Among them, TP represents true positive examples (the number of samples predicted by the model to be positive and actually are positive), TN represents true negative examples (the number of samples predicted by the model to be negative and actually are negative), FP represents false positive examples (the number of samples predicted by the model to be positive but actually are negative), and FN represents false negative examples (the number of samples predicted by the model to be negative but actually are positive).

[0023] (III) Beneficial effects

[0024] Compared with the prior art, the present invention provides a method for classifying student classroom behaviors based on an improved lightweight network, which has the following beneficial effects:

[0025] Compared with traditional classroom learning assessment methods, such as homework, questions, and observation, the present invention has significant advantages in terms of real-time assessment. Traditional assessment methods often have time lags and are difficult to capture students' learning dynamics in real time. The present invention can provide real-time feedback on students' learning status. This feature enables teachers to implement classroom intervention measures more promptly, thereby significantly improving students' learning quality and efficiency.

[0026] Previous classroom status recognition methods mostly relied on a single data source, which made it difficult to fully consider the individual differences of students and caused the problem of one-sided analysis. However, the present invention, by integrating multi-source data, can characterize students' classroom behavior and learning status from multiple angles, fully meeting the individual differences of students in the learning process and making the analysis method more comprehensive and scientific.

[0027] Based on the traditional MobileNet lightweight algorithm, the present invention cleverly incorporates another data modality. When describing complex classroom scenes, single-modal data inevitably has the limitation of incomplete information. By introducing a new data modality, the present invention achieves the complementarity of different modal data at the feature level, effectively making up for the shortcomings of single-modal data. Experimental verification shows that compared with the MobileNet baseline model, the model constructed by the present invention performs better in performance, fully demonstrating the practicability and effectiveness of the model in actual classroom applications, and providing strong technical support for the intelligent development of the education field. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of the method flow of the present invention;

[0029] Figure 2 It is a schematic diagram of the data processing flow of the present invention;

[0030] Figure 3 This is a framework diagram of an improved lightweight network model of the present invention;

[0031] Figure 4 A diagram showing comparative experimental results of various models in the experiments conducted by the present invention;

[0032] Figure 5 Evaluation index result diagram based on the improved lightweight model for each behavior category set for the present invention. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] Example

[0035] like Figure 1-5 As shown, a method for classifying student classroom behaviors based on an improved lightweight network in this embodiment is shown, and the identification method specifically includes the following steps:

[0036] Step 1: Data Collection:

[0037] Data collection consists of two parts:

[0038] Part of it is the collection of classroom video data. This embodiment uses the EZVIZ C6C 4K camera, which has a built-in traffic card and an SD card for storing videos. The SD card size is 64GB, which can meet the video storage requirements for 24 hours a day. The external data cable is connected to the classroom power supply to ensure that the camera has sufficient power.

[0039] During data collection, a camera with a built-in SD card was placed in the center of the classroom blackboard, and a magnet was attached under the camera to facilitate adsorption on objects such as the blackboard and easy movement. Video data of real classroom scenes of 24 students in a class of a technical secondary school throughout the semester was collected as the video data source, and the video was intercepted into video frame images;

[0040] The other part is the collection of classroom physiological information data;

[0041] The wristband used in this embodiment is the B2315+ series wristband developed by Shanghai Oufu Communication Technology Co., Ltd. The customized smart wristband has built-in different sensors for heart rate, body temperature, and blood pressure. During the video recording process, the background synchronously obtains the physiological information data of each student in real time, and uploads it to the cloud platform through the NBIOT network. The cloud platform is configured with the URL address and the corresponding receiving port where the computer terminal is located. The received physiological information data is sent to the server using the API interface to receive, parse and save it in a csv file. Researchers can use the cloud platform to send instructions to the smart wristband to modify the frequency of its data collection. Considering the class duration and analysis effect, the data transmission frequency is set to once per minute.

[0042] During the data collection period, all wristbands are charged before each use. The wristbands can be used for more than 48 hours when fully charged. The wristbands are equipped with adjustable buckles. Before each class, the wristbands are distributed to each student in the class to ensure that they are worn. They are also equipped with an app to observe the physiological signal monitoring status in real time to ensure the effective collection of data.

[0043] Step 2: Dataset construction:

[0044] Use OpenCV to split the video clips into 1s / frame and capture the image of a single person in each frame. OpenCV is a cross-platform computer vision and machine learning software library released under the Apache 2.0 license (open source). It can run on Linux, Windows, Android and Mac OS operating systems, and provides interfaces for languages ​​such as Python, Ruby, and MATLAB, which implements many common algorithms in image processing and computer vision.

[0045] Different Arabic numeral labels are used to represent different classroom behaviors that occur more frequently. Among the classroom behaviors that appear in the collected videos, listening, lowering the head, talking and sleeping are student classroom behaviors that occur more frequently. Therefore, the numbers 1-4 are used to represent these four types of actions, 1 represents raising the head to listen, 2 represents lowering the head, 3 represents talking, and 4 represents sleeping.

[0046] According to the timeline, the collected disordered wristband physiological data are arranged in the order of the time when the data is received, and the frame containing the wristband data timestamp is cropped. Then, the single image of the wristband wearer matching the IMEI number of each wristband is cropped, and the image is formatted and named containing the wristband physiological information. The image formatting name indicates that the wristband physiological data with the same timestamp as the cropped subject image is connected with &. The example is as follows:

[0047] heartbeat=76&bodyTemperature=36.3&wristTemperature=31.1&diastolic=75&shrink=114.png. Then, the students are classified according to the above-mentioned classroom behavior categories and organized into a multi-source fusion data set.

[0048] Step 3: Data Processing:

[0049] Read all the images in the dataset one by one by category, convert the images to RGB format, and then use OpenCV to unify the image size and resize them all to 224×224.

[0050] Then parse the formatting and naming of the image, convert the unified-size image, the parsed physiological information data (image formatting and naming), and the label (class folder name) data into NumPy array format, save them as NumPy array files, and store them in lists. Then export the three lists into three files in npy format: image_data.npy, labels.npy, and numeric_data.npy, which store the image, bracelet, and label data respectively.

[0051] Then read the above three npy files, use pytorch to define the data structure of the multi-source fusion dataset, align and encapsulate the images and bracelet data together to facilitate model training. Finally, define the dataset class and instantiate the dataset object, and generate train_DataLoader and val_DataLoader for model experiments according to the ratio of training set: validation set = 8:2.

[0052] Step 4: Network model construction:

[0053] MobileNetV1 is an efficient convolutional neural network designed for mobile and embedded vision applications. Deep separable convolution is the core structure of MobileNetV1, which consists of depthwise convolution (DW) and pointwise convolution (PW). When the standard convolution operation performs convolution on the input feature map, it will operate on all channels at the same time, which has a large number of parameters and calculations. The deep separable convolution decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution, which greatly reduces the number of parameters and calculations.

[0054] This embodiment uses MobileNetV1 as the baseline model and adopts a feature dimension alignment method to fuse the image data and the bracelet. Specifically, the five physiological data collected by the bracelet are first passed through a linear layer, as shown in the attached manual. Figure 3The FC5-16 and activation function layer (ReLu) in the bracelet data change the number of feature channels from 5 to 16, and then reshape the bracelet data into a tensor consistent with the image feature channels through the reshape operation, and then let the bracelet data pass a layer of depth-separable convolution, as shown in the attached manual. Figure 3 Dwconv3-512 in the network is used to fully learn the feature distribution of feature space images, and the bracelet data is added to the network as a feature of the image. The residual is pulled out after the first Dw convolution, and data is fused before the last Pw convolution. The bracelet data is used as the data of the residual branch and finally fused into the image data of the trunk. Finally, the final classification result is output through the fully connected layer.

[0055] Step 5: Measure the classification performance:

[0056] The classroom behavior classification effect is evaluated through accuracy, precision, recall rate and F1 value, so as to achieve comprehensive and accurate detection of students' classroom behavior.

[0057] Accuracy refers to the ratio of the number of samples predicted correctly by the model to the total number of samples. That is, the ratio of the number of correct predictions (including true positives and true negatives) to the number of all predictions (including true positives, true negatives, false positives, and false negatives). The calculation formula is as follows:

[0058]

[0059] Among them, TP represents true positive examples (the number of samples predicted by the model to be positive and actually are positive), TN represents true negative examples (the number of samples predicted by the model to be negative and actually are negative), FP represents false positive examples (the number of samples predicted by the model to be positive but actually are negative), and FN represents false negative examples (the number of samples predicted by the model to be negative but actually are positive).

[0060] Accuracy refers to the proportion of samples that are actually positive among all samples predicted by the model to be positive, also known as precision. It measures the reliability of the model's prediction of positive classes, and its calculation formula is as follows:

[0061]

[0062] Recall rate refers to the proportion of samples that are actually positive that are correctly predicted as positive by the model, also known as recall rate. It measures the ability of the model to find all positive samples. The calculation formula is as follows:

[0063]

[0064] The F1 value is the harmonic mean of precision and recall, which combines the two indicators of precision and recall and can more comprehensively reflect the performance of the model. Its calculation formula is as follows:

[0065]

[0066] The experimental results and the evaluation index results of each behavior category based on the improved lightweight model are shown in the attached Figure 4 , Figure 5 As shown. The MobileNet series of networks meet the needs of limited resources and fast classification in classroom scenarios. In the training of the classification model for the data set, the classification accuracy reached more than 79%, with the highest being 85.57%. At the same time, vgg16 was used for comparative experiments. The accuracy of the classification results was improved compared with MobileNet, but the difference was not much. Considering the practicality of classroom application scenarios, MobileNet is still a more suitable classification method. Compared with standard convolutional neural networks, the number of parameters of MobileNet can be reduced by an order of magnitude, which is crucial for mobile devices with limited storage resources. While the model has been simplified, it can still maintain relatively high accuracy in tasks such as image classification.

[0067] Based on the improved lightweight model, this paper achieved an average accuracy of 87.63% on the dataset, improving the classification accuracy of using a single data source of images. Among them, the accuracy of the listening category reached 88.37%, the accuracy of the head-down category reached 87.09%, the accuracy of the conversation category reached 81.79%, and the accuracy of the sleeping category reached 93.21%.

[0068] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for classifying student classroom behaviors based on an improved lightweight network, characterized in that: The steps include: Step 1: Place a camera with a built-in SD card in the center of the classroom blackboard, collect students' classroom videos as the video data source, and capture the videos into video frame images; before class, let students wear customized smart bracelets on their wrists, and collect students' dynamic physiological information data during video recording as the physiological information data source; Step 2: Divide the video clips into frames, capture the image of a single person in each frame, and use different Arabic numeral labels to represent different classroom behaviors with higher frequency; According to the timeline, the collected disordered wristband physiological data are arranged in the order of the time when the data is received, and the frames containing the wristband data timestamp are cropped. Then, the single image of the wristband wearer matching the IMEI number of each wristband is cropped, and the image is formatted and named containing the wristband physiological information. Then, it is classified according to the classroom behavior categories set above and organized into a multi-source fusion data set. Step 3: Read all images in the data set in order by category, convert the images into RGB format, unify the image size, parse the formatting and naming of the images, and perform data preprocessing on the unified size images, parsed physiological information data (image formatting and naming), and label (class folder name) data respectively; Step 4: Using MobileNetV1 as the baseline model, a feature dimension alignment method is used to fuse the image data and the bracelet. The residual is pulled out after the first Dw convolution, and the data is fused before the last Pw convolution. The bracelet data is used as the residual branch data and finally fused into the main image data. Finally, the final classification result is output through the fully connected layer. Step 5: Evaluate the effect of classroom behavior classification through accuracy, precision, recall rate and F1 value to achieve comprehensive and accurate detection of students' classroom behavior.

2. According to the method for classifying student classroom behaviors based on an improved lightweight network as described in claim 1, it is characterized in that: The customized smart bracelet in step 1 is equipped with different sensors for heart rate, body temperature and blood pressure. During the video recording process, the background synchronously obtains the physiological information data of each student in real time, and uploads it to the cloud platform through the NBIOT network. The cloud platform is configured with the URL address of the computer terminal and the corresponding receiving port. The API interface is used to send the received physiological information data to the server for reception, parsing and saving in a csv file. Considering the class duration and analysis effect, the data sending frequency is set to once per minute.

3. According to the method for classifying student classroom behaviors based on an improved lightweight network as described in claim 1, it is characterized in that: In step 2, OpenCV is used to divide the video clip into 1s / frame; among the classroom behaviors appearing in the video, listening, lowering the head, talking and sleeping are the students' classroom behaviors that appear more frequently, so the numbers 1-4 are used to represent these four types of actions, 1 represents raising the head to listen, 2 represents lowering the head, 3 represents talking, and 4 represents sleeping; the picture is formatted and named to include the wristband physiological information, indicating the wristband physiological data with the same timestamp as the cropped subject image, and each physiological data is connected with &.

4. According to claim 1, a method for classifying student classroom behaviors based on an improved lightweight network is characterized in that: In the step 3, OpenCV is used to unify the image size and adjust all of them to 224×224 size; then the formatting and naming of the image is parsed, and the images after the unified size, the parsed physiological information data (image formatting and naming) and the label (class folder name) data are converted into NumPy array format in turn, and saved as NumPy array files respectively, and stored in a list; then the three lists are exported into three files in npy format: image_data.npy, labels.npy, and numeric_data.npy, which store the image, bracelet and label data respectively; then the above three npy files are read, and pytorch is used to define the data structure of the multi-source fusion data set, and the image and bracelet data are aligned and encapsulated together for easy model training; finally, the class of the data set is defined and the data set object is instantiated, and train_DataLoader and val_DataLoader are generated according to the ratio of training set: validation set = 8:2 for model experiments.

5. The method for classifying student classroom behaviors based on an improved lightweight network according to claim 1 is characterized in that: The method for aligning the feature dimensions in step 4 is to first convert the five physiological data collected by the bracelet into 16 feature channels from 5 through a linear layer and an activation function layer (ReLu), and then reshape the bracelet data into a tensor consistent with the image feature channel through a reshape operation, and then let the bracelet data pass through a layer of depth-separable convolution to fully learn the feature distribution of the feature space image, and add the bracelet data to the network as a feature of the image.

6. The method for classifying student classroom behaviors based on an improved lightweight network according to claim 1 is characterized in that: The specific formulas for accuracy, precision, recall and F1 value in step 5 are as follows: Among them, TP represents true positive examples (the number of samples predicted by the model to be positive and actually are positive), TN represents true negative examples (the number of samples predicted by the model to be negative and actually are negative), FP represents false positive examples (the number of samples predicted by the model to be positive but actually are negative), and FN represents false negative examples (the number of samples predicted by the model to be negative but actually are positive).

Citation Information

Patent Citations

  • Student classroom behavior detection method based on deep learning

    CN113469001A

  • Disease and pest identification method based on multi-scale lightweight network

    CN115116054A

  • Student classroom state identification method based on multi-source information fusion

    CN118013389A

  • Facial expression recognition method and computer readable medium

    CN118711233A