Classroom student behavior detection and analysis system based on deep learning

By improving the YOLOv8s model and ByteTrack algorithm, combining adaptive noise adjustment and multi-head self-attention mechanism, the problem of student behavior detection error detection and missed detection in complex classroom environments is solved, and accurate behavior detection and tracking is achieved.

CN120496178AActive Publication Date: 2025-08-15HEFEI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510572525.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-15
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

Existing computer vision and deep learning models are difficult to accurately detect small targets and occluded objects in complex classroom environments. The target tracking algorithm is prone to mis-checking or missed detection during occlusion and movement changes, resulting in inaccurate detection of student behavior.

Method used

The improved YOLOv8s model and ByteTrack algorithm are adopted to improve feature extraction through the multi-head self-attention mechanism, combined with adaptive noise adjustment Kalman filtering and improved loss function, the object detection and tracking algorithm is optimized, and the detection and tracking capabilities are enhanced.

Benefits of technology

Accurately identify and continuously track student behavior in complex classroom environments, reduce mis-checking and missed inspections, provide a reliable data foundation, and support educational decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496178A_ABST
    Figure CN120496178A_ABST
Patent Text Reader

Abstract

The invention relates to the field of behavior detection, and discloses a classroom student behavior detection and analysis system based on deep learning, and the system comprises a data set construction module, a target detection model module, a target tracking algorithm module, a behavior detection and tracking module, and a data storage and analysis module. According to the invention, through construction of a classroom student behavior data set, an improved YOLOv8s model and a ByteTrack target tracking algorithm, accurate detection and tracking of classroom student behaviors can be realized under comprehensive cooperation; therefore, video data can be processed in real time, student behaviors can be automatically captured and classified and analyzed, a data report and a visual chart are generated, support is provided for education decision making, and meanwhile the system has the advantages of being high in accuracy, real-time, efficient, high in automation degree and capable of achieving data-driven decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of behavior detection, and in particular to a classroom student behavior detection and analysis system based on deep learning. Background Art

[0002] The existing computer vision technology uses classroom cameras to collect classroom data, cleans and pre-processes it, and then trains the YOLO model. The model is deployed to detect student behavior and output categories. However, there are still the following defects:

[0003] 1. Computer vision and deep learning models: In complex classroom environments, traditional target detection models (such as YOLOv5) have insufficient target detection accuracy. In low light, occlusion, overlapping of multiple people, and other conditions, small targets (such as looking down at a phone, turning around, or lying on the table) are easily missed. The dynamic response to changes in student posture (such as standing and turning around) is inaccurate, prone to false detection or missed detection. The YOLO series of algorithms has poor detection capabilities for small targets and is prone to missed detections when shooting from a distance in complex classroom scenes. The model also has insufficient processing of occlusion and overlapping objects, which can easily lead to incorrect identification or tracking when multiple targets are close or occluded.

[0004] 2. Target tracking algorithms: Existing target tracking algorithms do not adequately handle target occlusion. For example, ByteTrack is prone to tracking loss when the target is partially or permanently occluded. Tracking mismatch and drift are prone to occur when the target's appearance, posture, or motion pattern changes significantly, leading to incorrect target matching. The processing of target motion information primarily relies on appearance features, ignoring dynamic information such as motion trajectory and speed. This results in inaccurate tracking in fast-moving scenes and difficulty maintaining accuracy when target interactions are complex. Summary of the Invention

[0005] The purpose of the present invention is to provide a classroom student behavior detection and analysis system based on deep learning to solve at least one of the above technical problems.

[0006] The purpose of the present invention can be achieved through the following technical solutions:

[0007] The deep learning-based classroom student behavior detection and analysis system includes:

[0008] The dataset construction module is used to obtain real classroom videos, extract frames from the video to obtain image data, and use the MakeSense tool to annotate them to obtain a YOLO format dataset;

[0009] The object detection model module uses YOLOv8s as the baseline model, adds the multi-head self-attention mechanism MHSA to the backbone part, replaces the CIoU loss function with Focaler-DIoU, and adds a small object detection head to the head part. After training, a classroom student behavior detection model is obtained;

[0010] The target tracking algorithm module adopts the ByteTrack algorithm, introduces adaptive noise adjustment to the Kalman filter, and introduces improved CIoU and speed features when calculating the similarity metric and performs weighted merging calculations;

[0011] The behavior detection and tracking module uses the improved YOLOv8 model as the target detector to detect t frames of video to be detected, divides them into high and low frames according to the score threshold, and then uses the ByteTrack algorithm to perform target matching and trajectory prediction;

[0012] The data storage and analysis module integrates the detected data into a data package, including student ID, behavior category, coordinate location, confidence level, and timestamp, and stores it in the database. Python is used to perform quantitative statistics of student behaviors and time series statistics of individual student behaviors.

[0013] As a further technical solution, adaptive noise adjustment includes adaptive adjustment of measurement noise and adaptive adjustment of process noise;

[0014] In each filtering step, the innovation vector v is first calculated k , innovation vector v k It is the difference between the current predicted value of the Kalman filter and the actual measured value, which represents the error of the model; the calculation formula is as follows:

[0015] z k is the actual measurement value at time k, is the predicted value at time k, based on the state of the model at the previous time k-1 and the prediction result of the prediction model for the state at the current time, and H is the observation matrix;

[0016] According to the innovation vector v k The norm of the measurement noise covariance matrix R is dynamically adjusted:

[0017] R k =R k-1 +(1+||v k ||)·I4;

[0018] Among them, I4 is an identity matrix, ||v k || is the norm of the innovation vector; R k is the measurement noise covariance matrix at time k, R k-1 is the measurement noise covariance matrix at time k-1;

[0019] The noise covariance matrix Q of the process is adjusted according to the target motion state:

[0020]

[0021] β is the process noise adjustment factor, Q k The process noise covariance matrix at time k, Q k-1 is the process noise covariance matrix at time k-1, is the innovation vector v k The transpose of .

[0022] As a further technical solution, in the target detection model module, the Focaler-DIoU Loss loss function combines the dynamic weight mechanism of Focal Loss with the frame positioning optimization capability of DIoU Loss, and the calculation formula is:

[0023] Among them, Focaler increases the attention to difficult-to-detect samples by dynamically assigning weights to each sample. The calculation formula of the weight coefficient ω is: ω = (1-IoU) γ , γ is the adjustment factor, γ is greater than zero, IoU is the intersection over union ratio of the predicted box and the real box, (b,b gt ) is the center point b of the predicted box and the center point b of the real box gt The Euclidean distance between them is , and c is the diagonal length of the minimum bounding box between the predicted box and the true box.

[0024] As a further technical solution, in the target tracking algorithm module, the similarity calculation formula is: T s =γ·CIoU+δ·S s (v1, v2); γ and δ are weight coefficients, S s (v1, v2) is the similarity vector of velocity features;

[0025] in, CIoU is a similarity metric that takes into account the center distance, aspect ratio, and overlapping area of the target box; b and b gt are the coordinates of the center points of the prediction and detection boxes, v is the aspect ratio of the box, and α is the balance parameter.

[0026] As a further technical solution, Among them, w gt 、w gt are the width and height of the real box, w and h are the width and height of the predicted box respectively;

[0027] As a further technical solution, the speed characteristic S s The calculation is based on the change of the target's position information between different frames. The calculation formula is as follows:

[0028]

[0029] Where v1 and v2 are the velocity vectors of the historical trajectory target and the current detection box respectively; σ is a normalization parameter used to adjust the weight of the velocity feature.

[0030] As a further technical solution, in the target detection model module, MHSA is added between the ninth layer C2f module and the tenth layer SPPF of the Backbone part. After reshaping the feature map with a size of (20, 20, 512), it is converted into a query matrix, a key matrix, and a value matrix through three learnable projection matrices. The self-attention weight is calculated by scaling the dot product attention. The outputs of multiple attention heads are concatenated and mapped back to the original dimension through linear transformation. The output formula of the multi-head attention mechanism is as follows:

[0031] MultiHead(Q,K,V)=Concat(h1,h2......h h )·W;

[0032] Q, K, and V are the query, key, and value matrices shared by multiple heads in the multi-head attention mechanism, respectively. h1, h2...h h They represent the attention outputs calculated by each head in the multi-head attention mechanism, Concat represents the splicing operation, and W is the weight matrix, which is used to linearly transform the outputs of multiple heads and map them back to the original dimension.

[0033] As a further technical solution, the formula for calculating the self-attention weight by scaling the dot product attention is:

[0034]

[0035] Among them, Q i , K i 、V i Denote the query matrix, key matrix and value matrix respectively, d k Dimensions for queries or keys.

[0036] Beneficial effects of the present invention:

[0037] This paper combines the improved YOLOv8s model with the optimized ByteTrack target tracking algorithm, greatly enhancing the ability to detect and track student behavior in the classroom. In complex classroom environments, regardless of lighting changes, occlusions by people, or even subtle student movements, it can accurately identify and continuously track them. Compared with traditional technologies, it can more accurately judge various behaviors, effectively reducing false detections and missed detections, providing a reliable data foundation for subsequent analysis and ensuring that each student's behavioral performance in the classroom can be accurately grasped. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The present invention will be further described below with reference to the accompanying drawings.

[0039] Figure 1 Label a standards sheet for student behavior.

[0040] Figure 2 is the table of category labels.

[0041] Figure 3 This is the improved YOLOv8 model structure diagram.

[0042] Figure 4 This is the structure diagram of the multi-head self-attention mechanism (MHSA).

[0043] Figure 5 This is a comparison table of ablation experiment results.

[0044] Figure 6 This is the flowchart of the improved ByteTrack algorithm.

[0045] Figure 7 This is the CIoU structure diagram.

[0046] Figure 8 A comparison table of tracking results before and after improvement.

[0047] Figure 9 A score sheet for student behavior positivity. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0049] See also Figures 1-9 As shown, the present invention is a classroom student behavior detection and analysis system based on deep learning, including:

[0050] The dataset construction module is used to obtain real classroom videos, extract frames from the video to obtain image data, and use the MakeSense tool to annotate them to obtain a YOLO format dataset; in the dataset construction module, data collection is authorized to obtain real classroom videos of different schools and classroom layouts from multiple schools in China, and OpenCV is used to process the video files for frame extraction and save them as images; data collation and screening are carried out to select frontal angle images with diverse scenes, diverse behaviors and good quality, with a resolution of 1920×1080; data annotation is based on six types of behaviors: using mobile phones, sleeping on the table, looking at the blackboard / teacher, reading and writing, standing and turning around, and the bounding boxes are annotated one by one using the MakeSense annotation tool and saved in YOLO format.

[0051] The target detection model module uses YOLOv8s as the baseline model, adds a multi-head self-attention mechanism MHSA to the Backbone part, which can improve the feature extraction ability, replaces the loss function CIoU with Focaler-DIoU, and uses the Focaler-DIoU loss function to solve the inter-class imbalance and bounding box prediction problems; adds a small target detection head to the Head part, and adding a small target detection layer can improve the detection ability of long-distance and small-sized targets; after training, a classroom student behavior detection model is obtained; in the target detection model module, the resolution of the added small target detection layer P2 feature map is 1 / 4 of the input image size, that is, 160×160. By locating and extracting high-resolution feature maps in the backbone network, upsampling and concat splicing them with other layer feature maps for multi-scale feature fusion, target detection is performed.

[0052] The target tracking algorithm module uses the ByteTrack algorithm and introduces adaptive noise adjustment to the Kalman filter. This improves the Kalman filter in the tracking algorithm and dynamically adjusts the observation noise R and process noise Q based on innovations. The filter is dynamically optimized based on environmental and state changes, avoiding over-reliance on prediction and observation and improving overall accuracy. When calculating the similarity metric, improved CIoU and velocity features are introduced and weighted and combined to obtain more accurate matching results.

[0053] The behavior detection and tracking module uses the improved YOLOv8 model as the target detector to detect t frames of video to be detected, divides them into high and low frames according to the score threshold, and then uses the ByteTrack algorithm to perform target matching and trajectory prediction;

[0054] The data storage and analysis module integrates the detected data into a data package, including student ID, behavior category, coordinate location, confidence level, and timestamp, and stores it in a database. Python (Pandas, Matplotlib) is used to perform quantitative and temporal statistics of student behaviors. In this data storage and analysis module, the system can load videos in real time through a high-definition camera or select videos locally. The detection results are stored in a structured format in the database. Real-time single-frame and cumulative behavior statistics histograms are generated for category statistics. Student behaviors are scored for positivity, and time series diagrams are drawn in the form of dot-line graphs for individual student analysis.

[0055] In the target detection model module, MHSA is added between the ninth layer C2f module and the tenth layer SPPF of the Backbone part. After reshaping the feature map with a size of (20, 20, 512), it is converted into a query matrix, a key matrix, and a value matrix through three learnable projection matrices. The self-attention weight is calculated by scaled dot product attention. The outputs of multiple attention heads are concatenated and mapped back to the original dimension through linear transformation. The output formula of the multi-head attention mechanism is as follows:

[0056] MultiHead(Q,K,V)=Concat(h1,h2......h h )·W;

[0057] Q, K, and V are the query, key, and value matrices shared by multiple heads in the multi-head attention mechanism, respectively. h1, h2...h h They represent the attention outputs calculated by each head in the multi-head attention mechanism, Concat represents the splicing operation, and W is the weight matrix, which is used to linearly transform the outputs of multiple heads and map them back to the original dimension.

[0058] The calculation formula for calculating the self-attention weight by scaling the dot product attention is:

[0059]

[0060] Among them, Q i , K i 、V i Denote the query matrix, key matrix and value matrix respectively, d k Dimensions for queries or keys.

[0061] In this embodiment, the detection layers in the YOLOv8 basic model are the P3 layer with 8x downsampling, the P4 layer with 16x downsampling, and the P5 layer with 32x downsampling. After the training image is preprocessed by YOLOv8, the input image resolution is 640×640. The feature map resolutions of these three layers are 80×80, 40×40, and 20×20, respectively. The feature map resolution of the added small target detection layer P2 is the size of the input image, i.e., 160×160. First, the high-resolution feature map is located and extracted in the backbone network, and then multi-scale feature fusion is performed with the feature maps of other layers through upsampling and concat splicing. Target detection is performed on the multi-scale feature map, and the combined multi-scale result is output. The comparison of the experimental performance before and after the improvement is shown in the ablation experiment result table in the attached figure of the specification. It is not difficult to find from the various data that each indicator of the improved model has been significantly improved, which can balance the precision and recall rate, and improve the robustness and accuracy of the model in actual scenarios.

[0062] Adaptive noise adjustment includes adaptive adjustment of measurement noise and adaptive adjustment of process noise;

[0063] In each filtering step, the innovation vector v is first calculated k , innovation vector v k It is the difference between the current predicted value of the Kalman filter and the actual measured value, which represents the error of the model; the calculation formula is as follows:

[0064] z k is the actual measurement value at time k, is the predicted value at time k, based on the state of the model at the previous time k-1 and the prediction result of the prediction model for the state at the current time, and H is the observation matrix;

[0065] According to the innovation vector v k The norm of the measurement noise covariance matrix R is dynamically adjusted:

[0066] R k =R k-1 +(1+‖v k ‖)·I4;

[0067] Among them, I4 is an identity matrix, ||v k || is the norm of the innovation vector; R k is the measurement noise covariance matrix at time k, R k-1 is the measurement noise covariance matrix at time k-1;

[0068] The noise covariance matrix Q of the process is adjusted according to the target motion state:

[0069]

[0070] β is the process noise adjustment factor, Q k The process noise covariance matrix at time k, Q k-1 is the process noise covariance matrix at time k-1, is the innovation vector v k The transpose of .

[0071] In this embodiment, the goal of adaptive noise adjustment is to dynamically adjust the noise covariance matrix based on innovation, so that the Kalman filter can achieve better performance in different environments. It is mainly divided into two steps, and the process is as follows:

[0072] 1) Adaptive adjustment of measurement noise R

[0073] In each filtering step, the innovation vector v is first calculated k , which is the difference between the current predicted value of the Kalman filter and the actual measured value, indicating the error of the model. Its calculation formula is as follows:

[0074] Then calculate the norm of the innovation vector ||v k || and dynamically adjust the measurement noise covariance matrix R according to the size of this value. The measurement noise adjustment coefficient is:

[0075] R k =R k-1 +(1+||v k ||)·I4

[0076] Where I4 is the identity matrix, which is used to increase the diagonal elements of the measurement noise. The above formula shows that the innovation norm is the key to adjusting the noise level. A larger innovation vector indicates greater model error and, consequently, greater measurement noise, so R needs to be increased.

[0077] 2) Adaptive adjustment of process noise Q

[0078] Process noise Q is generally related to the target's motion state. The more uncertain the motion state, the greater the process noise should be. Based on this, it is proposed to use the size of the innovation to adjust the process noise during the Kalman filter process. If the norm of the innovation is small, it means that the model's prediction is more accurate, and the process noise Q can be appropriately reduced, thereby improving the system's tracking accuracy. The process noise adjustment formula is:

[0079]

[0080] Where β is the process noise adjustment factor, which controls the speed of noise update.

[0081] Combined with the following Kalman filter overall formula, the adaptive strategy optimizes the credibility of state prediction by dynamically adjusting Q; when the system is highly dynamic, increasing Q increases the uncertainty of the prediction to capture changes; in a stable environment, reducing Q improves the accuracy of the prediction. The Kalman gain K in the correction phase k Directly affected by R, after gain adjustment, the filter can suppress the influence of noise when the innovation is large, and make full use of the measurement value information when the innovation is small, thereby dynamically optimizing the accuracy and stability of state estimation.

[0082] The overall formula of the Kalman filter is as follows, which is used for state prediction and update in the improved ByteTrack target tracking algorithm to optimize the target tracking effect;

[0083]

[0084] State prediction equation: The function is: according to the state estimate value x at the previous moment (k-1) k-1|k-1 and the current control input u k , predict the state at the current time (k) A is the state transition matrix, which describes how the state changes over time; B is the control matrix, which determines the impact of control input on the state. For example, in a classroom student behavior detection scenario, if the student's position and speed are used as state quantities, the student's position and speed at the previous moment are known. Combined with the current possible movement control (for example, the teacher asks the student to stand up to answer a question, which will affect the student's position state), the student's state at the current moment can be predicted.

[0085] Covariance prediction equation: The role is to predict the covariance of the current state estimate Combined with the covariance P of the previous moment (k-1) state estimate k-1|k-1 , the state transition matrix A, and the process noise covariance Q. The covariance reflects the uncertainty of the state estimate, while Q represents process noise, reflecting the inherent uncertainty of the system. For example, when tracking student behavior, the student's movements may not be completely regular, and this uncertainty is represented by Q. This formula comprehensively considers the uncertainty of the previous moment and the uncertainty of the current process to predict the uncertainty of the state estimate at the current moment.

[0086] Kalman gain calculation equation: The role is to calculate the Kalman gain K k , determines how to combine the predicted value and the measured value to update the state estimate. k Affected by the predicted covariance The influence of the observation matrix H and the measurement noise covariance R. In practical applications, the measured values (such as the student position information collected by the camera) are noisy, and the predicted values are also uncertain. Kalman gain K k The key to balancing the two is to determine how much weight to give to the predicted value and the measured value based on the size of the predicted covariance and the measurement noise covariance to obtain a more accurate state estimate.

[0087] State update equation: The role is to use the Kalman gain K k , combined with the predicted value and the measured value z k The residual To update the current state estimate x k|k For example, when tracking student behavior, there may be a difference between the predicted and measured student positions. This difference is the residual. By weighting the residual using the Kalman gain and adding the predicted value, we can obtain a more realistic estimate of the student's state.

[0088] Covariance update equation: The role is: according to the Kalman gain K k and the predicted covariance Update the covariance P of the current state estimate k|k , used for the prediction and update calculations at the next moment. For example, after each state estimate update, the covariance needs to be updated accordingly to reflect the uncertainty of the new state estimate. I is the identity matrix. This formula adjusts the covariance by taking into account the Kalman gain and the prediction covariance, providing a more accurate uncertainty measure for the next moment's state prediction and update.

[0089] In the target detection model module, the Focaler-DIoU Loss loss function combines the dynamic weight mechanism of FocalLoss and the frame positioning optimization capability of DIoU Loss. The calculation formula is:

[0090] Focaler-DIoU combines the dynamic weighting mechanism of Focal Loss with the box positioning optimization capability of DIoU Loss to form a more effective loss function for object detection tasks. This formula adds a penalty term for center point distance to the IoU, solving the problem of insensitivity to center point offset when relying solely on IoU to evaluate box positioning.

[0091] Among them, Focaler increases the attention to difficult-to-detect samples by dynamically assigning weights to each sample. The calculation formula of the weight coefficient ω is: ω = (1-IoU) γ,γ is an adjustment factor used to control the dynamics of weight distribution. γ is greater than zero and is usually set to 2. IoU is the intersection over union ratio of the predicted box to the real box, (b,b gt ) is the center point b of the predicted box and the center point b of the real box gt The Euclidean distance between them is , and c is the diagonal length of the minimum bounding box between the predicted box and the true box.

[0092] In the target tracking algorithm module, the similarity calculation formula is: T s =γ·CIoU+δ·S s (v1, v2); γ and δ are weight coefficients, S s (v1, v2) is the similarity vector of velocity features;

[0093] in, CIoU is a similarity metric that takes into account the center distance, aspect ratio, and overlapping area of the target box; b and b gt are the coordinates of the center points of the prediction and detection boxes, v is the aspect ratio of the box, and α is the balance parameter; Among them, w gt 、w gt are the width and height of the real box, w and h are the width and height of the predicted box respectively; Speed characteristic S s The calculation is based on the change of the target's position information between different frames. The calculation formula is as follows:

[0094]

[0095] Where v1 and v2 are the velocity vectors of the historical trajectory target and the current detection box respectively; σ is a normalization parameter used to adjust the weight of the velocity feature.

[0096] In this embodiment, ByteTrack uses target location information as an association clue and adopts Intersection over Union (IoU) as a method to calculate the similarity metric between the detection box and the prediction box. However, it does not fully consider factors such as the target's shape, aspect ratio, and center point distance. CIoU considers the center distance, aspect ratio, and overlap area of the target box, thereby further optimizing the target matching based on IoU. CIoU can improve the calculation of IoU by considering the three aspects of center point distance, aspect ratio, and overlap.

[0097] The improved tracking algorithm's MOTA and MOTP values increased by 2.8% and 2.4% compared to the original algorithm, indicating greater accuracy in trajectory prediction and detection box assignment, with fewer mismatches. Increases in the IDP and IDR metrics reflect improved target identity consistency and higher tracking quality. A decrease in the IDSW evaluation metric indicates fewer target identity switching issues and more consistent tracking results, enabling more stable target identification and tracking in scenarios with frequent student interaction, occlusion, or large movements.

[0098] It should be noted that the calculation formulas and various parameters involved in the calculations in the present invention have been dimensionally processed in advance, and the process of dimensionless processing is well known in the industry and will not be described here.

[0099] The above is a detailed description of an embodiment of the present invention. However, the content described is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A classroom student behavior detection and analysis system based on deep learning, characterized by: include: The dataset construction module is used to obtain real classroom videos, extract frames from the video to obtain image data, and use the MakeSense tool to annotate them to obtain a YOLO format dataset; The object detection model module uses YOLOv8s as the baseline model, adds the multi-head self-attention mechanism MHSA to the backbone part, replaces the CIoU loss function with Focaler-DIoU, and adds a small object detection head to the head part. After training, a classroom student behavior detection model is obtained; The target tracking algorithm module adopts the ByteTrack algorithm, introduces adaptive noise adjustment to the Kalman filter, and introduces improved CIoU and speed features when calculating the similarity metric and performs weighted merging calculations; The behavior detection and tracking module uses the improved YOLOv8 model as the target detector to detect t frames of video to be detected, divides them into high and low frames according to the score threshold, and then uses the ByteTrack algorithm to perform target matching and trajectory prediction; The data storage and analysis module integrates the detected data into a data package, including student ID, behavior category, coordinate location, confidence level, and timestamp, and stores it in the database. Python is used to perform quantitative statistics of student behaviors and time series statistics of individual student behaviors.

2. The deep learning-based classroom student behavior detection and analysis system according to claim 1 is characterized in that: Adaptive noise adjustment includes adaptive adjustment of measurement noise and adaptive adjustment of process noise; In each filtering step, the innovation vector v is first calculated k , innovation vector v k It is the difference between the current predicted value of the Kalman filter and the actual measured value, which represents the error of the model; the calculation formula is as follows: z k is the actual measurement value at time k, is the predicted value at time k, based on the state of the model at the previous time k-1 and the prediction result of the prediction model for the state at the current time, and H is the observation matrix; According to the innovation vector v k The norm of the measurement noise covariance matrix R is dynamically adjusted: R k =R k-1 +(1+||v k ||)·I4; Among them, I4 is an identity matrix, ||v k || is the norm of the innovation vector; R k is the measurement noise covariance matrix at time k, R k-1 is the measurement noise covariance matrix at time k-1; The noise covariance matrix Q of the process is adjusted according to the target motion state: β is the process noise adjustment factor, Q k The process noise covariance matrix at time k, Q k-1 is the process noise covariance matrix at time k-1, is the innovation vector v k The transpose of .

3. The deep learning-based classroom student behavior detection and analysis system according to claim 2 is characterized in that: In the target detection model module, the Focaler-DIoU Loss loss function combines the dynamic weight mechanism of Focal Loss with the frame positioning optimization capability of DIoU Loss. The calculation formula is: Among them, Focaler increases the attention to difficult-to-detect samples by dynamically assigning weights to each sample. The calculation formula of the weight coefficient ω is: ω = (1-IoU) γ , γ is the adjustment factor, γ is greater than zero, IoU is the intersection over union ratio of the predicted box and the real box, (b,b gt ) is the center point b of the predicted box and the center point b of the real box gt The Euclidean distance between them is , and c is the diagonal length of the minimum bounding box between the predicted box and the true box.

4. The deep learning-based classroom student behavior detection and analysis system according to claim 3 is characterized in that: In the target tracking algorithm module, the similarity calculation formula is: T s =γ·CIoU+δ·S s (v1, v2); γ and δ are weight coefficients, S s (v1, v2) is the similarity vector of velocity features; in, CIoU is a similarity metric that takes into account the center distance, aspect ratio, and overlapping area of the target box; b and b gt are the coordinates of the center points of the prediction and detection boxes, v is the aspect ratio of the box, and α is the balance parameter.

5. The deep learning-based classroom student behavior detection and analysis system according to claim 4 is characterized in that: Among them, w gt 、w gt are the width and height of the real box, w and h are the width and height of the predicted box respectively; 6. The deep learning-based classroom student behavior detection and analysis system according to claim 4 is characterized in that: Speed characteristic S s The calculation is based on the change of the target's position information between different frames. The calculation formula is as follows: Where v1 and v2 are the velocity vectors of the historical trajectory target and the current detection box respectively; σ is a normalization parameter used to adjust the weight of the velocity feature.

7. The deep learning-based classroom student behavior detection and analysis system according to claim 1 is characterized in that: In the target detection model module, MHSA is added between the ninth layer C2f module and the tenth layer SPPF of the Backbone part. After reshaping the feature map with a size of (20, 20, 512), it is converted into a query matrix, a key matrix, and a value matrix through three learnable projection matrices. The self-attention weight is calculated by scaled dot product attention. The outputs of multiple attention heads are concatenated and mapped back to the original dimension through linear transformation. The output formula of the multi-head attention mechanism is as follows: MultiHead(Q,K,V)=Concat(h1,h2……h h )·W; Q, K, and V are the query, key, and value matrices shared by multiple heads in the multi-head attention mechanism, respectively. h1, h2...h h They represent the attention outputs calculated by each head in the multi-head attention mechanism, Concat represents the splicing operation, and W is the weight matrix, which is used to linearly transform the outputs of multiple heads and map them back to the original dimension.

8. The deep learning-based classroom student behavior detection and analysis system according to claim 7 is characterized in that: The calculation formula for calculating the self-attention weight by scaling the dot product attention is: Among them, Q i , K i 、V i Denote the query matrix, key matrix and value matrix respectively, d k Dimensions for queries or keys.

Citation Information

Patent Citations

  • Pedestrian small target detection method in video monitoring based on deep learning

    CN115240119A

  • Pedestrian detection method and device, equipment, storage medium and product

    CN119206853A

  • Self-adaptive multi-target tracking method based on stages

    CN119444802A

  • Student classroom behavior detection method based on deep learning

    CN119763179A

  • Multimodal data-based method and system for recognizing cognitive engagement in classroom

    US20250022314A1