Student behavior robust identification method for smart classroom

By improving the YOLOv11n model and combining it with multi-scale feature enhancement and deformable attention mechanism, the problems of multi-scale fusion and occlusion robustness in student target detection in smart classroom environments are solved, achieving efficient and accurate student behavior recognition.

CN120766342APending Publication Date: 2025-10-10CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510673568.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In the existing smart classroom environment, student target detection has problems such as insufficient multi-scale feature fusion capabilities, low accuracy in detecting occluded targets, and insufficient real-time performance of the model, making it difficult to meet real-time detection needs.

Method used

An improved YOLOv11n model is adopted to enhance the detection accuracy and real-time performance of the model in complex classroom environments by integrating the multi-scale feature enhancement mechanism and the deformable attention adaptation mechanism, including the C3K2_CSP_MSA structure block, the CSPDA structure block and the CGConv module.

Benefits of technology

It significantly improves the model's ability to integrate multi-scale features, enhances the robustness of occluded target detection, and improves the model's adaptability to complex scenarios while maintaining detection efficiency, solving the problem of student behavior detection in smart classroom environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766342A_ABST
    Figure CN120766342A_ABST
Patent Text Reader

Abstract

The invention provides a student behavior robust identification method for a smart classroom, and relates to the technical field of smart education, and the method comprises the steps: constructing a classroom behavior detection data set, and implementing three core improvements in a YOLOv11n framework: replacing C3K2 of a neck network with C3K2CSPMSA; replacing a part of Conv in the YOLOv11n with an efficient down-sampling module CGConv module designed based on a multi-scale context sensing principle; the deformable attention mechanism is fused into a backbone network C2PSA module to be reconstructed into a CSPDA module; performing iterative optimization training on the modified model by using a training set and a verification set to obtain an improved YOLOv11n classroom student behavior detection model; and inputting a test set into the improved YOLOv11n classroom student behavior detection model for detection, and outputting a student behavior detection result. According to the method, the omission ratio and the false detection rate in a complex shielding scene can be improved, and a high-precision and high-robustness real-time detection solution is provided for smart classroom behavior analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and behavior recognition technology, and in particular to a robust student behavior recognition method for smart classrooms. Background Art

[0002] Currently, smart classrooms are the main battlefield of teaching. They make full use of machine vision to empower teaching quality management, conduct intelligent analysis of massive data on students' learning process, optimize evaluation mechanisms, promote teaching reform, and provide intelligent support for school management decisions and education and teaching evaluation.

[0003] In today's education sector, smart classrooms, as the core vehicle for driving modern teaching reforms, have become a key platform for improving teaching quality and efficiency. With the rapid development of artificial intelligence technology, machine vision-based teaching quality management systems are gradually replacing traditional manual assessment methods. Through real-time monitoring and intelligent analysis of student classroom behavior, they provide teachers with accurate teaching feedback and provide educational administrators with a scientific basis for decision-making. However, current visual analysis technology in smart classroom environments still faces multiple technical bottlenecks:

[0004] In actual smart classroom environments, large image scale variations, densely packed student objects, and severe occlusions pose challenges to vision-based smart classroom analysis. Students in peripheral areas appear small due to their physical distance and perspective distortion. Furthermore, uneven lighting and occlusions from tables and chairs make them difficult for existing object detection technologies to accurately identify. Consequently, these students are often overlooked in teaching evaluations based on object detection technology. These technical and behavioral limitations lead to inaccurate classroom evaluations that fail to fully reflect the teaching situation.

[0005] Current mainstream detection models (such as early versions of the YOLO series) suffer from significant deficiencies in the feature extraction stage. These deficiencies include: a lack of multi-scale feature fusion capabilities; their fixed convolution kernel size cascade architecture struggles to effectively capture multi-scale features, from fine-grained to global scales; a rigid spatial attention mechanism; and a fixed sampling grid that fails to adapt to changes in target morphology in complex occlusion scenarios. Furthermore, the downsampling process suffers from severe information attenuation; traditional downsampling modules tend to lose small object features in edge regions when compressing feature maps, significantly degrading the model's detection performance for small-scale objects. Existing improvements often improve accuracy by stacking network layers. However, this brute-force approach leads to a surge in model parameters, making it difficult to meet the rigid real-time detection requirements of smart classroom scenarios. This further highlights the conflict between real-time and robustness. Enhancing the model's adaptability to complex scenarios while maintaining detection efficiency has become a pressing technical challenge. These technical bottlenecks severely restrict the practical application value of smart classroom analysis systems, necessitating the development of a lightweight solution that balances detection accuracy and real-time performance. SUMMARY

[0006] In view of the defects of dense occlusion interference, camera top view angle geometric distortion, multi-scale target recognition difficulty in student behavior detection in the smart classroom scene, low recognition accuracy, insufficient model robustness and difficulty in meeting real-time detection requirements in the prior art, the present application proposes a student behavior robust recognition method for smart classrooms. Based on the YOLO11n framework, the method significantly improves the detection accuracy and real-time performance of the model in complex classroom environments by fusing multi-scale feature enhancement mechanism and deformable attention adaptive mechanism.

[0007] The present application proposes a student behavior robust recognition method for smart classrooms, comprising the following steps:

[0008] S1: Obtain multi-view and multi-scene classroom image data, and divide it into a training set, a validation set and a test set;

[0009] S2: Construct an improved YOLOv11n model, the improved YOLOv11n model comprising a backbone network, a neck network and a detection head, and the following improvements are made:

[0010] S21: Replace the C3K2 structure block in the neck network of YOLOv11n with a C3K2_CSP_MSA structure block, the C3K2_CSP_MSA structure block fusing multi-scale features through an incremental multi-core convolution chain and a residual feature reorganization strategy;

[0011] S22: Replace the C2PSA structure block in the YOLOv11n model with a CSPDA structure block based on a deformable attention mechanism, the CSPDA structure block adjusting the attention range through a dynamic offset network;

[0012] S23: Replace the CONV module in the YOLOv11n model with a context-guided down-sampling module CGConv, the CGConv module preserving down-sampling features through a parallel local context branch, a surrounding context branch and a global attention mechanism;

[0013] S3: Use the training set to iteratively optimize and train the improved YOLOv11n model to generate a student behavior detection model;

[0014] S4: Input the classroom image to be tested into the student behavior detection model and output the student behavior detection result.

[0015] Further, the construction of the C3K2_CSP_MSA structure block in S21 comprises: replacing the BottleNeck structure in the C3K2 structure block of the neck network of YOLOv11n with a CSP_MSA module; the construction process of the CSP_MSA module is:

[0016] S21a: The input feature map is processed by a first convolutional layer with a convolution kernel size of 3 and a step size of 1, the input and output channel numbers remain unchanged, and the output feature map is equally divided into a first sub-feature map X1 and a second sub-feature map X2;

[0017] S21b: The second sub-feature map X2 is sequentially processed as follows:

[0018] S21b1: Extract features by a second convolutional layer with a convolution kernel size of 5 and a step size of 1, and equally divide the output feature map into a third sub-feature map X3 and a fourth sub-feature map X4;

[0019] S21b2: The fourth sub-feature map X4 is processed by a third convolutional layer with a convolution kernel size of 7 and a step size of 1 to generate a fifth sub-feature map U4;

[0020] S21c: The fifth sub-feature map U4, the third sub-feature map X3 and the first sub-feature map X1 are spliced in the channel dimension, wherein the fifth sub-feature map U4 is the output of 7x7 convolution, the third sub-feature map X3 is the output of 5x5 convolution, and the first sub-feature map X1 is the output of 3x3 convolution, forming a multi-scale fusion feature map;

[0021] S21d: Align the multi-scale fusion feature map in the channel dimension by a fourth convolutional layer with a convolution kernel size of 1x1 and a step size of 1 to generate an aligned feature map;

[0022] S21e: Add the aligned feature map and the input feature map of step S21a through residual connection, output the final feature, and complete the progressive multi-core feature fusion of cross-layer cascade. This helps gradient propagation, reduces computational complexity while maintaining multi-scale features, and compensates for receptive field loss and preserves bottom-level feature integrity. This makes the model have better occlusion robustness in high-density occlusion environment, thus helping to solve the problem of large image size and student target density in the smart classroom environment.

[0023] Further, the step S22 of replacing the C2PSA structure block in the YOLOv11n model with the CSPDA structure block based on the deformable attention mechanism includes: replacing the Attention block in the C2PSA block with a DAttention block, and the processing process of the CSPDA module includes:

[0024] S22a: The input feature map is sequentially processed by a fifth convolutional layer with a convolution kernel size of 1 and a step size of 1 for channel transformation, then normalized by BatchNorm to stabilize the data distribution, and then processed using a SiLU activation function to obtain an initial feature map;

[0025] S22b: Split the initial feature map into first branch features and second branch features in the channel dimension, wherein:

[0026] S22b1: The first branch feature is directly retained;

[0027] S22b2: The second branch feature is input to the DAModule for processing. The DAModule consists of a deformable attention module DAttention and a feedforward network FFN, and the DAttention output is added to the input feature through a residual connection;

[0028] S22c: Concatenate the processing results of the first branch feature and the second branch feature according to the channel dimension, and perform nonlinear transformation through the sixth convolution layer with a convolution kernel size of 1 and a step size of 1 to obtain the output features of the CSPDA structure block.

[0029] Furthermore, the processing of the DAttention module includes:

[0030] Input feature map Generate a uniformly distributed reference grid in Ration is the preset downsampling ratio, and the reference points are normalized;

[0031] The input feature map is converted into a query vector q=xW through a linear transformation operation q , while using the offset network θ offset To calculate the spatial position offset;

[0032] By performing feature sampling operations at the deformation point position, the key vectors are extracted respectively Sum vector

[0033] Finally, using the sampling function Get deformation feature map The sampling function uses bilinear interpolation, and the specific formula is as follows:

[0034]

[0035] In the formula Defined as:

[0036]

[0037] Where, g(a,b)=max(0,1-|ab|); (r x ,r y ) represents all positions of the feature map;

[0038] In the multi-head deformable attention mechanism, for the query vector q, the deformation key and deformation value Implement joint calculation; the output result corresponding to the mth attention head can be obtained by the following formula:

[0039]

[0040] Where, represents the position embedding matrix; R is the relative position deviation offset; σ represents the softmax activation function; d represents the key vector dimension;

[0041] After all attention heads are processed, they are spliced ​​together and passed through the projection layer to generate the final feature map Z, which is the output of the deformable attention mechanism.

[0042] Furthermore, the construction of the context-guided downsampling module CGConv in S23 includes the following specific steps:

[0043] S23a: The input feature map is downsampled through a convolution layer with a kernel size of 3 and a stride of 1. The height H and width W of the feature map are halved and the number of channels is doubled to generate the initial downsampled feature map.

[0044] S23b: Perform parallel branch processing on the initial down-sampling feature map:

[0045] S23b1: Local context branch f local : Extract local features from 8 adjacent feature vectors through a convolution layer with a convolution kernel size of 3 and a dilation of 1;

[0046] S23b2: surrounding context branch f surrounding : Extracting large-scale surrounding context features through a dilated convolution layer with a kernel size of 3 and a dilation of 3;

[0047] S23c: branch the local context f local The output features and the surrounding context branch f surrounding The output features are concatenated according to the channel dimension to form a joint feature map f joint ;

[0048] S23d: the joint feature map f joint Perform batch normalization and PReLU activation function processing in sequence, and then restore the number of channels through a convolution layer with a convolution kernel size of 1 and a stride of 1 to make it consistent with the number of channels of the initial downsampled feature map;

[0049] S23e: Enhanced module f through global attention globle The feature map outputted in step S23d is processed, including:

[0050] (e1) performing global average pooling on the feature map output by step S23d to generate a global feature vector for each channel;

[0051] (e2) constructing a channel attention mechanism through two fully connected layers, wherein the first fully connected layer performs nonlinear transformation using a ReLU activation function, and the second fully connected layer generates channel weight coefficients using a Sigmoid function;

[0052] (e3) multiplying the channel weight coefficients with the feature map output by step S23d channel by channel to output the final down-sampled features.

[0053] The classroom student behavior detection method and system based on the improved YOLOv11n proposed in the present application has significant comprehensive technical advantages compared with the prior art, mainly in three aspects of multi-scale feature fusion, occluded target detection capability and real-time robustness collaborative optimization.

[0054] In terms of multi-scale feature fusion, the prior art has difficulty in effectively capturing multi-scale features from fine granularity to global due to the use of a fixed convolution kernel size in a series structure, resulting in a decline in detection performance for small targets in the edge area of the classroom. The present application introduces a C3K2_CSP_MSA structure block in the neck network, uses a progressive multi-core convolution chain and a residual connection strategy, concatenates the output feature maps of different convolution kernels in the channel dimension, and compensates for the receptive field loss through residual connection. This progressive multi-core feature fusion method across layers not only synchronously retains the shallow details and deep semantics, but also significantly reduces the operation parameter amount, thereby providing multi-level and multi-scale feature representation for the subsequent detection head in the feature extraction stage, effectively alleviating the missed detection problem caused by the scale difference and dense distribution of student targets.

[0055] In terms of occluded target detection capability, the prior art cannot adapt to the shape changes of targets in complex occlusion scenarios due to the use of a fixed sampling grid attention mechanism, resulting in low detection accuracy for small targets with edge distortion. The present application replaces the C2PSA structure block with a CSPDA structure block based on a deformable attention mechanism, adjusts the attention range using a dynamic offset network, so that the model can dynamically adjust the receptive field according to the target shape, significantly enhancing the reconstruction capability for occluded targets. At the same time, the CGConv module effectively suppresses the information decay in the down-sampling process through parallel local context branches, surrounding context branches and global attention mechanisms, further strengthening the small target features. This cascading enhancement mechanism enables the CSPDA module and the CGConv module to cooperate with each other when processing high-dimensional feature maps, significantly improving the robustness of the model in dense occlusion scenarios.

[0056] In the aspect of real-time robust collaborative optimization, the prior art mainly improves the accuracy by stacking network layers, but causes the model parameter quantity to increase sharply, which is difficult to meet the real-time detection requirement. The present application optimizes the structure design of three core modules of C3K2_CSP_MSA, CSPDA and CGConv, significantly enhances the adaptability of the model to complex scenes while maintaining the detection efficiency. Specifically, the channel two-split strategy of the C3K2_CSP_MSA module reduces the operation parameter quantity, the efficient down-sampling operation of the CGConv module guarantees the quality of the high-dimensional feature map, and the dynamic attention mechanism of the CSPDA module improves the detection capability of the occluded target. The three are organically combined to form a closed-loop optimization mechanism of "deep layer guiding shallow layer and shallow layer correcting deep layer", which not only solves the problem of multi-scale dense student target detection in the smart classroom environment, but also significantly improves the real-time robustness of the model in complex scenes.

[0057] In summary, the present application has made a significant technical breakthrough in multi-scale feature fusion, occluded target detection capability and real-time robustness through innovative module design and collaborative optimization strategy, providing an efficient and accurate solution for student behavior detection in the smart classroom environment. BRIEF DESCRIPTION OF DRAWINGS

[0058] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, illustrate the present application together with the embodiments thereof, and explain the principles of the present application, and do not constitute a limitation of the present application.

[0059] Figure 1 A flowchart of a student behavior robust recognition method for a smart classroom according to the first embodiment of the present application;

[0060] Figure 2 A network structure diagram of a classroom student behavior detection model improved YOLOv11n according to the first embodiment of the present application;

[0061] Figure 3 A framework structure diagram of a classroom student behavior detection model improved YOLOv11n according to the first embodiment of the present application;

[0062] Figure 4 A structure diagram of a C3K2_CSP_MSA structure block according to the first embodiment of the present application;

[0063] Figure 5 A structure diagram of a CSP_MSA block according to the first embodiment of the present application;

[0064] Figure 6 A structure diagram of a CSPDA structure block according to the first embodiment of the present application;

[0065] Figure 7A structural diagram of DAttention proposed for the first embodiment of the present application;

[0066] Figure 8 A schematic diagram of a deformable attention mechanism proposed for the first embodiment of the present application;

[0067] Figure 9 A structural diagram of a CGConv structural block proposed for the first embodiment of the present application;

[0068] Figure 10 A behavior motion attention map of a student behavior robust recognition method for a smart classroom proposed for the first embodiment of the present application in a smart classroom environment. DETAILED DESCRIPTION

[0069] For the sake of understanding the technical solutions of the present application, the present application will be systematically described below in combination with the appended drawings. It should be noted that the illustrated cases are only a number of specific embodiments of the present application, and the actual application forms of the present application have diversity and are not limited to the examples listed in the specification. The fundamental purpose of providing these specific descriptions is to make the technical solutions of the present application present a more complete disclosure system through multi-dimensional interpretation.

[0070] It should be particularly noted that, except for special annotations, the technical terms used in the present application all follow the standard interpretation of conventional terms in the technical field. The specific terms involved in the specification only serve the technical description needs of the embodiments, and the use scope thereof should not be understood as a limitation on the claims of the present application.

[0071] The present application provides a student behavior robust recognition method for a smart classroom, comprising the following steps:

[0072] S1: Obtain multi-view and multi-scene classroom image data, and divide it into a training set, a validation set and a test set;

[0073] Illustratively, student class video is collected by using multi-angle and multi-scene classroom top camera equipment, N frames of video are extracted every N frames to generate 640x640 pixel images, and the student behavior in the images is rectangularly labeled, and the training set, the validation set and the test set are divided in a ratio of 6:2:2;

[0074] In terms of scene, the data set covers three education practice scenes of university classroom, middle school classroom and primary school classroom, enhancing the generalization ability of the model in actual application; specific scenes include three scenes of staircase classroom, small classroom and laboratory, and the shooting angles are various, further enhancing the diversity of the data set;

[0075] Optionally, the original Mosaic enhancement technology of YOLO algorithm is used for data enhancement through four-image splicing and scale transformation; dynamic geometric transformation enhancement is implemented, including but not limited to random horizontal or vertical flipping, rotation transformation operation;

[0076] It should be understood that the individual or combined application of the above-mentioned enhancement methods, as well as the equivalent adjustment of their technical parameters (such as modification of the flip probability threshold), all belong to the conventional extended implementation methods of the technical solution of this patent and are automatically included in the scope of protection of the claims.

[0077] S2: Build an improved YOLOv11n model, which includes a backbone network, a neck network, and a detection head, and make the following improvements:

[0078] S21: Replace the C3K2 block in the neck network of YOLOv11n with the C3K2_CSP_MSA block, which fuses multi-scale features through a progressive multi-kernel convolution chain and residual feature reconstruction strategy;

[0079] In YOLOv11n, the original C3K2 module's Bottleneck architecture uses convolutional layers with kernel sizes of 1 and 3 to improve computational efficiency. However, its limited receptive field prevents it from effectively fusing features from discontinuous occluded areas, resulting in reduced robustness to occluded features. When a student target is obscured by the front row (for example, only the shoulder or arm is visible), the original module struggles to effectively reconstruct the contextual information of the occluded area using hierarchical features. This is particularly true when recognizing similar postures (such as looking at a phone and reading a book), as the small differences in local features can easily lead to misjudgments.

[0080] To solve the above problems, a cross-stage partial multi-scale feature residual fusion module (Cross Stage Partial Multi-Scale Aggregation, CSP_MSA) is proposed, such as Figure 5 The CSP_MSA feature recombination mechanism and multi-level grouped convolution are the core of this module. By chaining and stacking grouped convolutions with increasing kernel sizes of 3, 5, and 7, it gradually extracts fine-grained features into coarse-grained features, thus overcoming the limitation of the single receptive field of the original Bottleneck structure and enhancing the model's multi-scale capabilities.

[0081] Furthermore, the construction of the C3K2_CSP_MSA block in S21 includes replacing the BottleNeck structure in the C3K2 block of the YOLOv11n neck network with the CSP_MSA module. Figure 4 .

[0082] For details, please refer to the attached Figure 5 , is the structural diagram of the CSP_MSA structural block. The construction process of the CSP_MSA module is:

[0083] S21a: The input feature map is processed by a first convolutional layer with a convolution kernel size of 3 and a step size of 1, the input and output channel numbers remain unchanged, and the output feature map is equally divided into a first sub-feature map X1 and a second sub-feature map X2;

[0084] S21b: The second sub-feature map X2 is sequentially processed as follows:

[0085] S21b1: Extract features by a second convolutional layer with a convolution kernel size of 5 and a step size of 1, and equally divide the output feature map into a third sub-feature map X3 and a fourth sub-feature map X4;

[0086] S21b2: The fourth sub-feature map X4 is processed by a third convolutional layer with a convolution kernel size of 7 and a step size of 1 to generate a fifth sub-feature map U4;

[0087] S21c: The fifth sub-feature map U4, the third sub-feature map X3 and the first sub-feature map X1 are spliced in the channel dimension, wherein the fifth sub-feature map U4 is the output of 7x7 convolution, the third sub-feature map X3 is the output of 5x5 convolution, and the first sub-feature map X1 is the output of 3x3 convolution, forming a multi-scale fusion feature map;

[0088] S21d: Align the multi-scale fusion feature map in the channel dimension by a fourth convolutional layer with a convolution kernel size of 1x1 and a step size of 1 to generate an aligned feature map;

[0089] S21e: Add the aligned feature map and the input feature map of step S21a through residual connection to output the final feature, completing the progressive multi-core feature fusion of cross-layer cascade. This helps gradient propagation, reduces computational complexity while maintaining the multi-scale nature of the feature, and compensates for the receptive field loss and maintains the integrity of the bottom features. This makes the model have better occlusion robustness in high-density occlusion environment, thus helping to solve the problem of large image size and dense student targets in the smart classroom environment.

[0090] In deep convolution, the size of the convolution kernel increases, the receptive field range expands, and the parameter quantity also increases. Therefore, with the help of CSPNet idea, the output of each convolutional layer is subjected to channel division operation, while the local features are retained and the remaining channels are passed to the next layer. Since the depth of the convolutional layer will cause the number of channels to decrease, the increase of the convolution kernel at this time will not cause the parameter quantity to increase greatly, thereby reducing the computational complexity while maintaining the multi-scale nature of the feature.

[0091] In summary, the application of the C3K2_CSP_MSA structure block helps gradient propagation, reduces computational complexity while maintaining feature multi-scale, compensates for receptive field loss and maintains the integrity of the bottom features, so that the model has better occlusion robustness in high-density occlusion environment, thus helping to solve the problem of large image scale change and dense student target in the smart classroom environment.

[0092] S22: replace the C2PSA structure block in the YOLOv11n model with a CSPDA structure block based on a deformable attention mechanism, the CSPDA structure block adjusts the attention range through a dynamic offset network;

[0093] It should be noted that in the classroom environment, the C2PSA module is difficult to capture the geometric changes of multi-scale targets through the original self-attention mechanism, resulting in limited spatial modeling capability in complex scenes. Due to the camera's top-down angle, the student target scale is distorted, and the scale difference between the front and back rows of students in the classroom is large, the original attention mechanism of C2PSA is rigid and cannot effectively adapt to these changes. This defect causes the model to focus on the image focus to deviate, thereby reducing the recognition effect of multi-scale targets.

[0094] To solve the above problems, deformable attention is introduced to replace the self-attention module in C2PSA, and a cross-stage partial deformable attention module CSPDA (Cross Stage Partial with Deformable Attention) is proposed. The CSPDA module effectively helps the model network to adaptively focus on the multi-scale target area, as shown in Figure 5 As shown, it inherits the main part of the C2PSA module, but simplifies the attention layer to 1 layer, thereby effectively reducing the model parameter amount and improving the calculation efficiency.

[0095] The specific processing process of the CSPDA module includes:

[0096] S22a: the input feature map is sequentially subjected to channel transformation through a fifth convolution layer with a convolution kernel size of 1 and a step size of 1, then subjected to BatchNorm normalization processing to stabilize the data distribution, and then subjected to SiLU activation function processing to obtain an initial feature map; This step aims to preliminarily adjust the channel representation of the feature map, while not changing the spatial size of the feature map. BatchNorm normalization processing helps to stabilize the data distribution, prevent gradient vanishing or explosion, and improve the stability of model training. The SiLU activation function enhances the expression ability of the model through its nonlinear characteristics.

[0097] S22b: split the initial feature map into a first branch feature and a second branch feature in the channel dimension, wherein:

[0098] S22b1: The first branch feature is directly retained;

[0099] S22b2: The second branch feature is input to the DAModule for processing. The DAModule is composed of a deformable attention module DAttention and a feedforward network FFN, and the DAttention output is added to the input feature through a residual connection, thereby improving the model learning ability.

[0100] DAModule is the core component of the CSPDA module, responsible for processing the second branch features. Figure 6 As shown in the figure, DAttention first feeds a feature map into an offset network, adaptively calculating the sampling network offset and generating a grid of reference points. The offset and reference points are combined to determine the final sampling positions. After calculating the sampling positions, bilinear interpolation is used to resample the features to obtain a new feature representation. Furthermore, a deep convolutional position encoding strategy is used to calculate relative position deviations and incorporate them into the multi-head attention calculation to enhance the model's perception of spatial position information.

[0101] The core advantage of the deformable attention-based DAttention building block lies in its ability to dynamically adjust the receptive field, enabling the model to flexibly focus on the visible portion of occluded objects and adaptively adjust its attention range based on different classroom environments. This mechanism effectively addresses the difficulty of detecting occluded objects in classroom settings and enhances the model's generalization capabilities across diverse classroom environments.

[0102] The deformable attention module DAttention can refer to the attached Figure 6 It should be noted that the tensor sizes shown in the figure are only used to understand the mechanism of this module. The actual size may be affected by parameter settings and is subject to actual conditions.

[0103] The deformable attention mechanism can be found in the attached Figure 8 : Feature map of the input First, construct a uniformly distributed grid as a reference point in Ration is the preset downsampling ratio, and the reference points are normalized. The feature map is converted into a query vector q = xW through a linear transformation operation. q , while using the offset network θ offset To calculate the spatial position offset. By performing feature sampling operations at the deformation point position, the key vectors are extracted respectively. Sum vector Finally, using the sampling function Get deformation feature map The sampling function uses bilinear interpolation, and the specific formula is as follows:

[0104]

[0105] wherein is defined as:

[0106]

[0107] wherein: g(a, b) = max(0, 1 - |a - b|); (r x y represent all positions of the feature map.

[0108] In the multi-head deformable attention mechanism, for the query vector q, the deformation key and the deformation value are implemented jointly; specifically, the output result corresponding to the mth attention head can be obtained by the following formula:

[0109]

[0110] wherein: represents a position embedding matrix, wherein R is a relative position bias offset; σ represents a softmax activation function; d represents a key vector dimension; after processing of all attention heads is completed, it is spliced together through a projection layer to generate a final feature map Z, that is, the output of the deformable attention mechanism.

[0111] S22c: the processing results of the first branch feature and the second branch feature are spliced in the channel dimension, and a sixth convolutional layer with a convolution kernel size of 1 and a step size of 1 is used for nonlinear transformation to obtain the output feature of the CSPDA structure block.

[0112] The CSPDA module adopts a channel segmentation strategy to ensure that local features are not lost due to DAModule module calculation while reducing the amount of calculation, and is suitable for real-time detection tasks in a classroom environment. Compared with the original self-attention mechanism, the deformable attention mechanism dynamically adjusts the receptive field through the deformable offset calculation module, so that the model can flexibly focus on the visible part of the occluded target, and adaptively adjust the attention range according to different classroom environments, solve the difficulty problem of occluded target detection in the classroom environment, and enhance the generalization ability of the model for different classroom environments.

[0113] ​S23: replace the CONV module in the YOLOv11n model with a context-guided down-sampling module CGConv, which retains down-sampling features through a parallel local context branch, a surrounding context branch, and a global attention mechanism; it should be noted that in the YOLO algorithm, the down-sampling process mainly extracts important information in the image through convolution and pooling operations, which can retain important semantic information in the image to a certain extent. However, due to the limitation of the receptive field, a single scale convolution cannot distinguish important information in the down-sampling process, and the loss of key feature information that can help the model identify the correct target. In the classroom environment, due to the position of the camera, the size of the students in the back row is small and there is mutual occlusion. Small targets have low resolution and less effective feature information, and are disturbed by surrounding invalid features. Such a single down-sampling operation often loses important information during feature abstraction, and small target information is severely lost in high-dimensional features, which is not conducive to detecting student targets in the edge area in the classroom environment. Therefore, the application proposes a context-guided down-sampling module CGConv (Context Guided Convolution). The CGConv module relies on context information to understand the scene, and then retains important information during down-sampling.

[0114] Specifically, the construction of the context-guided down-sampling module CGConv in S23 includes the following steps:

[0115] The construction of the context-guided down-sampling module CGConv in S23 includes the following specific steps:

[0116] S23a: input the feature map through a convolution layer with a convolution kernel size of 3 and a step size of 1 for down-sampling, halve the height H and width W of the feature map, double the number of channels, and generate an initial down-sampling feature map;

[0118] S23b: perform parallel branch processing on the initial down-sampling feature map:

[0119] S23b1: local context branch f local extract local features from adjacent 8 feature vectors through a convolution layer with a convolution kernel size of 3 and a dilation of 1;

[0120] S23b2: surrounding context branch f surrounding extract large-scale surrounding context features through a convolution layer with a convolution kernel size of 3 and a dilation of 3;

[0121] S23c: concatenate the features output by the local context branch f local and the features output by the surrounding context branch f surrounding in the channel dimension to form a joint feature map fjoint ;

[0121] S23d: processing the joint feature map f joint The batch normalization processing, PReLU activation function processing are sequentially performed, and then a convolution layer with a kernel size of 1 and a step of 1 is used to restore the channel number, so as to be consistent with the initial down-sampling feature map channel number; PReLU (Parametric Rectified Linear Unit) is a parameterized Relu, which is defined as:

[0122]

[0123] Obviously, the activation function is an improved activation function of the ReLU activation function. The main difference between PReLU and ReLU is that the slope on the negative half-axis is no longer fixed as 0 as a learnable parameter. PReLU allows the alpha parameter value to be adjusted adaptively during backpropagation to adapt to the needs of specific tasks and scenarios. At the same time, it relieves the problem of dead neurons, allowing negative inputs to pass through neurons, rather than the negative input in the ReLU activation function making the neuron output constant 0, which cannot update the weight (dead neuron).

[0124] S23e: enhancing the feature map f globle The feature map output by step S23d is processed, including:

[0125] (e1) performing global average pooling on the feature map output by step S23d to generate a global feature vector for each channel;

[0126] (e2) constructing a channel attention mechanism through two fully connected layers, wherein the first fully connected layer uses a ReLU activation function for nonlinear transformation, and the second fully connected layer uses a Sigmoid function to generate a channel weight coefficient; an activation function is a function added to a neural network to help the network learn complex patterns in the data. Similar to the neuron model in the biological body, the activation function decides the output to the next neuron according to a certain strategy according to the current neuron. In a neural network, the activation function of a node defines the output of the node under a given input or input set. Therefore, the activation function is a mathematical equation that determines the output of a neural network.

[0127] ReLU (Rectified Linear Unit) is a commonly used activation function in neural networks, which is defined as:

[0128]

[0129] For any input x, if x is less than or equal to 0, the output of ReLu is 0; if x is greater than 0, the output is itself. In a neural network, if only linear transformation is used, multiple linear layer combinations are still linear transformations, which cannot capture the nonlinear properties of the ReLu function. The nonlinear properties of the ReLu function can enable the neural network to represent and learn complex nonlinear mappings, improving the nonlinear representation ability of the model.

[0130] (e3) multiplying the channel weight coefficients with the feature maps output by step S23d channel by channel, to output final down-sampling features.

[0131] For details, please refer to the accompanying drawings Figure 9 Specifically, the feature map first passes through a convolutional layer with a step size of 1 and a convolution kernel size of 3, which halves the H and W of the input feature map and doubles the number of channels. Then, the feature map passes through two parallel branches of f local and f surrounding to capture the local and surrounding context of the image features. Specifically, the red area in figure (a) corresponds to f local , which is used to learn local features from 8 adjacent feature vectors; the green area in figure (b) corresponds to f surrounding , which is used for dilated convolution. The convolution kernel size of both parts is 3, f surrounding The dilation of the convolutional layer is 3, and the dilation of the f surrounding local convolutional layer is 1. Since the dilation of f local is larger, it has a larger receptive field, which enables it to better learn surrounding context information.

[0132] Further, f joint obtains joint feature information from the outputs of f surrounding and f local . f joint is a simple concatenation layer, followed by batch normalization and PReLU activation function, and finally through a convolutional layer with a kernel size of 1 and a step size of 1 to restore the number of channels to the same as the number of channels after the first step of down-sampling, maintaining the function of the down-sampling module.

[0133] Further, after f globle part, the global context corresponding to the purple area in figure (c) is aggregated with the output after f joint . f globle compresses the input feature map in the spatial dimension to a single global value for each channel through global average pooling, and then uses two consecutive fully connected layers to build a channel attention mechanism, which is nonlinearly transformed by a ReLU activation function and normalized by a Sigmoid function, to adaptively generate weight coefficients for each channel.

[0134] Further, the weight coefficients are multiplied by the original feature map, thereby realizing adaptive enhancement and inhibition of different channels, enabling the model to dynamically adjust the feature representation according to the global context, improving the expression ability and discriminability of the features. At the same time, the features are linked with the global context, thereby improving the effectiveness of the feature data.

[0135] The module refines the joint features of the local and surrounding context features using a global weight vector, helping the model to focus on the context during the downsampling process to preserve important local features, so that the model can focus on important features in the classroom environment and reduce irrelevant background interference. It helps to maintain effective information of students in the edge area and reduce the loss of small target features. Combining local and context features improves the small target detection capability. The small student target in the edge area has weak features, but the model can still enhance the representation with the help of surrounding information, thereby improving the detection effect. At the same time, spatial feature weighting is performed in the channel dimension, so that small targets in the edge area will not be ignored due to feature blurring.

[0136] S3: iteratively optimizing and training the improved YOLOv11n model using the training set to generate a student behavior detection model;

[0137] S4: inputting a to-be-tested classroom image into the student behavior detection model to output a student behavior detection result.

[0138] The application also provides a classroom student behavior detection system based on an improved YOLOv11n, comprising:

[0139] An image acquisition module is configured to acquire classroom image data in multiple perspectives and multiple scenes, and divide the data into a training set, a verification set and a test set;

[0140] A model construction module is configured to construct an improved YOLOv11n model, comprising:

[0141] A neck network replacement unit is configured to replace a C3K2 structure block in a YOLOv11n neck network with a C3K2_CSP_MSA structure block, wherein the C3K2_CSP_MSA structure block fuses multi-scale features through a progressive multi-core convolution chain and a residual feature reorganization strategy;

[0142] An attention mechanism replacement unit is configured to replace a C2PSA structure block in the YOLOv11n model with a CSPDA structure block based on a deformable attention mechanism, wherein the CSPDA structure block adjusts the attention range through a dynamic offset network;

[0143] A downsampling module replacement unit is configured to replace a CONV module in the YOLOv11n model with a context-guided downsampling module CGConv, wherein the CGConv module preserves the downsampling features through a parallel local context branch, a surrounding context branch and a global attention mechanism.

[0144] a training module configured to perform iterative optimization training on the improved YOLOv11n model using the training set to generate a student behavior detection model;

[0145] a detection execution module configured to input a to-be-detected classroom image into the student behavior detection model and output a student behavior detection result.

[0146] Experimental Example

[0147] The improved model was trained in an experimental environment shown in Table 1. The input image resolution was 640x640 pixels, the batch size was 16, the number of iterations was 300, the learning rate was 0.01, the number of working threads was 8, the weight decay was 0.0005, and the optimizer was SGD, all of which used the best results.

[0148] Table 1 Experimental Environment

[0149]

[0150] The data sets used were the public data sets SCB3-Dataset (Student Classroom Behavior Dataset Version 3) and POCO-Dataset, and the label classification data of the data sets is shown in Table 2.

[0151] The POCO-Dataset includes various different classroom environments and different camera perspectives. The training data set includes 1146 images and 83566 labels, the test data set includes 378 images and 26909 labels, and the verification data set includes 379 images and 27485 labels. A total of 10 label types are covered, including 8 common student classroom behaviors, i.e., looking up, looking down, turning head, playing mobile phone, reading, sleeping, standing, and bending, and 2 common classroom objects, i.e., mobile phone and book. The SCB-Dataset has a total of 3 label types, i.e., raising hand, reading, and writing. A total of 5015 images and 25810 labels are included.

[0152] It should be noted that the use of the above public data sets and hyperparameters is only for the convenience of reflecting the experimental results and the reproducibility of the experiments. The application of the present application to other data sets is within the protection scope of the present patent.

[0153] Table 2 Classification Data of Different Data Sets

[0154]

[0155]

[0156] The experiment uses parameters and gigaflops per second (GFLOPs) to evaluate the model speed performance, and uses precision (P), recall (R) and mean of average precision (mAP) as model precision evaluation indexes.

[0157]

[0158] In formula (4) and (5), TP is a correctly predicted positive sample, FP is a wrongly predicted positive sample, and FN is a wrongly predicted negative sample. That is, P is used to measure the reliability of the prediction result of the model, and R is used to measure the ability of the model to find real positive samples. In formula (6) and (7), AP is the detection accuracy of a single class, and mAP represents the average detection accuracy of all classes, which is the average of the sum of all-class AP. mAP@50 represents the average precision of all classes at an IoU threshold of 0.5, and is used to measure the detection performance of the model under loose positioning requirements; mAP@50-95 represents the average precision of all classes at an IoU threshold in the range of 0.5 to 0.95 (step size is 0.05), and is used to measure the comprehensive precision of the model under strict positioning requirements. In formula (8), N is the total number of model layers, Kernel is the size of the i-th convolution kernel; Channel is the input channel number; Height is the input feature map height; and Width is the input feature map width.

[0159] The present application is compared with the advanced target detection algorithm to illustrate that the baseline model selected by the present application and the improvement of the present application are optimal. The experimental environment and related parameters are the same, and the best results of each algorithm are selected for comparison. The experimental results are shown in Table 3.

[0160] Although Rtdetr-r18 has good advantages in precision, its parameter quantity and GFLOPs reach 19.8M and 57 respectively, which is larger than the lightweight model of the YOLO series algorithm. The DETR series algorithm requires a large amount of data, and may converge slowly on a small data set due to insufficient data diversity, and requires high computing resources, which does not meet the lightweight requirements of edge device deployment in a classroom environment. In the YOLO series model, yolov11n is the best in terms of detection accuracy and model parameters, and can balance the detection accuracy and computational efficiency. The better calculation accuracy and lower calculation parameter quantity can meet the real-time requirements of the classroom environment, and also have the ability to be deployed in embedded environments and edge system devices with limited computing resources.

[0161] Table 3 comparative experiment results

[0162]

[0163] To verify the effectiveness of the improved structure block, the experiment based on YOLOv11n model carried out progressive ablation experiment in POCO Dataset and SCB Dataset respectively, and the experimental results are shown in Table 3. All experiments use the same training strategy and environment configuration.

[0164] Table 4 ablation experiment results

[0165]

[0166] In order to show the independent effectiveness of each module, C3K2_CSP_MSA structure block, CGConv structure block and CSPDA module are introduced into yolov11n respectively for comparison; the above experiments show that a single module has a promoting effect on each detection accuracy index (mAP@50, mAP@50-95, P, R) of the model on POCO-Dataset, among which CSP_MSA has the most attention on the promotion of each detection accuracy index at the cost of slight promotion of parameter quantity and GFLOPs. Its mAP@50, mAP@50-95, P and R are improved by 2.0, 2.1, 0.5 and 2.1 percentage points respectively compared with the baseline model; in SCB-Dataset, each module has a slight decline in R index compared with the baseline model, and has a certain improvement in mAP@50, mAP@50-95 and P index compared with the baseline model, and the difference is that CGConv module has the most obvious promotion to the model; its mAP@50, mAP@50-95 and P are improved by 2.4, 1.8 and 3.5 percentage points respectively compared with the baseline model. Combining the different performances of single module on two datasets, it is not difficult to see that CGConv module has more promotion in model parameter quantity and GFLOPs; CGConv module is the main influencing factor that makes the parameter quantity and GFLOPs of the improved YOLOv11n algorithm higher. C3K2_CSP_MSA and CGConv module have the most obvious promotion on the accuracy evaluation index of two datasets respectively, which shows that the adaptability of these two modules in different classroom environments is different, and the organic integration of the three modules helps the model to better extract target features in different classroom environments, thereby improving the model generalization ability.

[0167] The synergistic effect of the improved modules is evaluated by a step-by-step integration strategy: in the SCBDataset, when the C3K2_CSP_MSA and CGConv modules are introduced at the same time, the two modules form a complement in the reconstruction of local features and the guidance of global context dimensions, respectively; this makes the model improve most obviously in the accuracy evaluation indicators except R, and the mAP@50, mAP@50-95, P of the model are improved by 2.9, 2.9 and 3.2 percentage points respectively compared with the baseline model, and the combination of other modules also improves the accuracy evaluation indicators except R. In the POCODataset, the introduction of the CSP_MSA and CGConv modules at the same time makes the model improve by 2.9, 2.9 and 3.2 percentage points in mAP@50, mAP@50-95, P and R indicators respectively compared with the baseline model.

[0168] When the C3K2_CSP_MSA, CGConv and CSPDA modules are added at the same time, whether in the POCODataset or the SCBDataset, the model reaches the highest in each detection accuracy indicator, and it is worth noting that in the SCBDataset, a single module or the fusion operation between two modules always makes the R indicator lower than the baseline model. After the fusion of the three modules, the R indicator is improved by 1.4 percentage points compared with the baseline model. This shows that the combination of the three modules can balance the performance and complexity. Specifically, the synergistic effect of the C3K2_CSP_MSA and CGConv modules improves the multi-scale feature extraction and fusion capability of the model, and at the same time improves the generalization ability of the model to adapt to different complex classroom environments. The effect of the CSPDA module alone is not obvious to the model, but under the multi-scale feature extraction capability of the previous two modules, the deformable attention mechanism effect is significant, so that all the accuracy indicators of the model reach the highest value.

[0169] In order to reflect the improved detection effect of the present application, images are selected from the dataset for visual analysis, and reference is made to the attached Figure 10 ; at the same time, this analysis can show the behavior attention distribution of the model in different image regions, so as to reflect the detection focus of the model in different classroom scenes.

[0170] Figure 10 (a) is the original image, Figure 10 (b) and Figure 10(c) YOLOv11n and the action of the application of the behavior attention map respectively. From left to right, in turn: (1) in the first picture, the improved model of the application can effectively separate the background and the effective target (such as the mobile phone at the bottom left) in the sparse target environment, reduce false positives and missed detection, and improve the overall performance of the model; (2) in the second and third pictures of the ultra-dense multi-scale target classroom environment, compared with YOLOv11n, the improved model can better capture multi-scale target boundary information and establish the relationship between targets, such as books and students, and better distinguish classroom student behavior. It is worth noting that in these two pictures, the improved model of the application pays high attention to the front row area. The reason is that, on the one hand, the front row area is an important area for the model to identify the classroom scene, so the improved model of the application has a more adaptive strategy to deal with different characteristics of different scenes and improves the generalization ability of the model; on the other hand, in the front row area, large-scale targets usually appear, which shows that the improved model of the application has a unique strategy to deal with the problem of difficult detection of large-scale targets.

[0171] The third aspect of the embodiment of the present application provides a computer device, which comprises a memory, a processor, and a processing program stored on the memory and executable on the processor, and the processing program is executed by the processor to implement the above-mentioned student behavior robust identification method for smart classrooms.

[0172] The fourth aspect of the embodiment of the present application provides a computer readable storage medium, which stores a processing program, and the processing program is executed by a processor to implement the above-mentioned student behavior robust identification method for smart classrooms.

[0173] It should be particularly stated that the embodiments of the present application are only used to illustrate the technical implementation principles of the present application, and the specific parameter configuration, structure combination mode and implementation scene selection do not constitute a limitation on the protection scope of the patent right. Any technical solution derived by adjusting the module connection relationship, optimizing the algorithm parameters or equivalent replacing the feature components of any person skilled in the art based on the basic principles of the present application, as long as the technical essence does not deviate from the innovative kernel defined in the claims, should be considered to fall within the protection scope of the present application.

[0174] It is further clarified that the enumeration of the embodiments of the present specification does not constitute an exclusive implementation of the technical solutions. In the premise of not violating the core design idea of the present application, for the network architecture adjustment, data processing flow optimization or equivalent replacement of module function, etc. Common technical improvements, which can be deduced by those skilled in the art without creative labor, belong to the implementation mode. The equivalent transformation or reorganization of elements of such technical solutions is limited to the protection scope recorded in the claims of the present application.

Claims

1. A robust student behavior recognition method for smart classrooms, characterized by: The following steps are involved: S1: Obtain multi-view and multi-scene classroom image data and divide it into training set, validation set and test set; S2: Build an improved YOLOv11n model, which includes a backbone network, a neck network, and a detection head, and make the following improvements: S21: Replace the C3K2 block in the neck network of YOLOv11n with the C3K2_CSP_MSA block, which fuses multi-scale features through a progressive multi-kernel convolution chain and residual feature reconstruction strategy; S22: Replace the C2PSA block in the YOLOv11n model with a CSPDA block based on a deformable attention mechanism, which adjusts the attention range through a dynamic offset network; S23: Replace the CONV module in the YOLOv11n model with a context-guided downsampling module CGConv, which preserves downsampling features through parallel local context branches, surrounding context branches, and a global attention mechanism; S3: Using the training set, the improved YOLOv11n model is iteratively optimized and trained to generate a student behavior detection model; S4: Input the classroom image to be tested into the student behavior detection model and output the student behavior detection result.

2. A method for robustly identifying student behavior in a smart classroom according to claim 1, characterized in that: The construction of the C3K2_CSP_MSA structure block in step S21 includes: replacing the BottleNeck structure in the C3K2 structure block of the YOLOv11n neck network with a CSP_MSA module; the processing process of the CSP_MSA module is as follows: S21a: The input feature map is processed by the first convolution layer with a convolution kernel size of 3 and a stride of 1, keeping the number of input and output channels unchanged, and the output feature map is equally divided into the first sub-feature map X1 and the second sub-feature map X2; S21b: The second sub-feature graph X2 undergoes the following processing in sequence: S21b1: Extract features through the second convolutional layer with a convolution kernel size of 5 and a stride of 1, and divide the output feature map equally into the third sub-feature map X3 and the fourth sub-feature map X4; S21b2: The fourth sub-feature map X4 is processed by a third convolution layer with a convolution kernel size of 7 and a stride of 1 to generate a fifth sub-feature map U4; S21c: Concatenate the fifth sub-feature map U4, the third sub-feature map X3, and the first sub-feature map X1 according to the channel dimension, wherein the fifth sub-feature map U4 is the output of 7×7 convolution, the third sub-feature map X3 is the output of 5×5 convolution, and the first sub-feature map X1 is the output of 3×3 convolution, to form a multi-scale fusion feature map; S21d: performing channel dimension alignment on the multi-scale fusion feature map through a fourth convolutional layer with a convolution kernel size of 1×1 and a stride of 1 to generate an aligned feature map; S21e: Add the aligned feature map to the input feature map of step S21a through a residual connection, output the final feature, and complete the progressive multi-core feature fusion of cross-layer cascade.

3. The method for robustly identifying student behavior in a smart classroom according to claim 1 is characterized in that: In step S22, replacing the C2PSA structure block in the YOLOv11n model with the CSPDA structure block based on the deformable attention mechanism includes: replacing the Attention block in the C2PSA block with the DAttention block. The processing process of the CSPDA module includes: S22a: The input feature map is sequentially transformed through the fifth convolution layer with a convolution kernel size of 1 and a stride of 1, and then BatchNorm normalization is performed to stabilize the data distribution. Subsequently, the SiLU activation function is used to obtain the initial feature map. S22b: Split the initial feature map into first branch features and second branch features in the channel dimension, wherein: S22b1: The first branch feature is directly retained; S22b2: The second branch feature is input to the DAModule for processing. The DAModule consists of a deformable attention module DAttention and a feedforward network FFN, and the DAttention output is added to the input feature through a residual connection; S22c: Concatenate the processing results of the first branch features and the second branch features according to the channel dimension, and perform nonlinear transformation through the sixth convolution layer with a convolution kernel size of 1 and a step size of 1 to obtain the output features of the CSPDA structure block.

4. The method for robustly identifying student behavior in a smart classroom according to claim 3 is characterized in that: The processing of the DAttention module includes: Input feature map Generate a uniformly distributed reference grid in Ration is the preset downsampling ratio, and the reference points are normalized; The input feature map is converted into a query vector q=xW through a linear transformation operation q , while using the offset network θ offset To calculate the spatial position offset; By performing feature sampling operations at the deformation point position, the key vectors are extracted respectively Sum vector Finally, using the sampling function Get deformation feature map The sampling function uses bilinear interpolation, and the specific formula is as follows: In the formula Defined as: Where, g(a,b)=max(0,1-|ab|); (r x ,r y ) represents all positions of the feature map; In the multi-head deformable attention mechanism, for the query vector q, the deformation key and deformation value Implement joint calculation; the output result corresponding to the mth attention head can be obtained by the following formula: Where, represents the position embedding matrix; R is the relative position deviation offset; σ represents the softmax activation function; d represents the key vector dimension; After all attention heads are processed, they are spliced ​​together and passed through the projection layer to generate the final feature map Z, which is the output of the deformable attention mechanism.

5. The method for robustly identifying student behavior in a smart classroom according to claim 1 is characterized in that: The construction of the context-guided downsampling module CGConv in S23 includes the following specific steps: S23a: The input feature map is downsampled through a convolution layer with a kernel size of 3 and a stride of 1. The height H and width W of the feature map are halved and the number of channels is doubled to generate the initial downsampled feature map. S23b: Perform parallel branch processing on the initial down-sampling feature map: S23b1: Local context branch f local : Extract local features from 8 adjacent feature vectors through a convolution layer with a convolution kernel size of 3 and a dilation of 1; S23b2: surrounding context branch f surrounding : Extracting large-scale surrounding context features through a dilated convolution layer with a kernel size of 3 and a dilation of 3; S23c: branch the local context f local The output features and the surrounding context branch f surrounding The output features are concatenated according to the channel dimension to form a joint feature map f joint ; S23d: the joint feature map f joint Perform batch normalization and PReLU activation function processing in sequence, and then restore the number of channels through a convolution layer with a convolution kernel size of 1 and a stride of 1 to make it consistent with the number of channels of the initial downsampled feature map; S23e: Enhanced module f through global attention globle The feature map outputted in step S23d is processed, including: (e1) performing global average pooling on the feature map output in step S23d to generate a global feature vector for each channel; (e2) Construct a channel attention mechanism through two fully connected layers, where the first fully connected layer uses the ReLU activation function for nonlinear transformation, and the second fully connected layer uses the Sigmoid function to generate channel weight coefficients; (e3) Multiply the channel weight coefficient by the feature map output in step S23d channel by channel to output the final downsampled feature.