Classroom behavior identification method based on improved YOLOv12 model

By introducing A2C2f_FRFN and C3k2_SAVSS modules in the YOLOv12 model, the feature expression and robustness of classroom behavior recognition are enhanced, and the detection accuracy problem of existing algorithms in occlusion and multi-scale change scenarios is solved, and efficient and accurate identification of classroom behavior is achieved.

CN120452061AActive Publication Date: 2025-08-08CHANGCHUN NORMAL UNIV

Patent Information

Application Number
CN202510524422.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-08
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing YOLO series algorithms lack the ability to capture global information in classroom behavior recognition tasks. The detection accuracy and robustness need to be improved when facing occlusion and multi-scale changes, especially in classroom scenarios with dense people, diverse behaviors and frequent occlusions.

Method used

In the improved YOLOv12 model, a P2 detection head is added, and the A2C2f_FRFN module combined with the feature refinement feedforward network FRFN module is designed at the backbone network and the neck network, and the C3k2_SAVSS module combined with the structure-aware visual state space module SAVSS is designed to enhance the attention to small movements or posture changes and capture details through multi-stage linear transformation, convolutional operation and gating mechanisms, while capturing long-range dependence and global features are captured through multi-directional scanning strategies.

Benefits of technology

It significantly improves the model's feature expression ability and recognition accuracy in complex scenarios, can accurately extract local action details in the classroom and correlate the previous and subsequent behaviors, adapt to scenes with different lighting and occlusion levels, and improves the accuracy and robustness of classroom behavior recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452061A_ABST
    Figure CN120452061A_ABST
Patent Text Reader

Abstract

The invention discloses a classroom behavior recognition method based on an improved YOLOv12 model, and the method comprises the steps: proposing the improved YOLOv12 model according to a classroom behavior recognition application scene, and designing and using an A2C2fFRFN module which combines A2C2f with a feature refinement feedforward network FRFN module at a backbone network and a neck network, by means of multi-stage linear transformation, convolution operation and a gating mechanism, the attention degree on tiny actions or posture changes can be enhanced in classroom behavior recognition, and the ability of capturing details can be improved. A C3k2SAVSS module combining C3k2 and a structure perception visual state space module SAVSS is designed and used at a backbone network and a neck network, long-range dependence and global features can be captured through direction scanning (horizontal, vertical and diagonal) and a structured state space model, and local details and global context can be comprehensively fused through combination of the long-range dependence and the global features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of image processing technology and education, and relates to a classroom behavior recognition method, and specifically to a classroom behavior recognition method based on an improved YOLOv12 model. Background Art

[0002] In modern education, with the in-depth development of educational informatization, the need for refined analysis and management of classroom teaching processes is becoming increasingly urgent. Classroom behavior recognition, as a key technology for achieving this goal, is extremely important. It can objectively and comprehensively capture student behavior in the classroom, providing strong support for teachers to adjust teaching strategies and optimize teaching methods, and for educational researchers to gain a deeper understanding of the teaching process.

[0003] Traditional classroom behavior recognition methods have numerous limitations. For example, relying on manual observation and recording of student behavior is not only labor-intensive and time-consuming, but also highly subjective, prone to omissions and biases, and unable to fully and accurately reflect students' actual classroom performance. With the rise of computer vision technology, some classroom behavior recognition methods based on traditional image processing and machine learning have begun to be applied. However, these methods perform poorly in feature extraction and model generalization. Faced with complex and changing classroom scenarios, such as varying lighting conditions, diverse student postures, and frequent occlusions, recognition accuracy is low, failing to meet the needs of practical applications.

[0004] In recent years, deep learning technology has achieved significant breakthroughs in object detection and image recognition, bringing new opportunities for classroom behavior recognition. Deep learning-based object detection algorithms are primarily categorized into one-stage and two-stage algorithms. While two-stage algorithms offer higher accuracy, they are slower and computationally more complex. One-stage algorithms require only a single pass for detection, requiring less computation and making them more suitable for real-time applications, such as classroom behavior recognition, which requires real-time analysis of video streams. Among one-stage algorithms, the YOLO family of algorithms has garnered widespread attention and application due to its excellent balance between accuracy and speed.

[0005] However, existing YOLO algorithms still face challenges when it comes to classroom behavior recognition. Classrooms are characterized by dense crowds, diverse behaviors, frequent occlusions, and large scale variations. Existing YOLO models are unable to capture global information, and their detection accuracy and robustness in the face of occlusion and multi-scale variations need improvement. For example, when multiple students partially occlude one another, the model is prone to missed or false detections. Furthermore, due to the large scale differences among students at varying distances from the camera, the model performs poorly for detecting small-scale objects. Summary of the Invention

[0006] In order to solve the above-mentioned problems existing in the existing YOLO series algorithms in processing classroom behavior recognition tasks, the present invention provides a classroom behavior recognition method based on an improved YOLOv12 model.

[0007] The purpose of the present invention is achieved through the following technical solutions:

[0008] A classroom behavior recognition method based on an improved YOLOv12 model includes the following steps:

[0009] Step 1: Obtain a dataset of classroom behavior images;

[0010] Step 2: Divide the dataset into validation set, training set, and test set, and process them into the network model recognition format of YOLOv12;

[0011] Step 3: Use the classroom behavior recognition detection model training dataset of the improved YOLOv12 model. The improved YOLOv12 model adds a P2 detection head to the detection layer, designs and uses the A2C2f_FRFN module that combines the A2C2f with the feature refinement feedforward network FRFN module in the backbone network and the neck network, and designs and uses the C3k2_SAVSS module that combines the C3k2 with the structure perception visual state space module SAVSS in the backbone network and the neck network;

[0012] Step 4: Use the training set and validation set to train the classroom behavior recognition detection model. After obtaining the optimal detection model, use the test set to evaluate the performance of the classroom behavior recognition detection model.

[0013] The training set and validation set are input into the classroom behavior recognition detection model of the improved YOLOv12 model, and the number of training times is set. As the number of training times increases, the loss function curve of the detection model gradually converges. When the loss function curve converges and stabilizes, the detection model is trained to the optimal state, and its optimal model weight file is saved. The image to be detected in the test set is input into the trained optimal behavior recognition detection model, and the detected image is output, where the detection image includes the type of each detection target, and the position of each target in the target detection image is marked;

[0014] Step 5: Combine monitoring with the optimal detection model to conduct real-time identification and detection of students' classroom behavior.

[0015] Compared with the prior art, the present invention has the following advantages:

[0016] The present invention proposes an improved YOLOv12 model based on the application scenario of classroom behavior recognition. The A2C2f_FRFN module that combines A2C2f with the feature refinement feedforward network FRFN module is innovatively designed and used in the backbone network and neck network. Through multi-stage linear transformation, convolution operation and gating mechanism, it can enhance the attention to small movements or posture changes and improve the ability to capture details in classroom behavior recognition. The C3k2_SAVSS module that combines C3k2 with the structure perception visual state space module SAVSS is innovatively designed and used in the backbone network and neck network. Through directional scanning (horizontal, vertical, diagonal) and structured state space model, it can capture long-range dependencies and global features. The combination of the two can fully integrate local details and global context. For example, in the classroom, it can accurately extract local action details such as students writing and raising their hands, and can also associate previous and subsequent behaviors to enhance the richness of feature expression. The multi-directional scanning strategy of SAVSS enables the module to capture multi-angle behaviors in the classroom more comprehensively, and can generate scenes that adapt to different lighting and occlusion levels through dynamic parameters. For example, when a student raises his hand sideways, multi-directional feature fusion can still accurately extract action features. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a flowchart of the classroom behavior recognition method based on the improved YOLOv12 model;

[0018] Figure 2 This is a detection model diagram of the classroom behavior recognition method based on the improved YOLOv12 model;

[0019] Figure 3 This is a schematic diagram of the A2C2f_FRFN module structure;

[0020] Figure 4 This is a schematic diagram of the ABlock_FRFN module structure;

[0021] Figure 5 This is a schematic diagram of the C3k2_SAVSS module structure;

[0022] Figure 6 This is a schematic diagram of the SAVSS module structure;

[0023] Figure 7 It is a schematic diagram of the GBC module structure;

[0024] Figure 8 For the FRFN2D module;

[0025] Figure 9 Schematic diagram of the PAF module structure. DETAILED DESCRIPTION

[0026] The technical solution of the present invention is further described below with reference to the accompanying drawings, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention that does not depart from the spirit and scope of the technical solution of the present invention should be included in the scope of protection of the present invention.

[0027] A classroom behavior recognition method based on an improved YOLOv12 model, the method comprising the following steps:

[0028] Step 1: Obtain a dataset of classroom behavior images.

[0029] Step 2: Divide the dataset into validation set, training set, and test set, and process them into the network model recognition format of YOLOv12.

[0030] Step 3. Use the classroom behavior recognition detection model training data set of the improved YOLOv12 model. The improved YOLOv12 model adds a P2 detection head on the detection layer, and designs and uses the A2C2f_FRFN module that combines the A2C2f with the feature refinement feedforward network FRFN module in the backbone network and the neck network. The C3k2_SAVSS module that combines the C3k2 with the structure perception visual state space module SAVSS is designed and used in the backbone network and the neck network.

[0031] In this step, the implementation process of the A2C2f_FRFN module is as follows: the input feature vector X is first fed into the initial convolutional layer built based on the parent class initialization logic. This layer performs channel compression on the input feature map through low-rank mapping. Using the sliding convolution of the convolution kernel in the feature map spatial domain, it performs a multiplication and accumulation operation on the weights and the elements at the corresponding positions, completing the dimensionality transformation from the input channel to the hidden channel and extracting feature information with basic representation capabilities. The feature map processed by this convolution is stored in a specific list structure. It then enters the loop processing flow composed of the ABlock_FRFN module, passing the end element of the list as input to each module sequence in sequence. The feature vector X2 generated by the ABlock_FRFN module is fused with the original input feature vector X through residual mapping, and the final output is the fused feature Y. Through the above-mentioned module architecture design and data processing flow, we can achieve efficient extraction, screening and fusion of complex input features, significantly enhance the discriminability and robustness of feature expression, and provide more discriminative and stable feature representations for high-level semantic tasks such as classroom behavior recognition. This effectively improves the model's ability to capture and process subtle feature changes in complex scenarios, ensuring the accuracy and reliability of recognition tasks from the feature level. Its forward propagation mathematical expression is:

[0032] Y=X+X2

[0033] The implementation process of the ABlock_FRFN module is as follows: the input feature vector X of the previous level is first processed by the regional attention module inherited from the ABlock class. The regional attention mechanism will assign weights to the features, highlight key information, and obtain the feature vector Y mlp1 , and performs the first residual connection with the original previous level input feature vector X to obtain X1. This operation is intended to alleviate the gradient vanishing problem and enhance the model's ability to capture key features. X1 is passed to the multi-layer perceptron module (a sequence consisting of two convolutional layers) in the parent class ABlock for feature transformation to obtain the feature vector Y mlp1 , and performs a second residual connection with X1 to obtain X2. X2 enters the FRFN2D module for processing, and the final feature vector is obtained and residually connected with X2 to obtain the output. Through a series of steps including initial convolution dimensionality reduction, regional attention mechanism, multi-layer perceptron transformation, and feature refinement of the FRFN2D module, the ABlock_FRFN module achieves multi-stage fine processing and deep fusion of input features, significantly enhancing the expressive power of features and providing more valuable feature representations for subsequent tasks. Its forward propagation mathematical expression is:

[0034] X1=X+Y att1

[0035] X2=X1+Y mlp 1

[0036] Output=X2+FRFN2D

[0037] The FRFN2D module is implemented as follows: The input feature vector X from the previous level is first split into two parts, X1 and X2, along the channel dimension. A 3×3 convolution is performed on X1. This step extracts local feature details through local convolution, enhancing the spatial correlation and local representation of features. It is then concatenated with X2 along the channel dimension. While preserving some of the original feature integrity, it also introduces localized, refined features, maintaining the overall structure and information diversity of the features. The concatenated features are then expanded by linear1 (which consists of a 1×1 convolution and an activation function). This expansion broadens the feature representation space by increasing the channel dimension, enabling the model to capture richer feature patterns and semantic information. The activation function further introduces nonlinearity, improving the model's ability to fit complex feature relationships. The features are then equally split into two parts, X3 and X4. A depthwise separable convolution is performed on X3, which reduces computational overhead while extracting local spatial information and inter-channel correlations. It is then element-wise multiplied with X4. A gating mechanism filters out features that are more critical to subsequent tasks, suppresses redundant or irrelevant features, and enhances the discriminative power and effectiveness of the features. Finally, the number of channels is restored to the input dimension dim through linear2 (1×1 convolution), so that the feature dimension adapts to the requirements of subsequent tasks. Without losing key information, the features are integrated and compressed to obtain the final output Output, providing a refined and highly expressive feature representation for subsequent model processing. Its forward propagation mathematical expression is:

[0038] X1=partialconv(X1)

[0039] X = Concat((X1, X2), dim = 1)

[0040] X=linear1(X)

[0041] X3,X4=X·chunk(2,dim=1)

[0042] X3=dwconv(X3)

[0043] X=X3X4

[0044] Output=linear2(X)

[0045] In this step, the implementation process of the C3k2_SAVSS module is as follows: the input feature X of the previous level is adjusted through the initial convolution layer to generate the hidden feature X h =Conv(X), where the number of hidden channels is c, and X h Input is sent to a sequence of n SAVSS modules, each SAVSS performs a specific transformation on the feature and outputs the intermediate feature F i (X h ); finally the final output Fn (X h ) is added to the original previous level input feature X to obtain the fusion result Y. Its forward propagation mathematical expression is:

[0046] X h =Conv(X)

[0047]

[0048] Y=X+F n (X h )

[0049] The implementation process of the SAVSS module is as follows: First, two GBC (Grouped Bottleneck Convolution) operations are performed on the previous level input feature X. This module significantly reduces the computational complexity through the grouped depth-separable convolution structure, and at the same time uses a multi-branch bottleneck design to enhance the local feature expression capability. It contains four parallel submodules: two 3×3 depth-wise convolution branches for capturing local spatial details, two 1×1 convolution branches for cross-channel information interaction, and group normalization (GN) and ReLU activation functions for efficient feature enhancement. Finally, residual connections are used to retain the original information to avoid gradient disappearance. The enhanced feature map X gbc It is expanded into a sequence form from the spatial dimension and fed into the SAVSS_2D module. The long-range dependency modeling is achieved through the structured state space model (SSM) combined with a multi-directional scanning strategy: horizontal scanning captures the pixel association within the row, vertical scanning establishes the dependency within the column, and dual diagonal scanning extracts the global diagonal features. The outputs of the four scanning paths are dynamically weighted and fused to ensure that the model perceives both local structure and global context. During the scanning process, the dynamically generated transfer matrix A and the content-aware projection weight B are adaptively adjusted according to the input features, making the state transfer process input-sensitive and significantly improving the modeling ability of complex spatial patterns. The serialized feature Y output by SSM ssm With X gbc Through deep fusion of the PAF (Position-Aware Fusion) module, the module first generates key-value pairs of base features and guide features through two-way feature transformation, calculates the spatial similarity graph S to quantify the position correlation, and then soft-assigns and fuses the original features and SSM features with S as the weight, which not only retains the underlying details but also injects global structural information. It is especially suitable for feature complementarity in occlusion or small target scenes. In the final stage, the training process is stabilized by linear transformation and residual connection: group normalization (GN) alleviates internal covariate offset, linear layer improves feature dimension adaptability, and jump connection ensures lossless information transmission and outputs Y. outThe entire module uses the three technologies of GBC local enhancement, SAVSS_2D multi-directional global modeling, and PAF dynamic fusion to significantly improve the mIoU index of dense prediction tasks compared to traditional convolutional modules while reducing the amount of computation. Its forward propagation mathematical expression is:

[0050] X gbc =GBC(GBC(X)))

[0051]

[0052] S=σ(Adapterr channel (Conv(X gbc ))⊙Adapter channel (Conv(Reshape(Y ssm ))))

[0053] Y paf =(1-S)⊙X gbc +S⊙Reshape(Y ssm )

[0054] Y out =Reshape(Linear(GN(Y paf )))+X gbc

[0055] Where σ is the Sigmoid function and ⊙ represents element-wise multiplication.

[0056] The implementation process of the GBC (GroupedBottleneckConvolution) module is as follows: First, the GBC operation is performed on the original previous level input feature X. This module can significantly reduce the computational complexity with the help of the grouped depth-separable convolution structure. It adopts a multi-branch bottleneck design internally and contains four parallel sub-modules: two 3×3 depth convolution branches, which are mainly used to capture local spatial details and mine subtle spatial structure information in the feature map. The two 1×1 convolution branches are used to interact cross-channel information and realize feature fusion and transformation between channels. These sub-modules cooperate with the group normalization (GN) operation to alleviate the problem of internal covariate offset, and then introduce nonlinearity through the ReLU activation function to achieve efficient feature enhancement. Finally, the original input information is added to the processed feature information through the residual connection to obtain the output, which can not only retain the original information and effectively avoid the gradient disappearance problem, but also help the model learn richer feature differences. Among them, X gbc is the enhanced feature map after two GBC operations. The first GBC operation obtains X gbc1 The mathematical expression of forward propagation is:

[0057] Xgbc1 =GBC(X)

[0058] X gbc =GBC(X gbc1 )

[0059] Y=X gbc1 +X

[0060] The implementation process of the SAVSS_2D module is as follows: First, the enhanced feature map X processed by the GBC module is gbc It is expanded into a sequence form from the spatial dimension as the input of the SAVSS_2D module. The SAVSS_2D module uses the structured state space model (SSM) combined with a multi-directional scanning strategy to process the input features. The multi-directional scanning strategy includes horizontal scanning, vertical scanning, and dual diagonal scanning. The horizontal scan processes the feature sequence row by row to capture the long-range dependencies between pixels in the row direction of the feature map; the vertical scan processes the feature sequence column by column to mine the long-range dependencies in the column direction; the dual diagonal scan scans the feature map from two diagonal directions to obtain the correlation information between pixels in the diagonal direction. During the scanning process, the module dynamically generates the transfer matrix A based on the input features. d and projection weight B d (d represents different scanning directions.) These parameters are adaptively adjusted according to the input features, so that the state transition process can better adapt to the changes in features, thereby improving the modeling ability of complex spatial patterns. For each scanning direction d, there is a corresponding structured state space model SSM d Based on the input feature sequence, transfer matrix and projection weight, the calculation is performed through the state transition rule. Finally, the outputs of the four scanning paths are dynamically weighted and fused. The weight α is calculated for the output of each scanning direction. d , and satisfy the condition that the weight sum is 1. Through weighted summation, the final serialized feature Y is obtained ssm . Its forward propagation mathematical expression is:

[0061]

[0062] z t,d =A d z t-1,d +B d x t,d

[0063] Y d =SSM d (Norm(Reshape(X gbc )),A d ,B d )

[0064]

[0065] The implementation process of the PAF module is as follows: The PAF module uses the serialized feature Y output by the SAVSS_2D module ssm And the original spatial feature X after processing by the GBC module gbc As input, through two-way feature transformation, X gbc Convert to base feature, Y ssm Convert it into a guide feature and generate a key-value pair. Based on this, the spatial similarity graph S is calculated. This process can quantify the correlation of features in spatial positions. Then, the original feature X is weighted with S. gbc and SSM feature Y ssm Perform soft assignment fusion and finally obtain the fused feature Y paf . It can retain X gbc The underlying detailed information can be injected into Y ssm The global structured information of is complementary to the features of small target behaviors and targets in classroom images. Its forward propagation mathematical expression is:

[0066] Y paf =(1-S)⊙X gbc +S⊙Reshape(Y ssm )

[0067] Step 4: Use the training set and validation set to train the classroom behavior recognition detection model, and use the test set to evaluate the performance of the classroom behavior recognition detection model to obtain the optimal detection model.

[0068] Step 5: Combine monitoring with the optimal detection model to conduct real-time identification and detection of students' classroom behavior. The specific steps are as follows:

[0069] Step 5: Place high-definition surveillance cameras in the classroom environment to capture key information about students' movements, postures, and facial expressions from different angles.

[0070] Step 52: Transmit the video stream data collected in real time by the surveillance camera to the computer processing terminal;

[0071] Step 53: After receiving the video stream data, the computer terminal uses the optimal detection model obtained in step 4 to extract and analyze the features of each input frame. Once the optimal classroom behavior detection model identifies that a student has typical classroom behavior, it will provide real-time feedback in a visual manner on the computer display interface.

[0072] Step 54: Use differentiated labeling methods for different behaviors.

[0073] Example:

[0074] This embodiment provides a classroom behavior recognition method based on an improved YOLOv12 model. Figure 1 As shown, the method includes the following steps:

[0075] Step 1: Obtain a dataset of classroom behavior images.

[0076] The Student Classroom Behavior Dataset (SCB-Dataset) provides authentic and accurate representations of student behavior. It contains 184,000 labels and 42,000 images, covering three behaviors: hand raising, reading, and writing. A subset of these images was selected for use.

[0077] Step 2: Use a computer to preprocess the dataset. Divide the dataset into a validation set, a training set, and a test set, and finally process it into the YOLOv12 network model recognition format.

[0078] Use Python code to divide the hand dataset into training set, validation set, and test set in a set ratio of 8:1:1. The training set will be used to train the model, the validation set will be used for evaluation during training, and the test set will be used to evaluate the performance of the model. The training set, validation set, and test set will be processed into the YOLOv12 network model recognition format.

[0079] Step 3: Use the improved YOLOv12 model to train the classroom behavior recognition detection model dataset. The improved YOLOv12 adds a P2 detection head on the detection layer, innovatively designs and uses the A2C2f_FRFN module that combines the A2C2f with the feature refinement feedforward network FRFN module in the backbone network and the neck network, and innovatively designs and uses the C3k2_SAVSS module that combines the C3k2 with the structure perception visual state space module SAVSS in the backbone network and the neck network. Figure 2 shown.

[0080] A P2 detection head is added to the detection layer, so that the hand image is passed into the Backbone network to extract four different feature maps. The four sizes of 120×120, 64×64, 32×32, and 16×16 are used to detect targets of four different sizes: tiny, small, medium, and large, respectively.

[0081] The A2C2f_FRFN module, designed in the backbone and neck networks, combines A2C2f with a feature refinement feedforward network (FRFN) module. This module enhances the model's spatial feature extraction capabilities by introducing a striped attention mechanism. Through multi-stage linear transformations, convolution operations, and gating mechanisms, it can enhance attention to subtle changes in movement or posture and improve its ability to capture details in classroom behavior recognition. The input tensor is first fed into an initial convolutional layer built based on the parent class's initialization logic. This layer applies channel compression to the input feature map using a low-rank mapping. Using sliding convolutions across the feature map's spatial domain, it performs a multiplication-accumulation operation between the weights and the corresponding elements, transforming the input channels into hidden channels and extracting feature information with basic representational capabilities. The convolutional feature map is stored in a specific list structure. This structure then enters a loop consisting of the ABlock_FRFN module, passing the last element of the list as input to each module in sequence. The feature vector generated by the ABlock_FRFN module is fused with the original input feature vector using residual mapping, ultimately outputting the fused feature. Through the above-mentioned module architecture design and data processing flow, we can achieve efficient extraction, screening and fusion of complex input features, significantly enhance the discriminability and robustness of feature expression, provide more discriminative and stable feature representation for high-level semantic tasks such as classroom behavior recognition, effectively improve the model's ability to capture and process subtle feature changes in complex scenarios, and ensure the accuracy and reliability of recognition tasks from the feature level. Figure 3 shown.

[0082] The ABlock_FRFN module helps the model focus on the most relevant parts of the input data. During action recognition, certain moments or actions contain more information than others. The attention mechanism allows the model to adaptively assign more weight to key frames or key actions. Input features are first processed by the regional attention module inherited from the ABlock class. The regional attention mechanism assigns weights to the features, highlighting key information. This produces a feature vector, the "regional attention feature vector," which is then residually connected with the original input features to produce the "first residual connection feature vector." This operation aims to alleviate the vanishing gradient problem and enhance the model's ability to capture key features. The "first residual connection feature vector" is then passed to the multi-layer perceptron module in the parent ABlock class for feature transformation, producing the "multi-layer perceptron feature vector." This feature vector is then residually connected with the "first residual connection feature vector" to produce the "second residual connection feature vector." The "second residual connection feature vector" is then processed by the FRFN2D module, ultimately producing a feature vector that is residually connected with the "second residual connection feature vector" to produce the output. Through a series of steps such as initial convolution dimensionality reduction, regional attention mechanism, multi-layer perceptron transformation, and feature refinement of the FRFN2D module, the ABlock_FRFN module achieves multi-stage fine processing and deep fusion of input features, significantly enhancing the feature expression ability and providing more valuable feature representation for subsequent tasks. Figure 4 shown.

[0083] The FRFN2D module utilizes partial convolution and a gate mechanism to enhance features, achieving higher accuracy in fine-grained behavior detection and real-time processing capabilities. When processing input features, they are first split into two parts along the channel dimension. A 3×3 convolution is then performed on one part, a local convolution operation that extracts local feature details, enhancing spatial correlation and local representation. This convolution-processed feature is then concatenated with the other part along the channel dimension. This approach preserves the integrity of some original features while introducing localized, refined features, maintaining the overall structure and information diversity of the features. The concatenated features enter the linear1 module. This module consists of a 1×1 convolution and an activation function. The 1×1 convolution expands the number of channels and broadens the spatial range of feature representation, allowing the model to capture richer feature patterns and semantic information. The activation function further introduces nonlinearity, improving the model's ability to fit complex feature relationships. The processed features are then split equally into two parts. A depth-wise separable convolution is performed on one part of it. This operation can extract the local spatial information of the features and the correlation between channels while reducing the amount of calculation. Subsequently, the convolved part of the features is multiplied element-by-element with the other part. This gating mechanism can filter out feature information that is more critical to subsequent tasks, suppress redundant or irrelevant features, and enhance the discriminability and effectiveness of features. Finally, the linear2 module (including 1×1 convolution) is used to restore the number of channels to the input dimension so that the feature dimension can adapt to the needs of subsequent tasks. Without losing key information, the features are integrated and compressed to obtain the final output result. This output result provides a refined and highly expressive feature representation for subsequent processing of the model. Figure 8 shown.

[0084] The C3k2_SAVSS module that combines C3k2 with the structure-aware visual state space module SAVSS is designed and used in the backbone network and the neck network. This module can capture long-range dependencies and global features through directional scanning (horizontal, vertical, diagonal) and structured state space models. The combination of the two can fully integrate local details and global context. For example, in the classroom, it can not only accurately extract local action details such as students writing and raising their hands, but also associate previous and subsequent behaviors to enhance the richness of feature expression. First, the input feature map will pass through the initial convolution layer to adjust the number of channels and generate hidden features. Then, the generated hidden features are input into a sequence of n SAVSS modules. Each SAVSS module performs a specific transformation on the features and outputs intermediate features. Finally, the final output features are added to the original input features to obtain the fusion result. As Figure 5 shown.

[0085] The SAVSS module leverages the synergy of GBC (local enhancement), SAVSS_2D (multi-directional global modeling), and PAF (dynamic fusion) to significantly improve the mean Intersection Over Union (MIoU) metric for dense prediction tasks while reducing computational overhead compared to traditional convolutional modules. First, two GBC (Grouped Bottleneck Convolution) operations are performed on the input feature map. This module significantly reduces computational complexity through a grouped depthwise separable convolutional architecture, while enhancing local feature representation through a multi-branch bottleneck design. It comprises four parallel submodules: two 3×3 depthwise convolutional branches for capturing local spatial details, and two 1×1 convolutional branches for cross-channel information exchange. Group normalization (GN) and ReLU activation functions are used for efficient feature enhancement. Finally, residual connections are used to preserve the original information and prevent gradient vanishing, resulting in an enhanced feature map. The enhanced feature map is then expanded from the spatial dimension into a sequence and fed into the SAVSS_2D module, where long-range dependencies are modeled using a structured state space model (SSM) combined with a multi-directional scanning strategy. Horizontal scanning captures pixel associations within a row, vertical scanning establishes dependencies within a column, and dual diagonal scanning extracts global diagonal features. The outputs of the four scanning paths are dynamically weighted and fused to ensure that the model perceives both local structure and global context. During the scanning process, the dynamically generated transfer matrix and content-aware projection weights are adaptively adjusted according to the input features, making the state transfer process input-sensitive and significantly improving the ability to model complex spatial patterns. The serialized features output by SSM are deeply fused with the original spatial features through the PAF (Position-Aware Fusion) module. This module first generates key-value pairs of base features and guide features through a two-way feature transformation, calculates a spatial similarity graph to quantify position correlation, and then softly assigns and fuses the original features and SSM features using the spatial similarity graph as weights, which not only retains the underlying details but also injects global structural information, and is especially suitable for feature complementarity under occlusion or small target scenes. In the final stage, the training process is stabilized through linear transformation and residual connection: group normalization (GN) alleviates internal covariate shift, the linear layer improves the adaptability of feature dimensions, and the jump connection ensures lossless transmission of information before outputting the final result. Figure 6 shown.

[0086] The GBC module enhances the ability to express local details through a multi-branch lightweight bottleneck convolution structure. This module can significantly reduce computational complexity with the help of a grouped depth-separable convolution structure. It adopts a multi-branch bottleneck design internally and contains four parallel sub-modules: two 3×3 depth convolution branches, which are mainly used to capture local spatial details and mine subtle spatial structure information in feature maps. The two 1×1 convolution branches are used to interact with cross-channel information and realize feature fusion and transformation between channels. These sub-modules cooperate with the group normalization (GN) operation to alleviate the problem of internal covariate shift, and then introduce nonlinearity through the ReLU activation function to achieve efficient feature enhancement. Finally, the original input information is added to the processed feature information through the residual connection to obtain the output. This not only retains the original information and effectively avoids the gradient disappearance problem, but also helps the model learn richer feature differences. Such as Figure 7 shown.

[0087] The SAVSS_2D module uses sequential modeling in different directions to enhance the model's directional awareness and improve the efficiency of long-distance information exchange. First, the enhanced feature map processed by the GBC module is spatially expanded into a sequence format, which serves as the input to the SAVSS_2D module. The SAVSS_2D module processes the input features using a structured state space model (SSM) combined with a multi-directional scanning strategy. This multi-directional scanning strategy includes horizontal scanning, vertical scanning, and dual diagonal scanning. Horizontal scanning processes the feature sequence row by row, aiming to capture long-range dependencies between pixels in the feature map along the rows. Vertical scanning processes the feature sequence column by column, mining long-range dependencies along the columns. Dual diagonal scanning scans the feature map in two diagonal directions to capture correlations between diagonal pixels. During the scanning process, the module dynamically generates a transfer matrix and projection weights based on the input features. These parameters are adaptively adjusted based on the input features, allowing the state transition process to better adapt to feature changes, thereby improving the ability to model complex spatial patterns. A corresponding structured state space model exists for each scanning direction. Based on the input feature sequence, the transition matrix, and the projection weights, calculations are performed using state transition rules. Finally, the outputs of the four scanning paths are dynamically weighted and fused. Weights are calculated for the outputs in each scanning direction, and the sum of the weights must be 1. The final sequenced features are obtained through weighted summation.

[0088] The PAF module improves the feature fusion performance through dynamic position-aware fusion. Taking the serialized features output by the SAVSS_2D module and the original spatial features processed by the GBC module as input, the original spatial features are converted into base features, the serialized features are converted into guide features, and key-value pairs are generated through two-way feature transformation. The spatial similarity graph is calculated based on these key-value pairs. This process can quantify the correlation of features in spatial positions. The original spatial features and SSM features are then soft-assigned and fused using the spatial similarity graph as the weight, and the fused features are finally obtained. This fusion method can not only retain the underlying detail information of the original spatial features, but also inject the global structural information of the serialized features. It is suitable for feature complementarity of small target behaviors in classroom images. Figure 9 shown.

[0089] Step 4: Use the training set and validation set to train the classroom behavior recognition detection model, use the test set to evaluate the performance of the classroom behavior recognition detection model, and obtain the optimal classroom behavior detection model.

[0090] The training and validation sets were fed into a classroom behavior recognition and detection model based on the improved YOLOv12 model. The number of training runs was set. As the number of training runs increased, the model's loss function curve gradually converged. When the loss function curve converged and stabilized, the model was trained to its optimal state, and its optimal model weight file was saved. The images to be tested from the test set were fed into the trained optimal classroom behavior recognition and detection model, which then output the optimal classroom behavior recognition detection images. The performance of the classroom behavior recognition and detection model based on the improved YOLOv12 model was evaluated based on the detection images to determine the optimal detection model.

[0091] Step 5: Combine monitoring and computer algorithms to conduct real-time identification and detection of students' classroom behavior.

[0092] High-definition surveillance cameras are strategically placed throughout the classroom environment to ensure comprehensive, comprehensive coverage of the student learning area, capturing key information such as student movements, posture, and facial expressions from various angles. Camera parameters such as resolution and frame rate are pre-optimized based on classroom size, lighting conditions, and the level of detail required to detect behaviors. This ensures image clarity and consistency, providing high-quality raw data for subsequent accurate recognition. Real-time video streams captured by surveillance cameras are transmitted stably and quickly to a preconfigured computer processing terminal via wired or wireless transmission. Upon receiving the video streams, the computer terminal immediately activates a trained, optimal classroom behavior detection model based on a modified YOLOv12, rapidly extracting and analyzing features from each input frame. Once the detection model identifies typical classroom behaviors, it provides real-time visual feedback on the computer display.

Claims

1. A classroom behavior recognition method based on an improved YOLOv12 model, characterized by The method comprises the following steps: Step 1: Obtain a dataset of classroom behavior images; Step 2: Divide the dataset into validation set, training set, and test set, and process them into the network model recognition format of YOLOv12; Step 3: Use the classroom behavior recognition detection model training dataset of the improved YOLOv12 model. The improved YOLOv12 model adds a P2 detection head to the detection layer, designs and uses the A2C2f_FRFN module that combines the A2C2f with the feature refinement feedforward network FRFN module in the backbone network and the neck network, and designs and uses the C3k2_SAVSS module that combines the C3k2 with the structure perception visual state space module SAVSS in the backbone network and the neck network; Step 4: Use the training set and validation set to train the classroom behavior recognition detection model. After obtaining the optimal detection model, use the test set to evaluate the performance of the classroom behavior recognition detection model. Step 5: Combine monitoring with the optimal detection model to conduct real-time identification and detection of students' classroom behavior.

2. The classroom behavior recognition method based on the improved YOLOv12 model according to claim 1 is characterized in that The implementation process of the A2C2f_FRFN module is as follows: the input feature vector X is first imported into the initial convolution layer built based on the parent class initialization logic. This layer implements channel compression on the input feature map through low-rank mapping, and uses the sliding convolution of the convolution kernel in the feature map space domain to perform multiplication and accumulation operations on the weights and the elements at the corresponding positions, completing the dimensionality transformation of the input channel to the hidden channel and extracting feature information with basic representation capabilities; then enters the loop processing flow composed of the ABlock_FRFN module, and passes the end element of the list as input to each module sequence in turn; the feature vector X2 generated by the ABlock_FRFN module is fused with the original input feature vector X by residual mapping, and finally outputs the fused feature Y. Its forward propagation mathematical expression is: Y=X+X2.

3. The classroom behavior recognition method based on the improved YOLOv12 model according to claim 2 is characterized in that The implementation process of the ABlock_FRFN module is as follows: the input feature vector X of the previous level is first processed by the regional attention module inherited from the ABlock class; the regional attention mechanism will assign weights to the features, highlight the key information, and obtain the feature vector Y mlp1 And the first residual connection is performed with the original upper level input feature vector X to obtain X1; X1 is passed to the multi-layer perceptron module in the parent class ABlock for feature transformation to obtain the feature vector Y mlp1 , and performs a second residual connection with X1 to obtain X2; X2 enters the FRFN2D module for processing, and finally obtains the feature vector and performs a residual connection with X2 to obtain the output Output. Its forward propagation mathematical expression is:

4. The classroom behavior recognition method based on the improved YOLOv12 model according to claim 3 is characterized in that The FRFN2D module is implemented as follows: the input feature vector X of the previous level is first divided into two parts X1 and X2 according to the channel dimension. X1 is convolved with 3×3 and then concatenated with X2 along the channel dimension. The concatenated features are expanded by linear1 to increase the number of channels. Subsequently, the features are evenly divided into two parts X3 and X4. X3 is convolved with depthwise separable convolution and then multiplied element-wise with X4. Finally, linear2 is used to restore the number of channels to the input dimension dim. Without losing key information, the features are integrated and compressed to obtain the final output Output. The forward propagation mathematical expression is:

5. The classroom behavior recognition method based on the improved YOLOv12 model according to claim 1 is characterized in that The implementation process of the C3k2_SAVSS module is as follows: the input feature X of the previous level is adjusted through the initial convolution layer to generate the hidden feature X h =Conv(X), where the number of hidden channels is c, and X h Input is sent to a sequence of n SAVSS modules, each SAVSS performs a specific transformation on the feature and outputs the intermediate feature F i (X h ); finally the final output F n (X h ) is added to the original previous level input feature X to obtain the fusion result Y, and its forward propagation mathematical expression is:

6. The classroom behavior recognition method based on the improved YOLOv12 model according to claim 5 is characterized in that The implementation process of the SAVSS module is as follows: First, two GBC operations are performed on the previous level input feature X. The module contains four parallel submodules: two 3×3 depth convolution branches are used to capture local spatial details, two 1×1 convolution branches are used for cross-channel information interaction, and group normalization and ReLU activation functions are used to achieve efficient feature enhancement. Finally, residual connections are used to retain the original information to avoid gradient disappearance; the enhanced feature map X is converted into gbc Expanded from the spatial dimension into a sequence form and fed into the SAVSS_2D module, the long-range dependency modeling is achieved by combining the structured state space model with a multi-directional scanning strategy; the serialized feature Y output by SSM ssm With X gbc Through deep fusion of the PAF module, the module first generates key-value pairs of base features and guide features through two-way feature transformation, calculates the spatial similarity graph S to quantify the position correlation, and then soft-assigns and fuses the original features and SSM features with S as the weight. In the final stage, the training process is stabilized by linear transformation and residual connection: group normalization alleviates internal covariate shift, linear layer improves feature dimension adaptability, and jump connection ensures lossless information transmission and outputs Y out , its forward propagation mathematical expression is: X gbc =GBC(GBC(X))) S=σ(Adapterr channel (Conv(X gbc ))⊙Adapter channel (Conv(Reshape(Y ssm ))))Y paf =(1-S)⊙X gbc +S⊙Reshape(Y ssm ) Y out =Reshape(Linear(GN(Y paf )))+X gbc Where σ is the Sigmoid function and ⊙ represents element-wise multiplication.

7. The classroom behavior recognition method based on the improved YOLOv12 model according to claim 6 is characterized in that The GBC module is implemented as follows: GBC is performed on the original previous-level input feature X. The module adopts a multi-branch bottleneck design and contains four parallel submodules: two 3×3 depthwise convolution branches for capturing local spatial details and mining subtle spatial structure information in the feature map, and two 1×1 convolution branches for cross-channel information exchange to achieve feature fusion and transformation between channels. These submodules cooperate with group normalization operations to alleviate the problem of internal covariate shift, and then introduce nonlinearity through the ReLU activation function to achieve efficient feature enhancement; finally, the original input information is added to the processed feature information through the residual connection to obtain the output, where X gbc is the enhanced feature map after two GBC operations. The first GBC operation obtains X gbc1 , the forward propagation mathematical expression is:

8. The classroom behavior recognition method based on the improved YOLOv12 model according to claim 6 is characterized in that The implementation process of the SAVSS_2D module is as follows: First, the enhanced feature map X processed by the GBC module is gbc Expanded from the spatial dimension into a sequence form as the input of the SAVSS_2D module; the SAVSS_2D module uses a structured state space model combined with a multi-directional scanning strategy to process the input features. The multi-directional scanning strategy includes horizontal scanning, vertical scanning and dual diagonal scanning. The horizontal scanning processes the feature sequence row by row to capture the long-range dependency between pixels in the row direction of the feature map; the vertical scanning processes the feature sequence column by column to mine the long-range dependency in the column direction; the dual diagonal scanning scans the feature map from two diagonal directions to obtain the correlation information between pixels in the diagonal direction; during the scanning process, the module dynamically generates the transfer matrix A based on the input features d and projection weight B d , for each scanning direction d, there is a corresponding structured state space model SSM d , based on the input feature sequence, transfer matrix and projection weight, the calculation is performed through the state transition rule; finally, the outputs of the four scanning paths are dynamically weighted and fused to calculate the weight α for the output of each scanning direction d , and satisfy the condition that the weight sum is 1, and the final serialized feature Y is obtained by weighted summation ssm , its forward propagation mathematical expression is:

9. The classroom behavior recognition method based on the improved YOLOv12 model according to claim 6 is characterized in that The PAF module implementation process is as follows: The PAF module uses the serialized feature Y output by the SAVSS_2D module ssm And the original spatial feature X after processing by the GBC module gbc As input, through two-way feature transformation, X gbc Convert to base feature, Y ssm Convert it into a guide feature and generate a key-value pair, based on which the spatial similarity graph S is calculated, and then the original feature X is weighted with S gbc and SSM feature Y ssm Perform soft assignment fusion and finally obtain the fused feature Y paf , its forward propagation mathematical expression is: AND paf =(1-S)⊙X gbc +S⊙Reshape(Y ssm )。 10. The classroom behavior recognition method based on the improved YOLOv12 model according to claim 1 is characterized in that The specific steps of step five are as follows: Step 5: Place high-definition surveillance cameras in the classroom environment to capture key information about students' movements, postures, and facial expressions from different angles. Step 52: Transmit the video stream data collected in real time by the surveillance camera to the computer processing terminal; Step 5.3: After the computer terminal receives the video stream data, it uses the optimal detection model obtained in step 4 to extract and analyze the features of each input frame image. Once the optimal classroom behavior detection model identifies that the student has typical classroom behavior, it will provide real-time feedback in a visual manner on the computer display interface.

Citation Information

Patent Citations

  • Classroom learning behavior identification method based on improved YOLOv8

    CN117671781A

  • Dangerous driving visual detection method based on YOLO-SGC

    CN118552939A

  • Student classroom behavior detection method based on Transform and task dynamic alignment

    CN119418243A

  • Student classroom behavior detection method based on deep learning

    CN119763179A

  • Behavior identification method and apparatus, and device and storage medium

    WO2023035891A1

Cited By

  • Image fine structure intelligent detection algorithm based on double-branch encoder

    CN121074418A