Computer vision-based intelligent recognition method and system for student behavior in classroom scene
By using a dual-stream parallel network structure and channel attention mechanism, combined with a frame buffer pool mechanism, the problem of feature fusion and real-time processing for student behavior recognition in classroom scenarios was solved, achieving efficient and accurate student behavior recognition and improving the level of intelligent teaching management.
Patent Information
- Application Number
- CN202510023908.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Existing student behavior recognition methods cannot effectively capture students' positional and action features in classroom scenarios, lack the utilization of temporal information from video sequences, and have high computational overhead, making it difficult to meet real-time processing requirements and affecting recognition accuracy and efficiency.
A dual-stream parallel network structure is adopted to extract position features and action features separately, which are then fused through a channel attention module. A frame buffer pool mechanism is used to control computational overhead. A multi-branch action feature extraction module and temporal position coding are designed to achieve efficient feature fusion and real-time processing.
It significantly improves the accuracy and real-time performance of student behavior recognition, enabling comprehensive monitoring of students' learning status, enhancing the intelligence level of teaching management, and meeting real-time processing needs.
Smart Images

Figure CN120071061B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, in particular to a classroom scene student behavior intelligent recognition method and system based on computer vision. BACKGROUND
[0002] With the rapid development of education informatization, using artificial intelligence technology to assist classroom teaching management has become a hot research topic. In the traditional classroom teaching process, it is difficult for teachers to pay attention to the learning state of each student in real time and comprehensively, which brings challenges to the improvement of teaching quality. Student behavior recognition and detection in the classroom scene have their particularity and complexity: first, the camera in the classroom environment is usually installed high in the front or back of the classroom, which limits the shooting angle and causes mutual occlusion between students; second, the number of students in the classroom is large, and the position distribution is dense, which increases the difficulty of target detection and behavior recognition; third, the behavior patterns of students in the classroom are complex and diverse, and there are subtle state transitions, which bring challenges to accurate behavior recognition. At present, the research on classroom student behavior recognition mainly exists the following problems:
[0003] (1) The existing student behavior recognition method often uses a single feature extraction method, which cannot effectively capture the position features and motion features of students at the same time, resulting in unsatisfactory recognition effect;
[0004] (2) The traditional method lacks effective use of time sequence information in video sequences, and it is difficult to accurately recognize the dynamic behavior features of students;
[0005] (3) The existing technology has deficiencies in feature fusion, and cannot effectively integrate position information and motion information, affecting the accuracy of recognition;
[0006] (4) The existing recognition algorithm has large computational overhead when processing multiple frames of video data, and it is difficult to meet the real-time processing demand.
[0007] Therefore, it is urgent to develop a computer vision method that can accurately and real-time recognize student behavior in the classroom scene, so as to improve the intelligent level of classroom teaching management. SUMMARY
[0008] The purpose of the present application is to overcome the deficiencies of the prior art and provide a classroom scene student behavior intelligent recognition method based on computer vision.
[0009] The purpose of the present application is achieved by the following technical solutions:
[0010] In a first aspect, the present application discloses a classroom scene student behavior intelligent recognition method based on computer vision, comprising the following steps:
[0011] S1, input video data of student behavior in a classroom scene;
[0012] S2, extract position features of each student in the current frame of the input video;
[0013] S3, set a buffer pool to store the current frame of the input video, and then extract the action features of each student from the time sequence frames stored in the buffer pool;
[0014] S4, fuse the position features and the action features through a channel attention module;
[0015] S5, process the fused features to output the final position and behavior information of the student.
[0016] Based on the first aspect, step S2 specifically includes: using a double-branch parallel structure to extract position features, wherein the first branch includes a 3x3 two-dimensional convolution layer, a standardization layer, a SiLU nonlinear layer, and a 1x1 convolution layer for feature compression, and finally two 3x3 convolution layers are used for feature extraction; the second branch includes a single 3x3 two-dimensional convolution layer for feature extraction and compression, and finally the features of the two branches are spliced to obtain complete position features.
[0017] Based on the first aspect, step S3 specifically includes: designing a five-branch parallel structure to extract action features, wherein the first to fourth branches use different three-dimensional convolution structures to extract multi-level action features in the video sequence; the fifth branch introduces a learnable time sequence position encoding, and after the features of the first to fourth branches are converged in the splicing layer, the time sequence position encoding of the fifth branch is added to obtain complete action features.
[0018] Based on the first aspect, step S3 further includes: the maximum number of current frames of the input video stored in the buffer pool is 20 frames.
[0019] Based on the first aspect, step S4 specifically includes: fusing the position features and the action features through a channel attention module, first performing average pooling compression on the position features and the action features, then performing feature abstraction through two fully connected layers and a ReLU nonlinear layer, and finally obtaining channel weights through a Sigmoid nonlinear layer to realize adaptive weighting of different feature channels.
[0020] Based on the first aspect, step S5 specifically includes: processing the fused features through two fully connected layers and a SiLU nonlinear layer to output the final student position and behavior information.
[0021] In a second aspect, the present application discloses a classroom scene student behavior intelligent recognition system based on computer vision, which is used for the classroom scene student behavior intelligent recognition method based on computer vision.
[0022] The position feature extraction module is used for extracting the position of each student in the current video frame.
[0023] The action feature extraction module is used for comprehensively judging the student action in a short time, so as to realize the fine-grained student behavior recognition.
[0024] The channel attention module is used for organically fusing the position feature and the action feature, and further improving the extraction ability of the position feature and the action feature of the target student.
[0025] The present application has the following beneficial effects:
[0026] 1) The present application provides a classroom scene student behavior intelligent recognition method based on computer vision, which extracts the position feature and the action feature through a double-flow parallel network structure, can effectively capture the static position information and the dynamic behavior feature of the student, and significantly improves the accuracy of the student behavior recognition in the classroom scene. In addition, the channel attention mechanism introduced in the present application realizes the effective fusion of the position feature and the action feature, solves the deficiency of the traditional method in the feature fusion, and further enhances the feature expression ability of the model.
[0027] 2) The present application can be used in an intelligent classroom management system, through the double-flow parallel network structure and the real-time processing optimization mechanism provided, the behavior state of the student in the classroom can be monitored and analyzed in real time, the teacher can more comprehensively understand the learning state of the student, and the teaching effect is improved. At the same time, the frame buffer pool mechanism adopted in the present application effectively controls the calculation cost, so that the system can meet the real-time processing demand, provides reliable technical support for intelligent teaching management, and promotes the development of education informatization. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 FIG. 1 is a step schematic diagram of a classroom scene student behavior intelligent recognition method based on computer vision according to an embodiment of the present application;
[0029] Figure 2 FIG. 3 is a specific structure schematic diagram of a classroom scene student behavior intelligent recognition system based on computer vision according to an embodiment of the present application;
[0030] Figure 3 FIG. 5 is a position feature extraction module schematic diagram of a classroom scene student behavior intelligent recognition system based on computer vision according to an embodiment of the present application;
[0031] Figure 4A motion feature extraction module schematic diagram of a computer vision-based classroom scene student behavior intelligent recognition system according to an embodiment of the present application;
[0032] Figure 5 A channel attention module schematic diagram of a computer vision-based classroom scene student behavior intelligent recognition system according to an embodiment of the present application;
[0033] Figure 6 A detection flow schematic diagram of a computer vision-based classroom scene student behavior intelligent recognition method according to an embodiment of the present application. DETAILED DESCRIPTION
[0034] The technical solutions of the present application will be described in detail below with reference to the embodiments, obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0035] The present application discloses a computer vision-based classroom scene student behavior intelligent recognition method, a double-flow parallel network structure is designed, the position features and motion features of students are extracted respectively, and the accuracy of recognition is improved; a multi-branch-based motion feature extraction module is proposed, the recognition ability of dynamic behavior is enhanced through multi-level feature extraction and time sequence position coding; the channel attention mechanism is introduced, the effective fusion of position features and motion features is realized, and the feature expression ability is improved; the frame buffer pool mechanism is used to control the number of processing frames, and through reasonable network structure design, the real-time processing requirement is realized. It aims to provide an efficient and accurate technical solution for student behavior intelligent recognition in a classroom scene, and provide strong technical support for intelligent teaching management. In view of the problems existing in student behavior recognition in the classroom scene, specifically, the multi-dimensional feature fusion idea is adopted in the present application, the position features of each student subject in the current video frame are captured through the position feature extraction branch, then the behavior motion features of each student subject in a short time are captured through the motion feature extraction branch, finally the features of the two aspects are organically fused, the complete position and behavior features of each student subject are obtained, so that more accurate student behavior recognition and target detection are realized. The specific steps of the method are shown in the schematic diagram as Figure 1 shown, including the following steps:
[0036] S1, inputting video data of student behavior in a classroom scene;
[0037] S2, extracting the position features of each student in the current frame of the input video;
[0038] S3, a buffer pool is set to store the current frame of the input video, and then the motion features of each student are extracted from the time sequence frames stored in the buffer pool;
[0039] S4, the position features and the motion features are fused through a channel attention module;
[0040] S5, the fused features are processed to output the final position and behavior information of the student.
[0041] Specifically, step S2 specifically includes: adopting a double-branch parallel structure to extract position features, wherein the first branch includes a 3x3 two-dimensional convolution layer, a standardization layer, a SiLU nonlinear layer, and a 1x1 convolution layer for feature compression, and finally two 3x3 convolution layers are used for feature extraction; the second branch includes a single 3x3 two-dimensional convolution layer for feature extraction and compression, and finally the features of the two branches are spliced to obtain complete position features.
[0042] Specifically, step S3 specifically includes: designing a five-branch parallel structure to extract motion features, wherein the first to fourth branches adopt different three-dimensional convolution structures to extract multi-level motion features in the video sequence; the fifth branch introduces a learnable time position encoding, and after the features of the first to fourth branches are converged in the splicing layer, the time position encoding of the fifth branch is added to obtain complete motion features.
[0043] Specifically, step S3 further includes: the maximum number of stored current frames of the input video in the buffer pool is 20 frames; the buffer pool is used to store the current frames of the input video, one frame at a time, when the maximum number of stored frames is reached, the current frame is added to the buffer pool as the newest frame, and the oldest frame in the buffer pool is discarded. The purpose of setting the buffer pool is to cache continuous actions for a period of time, to identify actions by integrating time sequence features, and to limit the calculation amount within a certain range. The time dimension of the motion features is compressed through the max-pooling layer, so that it can be spliced and fused with the position features.
[0044] Specifically, step S4 specifically includes: fusing the position features and the motion features through a channel attention module, first performing average pooling compression on the position features and the motion features, then performing feature abstraction through two fully connected layers and a ReLU nonlinear layer, and finally obtaining channel weights through a Sigmoid nonlinear layer to realize adaptive weighting of different feature channels.
[0045] Specifically, step S5 specifically includes: processing the fused features through two fully connected layers and a SiLU nonlinear layer to output the final position and behavior information of the student.
[0046] The application further discloses a classroom scene student behavior intelligent recognition system based on computer vision. Figure 2 As shown in a structural schematic diagram,
[0047] The position feature extraction module is used for extracting the position of each student in the current video frame.
[0048] The action feature extraction module is used for comprehensively judging the student action in a short time, so as to realize fine-grained student behavior recognition.
[0049] The channel attention module is used for organically fusing the position feature and the action feature, and further improving the extraction ability of the position feature and the action feature of the target student.
[0050] Specifically, the structure of the position feature extraction module is as shown in the figure. Figure 3 As can be seen, for a given input feature X, the feature is sent into two parallel branches for processing. In the left branch (first branch), the feature first undergoes a 3x3 two-dimensional convolution layer, a standardization layer and a nonlinear layer SiLU for preliminary feature extraction. Then in the following 1x1 two-dimensional convolution layer, feature compression is performed to halve the channel of the feature map, thereby reducing the calculation amount. Then two 3x3 two-dimensional convolution layers are used for further feature extraction, completing the feature extraction process of the left branch (first branch). In the right branch (second branch), a 3x3 two-dimensional convolution layer is used for feature extraction and feature compression to form a feature extraction mode different from the left branch (first branch), so as to obtain more rich features. At the same time, the existence of the right branch (second branch) is also conducive to the gradient transmission, so that the model is easier to converge. Finally, the features of the left branch (first branch) and the right branch (second branch) are spliced together to obtain a feature map for representing the target position information.
[0051] Specifically, the structure of the action feature extraction module is as shown in the figure. Figure 4 As can be seen, for a given input feature X, the feature is extracted in five parallel branches. The four different three-dimensional convolution structures on the left are responsible for extracting different levels of student action features contained between the sequential frames in the video stream. The multi-level features extracted by them will be converged in the splicing layer to form a mixed representation of multi-level features. In addition, in order to further improve the model's ability to understand the time sequence of multi-time video frames, we introduce a learnable time sequence position encoding in the rightmost branch. And finally, it is additively fused with the multi-level mixed representation obtained by the left four branches, which can enhance the model's ability to capture key action frames in multi-time video frames, thereby enhancing the model's action recognition accuracy.
[0052] Specifically, the structural diagram of the channel attention module is as shown in the figure. Figure 5 The module will first compress the input feature map through an average pooling layer, so that it is converted into a one-dimensional feature corresponding to the channel of the feature map. Then the one-dimensional feature will be abstracted in the combination of two fully connected layers and a ReLU nonlinear layer. Finally, after passing through a Sigmoid nonlinear layer, the high-level one-dimensional feature is converted into a weighted representation of different feature channels.
[0053] Specifically, Figure 6 The complete detection flowchart of the computer vision-based intelligent recognition method for student behavior in a classroom scene of the present application is shown in the figure. For the input video stream, the video frames in the video stream are sent into the position feature extraction module and the action feature extraction module for feature extraction, respectively. It is worth noting that in order to reduce the computational complexity of the model, a frame buffer pool is set before the action feature extraction module to control the maximum number of video frames sent into the action feature extraction module. In other words, the time span in short-time action recognition is controlled by using the frame buffer pool. In addition, since the action feature extraction module and the position feature extraction module have an additional time dimension, the features output by the action feature extraction module need to be compressed in the time dimension by a max-pooling layer, so that they can be further spliced and fused with the features output by the position feature extraction module. After obtaining the fused features formed by the two modules, the features are sent into the channel attention module for attention calculation, and the attention result is multiplied with the features to enhance the activation of key features. Finally, the activated features are converted into the final student position and behavior information through the simple combination of two fully connected layers and a nonlinear layer SiLU.
[0054] Exemplarily, the performance of the model can be changed by setting hyperparameters such as learning rate, training batch size, training times, frame buffer pool size, etc. In the present application, the initial learning rate is set to 0.01; the training batch size is set to 64; the training times are set to 300; the small batch stochastic gradient descent algorithm is used to optimize the model weight; the learning rate is reduced by 0.5 times at the 200th and 250th training; and the frame buffer pool size is set to 20 frames.
[0055] The above is only the preferred embodiment of the present application, and it should be understood that the present application is not limited to the form disclosed herein, and should not be considered as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein by the above-mentioned teaching or related art or knowledge. Any modification and change made by those skilled in the art without departing from the spirit and scope of the present application shall be within the protection scope of the claims of the present application.
Claims
1. A computer vision-based method for intelligent recognition of student behavior in a classroom scene, characterized in that, The method comprises the following steps: S1, inputting video data of student behavior in a classroom scene; S2, extracting position features of each student in the current frame of the input video; S3, setting a buffer pool to store the current frame of the input video, and then extracting motion features of each student from the time sequence frames stored in the buffer pool; S4, fusing the position features and the motion features through a channel attention module; S5, processing the fused features to output the final position and behavior information of the student; Step S2 specifically comprises: extracting position features by using a double-branch parallel structure, wherein the first branch comprises a 3*3 two-dimensional convolution layer, a standardization layer, a SiLU nonlinear layer, and a 1*1 convolution layer for feature compression, and finally two 3*3 convolution layers are used for feature extraction; the second branch comprises a single 3*3 two-dimensional convolution layer for feature extraction and compression, and finally the features of the two branches are spliced to obtain complete position features; Step S3 specifically comprises: designing a five-branch parallel structure to extract motion features, wherein the first to fourth branches use different three-dimensional convolution structures to extract multi-level motion features in the video sequence; the fifth branch introduces a learnable time sequence position encoding, and after the features of the first to fourth branches are converged in the splicing layer, the time sequence position encoding of the fifth branch is added to obtain complete motion features; Step S4 specifically comprises: fusing the position features and the motion features through the channel attention module, first performing average pooling compression on the position features and the motion features, then performing feature abstraction through two fully connected layers and a ReLU nonlinear layer, and finally obtaining channel weights through a Sigmoid nonlinear layer to realize adaptive weighting of different feature channels. 2.The computer vision-based intelligent classroom scene student behavior recognition method of claim 1, wherein, Step S3 further comprises: the maximum number of current frames of the input video stored in the buffer pool is 20 frames. 3.The computer vision-based method for intelligent recognition of student behavior in a classroom scene according to claim 2, characterized in that, Step S5 specifically comprises: processing the fused features through two fully connected layers and a SiLU nonlinear layer to output the final position and behavior information of the student.
4. A computer vision-based intelligent classroom student behavior recognition system for the computer vision-based intelligent classroom student behavior recognition method of any one of claims 1-3, wherein, Comprise: The position feature extraction module is used to extract the position of each student in the current video frame; The motion feature extraction module is used to comprehensively judge the student motion in a short time, so as to realize fine-grained student behavior recognition; The channel attention module is used to organically fuse the position features and the motion features, and further improve the extraction ability of the position features and the motion features of the target student.
Citation Information
Patent Citations
Classroom student behavior identification method based on computer vision
CN114170672A
Video action recognition method and device, electronic equipment and readable storage medium
CN116895038A