A Two-Stage Detection Method, Device and Medium for Classroom Teachers' Teaching Behaviors

Through the two-stage detection method, the YOLOv network and frequency adaptive expansion convolution FADConv are used to identify teachers' locations and teaching behaviors, which solves the problem of difficult to distinguish teachers from students in the existing technology, and achieves high-precision and robust teacher teaching behavior detection.

CN119851355BActive Publication Date: 2025-06-20XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510339423.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-20
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The existing technology is difficult to distinguish between teachers and students in classroom teacher teaching behavior testing, resulting in the misidentification of student behavior as teacher behavior, affecting the reliability of the detection.

Method used

The two-stage detection method is adopted. First, the teacher position is identified through the teacher position detection model based on the YOLOv network, and the frequency adaptive expansion convolution FADConv is used to improve the feature extraction ability; secondly, the teacher's teaching behavior is detected using the pre-trained teaching behavior recognition model.

Benefits of technology

Accurate detection of teachers' teaching behaviors is achieved, which reduces the interference of student behavior on detection and improves the accuracy and robustness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851355B_ABST
    Figure CN119851355B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and medium for detecting teachers' teaching behaviors in a two-stage classroom. The steps of the method include: receiving in real time the classroom video image data collected by a camera; inputting the received classroom video image data into a teacher position detection model to detect the position of the teacher as the detection result of the first stage; in the teacher position detection model, some convolutional operations adopt frequency-adaptive dilated convolution, and the frequency-adaptive dilated convolution performs convolution operations on the weighted features using an adaptive dilation rate and an adaptive convolution kernel through frequency selection to obtain an output feature layer; determining the teacher position area in the video image data according to the detection result of the first stage, and inputting the teacher position area and the classroom video image data into a teaching behavior recognition model to perform teaching behavior detection to obtain the behavior categories of the teacher. The present invention has the advantages of simple implementation method, low cost, high detection accuracy and strong robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent education, and particularly to a two-stage classroom teacher teaching behavior detection method, device and medium. Background Art

[0002] In traditional classroom teaching, the evaluation of teachers' classroom teaching behaviors is usually carried out by means of empirical descriptions such as classroom observation and behavior comparison, or by watching classroom videos. This can lead to problems such as heavy workload, poor timeliness, high labor costs, strong subjectivity, and the inability to achieve normalized evaluation. Using intelligent technology to detect teachers' classroom teaching behaviors and then conduct process intelligent evaluation of classroom teaching behaviors can help teachers improve their teaching behaviors and is of great significance for improving teaching quality.

[0003] Regarding the detection and recognition of classroom teachers' teaching behaviors, in the prior art, the recognition of single human body movement behaviors is usually based on artificial intelligence technology. However, in the classroom, there are also students besides teachers. The method of directly recognizing human behaviors cannot distinguish the identities of students and teachers, resulting in the possibility of misidentifying students' behaviors as teachers' behaviors, causing interference to the detection of teachers' classroom teaching behaviors and resulting in low reliability of the evaluation of classroom teaching behaviors. Summary of the Invention

[0004] The technical problem to be solved by the present invention lies in: aiming at the technical problems existing in the prior art, the present invention provides a two-stage classroom teacher teaching behavior detection method, device and medium with simple implementation method, low cost, high detection accuracy and strong robustness.

[0005] To solve the above technical problems, the technical solution proposed by the present invention is:

[0006] A two-stage classroom teacher teaching behavior detection method, the steps of which include:

[0007] Real-time receive the classroom video image data collected by the camera;

[0008] Input the received classroom video image data into a pre-trained teacher position detection model to detect the position of the teacher as the first-stage detection result;

[0009] The teacher position detection model is an object detection model constructed based on the YOLOv (You Only Look Once) network. In the YOLOv network, some convolutional operations adopt the frequency adaptive dilated convolution FADConv. The frequency adaptive dilated convolution FADConv obtains the corresponding frequency channel map by performing Fourier transform on the feature map, decomposes it into multiple frequency channel maps according to the frequency level, obtains the selection map of each frequency channel through convolutional operation on the feature map, multiplies the selection map by the corresponding frequency channel map to obtain the weighted feature to achieve frequency selection, and performs convolutional operation on the weighted feature using the adaptive dilation rate and the adaptive convolution kernel to obtain the output feature layer;

[0010] Determine the teacher position area in the video image data according to the first-stage detection result, and input the determined teacher position area and the classroom video image data into the pre-trained teaching behavior recognition model for teacher teaching behavior detection, and obtain the behavior category of the teacher as the detection result output of the second stage.

[0011] Furthermore, the teacher position detection model is constructed based on the YOLOV9 network. The backbone network in the YOLOV9 network includes the RepNCSPELAN4 module and multiple detection heads. The bottleneck layer in the RepNCSPELAN4 module and the convolutional operation in the YOLOV9 initial feature module adopt the frequency adaptive dilated convolution FADConv.

[0012] Furthermore, the adaptive dilation rate is determined for each feature pixel by using convolution and the Relu activation function. The convolution kernel parameters are decomposed into low-frequency convolution kernel parameters and high-frequency convolution kernel parameters. The ratio of the high-frequency convolution kernel weight to the low-frequency convolution kernel weight is adjusted according to the input context to form the adaptive convolution kernel. When performing average pooling operation in low-frequency convolution, the mean value is taken after summing average pooling and median pooling, and the weight of the channel is predicted through the first convolutional layer conv1, the Relu activation function, the MLP layer, and the sigmod function layer in sequence.

[0013] Furthermore, the teacher position detection model further includes a feature fusion module for fusing feature maps. The feature fusion module includes an ASSP module and a CA module. The ASSP module realizes semantic feature extraction of different scales by using convolution, average pooling, and dilated convolution with different dilation rates. The CA module fuses the feature maps extracted by the ASSP module through a residual module, and decomposes the channel attention into two one-dimensional feature encoding processes to aggregate features along two spatial directions respectively.

[0014] Further, the teaching behavior recognition model uses the Swin Transformer architecture as the feature extraction model. When the teaching behavior recognition model performs teaching behavior detection, it first uses a convolutional neural network to perform preliminary feature abstraction and block division on the input image, then uses the Swin Transformer architecture to extract features from the divided feature maps respectively, and finally obtains the detection results of the teaching behavior categories. The MLP modules in the W-MSA module and SW-MSA module of the Swin Transformer architecture are implemented using the KAN (Kolmogorov-Arnold Networks) module, and the KAN module is formed by combining spline curves and MLP.

[0015] Further, the representation structure of the KAN module is as follows:

[0016]

[0017] Where, , is the activation function , is the spline curve function, x represents the input feature, represents the output of the KAN module corresponding to the input feature x, L represents the number of network layers, represents the weight coefficient, represents the composite operation of the function.

[0018] Further, the loss function of the teacher position detection model includes the classification loss function and the position regression loss function , and the calculation expression is:

[0019]

[0020]

[0021]

[0022]

[0023] Where, represents the weight coefficient, represents the class weight focal loss function of the i-th class of identity person, is the class weight factor of students or teachers, where , m represents the number of identity categories, is the number of positive samples of the i-th class of identity person, is the predicted value of the i-th class of identity person, is the true label value for the i-th type of identity, is the balance factor, is the weight factor, represents the bounding box loss, represents the distribution focal loss, 、 represents the weight coefficient.

[0024] Furthermore, it also includes calculating a first metric based on the detection results of the first stage and calculating a second metric based on the results of the second stage, and evaluating the teaching behavior of the teacher according to the first metric and the second metric. The first metric includes the number of times the teacher is detected, and the second metric includes the total number of teaching behaviors of the teacher detected.

[0025] An electronic device includes a processor and a memory. The memory is used to store a computer program, and the processor is used to execute the computer program to perform the method as described above.

[0026] A computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the method as described above.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows: Starting from the perspective of fine-grained object recognition, the present invention uses a two-stage detection method to implement the detection of teachers' teaching behaviors. First, a teacher position detection model constructed based on the YOLOv network is used to detect the teacher's position, enabling the position of the teacher to be identified and located first. At the same time, frequency adaptive dilation convolution is used for some convolution operations in the network. By means of frequency selection, adaptive dilation rate, and adaptive convolution kernel, it is possible to balance the effective local frequency bandwidth and local receptive field of the feature map, so that under a small computational cost, the model's ability to represent local features of the feature map can be improved. Therefore, it is possible to obtain human features from a fine-grained perspective, improve the discrimination between teachers and students, and thus improve the recognition accuracy of teachers. After identifying the teacher's position, the detected teacher position is used to detect the teacher's behavior, which can accurately distinguish the behavior category of the teacher, reduce the interference of students' behaviors on the detection, and improve the detection accuracy and robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a schematic flowchart of the implementation of the two-stage classroom teacher teaching behavior detection method in Embodiment 1 of the present invention.

[0029] Figure 2 is a schematic structural principle diagram of the teacher position detection model used in Embodiment 1 of the present invention.

[0030] Figure 3It is a schematic diagram of the structural principle of the RepNCSPELAN4 module adopted in Embodiment 1 of the present invention.

[0031] Figure 4 It is a schematic diagram of the structural principle of the RepNCSP module adopted in Embodiment 1 of the present invention.

[0032] Figure 5 It is a schematic diagram of the structural principle of the FCBL module adopted in Embodiment 1 of the present invention.

[0033] Figure 6 It is a schematic diagram of the structural principle of the CBL module adopted in Embodiment 1 of the present invention.

[0034] Figure 7 It is a schematic diagram of the structural principle of the RepNBottleneck module implemented in Embodiment 1 of the present invention.

[0035] Figure 8 It is a schematic diagram of the structural principle of the RepACconvN module implemented in Embodiment 1 of the present invention.

[0036] Figure 9 It is a schematic diagram of the structural principle during the training of the ACconv convolutional module implemented in Embodiment 1 of the present invention.

[0037] Figure 10 It is a schematic diagram of the structural principle during the inference of the ACconv convolutional module implemented in Embodiment 1 of the present invention.

[0038] Figure 11 It is a schematic diagram of the structural principle of the frequency adaptive dilated convolution FADConv implemented in Embodiment 1 of the present invention.

[0039] Figure 12 It is a schematic diagram of the structural principle of the feature fusion module (ASSP + CA module) adopted in Embodiment 1 of the present invention.

[0040] Figure 13 It is a schematic diagram of the processing flow for realizing teaching behavior recognition in Embodiment 1 of the present invention.

[0041] Figure 14 It is a schematic diagram of the structural principle of the Swin Transformer architecture based on KAN in Embodiment 1 of the present invention. Detailed implementation manners

[0042] The present invention will be further described below in conjunction with the accompanying drawings of the specification and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.

[0043] Students and veteran teachers are the two main identities in the classroom. When identifying teachers' relevant teaching behaviors, students are interference factors, and when identifying students' relevant classroom behaviors, teachers are interference factors. Therefore, to achieve the accuracy and reliability of detecting teachers' classroom behaviors, it is necessary to correctly distinguish the identities of students and teachers first. Since the classroom is a relatively closed environment with little personnel flow, only students and teachers exist, and there are few interference items. Moreover, the teacher's position is generally in the front and back of the podium area and the pedestrian area between desks. When people distinguish teachers from students with their eyes, they usually identify them based on their positions, appearances and other characteristics. Starting from the perspective of fine-grained object recognition, the present invention adopts a two-stage detection method to detect teachers' teaching behaviors. First, a teacher position detection model constructed based on the YOLOv network is used to detect the teacher's position, so that the teacher's position can be identified and located first. At the same time, in some convolutional operations in the network, the frequency adaptive dilation convolution FADConv is adopted. By using the methods of frequency selection, adaptive dilation rate and adaptive convolution kernel, it is possible to balance the effective local frequency bandwidth and the local receptive field of the feature map, so that at a relatively small computational cost, the model's ability to represent local features of the feature map can be improved. Therefore, it is possible to obtain human features from a fine-grained perspective, improve the distinguishability between teachers and students, and thus improve the recognition accuracy of teachers. After identifying the teacher's position, the identified teacher's position is used to detect the teacher's behavior, which can accurately distinguish the category of the teacher's behavior, reduce the interference of students' behaviors on the detection, and improve the accuracy and robustness of the detection.

[0044] Embodiment 1:

[0045] As Figure 1 shown, the steps of the two-stage classroom teacher teaching behavior detection method in this embodiment include:

[0046] Step S01. Real-time receive the classroom video image data collected by the camera.

[0047] Specifically, cameras can be arranged in the classroom to collect video images during the teacher's teaching process in real time, and the real-time collected classroom teaching video image frames are uploaded to the data processor or control terminal, and the data processor or control terminal executes the subsequent steps S02 and S03 to achieve the detection of the teacher's classroom teaching behavior.

[0048] Step S02. Input the received classroom video image data into a pre-trained teacher position detection model to detect the position of the teacher as the first-stage detection result. The teacher position detection model is an object detection model constructed based on the YOLOv network, and in the YOLOv network, some convolutional operations use frequency-adaptive dilated convolution FADConv. The frequency-adaptive dilated convolution FADConv decomposes the feature map into multiple frequency channels, obtains the selection map of each frequency channel through convolutional operations, multiplies the selection map with the corresponding frequency channel to obtain weighted features to achieve frequency selection, and uses an adaptive dilation rate and an adaptive convolution kernel to perform convolutional operations on the weighted features to obtain the output feature layer.

[0049] The main purpose of teacher position detection is to correctly detect the teacher from the camera's field of view to exclude the interference of students in the field of view, thereby improving the accuracy of subsequent teacher teaching behavior detection. Since teachers and students belong to different identities within the same category, classroom teacher position detection can be attributed to fine-grained object detection. When designing a teacher position detection model, constructing a feature extraction and recognition network from a fine-grained perspective can improve the recognition accuracy of the model for teachers and students.

[0050] In this embodiment, the teacher position detection model is an object detection model constructed based on the YOLOv network. The YOLOv series of networks divides the input image into a grid of a fixed size, and each grid cell is responsible for detecting the objects within that area. If the center point of an object falls within a certain grid, then that grid is responsible for predicting the bounding box and category of the object, and has efficient real-time processing capabilities. As an alternative implementation, considering the balance between algorithm accuracy and real-time performance, the teacher position detection model can adopt an object detection model constructed based on the YOLOV9 network, that is, using the YOLOV9 network as the basic network architecture for teacher position detection. The YOLOV9 network further improves the accuracy of the detection task by introducing programmable gradient information (PGI) and generalized efficient layer aggregation network (GELAN).

[0051] Furthermore, on the basis of the YOLOV9 network, it is improved and optimized, such as Figure 2As shown in the figure, the YOLOV9 network includes a backbone network and an auxiliary branch. The backbone network consists of a RepNCSPELAN4 module and multiple detection heads (specifically 3 detection heads) to form the detection output. Among them, the RepNCSPELAN4 module is a key feature extraction and fusion module that integrates the CSPNet Block module of YOLOv5, the Rep module of YOLOv6, and the ELAN module of YOLOv7, enabling YOLOv9 to improve the detection accuracy while maintaining high computational efficiency. The 3 detection heads use feature layers with different resolutions as input features to enhance the detection ability of Yolov9 for multi-scale feature targets. The auxiliary branch is used to utilize multi-level auxiliary information to address the problem that the loss function cannot generate reliable gradients due to the information bottleneck caused by deepening the neural network. The multi-level auxiliary information is used to handle the error accumulation problem caused by deep supervision. During model training, both the backbone branch and the auxiliary branch participate in the training, while only the backbone network participates during inference. Among them, the CBLinear module is used to first adjust the channels of the feature layer and then equally divide it by channel; the CBFuse module performs the following processing: taking one scale as the standard, first interpolating the input feature layer, uniformly adjusting it to the appropriate size, and merging it by channel.

[0052] As Figures 3 - 9 shown, considering the real-time nature of the calculation, in this embodiment, an FCBL feature extraction module is respectively set at the input end (corresponding to the initial feature module) and the output end (corresponding to the bottleneck layer) of the RepNCSPELAN4 module to implement feature extraction using the frequency adaptive dilated convolution FADConv. Compared with the traditional convolution method, more fine-grained features of the target can be obtained by using the frequency adaptive dilated convolution FADConv, improving the model's ability to extract fine-grained features, enabling a balance between computational real-time and accuracy, and thus improving the teacher model's recognition ability for teachers. At the same time, only using FADConv convolution in the key feature extraction module can also achieve a balance between recognition accuracy and real-time.

[0053] Specifically, as Figure 3 shown, in this embodiment, the RepNCSPELAN4 module specifically includes an FCBL feature extraction module, a RepNCSP module, and a CBL module. Among them, the RepNCSP module includes multiple RepNBottleneck modules and performs feature fusion through a CSP (Cross-Stage Partial) structure. Each RepNBottleneck module is connected to a CBL module through a connection layer and then output to the FCBL feature abstraction module. As Figure 4As shown in the figure. The FCBL feature extraction module uses frequency adaptive dilated convolution FADConv to replace the traditional convolution operation, which can enhance the fine-grained feature extraction of the target and improve the feature recognition ability of the model. For example, Figure 5 As shown in the figure, where BN (Batch Normalization) is used to solve problems such as gradient vanishing, gradient explosion, and unstable training during the training process in deep learning, and SiLU represents the SiLU activation function. The CBL module uses Conv convolution operation for feature extraction to achieve feature abstraction. For example, Figure 6 As shown in the figure. The RepNBottleneck module is a standard bottleneck module composed of RepACconvN basic calculation blocks. For example, Figure 7 As shown in the figure. For example, Figure 8 As shown in the figure, the RepACconvN basic calculation block contains a 3×3 ACconv convolution + BN layer, a 1×1 convolution + BN layer, and a BN layer. The results of the three are added together through parallel connection.

[0054] In a specific application embodiment, the structure for implementing the ACconv convolution is as shown in Figure 9 、 10 As shown in the figure. During the training phase, the 3×3-BN structure is replaced by a multi-branch structure of 3×3-BN layer, 1×3-BN layer, and 3×1-BN layer for model training, and their outputs are summed up. For example, Figure 9 As shown in the figure. After training, during inference, each 1×3-BN layer and 3×1-BN layer asymmetric kernel is added to the skeleton of the 3×3-BN branch, that is, the intersection part of the square kernel. For example, Figure 10 As shown in the figure, the three parallel convolution models during training are converted back to the same structure as the original ordinary convolution. Through the above reparameterized convolution process, the ordinary convolution can enhance the feature recognition ability of the convolutional backbone part without increasing the calculation and changing the original structure, thereby further improving the overall feature representation ability of the model for the scene images of teachers and their behaviors.

[0055] Traditional convolutions usually act as high-pass filters, and the features they generate often exhibit a relatively high proportion of high-frequency components. This tendency leads to the use of a relatively small overall dilation rate to maintain a relatively high effective bandwidth, but it sacrifices the size of the receptive field. A smaller receptive field will reduce the range of connection between the target pixel and surrounding pixels, making it difficult to effectively extract the fine-grained features of the teacher model and reducing the model's recognition ability for the teacher. In this embodiment, by using the frequency-adaptive dilated convolution FADConv in the YOLOv network to replace the traditional convolution operation, the dilation rate of each position in the feature map can be dynamically adjusted in the feature map space according to the local frequency components of each position in the feature map, so that while the feature map as a whole can obtain a suitable dilation rate, a relatively high effective bandwidth can be maintained. Finally, the teacher position detection model can obtain more fine-grained features, improving the model's recognition ability for the teacher. Different from the traditional dilated convolution, the frequency-adaptive dilated convolution FADConv in this embodiment balances the effective bandwidth of the local frequency of the feature map and the local receptive field through three strategies: frequency selection (FreqSelect), adaptive dilation rate (AdaDR), and adaptive kernel (AdaKern), enabling the improvement of the model's ability to represent local features of the feature map at a relatively small computational cost.

[0056] As Figure 11 shown, the frequency selection module (FreqSelect) in this embodiment performs a Fourier transform on the feature map to obtain the corresponding frequency channel map, and performs frequency decomposition according to the frequency level to obtain different frequency channel maps (for example, Figure 11 is decomposed into 4 (configurable) frequency channel maps as decomposed features); then, a selection map for each frequency channel is obtained by performing a convolution operation on the feature map (for example, Figure 11 the selected mapping corresponds to the feature map in), and the selection map is multiplied by the corresponding frequency channel to re-weight the space of each frequency channel, which can balance the high-frequency and low-frequency components in the feature, enabling the frequency-adaptive dilated convolution to effectively learn a larger receptive field, so that the corresponding receptive field size can be increased through frequency selection (FreqSelect).

[0057] In the frequency-adaptive dilated convolution FADConv, the dilation rate is dynamically adjusted in a spatially variable manner through the adaptive dilation rate (AdaDR) to achieve a balance between the effective bandwidth and the receptive field. As Figure 11 shown, the adaptive dilation rate (AdaDR) assigns a suitable dilation rate to each feature pixel adaptively by using convolution (Conv) + Relu activation function on the weighted feature, that is, determines the adaptive dilation rate for each feature pixel by using convolution and Relu activation function, enabling a better trade-off between a large receptive field and an effective bandwidth.

[0058] Traditional convolutional kernel learning can capture features across different frequency bands, which is crucial for understanding complex visual patterns. However, once trained, they become static. To further enhance the effective bandwidth, in the Frequency Adaptive Dilated Convolution (FADConv) of this embodiment, more high-frequency components can be captured through the Adaptive Kernel (AdaKern), thereby increasing the effective bandwidth. In the FADConv of this embodiment, for the formation of low-frequency convolution, the mean of the sum of average pooling and median pooling is used to replace traditional average pooling, that is, (avg + med) / 2 (the mean of the sum of average pooling and median pooling) is used to replace traditional average pooling to reduce the impact of noise on low-frequency convolution. For the formation of high- and low-frequency convolution attention parameters, the dynamic weights (kernel weights) are predicted through a first convolutional layer conv1, a Relu activation function, an MLP layer, and a sigmod function layer in sequence, that is, the dynamic weights of the channels are predicted by conv1 + Relu + conv1 + sigmod. Compared with the traditional method of using the conv1 + Relu + conv1 + sigmod module to determine the channel weights, it can further integrate local semantics and global semantic information, improve the feature representation ability of attention parameters, and thus improve the generalization ability of the model.

[0059] Specifically, as Figure 11 shown, to implement the Adaptive Kernel (AdaKern), the FADConv in this embodiment decomposes the convolutional kernel parameters into low-frequency convolutional kernel parameters (low-frequency components) and high-frequency convolutional kernel parameters (high-frequency components). The adaptive convolutional kernel is formed by the low-frequency convolutional kernel and the high-frequency convolutional kernel, and dynamic weighting is introduced to adjust the frequency response. When performing average pooling operations in low-frequency convolution, the mean of the sum of average pooling and median pooling is used, that is, low-frequency convolution uses the mean of the sum of average pooling and median pooling to replace the original average pooling to reduce the impact of noise signals in the original convolution on low-frequency components. Global pooling and operations such as conv1, Relu, MLP layer, and sigmod are used to predict the weights of the channels to form high- and low-frequency convolution attention parameters, that is, the dynamic weights of the channels are predicted by a simple and lightweight global pooling plus conv1 + Relu + MLP + sigmod. Therefore, the kernel weights will change with the input, which can further integrate local semantics and global semantics and improve the feature representation ability of attention parameters. By dynamically adjusting the ratio of high-frequency convolutional kernel weights to low-frequency convolutional kernel weights according to the input context, the network can also focus on specific frequency bands and adapt to the complexity of visual patterns in features. Through the dynamic frequency adaptation method, the network's ability to capture low-frequency context and high-frequency local details can be enhanced, thereby increasing the effective bandwidth, and thus the performance can be improved in segmentation tasks that require extracting diverse features across different frequencies.

[0060] To further increase the context semantic information of the features, in this embodiment, the teacher position detection model further includes a feature fusion module for fusing the feature maps. The feature fusion module is used to further fuse the feature maps to enhance the semantic connection between each pixel in the feature map and its surrounding pixels, and further improve the model's ability to represent teacher features. The feature fusion module includes an ASPP module and a CA module. As Figure 12 shown, the ASPP module includes multiple feature extraction branches. Each branch uses convolution (Conv), global average pooling (Avgpool), and dilated convolution Dcon with different dilation rates rate to extract semantic features of different scales. By combining conv1×1, average pooling, and dilated convolution Dconv with different dilation rates, semantic features of different scales are extracted and merged by channels to strengthen the semantic information connection between the target pixel and its surrounding pixels.

[0061] As Figure 12 shown, the CA module fuses the feature maps extracted by the ASPP module through a residual structure to further enhance the semantic connection between pixels and their surrounding pixels, and decomposes the channel attention into two one-dimensional feature encoding processes to aggregate features along the X and Y spatial directions respectively. Specifically, for the output of the residual structure, pooling kernels with sizes (H,1) and (1,W) are first used to encode each channel along the horizontal and vertical coordinates respectively (X Avg pool, Y Avg pool). The outputs of X Avg pool and Y Avg pool are concatenated by channels, and feature extraction is performed on them through Conv2D convolution, and further feature abstraction is performed through BatchNorm (batch normalization) and Non - Linear (non - linear feature extraction layer). Then it is separated into components in the X and Y directions, and after convolution (Conv2d) and Sigmod operations respectively, Reweight (re - weighting) is performed. In this way, long - range dependencies can be captured along one spatial direction, while precise position information can be retained along the other spatial direction. Then the obtained feature maps are separately encoded into a pair of direction - aware and position - sensitive attention maps, which can be complementarily applied to the input feature map to enhance the representation of the teacher.

[0062] In this embodiment, feature fusion is achieved by adopting the ASPP + CA module. Compared with the traditional SPPELAN module, it can reduce the information loss of feature extraction while strengthening the semantic connection between the central pixel and its surrounding pixels, improving the fusion between semantic features of different scales, and thus further improving the fine - grained feature extraction ability of the model.

[0063] During the training process of this embodiment, the loss function of the teacher position detection model includes a classification loss function and the position regression loss function , and the calculation expression is:

[0064] (1)

[0065] where represents the weight coefficient.

[0066] Optionally, the classification loss function can be specifically calculated according to the following expression:

[0067] (2)

[0068] (3)

[0069] where represents the class-weighted focal loss function of the i-th class of identity person, is the class weight factor of the student or teacher, where , m represents the number of identity categories, is the number of positive samples of the i-th class of identity person, is the predicted value of the i-th class of identity person, is the true label value of the i-th class of identity person, is the balance factor, is the weight factor.

[0070] Optionally, the regression loss function can be specifically calculated according to the following expression:

[0071] (4)

[0072] where represents the bounding box loss, represents the distribution focal loss, 、 represent the weight coefficients.

[0073] Step S03. Determine the teacher position area in the video image data according to the first-stage detection result, and input the determined teacher position area and the classroom video image data into the pre-trained teaching behavior recognition model to perform teacher teaching behavior detection, and obtain the teacher's behavior category as the second-stage detection result for output.

[0074] After the teacher's position is accurately detected in the first stage, the second stage of detection is further carried out to identify the teacher's current teaching behavior and determine whether it is a teaching behavior. For example, it is identified whether it is a teaching behavior such as gestures, standing lectures, writing on the blackboard, pointing at the multimedia screen or the blackboard, operating teaching equipment, etc. The teacher's position detection and teaching behavior recognition use dedicated models for detection respectively to ensure the recognition accuracy and robustness of each task, so as to ensure the accuracy and reliability of teacher teaching behavior recognition.

[0075] Since the teaching behavior recognition involves the understanding of the semantic information of the whole picture, in order to further strengthen the semantic connection between pixels, improve the recognition accuracy of teaching behavior, and at the same time considering the global information relationship modeling ability of Transformer and the need to ensure a certain real-time performance of the model operation, this embodiment adopts the Swin Transformer architecture as the feature extraction model of the teacher behavior recognition model. When the teaching behavior recognition model conducts teaching behavior detection, it uses a convolutional neural network to perform preliminary feature abstraction on the input image and divides it into blocks, and then uses the Swin Transformer architecture to extract features from the divided feature maps respectively. Finally, the detection results of the teaching behavior categories are obtained. Among them, the MLP modules in the W-MSA module and SW-MSA module of the Swin Transformer architecture are implemented using the KAN module, that is, the original MLP layer is replaced by the KAN layer to implement the Transformer architecture based on KAN, which can improve the model's feature extraction ability for the teacher's actions. The KAN module is formed by combining a spline curve and an MLP. That is, KAN is essentially a combination of a spline curve and an MLP, so it can absorb the advantages of both and give full play to the advantages of the spline curve with high accuracy and strong interpretability in low dimensions, as well as the advantage of the MLP being convenient for dimension expansion.

[0076] Specifically, the steps for the teaching behavior recognition model to recognize teaching behavior in this embodiment are as follows:

[0077] Step S301. Image segmentation: Use a convolutional neural network based on frequency adaptive dilated convolution (FADConv) to perform preliminary feature abstraction on the teacher image to enhance the feature representation ability, and segment it to adapt to the input of the Swin Transformer architecture;

[0078] Step S302. Feature extraction: Use the Swin Transformer architecture to perform multiple feature extractions on the segmented feature maps respectively. After multiple executions, the balance between feature extraction performance and computational performance can be achieved.

[0079] Specifically, as Figure 13As shown in the figure, where x2 represents two such modules and x6 represents six such modules. After the input image passes through the FADConv layer, step 1 is executed (performing linear embedding and Swin Transformer Block (repeated twice)) for feature extraction. Linear embedding is achieved by linearly transforming the channel data of each pixel to abstract features. Then, step 2 is executed (including patch merging and Swin Transformer Block (repeated twice)) for further feature enhancement. Next, the features extracted above are subjected to step 3 (including patch merging and Swin Transformer Block (repeated 6 times)) for further feature extraction. Finally, the output of step 3 is subjected to step 4 (including patch merging and Swin Transformer Block (repeated 2 times)) for final feature extraction to obtain the final classification result. It can be understood that the number of executions of the Swin Transformer Block in the above steps can be configured according to actual needs.

[0080] Step S303. Classification and recognition: Identify the teacher's teaching behavior category based on the features extracted in step S302.

[0081] As Figure 14 shown, the Swin Transformer architecture based on KAN in this embodiment includes a KAN module, a W-MSA module, and an SW-MSA module. Using KAN to replace the MLP in the traditional Transformer architecture enables stronger function fitting ability to enhance the model's fitting ability for teacher behaviors and solve the problem of limited function fitting ability of the traditional MLP (multi-layer perceptron).

[0082] Specifically, the representation structure of the KAN module can be expressed as the following formula:

[0083] (5)

[0084] Where , is the activation function , is the spline curve function, x represents the input feature, represents the output of the KAN module corresponding to the input feature x, L represents the number of network layers, " " represents the composite operation of the function, represents the weight coefficient.

[0085] Considering that the teaching behavior recognition model is a classification model based on the Swim Transformer architecture and that a teacher may perform multiple teaching behaviors at the same time, during the training process of the teaching behavior recognition model, a multi-label classification loss function can be used Optimize the teaching behavior recognition model, and the calculation formula can adopt the same principle as formula (2) above.

[0086] Example 2:

[0087] This example is basically the same as Example 1, except that it further includes the step of evaluating the quality of classroom teaching: calculating the first index according to the detection results of the first stage and calculating the second index according to the results of the second stage, and evaluating the teaching behavior of the teacher according to the first index and the second index. The first index includes the number of times the teacher is detected, and the second index includes the total number of teaching behaviors of the teacher detected. That is, the steps of the method for detecting the teaching behavior of the classroom teacher in this example include:

[0088] Step S01. Receive the classroom video image data collected by the camera in real time;

[0089] Step S02. Input the received classroom video image data into a pre-trained teacher position detection model to detect the position of the teacher as the detection result of the first stage; the teacher position detection model is an object detection model constructed based on the YOLOv network, and in the YOLOv network, some convolution operations adopt the frequency adaptive dilation convolution FADConv. The frequency adaptive dilation convolution FADConv decomposes the feature map into multiple frequency channels, obtains the selection map of each frequency channel through convolution operations, multiplies the selection map with the corresponding frequency channel to obtain the weighted feature to achieve frequency selection, and performs convolution operations on the weighted feature using the adaptive dilation rate and the adaptive convolution kernel to obtain the output feature layer;

[0090] Step S03. Determine the teacher position area in the video image data according to the detection result of the first stage, and input the determined teacher position area and the classroom video image data into a pre-trained teaching behavior recognition model to detect the teaching behavior of the teacher, and obtain the behavior category of the teacher as the detection result of the second stage for output.

[0091] Step S04. Calculate the first index according to the detection result of the first stage and calculate the second index according to the result of the second stage, and evaluate the teaching behavior of the teacher according to the first index and the second index. The first index includes the number of times the teacher is detected, and the second index includes the total number of teaching behaviors of the teacher detected.

[0092] In this embodiment, a two-stage detection method is first adopted to sequentially detect the teacher's position and identify teaching behaviors. Each task uses a dedicated model for detection to ensure the recognition accuracy and robustness of each task, ensuring the accuracy of the teacher's position detection task and its teaching behavior recognition task, achieving accurate positioning of the teacher's identity in the classroom and accurately and robustly identifying their relevant teaching behaviors. Furthermore, by comprehensively combining the detection results of the two stages, accurate auxiliary evaluation of the teacher's teaching behavior can be realized, which is beneficial to improving the teaching quality.

[0093] In order to effectively evaluate the teacher's teaching behavior and conduct a quantitative evaluation of the teaching behavior, this embodiment evaluates the teaching behavior from two aspects: effective teaching duration and teaching initiative. By counting the effective teaching duration of the teacher in the classroom, some behaviors such as being late and leaving early can be identified, and the teacher's teaching behavior can be standardized. The teaching initiative can reflect the teacher's attitude towards teaching tasks.

[0094] Optionally, the effective teaching duration of the teacher in the classroom can be evaluated based on the number of times the teacher is detected , and the teaching initiative can be evaluated based on the total number of categories of the detected teacher's teaching behaviors , for example, the following calculation formulas can be respectively adopted:

[0095] (6)

[0096] (7)

[0097] Among them, represents the total number of detections, represents the total number of times the teacher is detected, m represents the total number of types of the teacher's teaching behaviors, represents the total number of times each teaching behavior is detected.

[0098] This embodiment further provides an electronic device, including a processor and a memory. The memory is used to store a computer program, and the processor is used to execute the computer program to execute the methods in Embodiments 1 and 2 as described above.

[0099] It can be understood that the above method of this embodiment can be executed by a single device, such as a computer or a server, etc., or can also be applied to a distributed scenario where multiple devices cooperate with each other to complete. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps of the above method of this embodiment, and the multiple devices interact with each other to complete the above method. The processor can be implemented in ways such as a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., and is used to execute relevant programs to implement the above method of this embodiment. The memory can be implemented in forms such as a read-only memory ROM, a random access memory RAM, a static storage device, and a dynamic storage device, etc. The memory can store an operating system and other application programs. When implementing the above method of this embodiment through software or firmware, the relevant program codes are stored in the memory and are called and executed by the processor.

[0100] This embodiment further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the methods of Embodiments 1 and 2 as described above.

[0101] Those skilled in the art should understand that the above embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or boxes Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one box or more boxes

[0102] The above is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above in a preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention

Claims

1. A two-stage classroom teacher teaching behavior detection method, characterized in that the steps include: Receive classroom video image data collected by the camera in real time; The received classroom video image data is input into the pre-trained teacher position detection model to detect the teacher's position as the first stage detection result; The teacher position detection model is a target detection model built based on the YOLOv network, and some convolution operations in the YOLOv network adopt frequency adaptive dilated convolution FADConv, which performs Fourier transform on the feature map to obtain the corresponding frequency channel map, and decomposes it into multiple frequency channel maps according to the frequency, obtains the selection map of each frequency channel by convolution operation on the feature map, multiplies the selection map with the corresponding frequency channel map to obtain weighted features to achieve frequency selection, and performs convolution operation on the weighted features using adaptive dilation rate and adaptive convolution kernel to obtain output feature layer; Determine the teacher's position area in the video image data according to the detection results of the first stage, input the determined teacher's position area and the classroom video image data into a pre-trained teaching behavior recognition model to perform teacher's teaching behavior detection, and obtain the teacher's behavior category as the detection result output of the second stage; The loss function of the teacher position detection model is Including classification loss function And the position regression loss function , the calculation expression is: in, represents the weight coefficient, represents the class-weighted focal loss function for the i-th class person, is the class weight factor of the student or teacher, where , represents the number of identity categories, is the number of positive samples of the i-th category person, is the predicted value of the i-th category person, is the true label value of the i-th category person, is the balance factor, is the weight factor, represents the bounding box loss, represents the distribution focal loss, , Represents the weight coefficient.

2. The double-stage classroom teacher teaching behavior detection method according to claim 1 is characterized in that: The adaptive expansion rate is determined for each feature pixel by using convolution and Relu activation function, the convolution kernel parameters are decomposed into low-frequency convolution kernel parameters and high-frequency convolution kernel parameters, and the ratio of the high-frequency convolution kernel weight to the low-frequency convolution kernel weight is adjusted according to the input context to form the adaptive convolution kernel. When the average pooling operation is performed in the low-frequency convolution, the average is taken by summing the average pooling and the median pooling, and the channel weight is predicted in sequence through the first convolution layer conv1, the Relu activation function, the MLP layer, and the sigmoid function layer.

3. The double-stage classroom teacher teaching behavior detection method according to claim 1 is characterized in that: The teacher position detection model is constructed based on the YOLOV9 network. The backbone network in the YOLOV9 network includes a RepNCSPELAN4 module and multiple detection heads. The bottleneck layer in the RepNCSPELAN4 module and the convolution operation in the YOLOV9 initial feature module adopt the frequency adaptive dilated convolution FADConv.

4. The double-stage classroom teacher teaching behavior detection method according to claim 1 is characterized in that: The teacher position detection model also includes a feature fusion module for fusing feature maps, and the feature fusion module includes an ASSP module and a CA module. The ASSP module realizes semantic feature extraction of different scales by adopting convolution, average pooling and dilated convolution with different expansion rates. The CA module fuses the feature maps extracted by the ASSP module through a residual module, and decomposes the channel attention into two one-dimensional feature encoding processes to aggregate features along two spatial directions respectively.

5. The double-stage classroom teacher teaching behavior detection method according to claim 1 is characterized in that: The teaching behavior recognition model is based on the Swin Transformer architecture as a feature extraction model. When the teaching behavior recognition model performs teaching behavior detection, the input image is preliminarily abstracted and divided into blocks using a convolutional neural network, and the feature maps after block division are respectively subjected to feature extraction using the Swin Transformer architecture, and finally the detection result of the teaching behavior category is obtained. The W-MSA module of the Swin Transformer architecture and the MLP module in the SW-MSA module are implemented using the KAN module, and the KAN module is formed by combining a spline curve and an MLP.

6. The double-stage classroom teacher teaching behavior detection method according to claim 5 is characterized in that: The representation structure of the KAN module is as follows: in, , is the activation function , is a spline function, x represents the input feature, represents the KAN module output corresponding to the input feature x, L represents the number of network layers, represents the weight coefficient, Represents a composite operation of a function.

7. The method for detecting the teaching behavior of teachers in a two-stage classroom according to any one of claims 1 to 6, characterized in that: It also includes calculating a first indicator based on the first stage detection results and calculating a second indicator based on the second stage results, and evaluating the teacher's teaching behavior based on the first indicator and the second indicator. The first indicator includes the number of times the teacher is detected, and the second indicator includes the total number of times the teacher's teaching behavior is detected.

8. An electronic device comprising a processor and a memory, wherein the memory is used to store a computer program, wherein: The processor is configured to execute the computer program to perform the method according to any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Classroom teaching quality auxiliary evaluation method, equipment and system based on cloud edge collaboration

    CN119418414A

  • Teaching assistance method and teaching assistance system using said method

    US20200175264A1