A multi-channel-based classroom teaching behavior recognition method and system
By recording videos of the podium and student area with dual cameras, combining face recognition and behavioral posture recognition, and using the Bayesian causal network to construct a classroom teaching behavior discriminant function, the objectivity and accuracy problems of classroom teaching behavior analysis in existing technologies are solved, and automated teaching behavior analysis is realized.
Patent Information
- Application Number
- CN202211530101.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-30
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2042-11-30
AI Technical Summary
Existing classroom teaching behavior analysis methods lack objectivity. Traditional methods require a lot of manpower and material resources, and the loss of video information results in incomplete and inaccurate analysis.
A multi-channel classroom teaching behavior recognition method is adopted. Dual cameras are used to record videos of the podium area and student area respectively. Combining face recognition and behavioral posture recognition, a classroom teaching behavior discriminant function is constructed through the Bayesian causal network to achieve automatic analysis.
It improves the objectivity, comprehensiveness and accuracy of classroom teaching behavior analysis, reduces the input of manpower and material resources, avoids the loss of video information, and realizes the automatic analysis of classroom teaching behavior.
Smart Images

Figure CN115719516B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of educational informatization technology, and more specifically, relates to a multi-channel based classroom teaching behavior recognition method and system. Background Art
[0002] By analyzing classroom teaching behaviors, we can obtain objective and effective classroom teaching evaluations, help transform teaching models, improve teachers' professional qualities, and optimize teaching quality.
[0003] Currently, most classroom teaching behavior analysis methods rely on questionnaires and observations. While valuable, these methods require both teachers and students to have a clear memory of the classroom teaching process, require observers to invest significant time and effort, and do not necessarily yield valid behavioral analysis information. Furthermore, self-evaluation of teachers' teaching behaviors and external observational evaluation of students' learning behaviors are often limited by factors such as teacher proficiency, student literacy, and cultural background, resulting in a lack of objectivity. Summary of the Invention
[0004] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a multi-channel classroom teaching behavior identification method and system, aiming to solve the problem of insufficient objectivity in conventional classroom teaching behavior analysis.
[0005] To achieve the above objectives, in a first aspect, the present invention provides a multi-channel classroom teaching behavior recognition method, comprising:
[0006] S101 obtains a first video frame and a second video frame captured by a first camera and a second camera at the same time, respectively; the first camera is used to capture video of the podium area of the classroom; the second camera is used to capture video of the student area of the classroom;
[0007] S102 performs face recognition and behavior posture recognition on the first video frame and the second video frame respectively to obtain identity information and corresponding behavior posture of each subject in the podium area and the student area;
[0008] S103 obtains the classroom teaching behavior corresponding to the same moment based on the identity information and corresponding behavior posture of each subject in the podium area and the student area, and a preset classroom teaching behavior discriminant function.
[0009] In an optional example, step S102 specifically includes:
[0010] Performing image fusion on the first video frame and the second video frame to obtain a classroom teaching image;
[0011] performing face target detection and human body target detection on the classroom teaching image to obtain each face region feature and each human body region feature in the classroom teaching image;
[0012] performing face recognition based on the each face region feature to obtain identity information corresponding to the each face region feature;
[0013] performing behavior and posture recognition based on the each human body region feature to obtain a behavior and posture corresponding to the each human body region feature;
[0014] matching and regionally dividing each identity information and each behavior and posture based on position information in the each face region feature and the each human body region feature to obtain identity information and a corresponding behavior and posture of each main character in a teacher's desk region and a student's region.
[0015] In an optional example, the performing behavior and posture recognition based on the each human body region feature to obtain a behavior and posture corresponding to the each human body region feature comprises:
[0016] performing human body key point extraction based on the each human body region feature to obtain a human body key point corresponding to the each human body region feature, and constructing a human body structure graph feature based on the human body key point;
[0017] performing behavior and posture classification based on the each human body region feature and the corresponding human body structure graph feature to obtain a behavior and posture corresponding to the each human body region feature.
[0018] In an optional example, the identity information of the main character includes a teacher and a student, and the behavior and posture of the main character includes reading / taking notes, listening to a class, turning a side, raising a hand, standing, and writing on a blackboard;
[0019] the classroom teaching behavior includes teacher's writing on a blackboard, teacher's teaching, teacher's asking questions, teacher's patrolling a classroom, student's writing on a blackboard, student's discussing, student's practicing, student's answering questions, and student's going on a stage.
[0020] In an optional example, the classroom teaching behavior discrimination function is constructed based on a Bayesian causal network.
[0021] In a second aspect, the present application provides a multi-channel based classroom teaching behavior recognition system, comprising:
[0022] a video frame acquisition module, configured to acquire a first video frame and a second video frame collected by a first camera and a second camera at the same time; the first camera is configured to collect a video of a teacher's desk region of a classroom; and the second camera is configured to collect a video of a student's region of the classroom;
[0023] A video frame recognition module, configured to perform face recognition and behavior posture recognition on the first video frame and the second video frame, respectively, and obtain identity information and corresponding behavior posture of each subject in the podium area and the student area;
[0024] The teaching behavior discrimination module is used to obtain the classroom teaching behavior corresponding to the same moment based on the identity information and corresponding behavior posture of each subject in the podium area and the student area, and a preset classroom teaching behavior discrimination function.
[0025] In an optional example, the video frame identification module includes:
[0026] An image fusion unit, configured to fuse the first video frame and the second video frame to obtain a classroom teaching image;
[0027] An object detection unit, configured to perform face object detection and human object detection on the classroom teaching image and obtain features of each face region and each human region in the classroom teaching image;
[0028] A face recognition unit, configured to perform face recognition based on the features of each face region and obtain identity information corresponding to the features of each face region;
[0029] a behavior posture recognition unit, configured to perform behavior posture recognition based on the features of each human body region and obtain the behavior posture corresponding to each human body region feature;
[0030] The matching and fusion unit is used to match and divide the identity information and behavior postures based on the position information in the facial area features and the body area features, and obtain the identity information and corresponding behavior postures of each subject in the podium area and the student area.
[0031] In an optional example, the behavior posture recognition unit is used to extract human body key points based on the human body area features to obtain the human body key points corresponding to the human body area features, and construct human body structure map features based on the human body key points, and perform behavior posture classification based on the human body area features and the corresponding human body structure map features to obtain the behavior posture corresponding to the human body area features.
[0032] In an optional example, the identity information of the subject determined by the video frame recognition module includes: a teacher and a student; the behavior postures of the subject determined include: reading / taking notes, listening to a lecture, leaning sideways, raising hands, standing, and writing on a blackboard;
[0033] The classroom teaching behaviors determined by the teaching behavior identification module include: teacher writing on the blackboard, teacher lecturing, teacher asking questions, teacher patrolling the classroom, students writing on the blackboard, students discussing, students practicing, students answering questions, and students going on stage.
[0034] In an optional example, the classroom teaching behavior discriminant function used by the teaching behavior discriminant module is constructed based on a Bayesian causal network.
[0035] In general, the above technical solutions conceived by the present invention have the following beneficial effects compared with the prior art:
[0036] The present invention provides a classroom teaching behavior recognition method and system based on multi-channel. By dividing the teaching area into a podium area and a student area, dual cameras are used to record videos of the podium area and the student area respectively, thereby avoiding the problem of loss of video information in the classroom teaching scene. Face recognition and behavior posture recognition are performed based on the dual video streams to obtain the identity information and corresponding behavior postures of all main characters in the podium area and the student area, and statistical analysis is performed based on this to finally determine the teaching behavior in the classroom teaching scene, thereby realizing automatic analysis of classroom teaching behavior and greatly improving the objectivity, comprehensiveness and accuracy of classroom teaching behavior analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 This is one of the flow charts of the multi-channel classroom teaching behavior recognition method provided by an embodiment of the present invention;
[0038] Figure 2 This is a flow chart of behavior posture classification and face recognition provided by an embodiment of the present invention;
[0039] Figure 3 Schematic diagram of the network structure of the behavior posture recognition network provided by an embodiment of the present invention;
[0040] Figure 4 This is the second flow chart of the multi-channel classroom teaching behavior recognition method provided by an embodiment of the present invention;
[0041] Figure 5 is a data alignment flow chart of dual video streams provided by an embodiment of the present invention;
[0042] Figure 6 This is a process framework diagram of the classroom subject identity-behavior posture representation model provided by an embodiment of the present invention;
[0043] Figure 7 This is an architecture diagram of a multi-channel classroom teaching behavior recognition system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0045] By analyzing classroom teaching behavior, we can obtain objective and effective classroom teaching evaluations, facilitate the transformation of teaching models, enhance teacher professionalism, and optimize teaching quality. Currently, domestic and international methods for analyzing classroom teaching behavior can be divided into two types based on the type of data studied: questionnaire- and observation-based methods, and video-based methods.
[0046] Common methods for measuring classroom teaching behavior are self-reflection reports and external observational evaluations. When conducting self-reflection or observer evaluations, evaluators use classroom teaching behavior standards or questionnaires to provide a rough summary of classroom teaching behavior. In the self-reflection method, teachers and students complete questionnaires on their individual teaching or learning behaviors. Questionnaires on teaching and learning behaviors typically require teachers and students to report on their interactions, conversations, and statements during the teaching process. Class participants rely solely on memory to complete the reports, making the data's validity difficult to verify.
[0047] In addition to teacher and student evaluation reports, some analytical methods rely on external observers (typically experienced teachers) to complete summative evaluation scales to analyze teacher and student classroom behavior. External observation and evaluation require observers to complete action classification scales, sample analyses, and case studies based on their observations of classroom processes. ST analysis (Student-Teacher Analysis) and Flanders Interaction Analysis are commonly used summative analysis scales.
[0048] While the aforementioned analytical methods have proven valuable in practice, their shortcomings are also clear: they require teachers and students to have a clear memory of the classroom teaching process, require observers to invest a significant amount of time and effort, and do not necessarily yield valid behavioral analysis information. Furthermore, both self-evaluation of teachers' teaching behaviors and external observation of students' learning behaviors require teachers, students, or observers to honestly and accurately assess their own or others' teaching behavior problems. Limited by factors such as teacher proficiency, student literacy, and cultural background, the results of these methods often lack objectivity.
[0049] With the widespread use of video cameras in school classrooms, the cost of acquiring classroom teaching videos has been significantly reduced, providing a user-friendly data acquisition method for teaching researchers. Furthermore, the rapid development of artificial intelligence (AI) technology has led to significant progress in image and video processing and classification techniques, enabling effective research on video-based classroom teaching behavior analysis. Currently, the general process for video-based teaching behavior analysis methods is as follows: First, classroom video recording is performed using classroom video equipment. Subsequently, an automated classroom teaching behavior discriminant function is constructed based on expert knowledge-based teaching behavior standards. Next, AI technology is used to sample and analyze the teaching videos and extract effective teaching behavior features. The classroom teaching behavior discriminant function is then used to encode the classroom teaching behavior of the current sample. Finally, by summarizing and statistically analyzing all classroom teaching behavior codes, a teaching behavior sequence is constructed to derive the classroom teaching model.
[0050] Video-based classroom teaching behavior analysis methods address the disruptions traditional methods often cause, saving manpower and time, and improving the objectivity of teaching behavior analysis. However, existing teaching behavior analysis systems often use multi-channel video switching, which results in a certain amount of video information loss, resulting in incomplete classroom teaching behavior analysis and a need to improve the accuracy of classroom teaching behavior recognition.
[0051] To address the shortcomings of existing teaching behavior analysis systems, this invention divides the teaching area into a podium area and a student area. Using dual, directional cameras, the cameras record each area separately, providing a dual-stream, multi-channel method for identifying classroom teaching behavior. Furthermore, by sampling the camera video streams at their frame rates, the method leverages the features inherent in the teaching videos to analyze classroom teaching behavior, improving the accuracy of classroom teaching behavior analysis. Figure 1 This is one of the flow charts of the multi-channel classroom teaching behavior recognition method provided by the embodiment of the present invention. Figure 1 As shown, the method includes:
[0052] Step S101, obtaining the first video frame and the second video frame captured by the first camera and the second camera respectively at the same time; the first camera is used to capture video of the podium area of the classroom; the second camera is used to capture video of the student area of the classroom.
[0053] Specifically, the first camera and the second camera can be installed at the front and back of the classroom, respectively. The first camera is used to capture video of the classroom's podium area, and the second camera is used to capture video of the classroom's student area. This allows for the complete capture of the classroom teaching scene using a non-switching method, avoiding the problem of video information loss. It should be noted that the terms "first" and "second" here are only used to distinguish between the two cameras and the videos they capture.
[0054] After the first camera and the second camera respectively capture videos of the podium area and the student area, video frames captured by the first camera and the second camera at the same time can be obtained, namely the first video frame of the podium area and the second video frame of the student area.
[0055] Step S102: performing face recognition and behavior posture recognition on the first video frame and the second video frame respectively to obtain the identity information and corresponding behavior posture of each subject in the podium area and the student area;
[0056] Step S103: Based on the identity information and corresponding behavior posture of each subject in the podium area and the student area, and a preset classroom teaching behavior discriminant function, the classroom teaching behavior corresponding to the moment is obtained.
[0057] Specifically, on the basis of obtaining the video frames of the podium area and the student area collected at the same time, it is also taken into consideration that a single behavioral posture cannot accurately reflect the actual teaching behavior in the entire classroom scene. For example, when the teacher is asking questions, some students are listening and some students are raising their hands. If only a single behavioral posture is used to judge, the final teaching behavior may be that the teacher is lecturing, rather than the teacher asking questions; for example, the teacher asks one of the students to go to the podium to write on the blackboard, the teacher is standing, and most of the other students are listening. If only a single behavioral posture is used to judge, the final teaching behavior may also be that the teacher is lecturing, rather than the student writing on the blackboard.
[0058] To address the above issues, an embodiment of the present invention performs face recognition and behavioral posture recognition on the first video frame, and performs face recognition and behavioral posture recognition on the second video frame, thereby obtaining the identity information and corresponding behavioral posture of each subject in the podium area, as well as the identity information and corresponding behavioral posture of each subject in the student area. Subsequently, the identity information and corresponding behavioral posture of all subjects in the podium area, as well as the identity information and corresponding behavioral posture of all subjects in the student area, are statistically analyzed. The classroom teaching behavior corresponding to the moment indicated in step S101 is then determined based on a pre-set classroom teaching behavior discriminant function. For example, if statistics show that the proportion of students in the student area who are in a raised-hand posture is greater than a preset threshold, and no students are currently in a standing posture, then the classroom teaching behavior corresponding to that moment can be determined to be the teacher asking a question. If statistics show that the identity information of a subject in the student area is a teacher, and the number of subjects in the podium area who are in a blackboard writing posture is 1, then the classroom teaching behavior corresponding to that moment can be determined to be the teacher asking a question.
[0059] Here, the embodiment of the present invention does not specifically limit the order in which face recognition and behavioral posture recognition are performed on the first and second video frames. For example, face recognition and behavioral posture recognition can be performed on the first video frame first, followed by face recognition and behavioral posture recognition on the second video frame, or on the second video frame first, followed by face recognition and behavioral posture recognition on the first video frame. Alternatively, the first and second video frames can be fused into a single image, and then face recognition and behavioral posture recognition can be performed on the fused image. The recognition results are then divided into the identity information and corresponding behavioral posture of each subject in the podium area and the student area by region. The classroom teaching behavior discriminant function is used to summarize and analyze the identity-behavioral posture information in the podium area and the student area, thereby determining the category of the classroom teaching behavior corresponding to that moment.
[0060] The method provided by the embodiment of the present invention divides the teaching area into a podium area and a student area, and uses dual cameras to record videos of the podium area and the student area respectively, thereby avoiding the problem of loss of video information in the classroom teaching scene, and performs face recognition and behavior posture recognition based on the dual video streams to obtain the identity information and corresponding behavior postures of all main characters in the podium area and the student area, and performs statistical analysis based on this, and finally determines the teaching behavior in the classroom teaching scene, thereby realizing automatic analysis of classroom teaching behavior, and greatly improving the objectivity, comprehensiveness and accuracy of classroom teaching behavior analysis.
[0061] Based on the above embodiment, step S102 specifically includes:
[0062] Performing image fusion on the first video frame and the second video frame to obtain a classroom teaching image;
[0063] Performing face target detection and human target detection on classroom teaching images and obtaining features of each face region and each human region in the classroom teaching images;
[0064] Performing face recognition based on the features of each facial region and obtaining identity information corresponding to each facial region feature;
[0065] Performing behavioral posture recognition based on the characteristics of each human body region and obtaining the behavioral posture corresponding to each human body region feature;
[0066] Based on the position information of each face region feature and each body region feature, each identity information and each behavior posture are matched and divided into regions, and the identity information and corresponding behavior posture of each subject in the podium area and the student area are obtained.
[0067] Specifically, in order to improve the efficiency of face recognition and behavior posture recognition in the video frames of the two regions and reduce the consumption of system resources, an embodiment of the present invention first fuses the first video frame of the podium region and the second video frame of the student region into one image, namely, a classroom teaching image, and then inputs the classroom teaching image into a target detection network. The target detection network locates the human body region and the face region in the classroom teaching image, thereby obtaining the features of each face region and each body region in the classroom teaching image. Here, each face region feature and each body region feature may include the category information cls of whether the target of the current region is a face or a body, and may also include the location information loc of the current region in the image. Segmentation can be performed based on the location information to obtain each face region image and each body region image in the classroom teaching image.
[0068] Then, the feature set can be divided according to cls so that R = {R body ,R face}, R body is a set of features of each human body region, R face is a set of features of each face region, and R body With R face Input into the behavior posture recognition network and the face recognition network respectively, that is, R face Input the face recognition network to obtain the subject identity of the face in the current face area, thereby obtaining the identity information corresponding to the features of each face area, and convert R body Input into the behavior posture recognition network to obtain the behavior posture corresponding to each human body area feature.
[0069] Subsequently, based on the location information (loc) from each facial region feature and body region feature obtained by the object detection network, the face recognition results are matched with the classroom behavior and posture recognition results through multi-model feature fusion. This establishes a correspondence between the subject's identity information and behavior and posture, thereby obtaining the identity-behavior-posture information of all subjects in the classroom. Finally, the area to which loc belongs is divided according to whether it is the podium area or the student area, thereby obtaining the identity-behavior-posture information of all subjects in the podium area and the identity-behavior-posture information of all subjects in the student area.
[0070] It should be noted that, considering that if a teacher or student is writing on the blackboard, it may not be possible to obtain the facial image of the subject, and thus the identity information of the subject cannot be obtained. For similar situations, when matching the identity information and behavioral posture of the subject, if the identity information corresponding to the behavioral posture cannot be matched based on the position information, the identity information corresponding to the behavioral posture can be set to empty.
[0071] Furthermore, Figure 2This is a flowchart of behavior posture classification and face recognition provided by an embodiment of the present invention. Considering that in the conventional behavior posture classification and face recognition process, using two models at the same time will lead to an increase in system resources or analysis time, the embodiment of the present invention trains a target detection network FP-YOLO (You Only Look Once Advanced by Fully Pre-training) based on YOLO, which can simultaneously obtain the positions of the human body and the face. It has the advantages of being fast and consuming less resources. Figure 2 .c. Figure 2 .a is a serialized target detection method, that is, human target detection is performed first and then face target detection. Figure 2 .b is a separate concurrent target detection method, that is, two models are used for human target detection and face target detection respectively. Figure 2 .a、 Figure 2 .b Increase the system's time consumption and resource consumption respectively.
[0072] Based on any of the above embodiments, performing behavior posture recognition based on the characteristics of each human body region and obtaining the behavior posture corresponding to each human body region feature includes:
[0073] Extracting key points of the human body based on the features of each human body region, obtaining the key points of the human body corresponding to the features of each human body region, and constructing the features of the human body structure map based on the key points of the human body;
[0074] Behavior posture classification is performed based on the characteristics of each human body region and the corresponding human body structure map characteristics, and the behavior posture corresponding to each human body region feature is obtained.
[0075] Specifically, during classroom behavior and posture recognition, there are issues such as cluttered backgrounds, large differences in the shapes of the subjects, and small differences in the subjects' behavior and posture. Therefore, the recognition rate of general image classification networks in classroom teaching is not very outstanding. To address these issues and improve the accuracy of behavior and posture recognition, the behavior and posture recognition network in the embodiment of the present invention includes a human body key point extraction subnetwork and a behavior and posture classification subnetwork. That is, behavior and posture recognition can be divided into two steps: human body key point extraction and behavior and posture classification. In the human body key point extraction stage, the human body key point extraction subnetwork quickly obtains human body key points in each human body region based on the features of each human body region, and constructs corresponding human body structure map features based on the human body key points. Here, the human body structure map features can be obtained by connecting the human body key points and are used to represent the human limb structure. In the behavior and posture classification stage, the behavior and posture classification subnetwork combines the human body structure map features with the original human body region features to classify the human body behavior and posture in the human body region, thereby obtaining the behavior and posture corresponding to each human body region feature.
[0076] Furthermore, the specific network structure of the behavior posture recognition network provided by the embodiment of the present invention is as follows: Figure 3 As shown, the behavior posture recognition network includes a key point extraction subnetwork (i.e. Figure 3 The human body region image obtained after target detection is input into OpenPose to obtain the key points of the human body in the image and construct the human body structure map feature. The human body structure map feature is processed by GAT and GCN to obtain the OpenPose g Input into the attention residual network, and the original human body area image is obtained after CNN O c Input into the attention residual network, the attention residual network applies the attention mechanism to associate the above two inputs, where Q g , K c and V c is the weight matrix of the attention mechanism, and the final features are used for subsequent behavior posture classification. It should be noted that the multi-feature behavior posture classification method based on the attention mechanism provided by the embodiment of the present invention greatly improves the classification accuracy compared with the traditional classroom behavior posture classification method.
[0077] Based on any of the above embodiments, the identity information of the subject includes: teacher and student; the behavior of the subject includes: reading / taking notes, listening to a lecture, leaning sideways, raising hands, standing, and writing on the blackboard;
[0078] Classroom teaching behaviors include: teacher writing on the blackboard, teacher lecturing, teacher asking questions, teacher patrolling the classroom, students writing on the blackboard, student discussions, student exercises, students answering questions, and students going on stage.
[0079] Based on any of the above embodiments, the classroom teaching behavior discriminant function is constructed based on the Bayesian causal network.
[0080] Specifically, an embodiment of the present invention proposes an ST teaching analysis method based on a Bayesian causal network for dual video streams. The identity information and corresponding behavioral posture of each subject in the podium area and the student area are used as conditional information. The identity information and behavioral posture in the podium area, as well as the identity information and behavioral posture in the student area are statistically analyzed. Then, a classroom teaching behavior discriminant function is constructed based on the Bayesian causal network to determine the category to which the classroom teaching behavior at the current moment belongs.
[0081] Furthermore, the specific classroom teaching behavior discriminant function provided by the embodiment of the present invention is shown in Table 1. When the number of subjects in the blackboard writing posture in the podium area is equal to 1 and the number of subjects with identity information of teachers is equal to 1 (T=1), or the number of subjects in the blackboard writing posture in the podium area is equal to 1, and the face detection in the student area has a teacher identity or the number of subjects in the standing posture in the student area is equal to 1 and its corresponding identity information is not a student (E=1andIE≠S), or the number of subjects in the blackboard writing posture in the podium area is greater than 1, it indicates that the current teaching state is the student writing state, that is, the classroom teaching behavior is behavior 5. When the proportion of students in the student area in the sideways behavior posture is at least 30% of the total number of students, it indicates that the current teaching state is the student discussion state, that is, the classroom teaching behavior is behavior 6. When a student appears in the podium area in a standing behavior posture, it indicates that the current state is the student going on stage, that is, the classroom teaching behavior is behavior 9.
[0082] In addition to the above behaviors, when the number of students in the student area listening to a lecture exceeds 30% of the total student population, the current teaching state is assumed to be the teacher lecturing, i.e., Behavior 2. Under Behavior 2, if the number of subjects in the student area standing is 0, no teacher face is detected in the student area, and the number of subjects in the podium area writing on the blackboard is 1, the current teaching state is the teacher writing on the blackboard, i.e., Behavior 1. If the number of students in the student area raising their hands exceeds 10% of the total student population and no students are currently standing, the current teaching state is the teacher asking questions, i.e., Behavior 3. If the number of students in the student area standing is 1 or more, the current teaching state is the student answering questions, i.e., Behavior 8. If the proportion of students in the student area reading / note-taking is greater than 30%, the current teaching state is assumed to be the student practice, i.e., Behavior 7. Under Behavior 7, and the teacher is not in the podium area, the current teaching state is the teacher patrolling the classroom, i.e., Behavior 4.
[0083] Table 1
[0084]
[0085] In this way, we can obtain classroom teaching behaviors corresponding to multiple moments, and by connecting multiple moments in chronological order, we can obtain a classroom teaching behavior sequence.
[0086] Based on any of the above embodiments, machine learning and deep learning are currently hot topics in artificial intelligence research and have achieved good results in many fields. Applying artificial intelligence technology to the classroom and constructing an automated classroom teaching behavior analysis system based on the ST teaching analysis method greatly reduces the manpower and material resources required for classroom teaching behavior analysis, avoids the drawbacks of traditional teaching behavior analysis methods, and improves the objectivity of classroom teaching behavior analysis. However, the current classroom teaching behavior analysis system has the following problems: the existing analysis methods mostly use multi-channel lens switching videos, which require special personnel to shoot, and cannot obtain the overall situation of the classroom. There is a certain amount of loss in the amount of video information, resulting in incomplete classroom teaching behavior analysis and the accuracy of classroom teaching behavior recognition needs to be improved; and the existing methods cannot perform statistical analysis on individual classroom teaching behaviors.
[0087] To address the above problems, the present invention provides a classroom teaching behavior recognition method based on dual video streams and multiple channels (Classroom Teaching Behavior Analysis based on Dual Video Streams, DVS-TBA). Figure 4 This is the second flow chart of the multi-channel classroom teaching behavior recognition method provided by the embodiment of the present invention. Figure 4 As shown in Figure 2, the method can be divided into four stages: dual video stream image fusion, image target detection, image representation extraction, and classroom teaching behavior encoding, as follows:
[0088] Step 1: Image fusion of dual video streams
[0089] The image fusion of dual video streams is divided into data alignment of dual video streams and data fusion of dual video streams. The dual video streams here come from the first camera and the second camera respectively.
[0090] The data alignment process of the dual video streams provided by the embodiment of the present invention is as follows: Figure 5 As shown, first, the two cameras are synchronized via a time server, maintaining the same timeline. To minimize the synchronization error between the two cameras, the time server and the cameras must be on the same local area network. Next, the parsed video stream data I is time-aligned to generate data I′ = {P1, P2, T}. P1 represents the image from the first camera, i.e., the first video frame of the podium area; P2 represents the image from the second camera, i.e., the second video frame of the student area; and T represents the time corresponding to the current image.
[0091] On the basis of data alignment, the two images are spliced into one image, i.e. a classroom teaching image (for example, the podium area is in the upper 1 / 2 area, the student area is in the lower 1 / 2 area, and a narrow black bar is added in the middle to distinguish the two areas to avoid recognition errors), which is used for subsequent extraction of identity-behavioral posture information of the main person in the classroom.
[0092] Step 2: image target detection and image feature extraction
[0093] By investigating traditional classroom teaching behavior analysis methods and analyzing the current more advanced technologies, the embodiment of the present application proposes a classroom main body identity-behavioral posture feature model based on multiple models, and the overall process framework of the model is as shown in Figure 6 On the basis of step 1, the FP-YOLO algorithm is used to locate the human body area and the face area in the classroom teaching image, and a human body-face area positioning feature set R = {r1, r2, … r i ,…,r n}, r i = {cls, loc}, cls ∈ {body, face}. Wherein, cls represents the class of the image in the current area; loc = [x, y, w, h], x and y represent the horizontal and vertical coordinates of the center point of the area, w represents the width of the area, and h represents the height of the area.
[0094] Then, the region division feature data makes R = {R body ,R face}, and R body and R face are respectively input into the behavior posture recognition network and the face recognition network for further analysis, as in the second step of Figure 6 The behavior posture recognition is divided into two steps, i.e. human key point extraction and behavior posture classification. In the human key point extraction stage, the key point extraction sub-network is based on the performance of human posture estimation of OpenPose, and quickly obtains the human key points in the human body area image, and constructs a human structure graph feature. In the behavior posture classification stage, the human structure graph feature and the original human body area feature are input into the MCG-CPR (Classroom Pose Recognition based on Mixed Cnn and Gnn neural networks) network, the MCG-CPR applies attention mechanism to associate the two features, and finally obtains the behavior posture classification result A = {a1, a2, …, a m}, a m={cls,loc,action}. In the face recognition network, the subject identity represented by the face in the current face area is obtained according to the face area image, and the face recognition result D = {d1,d2,…,d k},d k ={cls,loc,identify}, where identify represents the subject identity information corresponding to the face in the current area.
[0095] Finally, based on the loc features obtained by the FP-YOLO model, the face recognition results are matched and fused with the classroom behavior posture recognition results through multi-model feature fusion to construct the identity-behavior posture information H = {h1,h2,…,h l},h l ={identify,action},l=max(m,k).
[0096] Step 3: Coding classroom teaching behaviors
[0097] Based on step 2, the identity-behavior posture information of each subject in the classroom is first input into the classroom subject identity-behavior posture information segmentation module of the dual video stream, so that all the subject identity-behavior posture information H all ={H platform ,H student}. Among them, H platform represents the identity-behavior posture information extracted from the podium area, H student represents the identity-behavior posture information extracted from the student area. all As conditional information, determine the current classroom teaching behavior. all The subject's behavior and posture information includes reading / notetaking, listening, leaning sideways, raising hands, standing, and writing on the blackboard. The Bayesian causal network summarizes and analyzes the identity-behavior posture information in the podium and student areas, encoding the current moment's classroom teaching behavior. By concatenating multiple moments in chronological order, a classroom teaching behavior sequence is generated.
[0098] The dual-video stream ST teaching analysis method based on the Bayesian causal network proposed in the embodiment of the present invention can effectively obtain classroom teaching behavior coding by utilizing the classroom subject identity-behavior posture information of the dual video streams.
[0099] Based on any of the above embodiments, the embodiments of the present invention construct a classroom teaching behavior analysis system that does not require human intervention and can fully use teaching information to analyze classroom and individual teaching behaviors. The method includes data alignment and fusion methods for dual video streams, a classroom subject identity-behavior posture representation model based on multiple models, a classroom subject identity-behavior posture information segmentation method for dual video streams, and a ST teaching analysis method based on a Bayesian causal network for dual video streams, as follows:
[0100] 1. To address the problem of information loss caused by traditional multi-channel lens switching videos, the embodiment of the present invention divides the teaching area into a podium area and a student area, and records the podium area and the student area separately through directional dual cameras. It also proposes a data alignment and fusion method for dual video streams, so that the dual video stream data can be used for information extraction based on a multi-model classroom subject identity-behavior posture representation model.
[0101] In order to facilitate the subsequent characterization and extraction model to extract information, it is necessary to preprocess the data of the dual video streams. First, the video stream is parsed to obtain the time information and image data contained in the video stream. Then, the time information is used to align the image data with the time (such as Figure 5 As shown), image fusion constructs a fused image, namely the classroom teaching image.
[0102] 2. In order to encode classroom teaching behaviors through non-switched dual-channel video streams, an embodiment of the present invention obtains all behavioral posture features contained in the classroom teaching video stream based on a multi-model classroom subject identity-behavioral posture representation model, and confirms the identity information corresponding to the behavioral posture features, that is, extracts the classroom subject identity-behavioral posture representation.
[0103] The embodiment of the present invention proposes a multi-model-based classroom subject identity-behavior posture representation model to extract information from the video and obtain classroom subject identity-behavior posture information, which can realize the processing of classroom video data and improve the accuracy and speed of recognition.
[0104] 3. In the process of obtaining the identity of the classroom subject - the characterization of behavior and posture, target detection is first required, that is, locating the position of the human body and the face, and then performing behavior and posture classification and face recognition. Considering that in the conventional behavior and posture classification and face recognition process, using two models at the same time will lead to an increase in system resources or analysis time, the embodiment of the present invention trains a target detection network FP-YOLO based on YOLO that can simultaneously obtain the position of the human body and the face, which has the advantages of being fast and consuming less resources, such as Figure 2 .c.
[0105] 4. Classification of classroom behavior and posture poses often suffers from background clutter, significant differences in subject shape, and minimal differences in subject behavior and posture. Consequently, conventional image classification networks are not very effective in classroom teaching. To improve the accuracy of classroom behavior and posture classification, the present invention investigates the particularities of classroom teaching scenarios and proposes a multi-layered neural network architecture, the attention-based multi-feature classroom behavior and posture classification network (MCG-CPR).
[0106] By analyzing the characteristics of classroom teaching video data and based on the relevant information of the survey, the original data is preprocessed and the corresponding features are extracted to construct a human body structure feature. Combined with the original image classification model, multi-feature classroom teaching behavior posture classification is performed. The specific network structure is as follows Figure 3 As shown in the figure, the human body area image obtained after target detection is input into OpenPose to obtain the human body key points in the image, and the human body structure map features are constructed. The human body structure map features are input into the attention residual network after passing through GAT and GCN. At the same time, the original human body area image is input into the attention residual network after passing through CNN. The combined features of the two are then used for behavior posture classification.
[0107] 5. In order to obtain classroom teaching behavior coding using the classroom subject identity-behavior posture representation extracted from the dual video streams, the embodiment of the present invention inputs the extracted identity-posture features of each subject in the classroom into the classroom subject identity-behavior posture information segmentation module of the dual video streams to obtain the identity-behavior posture information extracted from the podium area and the student area respectively. On this basis, a ST teaching analysis method based on the Bayesian causal network is proposed for the dual video streams to encode the classroom teaching behavior at the current moment.
[0108] Based on any of the above embodiments, Figure 7 This is an architecture diagram of a multi-channel classroom teaching behavior recognition system provided by an embodiment of the present invention. Figure 7 As shown, the system includes:
[0109] The video frame acquisition module 710 is configured to acquire a first video frame and a second video frame captured by a first camera and a second camera at the same time. The first camera is configured to capture video of the podium area of the classroom; the second camera is configured to capture video of the student area of the classroom.
[0110] The video frame recognition module 720 is used to perform face recognition and behavior posture recognition on the first video frame and the second video frame respectively, and obtain the identity information and corresponding behavior posture of each subject in the podium area and the student area;
[0111] The teaching behavior discrimination module 730 is used to obtain the classroom teaching behavior corresponding to the moment based on the identity information and corresponding behavior posture of each subject in the podium area and the student area, as well as the preset classroom teaching behavior discrimination function.
[0112] The system provided by the embodiment of the present invention divides the teaching area into a podium area and a student area, and uses dual cameras to record videos of the podium area and the student area respectively, thereby avoiding the problem of loss of video information in the classroom teaching scene. Face recognition and behavioral posture recognition are performed based on the dual video streams to obtain the identity information and corresponding behavioral postures of all main characters in the podium area and the student area, and statistical analysis is performed based on this to finally determine the teaching behavior in the classroom teaching scene, thereby realizing automatic analysis of classroom teaching behavior and greatly improving the objectivity, comprehensiveness and accuracy of classroom teaching behavior analysis.
[0113] It is understandable that the detailed functional implementation of each of the above modules can be found in the introduction of the aforementioned method embodiment, and will not be repeated here.
[0114] In addition, an embodiment of the present invention provides another multi-channel-based classroom teaching behavior recognition device, which includes: a memory and a processor;
[0115] The memory is used to store computer programs;
[0116] The processor is configured to implement the method in the above embodiment when executing the computer program.
[0117] In addition, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method in the above embodiment is implemented.
[0118] Based on the method in the above embodiment, an embodiment of the present invention provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.
[0119] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-channel classroom teaching behavior recognition method, characterized in that: include: S101 obtains a first video frame and a second video frame captured by a first camera and a second camera at the same time; The first camera is used to capture video of the podium area of the classroom; The second camera is used to capture video of the student area of the classroom; S102 performs face recognition and behavior posture recognition on the first video frame and the second video frame respectively to obtain identity information and corresponding behavior posture of each subject in the podium area and the student area; S103 obtains the classroom teaching behavior corresponding to the same moment based on the identity information and corresponding behavior posture of each subject in the podium area and the student area, and a preset classroom teaching behavior discriminant function; Step S102 specifically includes: Performing image fusion on the first video frame and the second video frame to obtain a classroom teaching image; Performing face target detection and human target detection on the classroom teaching image and obtaining features of each face region and features of each human region in the classroom teaching image; Performing face recognition based on the facial region features and obtaining identity information corresponding to the facial region features; Performing behavior posture recognition based on the human body region features and obtaining the behavior posture corresponding to the human body region features; Based on the position information of each face region feature and each body region feature, each identity information and each behavior posture are matched and divided into regions to obtain the identity information and corresponding behavior posture of each subject in the podium area and the student area; The performing behavior posture recognition based on the human body region features and obtaining the behavior posture corresponding to the human body region features includes: Extracting human body key points based on the human body region features to obtain human body key points corresponding to the human body region features, and constructing human body structure map features based on the human body key points; Behavior posture classification is performed based on the characteristics of each human body region and the corresponding human body structure diagram characteristics to obtain the behavior posture corresponding to each human body region feature.
2. The classroom teaching behavior recognition method according to claim 1, characterized in that: The subject's identity information includes: teacher and student; the subject's behavior includes: reading / taking notes, listening to a lecture, leaning sideways, raising hands, standing, and writing on the blackboard; The classroom teaching behaviors include: teacher writing on the blackboard, teacher lecturing, teacher asking questions, teacher patrolling the classroom, students writing on the blackboard, students discussing, students practicing, students answering questions and students going on stage.
3. The classroom teaching behavior recognition method according to any one of claims 1 to 2, characterized in that: The classroom teaching behavior discriminant function is constructed based on the Bayesian causal network.
4. A multi-channel classroom teaching behavior recognition system, characterized by: include: A video frame acquisition module, configured to acquire a first video frame and a second video frame captured by the first camera and the second camera respectively at the same time; The first camera is used to capture video of the podium area of the classroom; The second camera is used to capture video of the student area of the classroom; A video frame recognition module, configured to perform face recognition and behavior posture recognition on the first video frame and the second video frame, respectively, and obtain identity information and corresponding behavior posture of each subject in the podium area and the student area; A teaching behavior discrimination module is used to obtain the classroom teaching behavior corresponding to the same moment based on the identity information and corresponding behavior posture of each subject in the podium area and the student area, and a preset classroom teaching behavior discrimination function; The video frame recognition module includes: an image fusion unit, configured to fuse the first video frame and the second video frame to obtain a classroom teaching image; An object detection unit, configured to perform face object detection and human object detection on the classroom teaching image and obtain features of each face region and each human region in the classroom teaching image; A face recognition unit, configured to perform face recognition based on the features of each face region and obtain identity information corresponding to the features of each face region; a behavior posture recognition unit, configured to perform behavior posture recognition based on the features of each human body region and obtain the behavior posture corresponding to each human body region feature; A matching and fusion unit is configured to match and divide the identity information and behavior postures based on the position information of the facial region features and the body region features, and obtain the identity information and corresponding behavior posture of each subject in the podium area and the student area; The behavior posture recognition unit is used to extract human body key points based on the human body region features to obtain human body key points corresponding to the human body region features, construct human body structure map features based on the human body key points, and classify behavior postures based on the human body region features and the corresponding human body structure map features to obtain behavior postures corresponding to the human body region features.
5. The classroom teaching behavior recognition system according to claim 4 is characterized in that: The identity information of the subject determined by the video frame recognition module includes: teachers and students; the behavior postures of the subject determined include: reading / taking notes, listening to a lecture, leaning sideways, raising hands, standing, and writing on the blackboard; The classroom teaching behaviors determined by the teaching behavior identification module include: teacher writing on the blackboard, teacher lecturing, teacher asking questions, teacher patrolling the classroom, students writing on the blackboard, students discussing, students practicing, students answering questions, and students going on stage.
6. The classroom teaching behavior recognition system according to any one of claims 4 to 5, characterized in that: The classroom teaching behavior discriminant function used by the teaching behavior discriminant module is constructed based on the Bayesian causal network.