Classroom behavior analysis method and system, electronic equipment and storage medium

By preprocessing classroom video and audio data and extracting deep learning features, a multi-dimensional classroom behavior index system is built, which solves the problems of low accuracy and single dimensions of classroom behavior analysis in the existing technology, and realizes high-precision comprehensive evaluation of multi-dimensional classroom behavior and teaching effect correlation.

CN120180388AActive Publication Date: 2025-06-20NANJING LANZHONG INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510661711.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The existing technology has problems in classroom behavior analysis with limited identification accuracy, poor data processing capabilities and single analysis dimensions, making it difficult to achieve high-precision, multi-dimensional classroom behavior comprehensive evaluation.

Method used

By obtaining classroom video stream data and teacher and student audio stream data, image enhancement, noise filtering and video frame segmentation are performed to generate a standardized image sequence; noise reduction, sound source positioning and speech segmentation are performed on the audio stream data to generate a standardized speech sequence. CNN is used for feature extraction, combined with the MTCNN cascade network and multi-synchronous node head pose estimation model, to identify students' facial key points and head poses; two-way LSTM network is used for timing modeling to obtain the length and frequency of speeches of teachers and students. Based on this information, a system of indicators for students' individual behavior and teacher-student interaction behavior is constructed, a multi-dimensional classroom behavior portrait is generated, and it is related to the teaching effect.

Benefits of technology

It improves the accuracy of classroom behavior analysis, enhances data processing capabilities, realizes multi-dimensional comprehensive classroom behavior evaluation, and can guide teaching intervention and optimization more scientifically.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180388A_ABST
    Figure CN120180388A_ABST
Patent Text Reader

Abstract

The invention provides a classroom behavior analysis method and system, electronic equipment and a storage medium, and relates to the technical field of educational informationization, and the method comprises the steps: collecting videos and audios, carrying out image enhancement, noise filtering and frame segmentation on the videos, carrying out noise reduction, sound source positioning and voice segmentation on the audios, and generating standardized images and voice sequences; an improved MTCNN cascade network is combined with a posture estimation technology, facial key points of students are extracted from videos, class arrival states, head postures and facial micro-expressions are recognized, audios are analyzed through a bidirectional LSTM network, and speaking duration and frequency are obtained; constructing individual behavior indexes according to attendance, head actions and expressions of the students; constructing an interaction index according to the teacher and student speaking data; the individual and interaction indexes are summarized to generate the multi-dimensional classroom behavior portrait, and the relation between the portrait and the teaching effect is analyzed, so that the classroom behavior analysis precision can be improved, the data processing capability can be enhanced, and the multi-dimensional classroom behavior comprehensive evaluation can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of educational informatization technology, and particularly to a classroom behavior analysis method, system, electronic device, and storage medium. Background Art

[0002] In the field of educational informatization, with the development of intelligent devices and big data analysis technology, classroom behavior analysis has become an important means to improve teaching quality and optimize teaching management.

[0003] In the prior art, classroom behavior analysis mainly captures students' facial expressions, body movements, etc. through a video monitoring system, and combines machine learning algorithms for behavior recognition to evaluate students' learning status and participation. However, traditional classroom behavior analysis methods often rely on manual observation and subjective judgment, and have many defects. Specifically, first, the recognition accuracy is limited, and the recognition effect drops significantly under complex lighting conditions or when the students' faces are blocked; second, the data processing ability is poor, and it is difficult to achieve real-time analysis of large-scale data; third, the analysis dimension is single, and there is a lack of comprehensive evaluation of the quality of teacher-student interaction.

[0004] Therefore, it is necessary to provide a classroom behavior analysis method, system, electronic device, and storage medium to solve the above technical problems. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a classroom behavior analysis method, system, electronic device, and storage medium, which are used to solve the problems that the prior art cannot effectively improve the accuracy of classroom behavior analysis, has poor data processing ability, and cannot achieve multi-dimensional comprehensive evaluation of classroom behavior.

[0006] The classroom behavior analysis method provided by the present invention includes: Obtain classroom video stream data and teacher-student audio stream data, perform image enhancement, noise filtering, and video frame segmentation on the classroom video stream data to generate a standardized image sequence, and perform noise reduction, sound source localization, and speech segmentation on the teacher-student audio stream data to generate a standardized speech sequence; Use CNN to extract features from the standardized image sequence, that is, locate the key points of the students' faces through an improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, combine with a multi-joint nodding head pose estimation model to identify the students' attendance status, head pose, and facial micro-expressions, and use a bidirectional LSTM network to perform temporal modeling on the standardized speech sequence to obtain the speaking duration and frequency of the teacher and the speaking duration and frequency of the students; Construct a student individual behavior index system based on the student attendance status, the head posture, and the facial micro - expressions, including student attendance rate, head - up rate, head - down duration, and expression change frequency, and construct a teacher - student interaction behavior index system based on the teacher's speaking duration and frequency and the student's speaking duration and frequency, including the proportion of teacher - student speech, student question frequency, and teacher feedback duration; Summarize the student individual behavior index system and the teacher - student interaction behavior index system to generate a multi - dimensional classroom behavior portrait, and correlate the multi - dimensional classroom behavior portrait with the classroom teaching effect.

[0007] Preferably, obtain the classroom video stream data and the teacher - student audio stream data, perform image enhancement, noise filtering, and video frame segmentation on the classroom video stream data to generate a standardized image sequence, and perform noise reduction, sound source localization, and speech segmentation on the teacher - student audio stream data to generate a standardized speech sequence. Specifically, it includes: Deploy multi - angle cameras and array microphones in the target teaching area to form a distributed sensor network; Based on the distributed sensor network, capture the classroom video stream data with a resolution above 720P in real - time through the RTSP protocol, and synchronously collect the classroom audio stream data with a sampling rate of 48kHz to complete the spatio - temporal alignment acquisition of the classroom video stream data and the classroom audio stream data; Perform image enhancement, noise filtering, and video frame segmentation on the classroom video stream data, that is, improve the scene contrast through the adaptive histogram equalization method; based on the object detection model, prior knowledge of seat layout, and bilateral filtering algorithm, remove the teacher's image noise and retain the student's image information to generate the standardized image sequence; Perform noise reduction, sound source localization, and speech segmentation on the classroom audio stream data, that is, use the Wiener filtering algorithm to remove environmental noise, and combine the short - time energy detection method to filter out the silent segments; based on the beamforming technology, locate the sound source direction, train the teacher - student voiceprint features through the GMM - UBM model, and automatically separate the teacher's speech signal and the student's speech signal; use the speech activity detection algorithm to segment the effective speech segments to generate the standardized speech sequence with timestamps.

[0008] Preferably, the method of removing the teacher's image noise and retaining the student's image information based on the object detection model, prior knowledge of seat layout, and bilateral filtering algorithm to generate the standardized image sequence specifically includes: For the classroom video stream data, identify the teacher's human body area based on the object detection model to generate a teacher mask ; Based on the prior knowledge of seat layout, combined with the background difference method, extract the student's human body area to generate a student mask , complete the spatial separation of the teacher and student images in the classroom video stream data; Use the bilateral filtering algorithm for the teacher mask to perform denoising processing to generate a denoised teacher area, and the corresponding calculation formula is as follows: In the formula, represents the pixel p in the denoised teacher area; represents the normalization factor at pixel p; N(p) represents the spatial neighborhood of pixel p; represents the spatial standard deviation. If the details in the teacher mask are to be retained, then ; if enhanced denoising is performed on the teacher mask , then ; represents the spatial domain Gaussian weight function, which is used to control the spatial influence range of neighboring pixels, and ; represents the range standard deviation, , represents the standard deviation of the pixel values in the teacher's face area; represents the pixel p in the teacher's original image area; represents the pixel q in the teacher's original image area; represents the range Gaussian weight function, which is used to control the influence of color similarity of neighboring pixels, and ; Directly retain the student original image area covered by the student mask , and fuse the denoised teacher area with the student original image area according to the corresponding mask to generate the standardized image sequence, and the corresponding calculation formula is as follows: In the formula, represents the standardized image sequence; represents the teacher mask; represents the denoised teacher area; represents the student mask; represents the student original image area.

[0009] Preferably, during the process of using the bilateral filtering algorithm to perform denoising processing on the teacher mask , a motion compensation mechanism is introduced, which specifically includes: For the motion blur noise in the teacher mask , a motion compensation term is added to the spatial domain Gaussian weight function, where denote the motion vector of pixel p; denote the motion vector of pixel q, then the updated spatial domain Gaussian weight function is .

[0010] Preferably, the CNN is used to extract features from the standardized image sequence, that is, the improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion is used to locate the key points of the student's face, and combined with the multi-joint nodding head pose estimation model, the student's attendance status, head pose and facial micro-expression are recognized, specifically including: Through the improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, according to the complexity of the facial region in the standardized image sequence, the weights of each layer of the network are dynamically allocated, and the facial feature information at different scales is fused to locate 68 key points of the student's face; Synchronously access the multi-joint nodding head pose estimation model, and calculate the three-dimensional pose parameters of the student's head in real time, and the multi-joint nodding head pose estimation model at least includes a bone point detection module and an angle calculation module based on a convolutional neural network; Establish a double-layer detection mechanism for the facial presence state and the head pose matrix, and based on the key points of the student's face and the three-dimensional pose parameters of the student's head, recognize the student's attendance status, the head pose and the facial micro-expression.

[0011] Preferably, the bidirectional LSTM network is used to perform temporal modeling on the standardized speech sequence to obtain the teacher's speech duration and frequency and the student's speech duration and frequency, specifically including: Convert the standardized speech sequence into a Mel-frequency cepstral coefficient sequence, and input the Mel-frequency cepstral coefficient sequence into the bidirectional LSTM network for temporal modeling; Extract the speech energy, zero-crossing rate and fundamental frequency features in the Mel-frequency cepstral coefficient sequence through a preset sliding window, and combine the attention mechanism to obtain the key speech period; Construct a teacher-student speech classifier to distinguish between the teacher's teaching period and the student's interaction period in the key speech period, and output the teacher's speech duration and frequency and the student's speech duration and frequency.

[0012] Preferably, the student individual behavior index system and the teacher-student interaction behavior index system are summarized to generate a multi-dimensional classroom behavior portrait, and the multi-dimensional classroom behavior portrait is associated with the classroom teaching effect, specifically including: Perform feature normalization processing on the student individual behavior index system and the teacher-student interaction behavior index system to generate a classroom behavior index set X; Obtain the preset number of clusters, use the K-means algorithm to group the set of classroom behavior indicators X to obtain multiple clusters, calculate the central vectors of each cluster, and construct the multi-dimensional classroom behavior portrait Y; Establish a set of classroom teaching effect indicators E, use a multiple linear regression model to fit the linear relationship between the multi-dimensional classroom behavior portrait Y and the set of classroom teaching effect indicators E, and screen out significantly associated indicators through t-tests; Introduce SHAP values to calculate the contribution degree of each indicator in the set of classroom behavior indicators X to the classroom teaching effect, sort the indicators in the set of classroom behavior indicators X from largest to smallest according to the contribution degree, and select the top A indicators as key classroom behavior indicators, where A represents the preset number of key indicators; Calculate the Pearson correlation coefficient between the key classroom behavior indicators and the classroom teaching effect and construct a behavior-effect association heat map.

[0013] A classroom behavior analysis system, the analysis system includes: A data acquisition module, used to acquire classroom video stream data and teacher-student audio stream data, perform image enhancement, noise filtering, and video frame segmentation on the classroom video stream data to generate a standardized image sequence, and perform noise reduction, sound source localization, and speech segmentation on the teacher-student audio stream data to generate a standardized speech sequence; A data analysis module, used to extract features from the standardized image sequence using CNN, that is, locate the key points of the student's face through an improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, combine a multi-joint nodding head pose estimation model to identify the student's attendance status, head pose, and facial micro-expressions, and use a bidirectional LSTM network to perform temporal modeling on the standardized speech sequence to obtain the teacher's speaking duration and frequency as well as the student's speaking duration and frequency; An indicator construction module, used to construct a student individual behavior indicator system based on the student's attendance status, head pose, and facial micro-expressions, including student attendance rate, head-up rate, head-down duration, and expression change frequency, and construct a teacher-student interaction behavior indicator system based on the teacher's speaking duration and frequency as well as the student's speaking duration and frequency, including the teacher-student speaking ratio, student question frequency, and teacher feedback duration; A summary association module, used to summarize the student individual behavior indicator system and the teacher-student interaction behavior indicator system to generate a multi-dimensional classroom behavior portrait, and associate the multi-dimensional classroom behavior portrait with the classroom teaching effect.

[0014] An electronic device includes a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the steps of the classroom behavior analysis method described in any one of the above.

[0015] A readable storage medium stores a computer program. When the computer program is executed by a processor, it is used to implement the steps of the classroom behavior analysis method described in any one of the above.

[0016] Compared with the related art, the classroom behavior analysis method, system, electronic device and storage medium provided by the present invention have the following beneficial effects: The present invention obtains classroom video stream data and teacher-student audio stream data, performs image enhancement, noise filtering and video frame segmentation on the classroom video stream data to generate a standardized image sequence, performs noise reduction, sound source localization and speech segmentation on the teacher-student audio stream data to generate a standardized speech sequence; uses CNN to extract features from the standardized image sequence, that is, locates the key points of the student's face through an improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, combines a multi-joint nodding head pose estimation model to identify the student's attendance status, head pose and facial micro-expression, and uses a bidirectional LSTM network to perform temporal modeling on the standardized speech sequence to obtain the teacher's speaking duration and frequency and the student's speaking duration and frequency; constructs a student individual behavior index system based on the student's attendance status, head pose and facial micro-expression, including student attendance rate, head-up rate, head-down duration and expression change frequency, and constructs a teacher-student interaction behavior index system based on the teacher's speaking duration and frequency and the student's speaking duration and frequency, including the proportion of teacher-student speech, student question frequency and teacher feedback duration; summarizes the student individual behavior index system and the teacher-student interaction behavior index system to generate a multi-dimensional classroom behavior portrait, and correlates the multi-dimensional classroom behavior portrait with the classroom teaching effect, so as to effectively improve the accuracy of classroom behavior analysis, enhance the data processing ability, and realize the multi-dimensional comprehensive evaluation of classroom behavior.

[0017] The present invention improves the MTCNN cascade network, combines pose estimation technology, and constructs a two-layer mechanism for face detection and pose assistance, which can effectively solve the recognition problems in complex scenarios such as low light and face occlusion, and provide a reliable basis for accurately evaluating the learning state. The present invention combines a bilateral filtering algorithm and a motion compensation mechanism through a distributed sensor network, which can achieve high-quality acquisition and preprocessing of classroom data, and at the same time support real-time streaming processing and batch analysis of classroom data, improve the data processing efficiency, and meet the low-latency requirements of large-scale education scenarios. The present invention can construct a multi-dimensional index system including individual behaviors and teacher-student interactions, generate student behavior portraits through clustering analysis, and combine statistical modeling to quantify the relationship between classroom behaviors and teaching effects, providing scientific guidance for teaching intervention. The present invention adopts a full-process automated system design, supports rapid deployment in multiple scenarios such as smart classrooms and online education, is compatible with multi-terminal display and distributed computing mechanisms, and can adapt to different classroom environments without manual calibration, significantly improving the intelligent level and resource adaptation ability of education management. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of the classroom behavior analysis method of the present invention; Figure 2 is a system block diagram of the classroom behavior analysis system of the present invention; Figure 3 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] As Figure 1 shown, it is a flowchart of the classroom behavior analysis method provided by an embodiment of the present invention, Figure 1The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include but is not limited to at least one of the following: user equipment, network equipment, etc. Among them, the user equipment can include but is not limited to computers, smartphones, personal digital assistants (Personal Digital Assistant, abbreviated as: PDA), and the electronic devices mentioned above. The network equipment can include but is not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing. Among them, cloud computing is a type of distributed computing, which is composed of a group of loosely coupled computers to form a super virtual computer. This embodiment does not limit this. It includes steps S1 to S4, specifically as follows: S1, obtain classroom video stream data and teacher-student audio stream data, perform image enhancement, noise filtering, and video frame segmentation on the classroom video stream data to generate a standardized image sequence, and perform noise reduction, sound source localization, and speech segmentation on the teacher-student audio stream data to generate a standardized speech sequence; Among them, the classroom video stream data refers to the continuous dynamic video information in the classroom scene, which exists in the form of a data stream and contains various visual elements in the classroom, such as the actions and expressions of teachers and students, the classroom environment, etc. The teacher-student audio stream data refers to the continuous audio information in the classroom scene, which is presented in the form of a data stream and covers the teaching voices of teachers, the speaking voices of students, and various audio contents such as voice interactions in the classroom. These audio stream data can reflect the speech interaction situation between teachers and students in the classroom and provide a data basis for analyzing aspects such as language expression and teacher-student communication in the classroom.

[0021] In addition, the standardized image sequence refers to a series of images with unified specifications and standards generated after performing image enhancement, noise filtering, and video frame segmentation processing on the obtained classroom video stream data. This standardized processing makes the images consistent in terms of resolution, format, brightness, contrast, etc., facilitating subsequent analysis and processing of information in the images, such as the facial features and head postures of students, using relevant algorithms and models, and improving the accuracy and efficiency of analysis. The standardized speech sequence refers to a standardized and ordered speech data sequence obtained after performing operations such as noise reduction, sound source localization, and speech segmentation on the teacher-student audio stream data. The standardization process makes the speech data consistent in terms of sampling rate, quantization accuracy, duration segmentation, etc., which helps to perform temporal modeling on the speech data using a bidirectional LSTM network, so that valuable information such as the speaking duration and frequency of teachers and students can be obtained, providing a clear and standardized data basis for analyzing the speech interaction situation in the classroom.

[0022] In practical applications, for classroom video stream data, image enhancement algorithms can be used to improve visual quality such as image contrast and sharpness, noise filtering algorithms can be used to remove interferences such as Gaussian noise and salt-and-pepper noise, and video frame segmentation technology can be adopted to decompose continuous video streams into discrete image frames, generating a standardized image sequence with a unified resolution and format. For teacher-student audio stream data, noise reduction algorithms can be used to suppress environmental noise, sound source localization technology can be used to distinguish the sound source positions of teachers and students, and speech segmentation algorithms can be used to segment continuous speech into meaningful segments, forming a standardized speech sequence.

[0023] S2, use CNN to extract features from the standardized image sequence, that is, use an improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion to locate the key points of the student's face, combine it with a multi-joint nodding head pose estimation model to identify the student's attendance status, head pose and facial micro-expression, and use a bidirectional LSTM network to perform temporal modeling on the standardized speech sequence to obtain the teacher's speaking duration and frequency as well as the student's speaking duration and frequency; It can be understood that MTCNN (Multi-Task Cascaded Convolutional Networks) is a deep learning model commonly used for face detection and related key point localization, which is composed of multiple cascaded convolutional neural networks. It can quickly and accurately detect faces in images and locate the key features of faces, such as eyes, nose, mouth, etc.

[0024] Among them, the adaptive weight adjustment operation can automatically adjust the weight parameters of each part in the network according to different input data and task requirements, so that the network can be more flexible and accurate when dealing with different situations, improving the performance and adaptability of the network. For example, when facing face images with different illuminations, angles and poses, the adaptive weight adjustment operation can better capture face features. The multi-scale feature fusion operation can consider the situation that the faces in the image may have different sizes. By fusing feature information at different scales, the network can better process faces of different sizes. For example, when detecting a smaller face, the multi-scale feature fusion operation can comprehensively utilize features at different resolutions to avoid missed detection or false detection caused by scale differences, thereby improving the accuracy of face detection and key point localization. This improved MTCNN cascade network has higher accuracy and robustness in locating the key points of the student's face.

[0025] It should be noted that the key points on the student's face refer to the representative key position points on the student's face, such as the center of the eyes, the corners of the eyes, the tip of the nose, the corners of the mouth, the endpoints of the eyebrows, etc. These key points can reflect the basic structure and expression characteristics of the student's face. By locating these key points, the changes in the student's facial expressions, such as smiling and frowning, can be analyzed, and the attention state of the student can also be judged. For example, when the eyes and head posture of the student change, the positions of the facial key points will also change accordingly, providing an important basis for the analysis of the student's behavior and emotions.

[0026] For the multi-joint head pose estimation model, the joint points here refer to some key parts of the head, such as the top of the head, the connection point between the neck and the head, etc. These joint points can help determine the overall posture of the head. The head pose estimation model is a model that judges the head pose by analyzing the head features and position information in an image or video. Combining the information of multiple joint points, this model can more accurately estimate the head pose of the student, such as actions like looking up, looking down, and turning the head. For example, by analyzing the relative positions and movement conditions of the head joint points, it can be judged whether the student is concentrating on listening, such as looking up and the head being relatively stable, or doing other things, such as looking down or turning the head frequently.

[0027] For the bidirectional LSTM (Long Short-Term Memory) network, the LSTM network is a special type of recurrent neural network (RNN) that can effectively handle the long-term dependence problem in sequential data. When processing time-series data such as speech, the LSTM network can save and update information through memory units, avoiding the problem of gradient disappearance or gradient explosion that may occur in traditional RNNs. Bidirectional means that the network processes information not only from the start to the end of the sequence but also from the end to the start. Through this bidirectional processing, the context information in the sequential data can be captured more comprehensively. When processing a standardized speech sequence, the bidirectional LSTM network can consider the information before and after the speech simultaneously, more accurately model the temporal characteristics of the speech, and thus obtain information such as the speaking duration and frequency of the teacher and the speaking duration and frequency of the student.

[0028] Specifically, in terms of image feature extraction, a convolutional neural network (CNN) can be used to process the standardized image sequence. Through the multi-task cascaded convolutional neural network (MTCNN) improved based on the adaptive weight adjustment mechanism and multi-scale feature fusion strategy, multiple convolutional network modules are cascaded to accurately locate the key points of the student's face, such as the corners of the eyes, the tip of the nose, the corners of the mouth, etc. Combining with the multi-joint head pose estimation model, using computer vision technology, analyze the relative position changes of head joints from multiple angles to realize the recognition of students' attendance status, head pose detection and facial micro-expression analysis. Among them, the head pose can be looking up, looking down, turning sideways, etc., and the facial micro-expression can be smiling, frowning, etc. In terms of speech feature extraction, a bidirectional long short-term memory network (Bi-LSTM) is used to perform temporal modeling on the standardized speech sequence, and its bidirectional processing ability is used to capture the speech context information to accurately count the speaking duration and frequency of teachers and students.

[0029] S3. Construct a student individual behavior index system based on the student attendance status, the head pose, and the facial micro-expression, including the student attendance rate, the looking-up rate, the looking-down duration, and the expression change frequency, and construct a teacher-student interaction behavior index system based on the teacher's speaking duration and frequency and the student's speaking duration and frequency, including the teacher-student speaking ratio, the student question frequency, and the teacher feedback duration; Among them, the student individual behavior index system refers to a set of indexes used to measure and evaluate the behavior performance of students individually in the classroom. The teacher-student interaction behavior index system refers to a series of indexes used to describe and analyze the interaction situation between teachers and students in the classroom, which can reflect the dynamic process of classroom teaching from the perspective of verbal communication and interaction.

[0030] In practical applications, the student attendance rate is obtained by counting the ratio of the actual number of students present to the number of students supposed to be present; the looking-up rate is measured by calculating the ratio of the student's looking-up time to the total classroom duration; the looking-down duration is used to record the duration of the student's continuous looking down; the expression change frequency is determined by analyzing the number of dynamic changes of facial micro-expressions. The teacher-student speaking ratio is the ratio of the teacher's and student's speaking durations to the total classroom duration; the student question frequency is determined by counting the number of times the student actively asks questions; the teacher feedback duration is the time length of the teacher's response to the student's speech or question.

[0031] S4. Summarize the student individual behavior index system and the teacher-student interaction behavior index system to generate a multi-dimensional classroom behavior portrait, and associate the multi-dimensional classroom behavior portrait with the classroom teaching effect.

[0032] It should be noted that the multi-dimensional classroom behavior portrait is a portrait that comprehensively depicts and presents the behaviors of teachers and students in the classroom. The classroom teaching effect refers to the actual results achieved by teachers guiding students to learn during the classroom teaching process.

[0033] It is understandable that the individual student behavior index system and the teacher-student interaction behavior index system can be data-fused to generate a multi-dimensional classroom behavior portrait covering the individual student behavior characteristics and the teacher-student interaction pattern. Through statistical methods such as correlation analysis and regression analysis, an association model between the classroom behavior portrait and the classroom teaching effect is established to provide data support for teaching quality evaluation and teaching strategy optimization. Among them, the classroom teaching effect includes the degree of students' knowledge mastery, the effectiveness of ability cultivation, and the changes in emotional attitudes, etc.

[0034] In the specific implementation process, obtaining the classroom video stream data and the teacher-student audio stream data, and performing image enhancement, noise filtering, and video frame segmentation on the classroom video stream data to generate a standardized image sequence, and performing noise reduction, sound source localization, and speech segmentation on the teacher-student audio stream data to generate a standardized speech sequence specifically includes: Deploy multi-angle cameras and array microphones in the target teaching area to form a distributed sensor network; Based on the distributed sensor network, the classroom video stream data with a resolution of more than 720P is captured in real time through the RTSP protocol, and the classroom audio stream data with a sampling rate of 48kHz is synchronously collected to complete the spatio-temporal alignment acquisition of the classroom video stream data and the classroom audio stream data; Perform image enhancement, noise filtering, and video frame segmentation on the classroom video stream data, that is, improve the scene contrast through the adaptive histogram equalization method; based on the object detection model, the prior knowledge of seat layout, and the bilateral filtering algorithm, remove the teacher image noise and retain the student image information to generate the standardized image sequence; Perform noise reduction, sound source localization, and speech segmentation on the classroom audio stream data, that is, use the Wiener filtering algorithm to remove the environmental noise, and combine the short-time energy detection method to filter out the silent segments; based on the beamforming technology, locate the sound source direction, train the teacher-student voiceprint features through the GMM-UBM model, and automatically separate the teacher voice signal and the student voice signal; use the voice activity detection algorithm to segment the effective speech segments to generate the standardized speech sequence with timestamps.

[0035] In the actual teaching area, a distributed sensor network can be constructed, and multiple high-definition cameras and array microphones with different perspectives are deployed. Among them, the multi-angle cameras can cover the entire classroom panorama to ensure the complete capture of the teacher-student behavior; the array microphones adopt a multi-channel sound pickup design, and the spatial array layout enhances the directivity and signal-to-noise ratio of sound collection, providing a basis for subsequent sound source localization and speech separation.

[0036] Furthermore, based on a distributed sensor network, the real-time capture of classroom video stream data can be achieved using the Real-Time Streaming Protocol (RTSP). At this time, the acquisition resolution is not lower than 720P to ensure that the details of the picture are clearly distinguishable. Meanwhile, the classroom audio stream data is synchronously acquired at a sampling rate of 48 kHz, which meets the CD audio quality standard and can completely retain the high-frequency details of the voice signal. Through the hardware timestamp synchronization technology and the time calibration algorithm, the video and audio data streams can be accurately aligned in space and time to ensure that they strictly correspond in the time dimension and avoid the problem of out-of-sync audio and video caused by transmission delay, laying a foundation for subsequent joint analysis.

[0037] During the preprocessing of the video stream data, for image enhancement processing, the Contrast Limited Adaptive Histogram Equalization (CLAHE) method is used to process the video frames. This algorithm dynamically adjusts the image contrast through local histogram equalization, effectively improving the picture clarity in scenes with uneven illumination and enhancing the visual distinguishability between people and the background. For noise filtering processing, based on a deep learning object detection model such as the YOLO series, the teacher and student areas in the video frames are identified, and combined with the prior knowledge of the seat layout, the student positions are located. Through the bilateral filtering algorithm, noise reduction processing can be performed on the area where the teacher is located. This algorithm can retain the image edge information while smoothing the noise, avoiding excessive blurring that affects the recognition of students' facial features. Through selective noise reduction, while removing the noise interference in the teacher's image, the complete information of the student image can be retained to the greatest extent. For video frame segmentation processing, the continuous video stream can be segmented into independent image frames at a fixed frame rate, such as 30 fps, so as to unify the image size and format and generate a standardized image sequence, providing a standardized data input for subsequent computer vision analysis.

[0038] In the process of preprocessing audio stream data, noise reduction processing uses the Wiener filtering algorithm to reduce noise in the audio stream. This algorithm is based on the minimum mean square error criterion and dynamically adjusts the filtering parameters according to the statistical characteristics of the signal and noise, effectively suppressing environmental background noise. Through the short-time energy detection method, silent segments in the audio can be automatically identified and removed, reducing the amount of invalid data. Sound source localization processing uses beamforming technology to perform spatial processing on the audio signal. By adjusting the weighting coefficients of each channel of the array microphone, a beam pointing in a specific direction is formed to achieve accurate localization of the sound source direction. Based on the Gaussian mixture model-universal background model (GMM-UBM), voiceprint features of teachers and students are trained to construct a personalized acoustic model to achieve automatic separation of voice signals from different sound sources. Voice segmentation processing uses the voice activity detection (VAD) algorithm. By analyzing features such as the short-time energy and zero-crossing rate of the audio signal, the start and end boundaries of the voice are automatically detected, and the continuous voice stream is segmented into independent valid voice segments. An accurate timestamp is added to each voice segment to generate a standardized voice sequence, which is convenient for subsequent time series modeling and interaction analysis.

[0039] Through the above multi-step and multi-technology fusion processing flow, high-quality acquisition and standardized preprocessing of classroom video stream and teacher-student audio stream data can be achieved, providing a reliable data basis for subsequent classroom behavior analysis based on deep learning.

[0040] Removing the teacher image noise based on the object detection model, prior knowledge of seat layout and bilateral filtering algorithm, retaining the student image information, and generating the standardized image sequence specifically includes: For the classroom video stream data, based on the object detection model, the teacher's human body area is identified to generate a teacher mask ; Based on the prior knowledge of the seat layout and combined with the background difference method, the student's human body area is extracted to generate a student mask , completing the spatial separation of the teacher and student images in the classroom video stream data; Using the bilateral filtering algorithm to perform denoising processing on the teacher mask , generating the denoised teacher area, and the corresponding calculation formula is as follows: In the formula, represents the pixel p in the denoised teacher area; represents the normalization factor at pixel p; N(p) represents the spatial neighborhood of pixel p; represents the spatial standard deviation. If the details in the teacher mask are retained, then ; if enhanced denoising is performed on the teacher mask , then ; represents the Gaussian weight function in the spatial domain, which is used to control the spatial influence range of neighboring pixels, and ; represents the standard deviation of the value range, , represents the standard deviation of the pixel values in the teacher's face region; represents the pixel p in the teacher's original image region; represents the pixel q in the teacher's original image region; represents the Gaussian weight function in the value range, which is used to control the influence of color similarity of neighboring pixels, and ; directly retain the student mask covered student original image region , and the denoised teacher region and the student original image region are fused according to the corresponding masks to generate the standardized image sequence. The corresponding calculation formula is as follows: In the formula, represents the standardized image sequence; represents the teacher mask; represents the denoised teacher region; represents the student mask; represents the student original image region.

[0041] Among them, first, an object detection model can be used to identify the teacher's human body region in the classroom video stream data, generating a teacher mask that can mark the teacher's position. Secondly, based on the pre-known seat layout information and combined with the background difference method, the student region can be extracted to form a student mask, thereby realizing the spatial distinction between the teacher and student images.

[0042] Then, for the teacher mask, a bilateral filtering algorithm can be used to perform denoising operations. The bilateral filtering algorithm considers the spatial distance and color similarity between pixels through two Gaussian weight functions, and sets the spatial standard deviation according to different requirements. If the details of the teacher mask are to be retained, the spatial standard deviation is set to a smaller value; if the denoising effect needs to be enhanced, the spatial standard deviation is set to a larger value. The standard deviation of the value range is determined according to the standard deviation of the pixel values in the teacher's face region, and the denoised teacher region is obtained through formula calculation.

[0043] Finally, the student original image region covered by the student mask can be directly retained, and the denoised teacher region and the student original image region are accurately fused according to the teacher mask and the student mask, finally generating a standardized image sequence, providing a standardized and clear data basis for subsequent classroom behavior analysis.

[0044] In the process of denoising the teacher mask using the bilateral filtering algorithm, a motion compensation mechanism is introduced, which specifically includes: During the process of denoising the using the bilateral filtering algorithm, a motion compensation mechanism is introduced, specifically including: For the motion blur noise in the teacher mask a motion compensation term is added to the spatial domain Gaussian weight function , where represents the motion vector of pixel p; represents the motion vector of pixel q, then the updated spatial domain Gaussian weight function is .

[0045] It can be understood that when using the bilateral filtering algorithm to denoise the teacher mask, in order to effectively process the motion blur noise therein, a motion compensation mechanism can be introduced. The bilateral filtering originally realizes denoising by comprehensively considering the spatial distance and color similarity between pixels through the Gaussian weight functions in the spatial domain and the value domain. The motion compensation mechanism adds a motion compensation term to the spatial domain Gaussian weight function, and this compensation term is constructed based on the motion vectors of pixels.

[0046] In addition, the motion vector of a pixel represents the motion direction and distance of the pixel between video frames. By incorporating the motion vector of the pixel into the calculation of the spatial domain Gaussian weight function, the updated function can perceive and utilize the motion information of the pixel. When processing the teacher mask with motion blur noise, the spatial domain Gaussian weight function considering the motion factor can more accurately weight adjacent pixels, better retain the details and edges of the image while removing noise, and improve the denoising effect and image quality.

[0047] The CNN is used to extract features from the standardized image sequence, that is, the improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion is used to locate the key points of the student's face, and combined with the multi-joint nodding head pose estimation model, the student's attendance status, head pose and facial micro-expression are recognized, specifically including: Through the improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, according to the complexity of the facial region in the standardized image sequence, the weights of each layer of the network are dynamically allocated, and at the same time, the facial feature information at different scales is fused to locate 68 key points of the student's face; The multi-joint nodding head pose estimation model is synchronously accessed to calculate the three-dimensional pose parameters of the student's head in real time, and the multi-joint nodding head pose estimation model at least includes a bone point detection module and an angle calculation module based on a convolutional neural network; A two-layer detection mechanism of the facial presence state and the head pose matrix is established, and based on the key points of the student's face and the three-dimensional pose parameters of the student's head, the student's attendance status, the head pose and the facial micro-expression are recognized.

[0048] In terms of student facial key point localization, an improved multi-task cascaded convolutional neural network based on adaptive weight adjustment and multi-scale feature fusion can be adopted. When the traditional multi-task cascaded convolutional neural network processes facial images of different complexities, the weights of each layer of the network are fixed, which may lead to poor feature extraction effects. The improved network has an adaptive ability and can dynamically allocate weights to each layer of the network according to the complexity of the facial region in the standardized image sequence, such as different lighting conditions, occlusion situations, or pose changes. When the lighting is strong, resulting in complex light and dark contrast on the face, or there are occlusions such as glasses or hair, the network will automatically increase the weights of the key layers to better extract facial features. At the same time, the network fuses facial feature information at different scales. At a small scale, it can capture the overall contour of the face, and at a large scale, it can obtain more detailed information. Through multi-scale feature fusion, whether it is a small far-shot face or a large near-shot face, 68 key points of the student's face can be accurately located, including the key coordinate points of parts such as eyes, eyebrows, nose, and mouth.

[0049] When estimating the head pose, a multi-joint head pose estimation model can be synchronously accessed. This model includes a bone point detection module and an angle calculation module based on a convolutional neural network. Specifically, the bone point detection module uses the powerful feature extraction ability of the convolutional neural network to identify the position information of the key points of the head bones; the angle calculation module then calculates the three-dimensional pose parameters of the student's head in real time based on the detected bone points, such as pitch angle, yaw angle, and roll angle, so as to determine the pose of the head in three-dimensional space.

[0050] In order to accurately identify the student's attendance status, head pose, and facial micro-expression, a double-layer detection mechanism of the facial presence status and the head pose matrix can be established. Through this mechanism, on the one hand, based on the located key points of the student's face, it can be judged whether the student's face is completely presented in the picture to identify the student's attendance status; on the other hand, combined with the three-dimensional pose parameters of the student's head and the facial key point information, the head pose can be analyzed. For example, when the pitch angle of the head exceeds a certain threshold, it is judged as a bowed head state, and when it is lower than this threshold, it is a raised head state. At the same time, by observing the relative position changes between the facial key points, such as the corners of the mouth rising or the eyebrows frowning, the facial micro-expression can be identified, and thus the behavior and emotional state of the student in the classroom can be comprehensively understood, providing rich and accurate feature data for classroom behavior analysis.

[0051] The use of a bidirectional LSTM network to perform temporal modeling on the standardized speech sequence to obtain the teacher's speaking duration and frequency as well as the student's speaking duration and frequency specifically includes: Convert the standardized speech sequence into a Mel Frequency Cepstral Coefficient (MFCC) sequence, and input the MFCC sequence into the bidirectional LSTM network for temporal modeling; Extract the speech energy, zero-crossing rate, and fundamental frequency features from the MFCC sequence through a preset sliding window, and combine the attention mechanism to obtain the key speaking periods; Construct a teacher-student speech classifier to distinguish between the teacher's lecture periods and the students' interaction periods during the key speaking periods, and output the teacher's speaking duration and frequency, as well as the students' speaking duration and frequency.

[0052] Among them, the Mel Frequency Cepstral Coefficient is a feature parameter widely used in the field of speech processing. It performs frequency-domain analysis of the speech signal on the Mel scale by simulating the characteristics of the human auditory system, and then obtains the cepstrum through cepstrum calculation. This conversion can effectively extract the important features related to human ear perception in the speech signal and remove redundant information.

[0053] After the conversion, the MFCC sequence can be input into a bidirectional long short-term memory network for temporal modeling. As a special type of recurrent neural network, the bidirectional long short-term memory network is unique in that it can process information from both the start and end directions of the speech sequence, thereby capturing the context information before and after the speech more comprehensively. It overcomes the problem of gradient disappearance or gradient explosion that traditional recurrent neural networks are prone to when dealing with long sequence data, providing strong support for accurately analyzing the temporal features of speech.

[0054] Secondly, the speech energy, zero-crossing rate, and fundamental frequency features can be extracted from the MFCC sequence through a preset sliding window. The sliding window slides on the sequence data at a fixed step size, and each time a segment of data is intercepted for feature extraction. Specifically, the speech energy reflects the strength of the speech signal, the zero-crossing rate represents the number of times the speech signal crosses the zero axis per unit time, and the fundamental frequency is related to the pitch of the speech. These features describe the characteristics of the speech signal from different perspectives. Combining the attention mechanism, the bidirectional LSTM network can automatically focus on the key speaking periods that are important for analyzing the teacher-student speech, ignoring irrelevant or secondary speech segments, and improving the model processing efficiency and analysis accuracy. For example, in a speech sequence containing teacher-student conversations and classroom background noise, the attention mechanism can help the network identify the truly informative teacher-student speech parts and exclude noise interference.

[0055] Finally, a teacher-student speech classifier can be constructed to classify key speech periods. This classifier is trained based on deep learning or machine learning algorithms and can accurately distinguish between the teacher's teaching period and the student interaction period during the key speech period according to the extracted speech features. When the teacher is teaching, the speech features may have specific rhythms and pitch patterns; the speech features during the student interaction period are different in terms of volume, speech rate, etc. By learning these differences, the classifier can achieve accurate classification and output the teacher's speech duration and frequency as well as the student's speech duration and frequency. These data provide a quantitative basis for analyzing the teacher-student interaction pattern and teaching participation in the classroom, which helps to deeply understand the dynamics of classroom teaching.

[0056] Summarize the individual student behavior index system and the teacher-student interaction behavior index system to generate a multi-dimensional classroom behavior portrait, and correlate the multi-dimensional classroom behavior portrait with the classroom teaching effect, specifically including: Perform feature normalization on the individual student behavior index system and the teacher-student interaction behavior index system to generate a classroom behavior index set X; Obtain the preset number of clusters, use the K-means algorithm to group the classroom behavior index set X to obtain multiple clusters, calculate the center vectors of each cluster and construct the multi-dimensional classroom behavior portrait Y; Establish a classroom teaching effect index set E, use a multiple linear regression model to fit the linear relationship between the multi-dimensional classroom behavior portrait Y and the classroom teaching effect index set E, and screen out significant correlation indicators through t-tests; Introduce the SHAP value to calculate the contribution degree of each index in the classroom behavior index set X to the classroom teaching effect, sort the indexes in the classroom behavior index set X from largest to smallest according to the contribution degree, and select the top A indexes as the key classroom behavior indexes, where A represents the preset number of key indexes; Calculate the Pearson correlation coefficient between the key classroom behavior indexes and the classroom teaching effect and construct a behavior-effect correlation heat map.

[0057] In practical applications, due to the differences in the value ranges and dimensions of different indexes, for example, the student attendance rate is a percentage value, while the teacher feedback duration is measured in minutes, directly using the original data will affect the accuracy of analysis and the performance of the model. Therefore, through normalization processing, all indexes can be converted into a unified numerical interval to eliminate the influence of dimensions and generate a classroom behavior index set containing various classroom behavior data.

[0058] Subsequently, the K-means algorithm can be used to perform clustering analysis on the set of classroom behavior indicators. Before the clustering operation, the number of clusters needs to be preset, which is determined based on the complexity of the classroom behavior data and the analysis purpose. The K-means algorithm uses distance as the metric standard, divides similar classroom behavior data into the same cluster, and finally obtains multiple groups with similar behavior characteristics by continuously iteratively calculating the center vectors of each cluster. These groups and their center vectors jointly construct a multi-dimensional classroom behavior portrait, which can visually present different behavior patterns and feature distributions in the classroom.

[0059] Subsequently, a set of classroom teaching effect indicators can be established, which covers teaching effect evaluation indicators in multiple dimensions such as knowledge mastery, ability cultivation, and emotional attitude changes. Through a multiple linear regression model, the linear relationship between the multi-dimensional classroom behavior portrait and the set of classroom teaching effect indicators can be fitted, and the correlation trend between the two can be described through model parameter estimation. At the same time, a t-test can be used to conduct a significance test on each indicator in the regression model, screen out the classroom behavior indicators that have a significant impact on the classroom teaching effect, and exclude the interference of irrelevant or less influential indicators.

[0060] Furthermore, the SHAP value can be introduced to quantify the contribution degree of each indicator in the set of classroom behavior indicators to the classroom teaching effect. It should be noted that the SHAP value is based on game theory principles and can fairly allocate the importance of each feature in model prediction. By calculating the SHAP values of each indicator and sorting the indicators in the set of classroom behavior indicators from largest to smallest according to the contribution degree, a preset number of indicators are selected as key classroom behavior indicators. These key indicators play an important role in influencing the classroom teaching effect.

[0061] Finally, the Pearson correlation coefficient between the key classroom behavior indicators and the classroom teaching effect can be calculated, which is used to measure the linear correlation degree between two variables. Based on the calculated correlation coefficient, a behavior-effect correlation heat map can be constructed to visually display the correlation strength and direction between the key classroom behavior indicators and the classroom teaching effect. And the depth of color in the heat map represents the size of the correlation coefficient, the darker the color, the stronger the correlation, so that the classroom behaviors that have a greater impact on the teaching effect can be intuitively understood, providing a scientific basis for teaching improvement and optimization.

[0062] As Figure 2 shown, it is the system block diagram of the classroom behavior analysis system provided by the embodiment of the present invention. The analysis system includes: A data acquisition module, configured to acquire classroom video stream data and teacher-student audio stream data, perform image enhancement, noise filtering, and video frame segmentation on the classroom video stream data to generate a standardized image sequence, and perform noise reduction, sound source localization, and speech segmentation on the teacher-student audio stream data to generate a standardized speech sequence; A data analysis module, configured to extract features from the standardized image sequence using a CNN, that is, locate key points of students' faces through an improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, combine with a multi-joint nodding head pose estimation model to identify students' attendance status, head poses, and facial micro-expressions, and use a bidirectional LSTM network to perform temporal modeling on the standardized speech sequence to obtain the speaking duration and frequency of teachers and the speaking duration and frequency of students; An index construction module, configured to construct a student individual behavior index system based on the students' attendance status, head poses, and facial micro-expressions, including student attendance rate, head-up rate, head-down duration, and expression change frequency, and construct a teacher-student interaction behavior index system based on the speaking duration and frequency of teachers and the speaking duration and frequency of students, including the proportion of teacher-student speeches, the frequency of students' questions, and the feedback duration of teachers; A summary and association module, configured to summarize the student individual behavior index system and the teacher-student interaction behavior index system to generate a multi-dimensional classroom behavior portrait, and associate the multi-dimensional classroom behavior portrait with the classroom teaching effect.

[0063] Figure 2 The device of the illustrated embodiment can correspondingly be used to execute Figure 1 the steps in the method embodiment shown, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0064] An electronic device includes a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the steps of the classroom behavior analysis method described in any one of the above.

[0065] As Figure 3 shown, it is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. The electronic device 30 includes: a processor 31, a memory 32, and a computer program; wherein The memory 32 is used to store the computer program, and the memory can also be a flash memory. The computer program is, for example, an application program, a functional module, etc. that implement the above method.

[0066] The processor 31 is used to execute the computer program stored in the memory to implement each step executed by the device in the above method. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.

[0067] Optionally, the memory 32 can be either independent or integrated with the processor 31.

[0068] When the memory 32 is a device independent of the processor 31, the device may further include: A bus 33 for connecting the memory 32 and the processor 31.

[0069] A readable storage medium stores a computer program, and when the computer program is executed by a processor, it is used to implement the steps of the classroom behavior analysis method described in any one of the above.

[0070] Among them, the readable storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transfer of a computer program from one place to another. The computer storage medium can be any available medium that can be accessed by a general or special purpose computer. For example, the readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). In addition, the ASIC can be located in the user equipment. Of course, the processor and the readable storage medium can also exist as discrete components in the communication device. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0071] The present invention also provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of the device can read the execution instructions from the readable storage medium, and the execution of the execution instructions by at least one processor causes the device to implement the methods provided by the above various embodiments.

[0072] In an embodiment of the above device, it should be understood that the processor may be a central processing unit (CPU for short), or other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in conjunction with the present invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0073] Through the introduction of the above embodiments, the present invention can obtain classroom video stream data and teacher-student audio stream data through a classroom behavior analysis method, system, electronic device, and storage medium, and perform image enhancement, noise filtering, and video frame segmentation on the classroom video stream data to generate a standardized image sequence, and perform noise reduction, sound source localization, and speech segmentation on the teacher-student audio stream data to generate a standardized speech sequence; use CNN to extract features from the standardized image sequence, that is, locate the key points of the students' faces through an improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, and combine a multi-joint nodding head pose estimation model to identify the students' attendance status, head pose, and facial micro-expressions, and use a bidirectional LSTM network to perform temporal modeling on the standardized speech sequence to obtain the teacher's speaking duration and frequency, as well as the students' speaking duration and frequency; construct a student individual behavior index system based on the students' attendance status, head pose, and facial micro-expressions, including the student attendance rate, head-up rate, head-down duration, and expression change frequency, and construct a teacher-student interaction behavior index system based on the teacher's speaking duration and frequency, as well as the students' speaking duration and frequency, including the teacher-student speaking ratio, student question frequency, and teacher feedback duration; summarize the student individual behavior index system and the teacher-student interaction behavior index system to generate a multi-dimensional classroom behavior portrait, and associate the multi-dimensional classroom behavior portrait with the classroom teaching effect, so as to effectively improve the accuracy of classroom behavior analysis, enhance the data processing ability, and realize the comprehensive evaluation of multi-dimensional classroom behavior.

[0074] The present invention improves the MTCNN cascade network, combines pose estimation technology, and constructs a two-layer mechanism of face detection and pose assistance, which can effectively solve the recognition problems in complex scenarios such as low light and face occlusion, and provide a reliable basis for accurately evaluating the learning state. Through a distributed sensor network, combined with a bilateral filtering algorithm and a motion compensation mechanism, the present invention can achieve high-quality acquisition and preprocessing of classroom data, while supporting real-time streaming processing and batch analysis of classroom data, improving the data processing efficiency and meeting the low-latency requirements of large-scale education scenarios. The present invention can construct a multi-dimensional index system including individual behaviors and teacher-student interactions, generate student behavior portraits through clustering analysis, and quantify the correlation between classroom behaviors and teaching effects by combining statistical modeling, providing scientific guidance for teaching intervention. The present invention adopts a full-process automated system design, supports rapid deployment in multiple scenarios such as smart classrooms and online education, is compatible with multi-terminal display and distributed computing mechanisms, and can adapt to different classroom environments without manual calibration, significantly improving the intelligent level and resource adaptation ability of education management.

[0075] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0076] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other medium that can be used to carry or store data and is readable by a computer.

[0077] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent in such a process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, commodity or device comprising the element.

Claims

1. A method for analyzing classroom behavior, characterized in that, The analysis method includes: Obtaining classroom video stream data and teacher-student audio stream data, performing image enhancement, noise filtering, and video frame segmentation on the classroom video stream data to generate a standardized image sequence, and performing noise reduction, sound source localization, and speech segmentation on the teacher-student audio stream data to generate a standardized speech sequence; Using a CNN to extract features from the standardized image sequence, that is, locating the key points of students' faces through an improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, combining with a multi-joint nodding head pose estimation model to identify the arrival status, head pose, and facial micro-expressions of students, and using a bidirectional LSTM network to perform temporal modeling on the standardized speech sequence to obtain the speaking duration and frequency of teachers and the speaking duration and frequency of students; Constructing a student individual behavior index system based on the student arrival status, the head pose, and the facial micro-expressions, including student attendance rate, head-up rate, head-down duration, and expression change frequency, and constructing a teacher-student interaction behavior index system based on the speaking duration and frequency of teachers and the speaking duration and frequency of students, including the proportion of teacher-student speech, the frequency of student questions, and the teacher feedback duration; Summarizing the student individual behavior index system and the teacher-student interaction behavior index system to generate a multi-dimensional classroom behavior portrait, and correlating the multi-dimensional classroom behavior portrait with the classroom teaching effect.

2. The method for analyzing classroom behavior according to claim 1, characterized in that, The obtaining of the classroom video stream data and the teacher-student audio stream data, and performing image enhancement, noise filtering, and video frame segmentation on the classroom video stream data to generate a standardized image sequence, and performing noise reduction, sound source localization, and speech segmentation on the teacher-student audio stream data to generate a standardized speech sequence specifically includes: Deploying multi-angle cameras and array microphones in the target teaching area to form a distributed sensor network; Based on the distributed sensor network, capturing the classroom video stream data with a resolution above 720P in real time through the RTSP protocol, and synchronously collecting the classroom audio stream data with a sampling rate of 48kHz to complete the spatio-temporal aligned acquisition of the classroom video stream data and the classroom audio stream data; Performing image enhancement, noise filtering, and video frame segmentation on the classroom video stream data, that is, improving the scene contrast through the adaptive histogram equalization method; removing teacher image noise based on the object detection model, prior knowledge of seat layout, and bilateral filtering algorithm, and retaining student image information to generate the standardized image sequence; Performing noise reduction, sound source localization, and speech segmentation on the classroom audio stream data, that is, using the Wiener filtering algorithm to remove environmental noise, combining with the short-time energy detection method to filter out silent segments; locating the sound source direction based on beamforming technology, training the voiceprint features of teachers and students through the GMM-UBM model, and automatically separating the teacher speech signal and the student speech signal; using the speech activity detection algorithm to segment the effective speech segments to generate the standardized speech sequence with timestamps.

3. The method for analyzing classroom behavior according to claim 2, characterized in that, The removing of teacher image noise based on the object detection model, prior knowledge of seat layout, and bilateral filtering algorithm, and retaining student image information to generate the standardized image sequence specifically includes: For the classroom video stream data, identify the teacher's human body region based on the object detection model to generate a teacher mask ; Based on the prior knowledge of the seat layout and combined with the background difference method, extract the student body region and generate a student mask , and complete the spatial separation of the teacher and student images in the classroom video stream data; Apply the bilateral filtering algorithm to the teacher mask for denoising to generate a denoised teacher region, and the corresponding calculation formula is as follows: In the formula, represents pixel p in the denoised teacher region; represents the normalization factor at pixel p; N(p) represents the spatial neighborhood of pixel p; represents the spatial standard deviation. If the details in the teacher mask are retained, then ; if enhanced denoising is performed on the teacher mask , then ; represents the spatial domain Gaussian weight function, which is used to control the spatial influence range of neighboring pixels, and ; represents the value domain standard deviation, , represents the standard deviation of the pixel values in the teacher's face region; represents pixel p in the teacher's original image region; represents pixel q in the teacher's original image region; represents the value domain Gaussian weight function, which is used to control the influence of color similarity of neighboring pixels, and ; Directly retain the student mask Covered original student image area , and fuse the denoised teacher area with the original student image area according to the corresponding mask to generate the standardized image sequence, and the corresponding calculation formula is as follows: In the formula, represents the standardized image sequence; represents the teacher mask; represents the denoised teacher region; represents the student mask; represents the student's original image region.

4. The method for analyzing classroom behavior according to claim 3, characterized in that, When using the bilateral filtering algorithm to denoise the teacher mask During the process, a motion compensation mechanism is introduced, specifically including: For the motion blur noise in the teacher mask a motion compensation term is added to the spatial domain Gaussian weight function , where represents the motion vector of pixel p; represents the motion vector of pixel q, then the updated spatial domain Gaussian weight function is .

5. The method for analyzing classroom behavior according to claim 1, characterized in that, Using CNN to extract features from the standardized image sequence, that is, locating the key points of the student's face through an improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, and combining with a multi-joint nodding head pose estimation model to identify the student's attendance status, head pose, and facial micro-expression, specifically including: Through the improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, dynamically allocate the weights of each layer of the network according to the complexity of the facial region in the standardized image sequence, and at the same time fuse the facial feature information at different scales to locate 68 key points of the student's face; Synchronously access the multi-joint nodding head pose estimation model to calculate the three-dimensional pose parameters of the student's head in real time, and the multi-joint nodding head pose estimation model at least includes a bone point detection module and an angle calculation module based on a convolutional neural network; Establish a double-layer detection mechanism for the facial presence status and the head pose matrix, and identify the student's attendance status, the head pose, and the facial micro-expression based on the key points of the student's face and the three-dimensional pose parameters of the student's head.

6. The method for analyzing classroom behavior according to claim 1, characterized in that, Using a bidirectional LSTM network to perform temporal modeling on the standardized speech sequence to obtain the teacher's speaking duration and frequency, as well as the student's speaking duration and frequency, specifically including: Convert the standardized speech sequence into a Mel frequency cepstral coefficient sequence, and input the Mel frequency cepstral coefficient sequence into the bidirectional LSTM network for temporal modeling; Extract the speech energy, zero-crossing rate, and fundamental frequency features in the Mel frequency cepstral coefficient sequence through a preset sliding window, and combine the attention mechanism to obtain the key speaking period; Construct a teacher-student speech classifier to distinguish the teacher's teaching period and the student's interaction period in the key speaking period, and output the teacher's speaking duration and frequency, as well as the student's speaking duration and frequency.

7. The method for analyzing classroom behavior according to claim 1, characterized in that, Summarize the student individual behavior index system and the teacher-student interaction behavior index system to generate a multi-dimensional classroom behavior portrait, and correlate the multi-dimensional classroom behavior portrait with the classroom teaching effect, specifically including: Perform feature normalization processing on the student individual behavior index system and the teacher-student interaction behavior index system to generate a classroom behavior index set X; Obtain the preset number of clusters, use the K-means algorithm to group the classroom behavior index set X to obtain multiple clusters, calculate the center vectors of each cluster, and construct the multi-dimensional classroom behavior portrait Y; Establish a classroom teaching effect index set E, use a multiple linear regression model to fit the linear relationship between the multi-dimensional classroom behavior portrait Y and the classroom teaching effect index set E, and screen out significant correlation indicators through t-tests; Introduce SHAP values to calculate the contribution degree of each index in the classroom behavior index set X to the classroom teaching effect, sort the indexes in the classroom behavior index set X from largest to smallest according to the contribution degree, and select the top A indexes as the key classroom behavior indexes, where A represents the preset number of key indexes; Calculate the Pearson correlation coefficient between the key classroom behavior indexes and the classroom teaching effect, and construct a behavior-effect correlation heat map.

8. A classroom behavior analysis system, applied to the classroom behavior analysis method according to any one of claims 1-7, characterized in that, The analysis system includes: A data acquisition module, configured to acquire classroom video stream data and teacher-student audio stream data, perform image enhancement, noise filtering, and video frame segmentation on the classroom video stream data to generate a standardized image sequence, and perform noise reduction, sound source localization, and speech segmentation on the teacher-student audio stream data to generate a standardized speech sequence; A data analysis module, configured to use a CNN to extract features from the standardized image sequence, that is, locate the key points of the students' faces through an improved MTCNN cascade network based on adaptive weight adjustment and multi-scale feature fusion, combine a multi-joint nodding head pose estimation model to identify the students' attendance status, head pose, and facial micro-expressions, and use a bidirectional LSTM network to perform temporal modeling on the standardized speech sequence to obtain the speaking duration and frequency of the teacher and the speaking duration and frequency of the students; An index construction module, configured to construct a student individual behavior index system based on the students' attendance status, head pose, and facial micro-expressions, including student attendance rate, head-up rate, head-down duration, and expression change frequency, and construct a teacher-student interaction behavior index system based on the teacher's speaking duration and frequency and the students' speaking duration and frequency, including teacher-student speaking ratio, student question frequency, and teacher feedback duration; A summary and association module, configured to summarize the student individual behavior index system and the teacher-student interaction behavior index system to generate a multi-dimensional classroom behavior portrait, and associate the multi-dimensional classroom behavior portrait with the classroom teaching effect.

9. An electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the processor runs the computer program stored in the memory, the processor executes the steps of the classroom behavior analysis method according to any one of claims 1-7.

10. A readable storage medium, wherein a computer program is stored in the readable storage medium, characterized in that, When the computer program is executed by the processor, it is used to implement the steps of the classroom behavior analysis method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Virtual teaching management method and device based on artificial intelligence

    CN115936944A

  • Bimodal astronaut emotion recognition method based on facial expressions and voices in video

    CN116386101A

  • Classroom video-based AI multi-dimensional teaching behavior analysis method and system

    CN118658128A

  • Classroom teaching effect evaluation system based on voice multi-feature progressive embedding

    CN118782096A

  • Intelligent evaluation method for real-time learning and teaching effect feedback

    CN119941053A

Cited By

  • Multimedia conference room sound system based on artificial intelligence

    CN120812477A

  • Student classroom performance evaluation method and system based on multi-modal audio and video

    CN120953873A

  • Intelligent teaching evaluation and diagnosis system and method based on multi-mode audio and video analysis

    CN120997010A

  • Intelligent teaching evaluation and diagnosis system and method based on multi-modal audio and video analysis

    CN120997010B

  • Classroom teaching interaction method and device, electronic equipment and storage medium

    CN121301425A