Classroom abnormal behavior intelligent detection and teaching intervention system
Through multimodal data acquisition and processing technology, combined with image and sound signals, accurate detection and teaching intervention in students' classroom behavior is achieved, and the problems of strong subjectivity and environmental interference in traditional classroom management are solved, and teaching efficiency and management effect are improved.
Patent Information
- Application Number
- CN202510405470.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional classroom behavior assessment is highly subjective, difficult to pay attention to multiple students at the same time, is susceptible to environmental interference, and fails to effectively deal with students' persistent abnormal behavior, resulting in low teaching management efficiency.
Multimodal data acquisition and processing technology is adopted, and time synchronization and feature extraction are performed through the combination of image and sound signals, abnormal behavior detection is performed using behavioral state classification model, and teaching intervention instructions are generated.
Quantitative evaluation and precise classification of students' classroom behaviors have been achieved, timely reminders or interventions have been made, and classroom teaching efficiency and management effectiveness have been improved.
Smart Images

Figure CN120279596A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and particularly relates to an intelligent detection and teaching intervention system for classroom abnormal behaviors. Background Art
[0002] With the development of image recognition technology, classroom behavior analysis technology has emerged. It can capture students' expressions in real time through computer vision and deep learning, and analyze them through intelligent algorithms to identify students' classroom states, such as concentration, distraction, and dozing, providing new possibilities for intelligent classroom management and personalized teaching.
[0003] In traditional technologies, teachers mainly rely on their own observations and experiences for classroom discipline management and teaching adjustment. There may be significant differences in the management styles and judgment criteria of different teachers, resulting in strong subjectivity in classroom behavior evaluation. It is difficult for teachers to simultaneously pay attention to the behaviors of multiple students. Especially in a large-class teaching environment, it is easy to miss some behaviors, and relying on teacher reminders is likely to interrupt the classroom teaching rhythm.
[0004] Existing classroom analysis technologies usually analyze students' behaviors based on a single modality. Single-modality methods are easily interfered by the environment. For example, simple facial expression detection may misjudge students' concentration states, and single speech analysis may not be able to distinguish normal discussions from disturbing speeches. Moreover, they only classify instantaneous behaviors and do not consider the persistence of students' behaviors, and cannot take reasonable reminder or guiding teacher intervention measures based on students' abnormal behaviors. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide an intelligent detection and teaching intervention system for classroom abnormal behaviors that can detect and remind classroom abnormal behaviors by using the temporal features of multi-modal data.
[0006] In the first aspect, the present application provides an intelligent detection and teaching intervention system for classroom abnormal behaviors, including:
[0007] A data acquisition layer for obtaining multi-modal data in the classroom of students; the multi-modal data includes a set of images of students' classroom performances and sound signals corresponding to the time.
[0008] A data standardization layer for preprocessing and time synchronization of the multi-modal data to obtain standardized multi-modal data.
[0009] An abnormal behavior detection layer for inputting the standardized multi-modal data into a behavior state classification model to obtain a behavior classification result; the behavior classification result includes behavior type, behavior duration, and intervention level label.
[0010] An intervention decision-making layer, which is used to generate teaching intervention instructions based on the intervention decision-making logic using the behavior classification results; the teaching intervention instructions include system prompt instructions and teacher intervention reminder instructions.
[0011] In one embodiment, the multi-modal data is preprocessed and time-synchronized to obtain standardized multi-modal data, including:
[0012] Use a face detection algorithm to locate the facial regions in the set of student classroom performance images, and segment the located facial region images to obtain a facial region image matrix;
[0013] Use a pose estimation algorithm to extract the skeleton key points from the set of student classroom performance images to obtain the skeleton key point coordinates;
[0014] Filter and frame the sound signal to obtain a denoised sound signal;
[0015] Based on the time stamps, perform time alignment of the cross-modal data for the facial region image matrix, skeleton key point coordinates, and denoised sound signal to obtain standardized multi-modal data; the standardized multi-modal data includes multiple sets of facial region image matrices, skeleton key point coordinates, and denoised sound signals that match within the same time window.
[0016] In one embodiment, the abnormal behavior detection layer includes a feature extraction and fusion unit and a behavior state classification unit;
[0017] The feature extraction and fusion unit is used to extract features from the standardized multi-modal data and fuse the extracted features to obtain a behavior state feature representation;
[0018] The behavior state classification unit is used to obtain a behavior classification result based on the behavior state feature representation.
[0019] In one embodiment, extracting features from the standardized multi-modal data and fusing the extracted features to obtain a behavior state feature representation includes:
[0020] Perform temporal change extraction and facial detail extraction on the facial region image matrix to obtain facial features;
[0021] Extract the time-domain and frequency-domain features of the denoised sound signal to obtain sound features;
[0022] Extract the limb movement amplitude, frequency, and change trend based on the skeleton key point coordinates to obtain skeleton motion features;
[0023] Perform multi-modal feature fusion on the facial features, sound features, and skeleton motion features to obtain a behavior state feature representation.
[0024] In one embodiment, facial features, voice features, and skeleton motion features are fused in a multi-modal manner to obtain a representation of the behavior state, including:
[0025] Use a sliding time window to segment the facial features, voice features, and skeleton motion features, and calculate feature vectors for the facial features, voice features, and skeleton motion features within each sliding time window to obtain multi-modal raw features;
[0026] Concatenate multiple multi-modal raw features in the order of the sliding time window and perform preliminary fusion through a fully connected layer to obtain multi-modal preliminary fusion features;
[0027] Perform weighted fusion on the multi-modal preliminary fusion according to a preset weight strategy to obtain a representation of the behavior state.
[0028] In one embodiment, the behavior state classification model is obtained in the following manner:
[0029] Obtain a multi-modal dataset, which includes multiple modal data, as well as behavior types and intervention level labels corresponding one-to-one to the multi-modal data; the behavior types include normal class listening, short-term distraction, continuous distraction, dozing off, short-term speech, and continuous disruptive speech; the intervention level labels include no intervention, system automatic reminder, and prompt teacher intervention;
[0030] Extract and fuse features from the multi-modal data to obtain a time series feature vector;
[0031] Divide the time series feature vector according to a preset time window length to obtain a multi-modal feature vector sequence; the multi-modal feature vector sequence includes time windows and their corresponding multi-modal feature vectors;
[0032] Mark the behavior type and intervention level within the time window in the multi-modal feature vector sequence to obtain a training set;
[0033] Based on the self-attention mechanism, iteratively learn the mapping relationship between the multi-modal feature vectors and the behavior classification results in the training set to obtain model parameters;
[0034] Construct a state classification model according to the model parameters.
[0035] In one embodiment, based on the intervention decision logic, generate a teaching intervention instruction using the behavior classification result, including:
[0036] Calculate the abnormal behavior index according to the behavior type, behavior duration, and intervention level label;
[0037] When the abnormal behavior index reaches the first intervention range, generate a system intervention instruction; the system intervention instruction is used to instruct the system to send a reminder to the student terminal corresponding to the abnormal behavior;
[0038] When the abnormal behavior index reaches the second intervention range, a teacher intervention reminder instruction is generated; the teacher intervention reminder instruction is used to instruct the system to send the student information corresponding to the abnormal behavior and the behavior type to the corresponding teacher terminal.
[0039] In one embodiment, the system further includes a feedback collection and model adaptive adjustment layer;
[0040] The feedback collection and model adaptive adjustment layer is used to collect teacher intervention feedback, and based on a learning algorithm, use the teacher intervention feedback to adjust the weight factors in the weight strategy and the model parameters of the behavior state classification model; the teacher intervention feedback includes accepting the intervention reminder and ignoring the intervention reminder.
[0041] In one embodiment, the system further includes an information security and encrypted transmission layer;
[0042] The information security and encrypted transmission layer is used to encrypt the data for the acquisition, transmission, and storage of multimodal data.
[0043] In a second aspect, the present application also provides a method for intelligent detection of classroom abnormal behavior and teaching intervention, including:
[0044] Obtain multimodal data of students in the classroom;
[0045] Perform preprocessing and time synchronization on the multimodal data to obtain standardized multimodal data;
[0046] Input the standardized multimodal data into the behavior state classification model to obtain a behavior classification result; the behavior classification result includes the behavior type, behavior duration, and intervention level label;
[0047] Generate a teaching intervention instruction based on the intervention decision logic using the behavior classification result.
[0048] The above-mentioned intelligent detection and teaching intervention system for classroom abnormal behavior obtains multimodal data such as a set of student classroom performance images and sound signals at corresponding times through the data acquisition layer, and realizes the time synchronization of multimodal data through the data standardization layer, ensuring that in subsequent feature extraction and behavior analysis, data of different modalities can correspond to the same time point. The behavior state classification model is used to analyze the standardized multimodal data to obtain the behavior type, behavior duration, and intervention level label, so as to quantify and classify the behavior of students. System prompt instructions and teacher intervention reminder instructions are generated according to the behavior classification result, which can perform targeted interventions according to the severity and duration of students' behaviors, help teachers discover and handle students' abnormal behaviors in a timely manner, reduce the interference of students' abnormal behaviors on the classroom, improve classroom teaching efficiency, and remind students to quickly return to the normal learning state. Description of the Drawings
[0049] To more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a composition structure diagram of a classroom abnormal behavior intelligent detection and teaching intervention system of the present invention;
[0051] Figure 2 It is a flow schematic diagram of a classroom abnormal behavior intelligent detection and teaching intervention method of the present invention. Detailed implementation manners
[0052] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0053] In one embodiment, as Figure 1 shown, a classroom abnormal behavior intelligent detection and teaching intervention system is provided. In this embodiment, taking the application of the system to a terminal as an example, it can be understood that the system can also be applied to a server, and can also be applied to include a terminal and a server, and is realized through the interaction between the terminal and the server. In this embodiment, it includes:
[0054] A data acquisition layer, used to obtain multi-modal data in the student classroom; the multi-modal data includes a set of student classroom performance images and the sound signal corresponding to the time.
[0055] The data acquisition layer is used to obtain the acquisition data of various hardware devices deployed in the classroom. Schematically, a high-definition camera is used to collect images of students' classroom performance. The image sensor in the camera converts the optical signal into an electrical signal. In the classroom scenario, the camera continuously captures images at a certain frame rate. Each image is a frame of image, which is stored in the form of a digital matrix. Each element in the matrix corresponds to a pixel point of the image and contains the values of the red, green, and blue (RGB) color channels to record the color and brightness information of the image. A microphone array is used as a sound acquisition device to collect sound signals. The microphone array consists of multiple microphones, and multiple microphones are distributed in multiple positions in the classroom. The microphones convert the mechanical vibration of sound into an electrical signal through the built-in transducer. There are differences in the time and intensity of the sound signals received by microphones at different positions. This difference can be used to locate and enhance the sound. The sound signal is output in the form of an analog electrical signal and becomes a digital signal after analog-to-digital conversion. It is digitized according to a certain sampling rate and quantization accuracy and stored as audio data. To ensure that the collected images and sound signals correspond to the same time, a high-precision clock chip can be added to the acquisition device. When collecting data, the clock chip adds an accurate timestamp to each frame of image and each segment of audio data. Subsequently, in the data processing process, the multi-modal data is aligned and matched according to the timestamps to ensure the time consistency of the data
[0056] The data standardization layer is used to preprocess and synchronize the time of multi-modal data to obtain standardized multi-modal data.
[0057] Schematically, the collected image data may have problems such as noise, uneven illumination, and inconsistent sizes. Filtering algorithms, illumination equalization techniques, and normalization can be used to process the set of images of students' classroom performance using filtering algorithms. Exemplarily, Gaussian filtering is used to denoise the image, and the image is smoothed by calculating the weighted average within the pixel neighborhood to remove random noise. Histogram equalization technology is used for illumination equalization. By adjusting the gray histogram of the image, the gray values of the image are redistributed to make the brightness distribution of the image more uniform and enhance the contrast of the image for subsequent feature extraction. The images are normalized by uniformly adjusting all images to a specific size and using algorithms such as bilinear interpolation to ensure the quality of the image after scaling and reduce information loss.
[0058] The sound signal is easily interfered by environmental noise. An adaptive filtering algorithm is used to eliminate the noise, effectively removing the background noise, and the sound signal is standardized to unify the sampling rate and duration so that different sound signals are comparable in subsequent processing.
[0059] The abnormal behavior detection layer inputs the standardized multimodal data into the behavior state classification model to obtain the behavior classification results; the behavior classification results include behavior type, behavior duration and intervention level labels.
[0060] Standardized multimodal data includes image, sound information and skeleton key point coordinates. For example, for image data and skeleton key point coordinates, convolutional neural network is used to extract students' posture, expression and other features. For sound data, recurrent neural network and its variant long time series neural network are used to process sequence features and extract features such as volume and spectrum. The extracted image and sound features are fused to obtain multimodal features, and the behavior state classification model is used to analyze the multimodal features to obtain the behavior classification results. Among them, the input multimodal features are matched with the patterns learned by the model through the behavior type mapping features of the pre-trained state classification model to determine the behavior type. The timestamps of the start and end of the behavior type are recorded, and the time difference between the two is calculated to obtain the behavior duration of the behavior. The intervention level rules corresponding to different behavior types and durations are pre-defined, and the intervention level labels are determined according to the behavior type and behavior duration.
[0061] The intervention decision layer is used to generate teaching intervention instructions based on the intervention decision logic and the behavior classification results; the teaching intervention instructions include system prompt instructions and teacher intervention reminder instructions.
[0062] The intervention decision layer analyzes the behavior classification results output by the abnormal behavior detection layer according to the pre-set rules and strategies, and generates corresponding teaching intervention instructions according to the judgment results of the intervention decision logic. For example, for minor inattention behaviors such as short-term distraction, the duration is lower than the threshold and is set to a low intervention level; if the duration is higher than the threshold, it is raised to a medium intervention level. For behaviors that seriously disrupt classroom order, such as loud noises corresponding to a volume higher than the threshold, regardless of the duration, a high intervention level is set. The system prompt instruction attracts the student's attention through the student's terminal device by vibration, flashing prompt lights or short text reminders, and guides him back to a normal learning state. The student's terminal device can be a smart bracelet worn by the student, a classroom electronic screen or an interactive electronic device in a smart classroom. The teacher intervention reminder instruction notifies the teacher to intervene in the form of pop-up windows, message push, etc. to the teacher's teaching management platform, and attaches detailed information on the student's abnormal behavior, such as the type of behavior, duration, student's name and location, etc., to help teachers quickly understand the situation and make corresponding handling.
[0063] In the above intelligent detection and teaching intervention system for classroom abnormal behaviors, by collecting various classroom performance data of students, combining multi-modal feature extraction, comprehensively analyzing image and sound information, the behavior type can be accurately judged. The duration of the behavior is calculated by recording the timestamp, and the intervention level label is determined according to the predefined rules, realizing the quantitative evaluation and accurate classification of students' behaviors. Teaching intervention instructions are generated according to the behavior classification results and preset rules, and different methods are adopted for different intervention levels. The system prompt instruction can timely remind students with minor abnormal behaviors and guide them to correct themselves; the teacher intervention reminder instruction enables teachers to quickly obtain detailed information, effectively intervene in serious abnormal behaviors, maintain classroom order, ensure the smooth progress of teaching activities, and improve the efficiency of classroom management and teaching quality.
[0064] In one of the embodiments, the multi-modal data is preprocessed and time-synchronized to obtain standardized multi-modal data, including:
[0065] S11. Use a face detection algorithm to locate the facial region in the set of students' classroom performance images, and segment the located facial region image to obtain a facial region image matrix.
[0066] The face detection algorithm can identify the position and boundary of the face in the image. Optionally, the Haar cascade algorithm is adopted, and a classifier is trained based on Haar features and the Adaboost algorithm. The classifier slides a window on the image and calculates the Haar feature values of windows at different positions. When the feature values of the window meet the requirements of the classifier, it is determined that there is a face within the window.
[0067] After the facial region is located, it is segmented from the original image. Exemplarily, a threshold-based segmentation method can be adopted to distinguish the pixel values of the facial region from the background pixel values. Specifically, according to the distribution characteristics of the facial color in the RGB color space, a suitable color threshold range is set, and the pixel points that meet the threshold range are extracted to form a facial region image. A contour detection-based method can also be used, by detecting the edge contour of the facial region, and segmenting the internal pixels. The obtained facial region image is stored in matrix form, and the elements in the matrix correspond to the pixel values of the image, and each pixel contains RGB color information.
[0068] S12. Adopt a pose estimation algorithm to extract the skeleton key points from the set of students' classroom performance images to obtain the skeleton key point coordinates.
[0069] The pose estimation algorithm can determine the pose of a human body in an image by extracting the skeleton key points of the human body. Exemplarily, the OpenPose algorithm based on a convolutional neural network and a multi-stage detection strategy is adopted. The convolutional neural network is used to extract features from the student classroom performance image, and various feature representations of the human body in the image are learned. Further, a specific branch network is used to predict the positions of the key points of the human body. The key points include the head, neck, shoulders, elbows, wrists, etc. During the prediction process, the network extracts the coordinates of each key point in the image not only according to the position coordinates of each key point, but also according to the spatial relationship and context information between the key points, forming the skeleton key point coordinates to describe the pose and actions of the human body.
[0070] S13. Filter and frame the voice signal to obtain a denoised voice signal.
[0071] Schematically, an adaptive filtering algorithm is adopted to adaptively adjust the parameters of the filter according to the statistical characteristics of the noise. By comparing the input noisy voice signal with the output of the filter, an error signal is obtained, and then the weights of the filter are adjusted according to the error signal, so that the output of the filter is as close as possible to the original pure voice signal. As the signal is continuously input, the filter can gradually adapt to the change of the noise and effectively remove the noise. Optionally, the pre-emphasis technique is further used to enhance the intensity of the high-frequency signal, and the continuous voice signal is segmented into shorter frames by framing and window function processing for subsequent independent analysis of each frame of the signal. Exemplarily, a suitable frame length is selected for framing, and the window slides on the voice signal with a certain step size, and the signal within the window is taken as a frame. Therefore, the continuous voice signal is segmented into multiple short frames, and each short frame can be independently subjected to feature extraction and analysis. After filtering and framing processing, the obtained denoised voice signal can more clearly reflect the speech content and provide a better data basis for subsequent speech analysis. After framing and window function processing, basic features such as the short-time energy, spectral envelope, and MFCC (Mel Frequency Cepstral Coefficients) of the voice signal are extracted to reflect important information such as the energy distribution and frequency characteristics of the voice signal.
[0072] S14. Based on the time stamp, perform time alignment of the cross-modal data of the facial region image matrix, the skeleton key point coordinates, and the denoised voice signal to obtain standardized multi-modal data; the standardized multi-modal data includes multiple sets of facial region image matrices, skeleton key point coordinates, and denoised voice signals that are matched within the same time window.
[0073] During data acquisition, due to differences in acquisition devices and transmission processes, there may be certain errors in timestamps. Using a time synchronization algorithm, with the time of a certain device as the reference, the timestamps of the data collected by other devices are calibrated. By calculating the time difference and adjusting the data according to the time difference, it is ensured that at the same time point, the image data and sound data can be precisely aligned. In subsequent data processing, the multi-modal data is processed in the calibrated time order to ensure the temporal consistency of the data.
[0074] In one embodiment, the abnormal behavior detection layer includes a feature extraction and fusion unit and a behavior state classification unit.
[0075] The feature extraction and fusion unit is used to extract features from the standardized multi-modal data and fuse the extracted features to obtain a behavior state feature representation.
[0076] Schematically, feature extraction is respectively performed on the facial region image matrix, skeleton key point coordinates, and denoised sound signals, and the features extracted from different modal data are fused to obtain a behavior state feature representation. Exemplarily, for the facial region image matrix, a convolutional neural network (CNN) is used for feature extraction. The CNN consists of a convolutional layer, a pooling layer, and a fully connected layer. The convolutional layer slides a convolutional kernel over the image matrix to perform a convolution operation to extract local features of the image. The pooling layer downsamples the output of the convolutional layer to reduce the data volume, while retaining important feature information and enhancing the robustness of the features. The fully connected layer integrates the extracted features to obtain a high-level feature representation of the image. Further, pre-trained models such as VGG and ResNet can be directly applied to the feature extraction of the facial region image.
[0077] For the skeleton key point coordinates, features such as the relative position relationship, angle, and distance between key points can be calculated. Exemplarily, calculating the distance between adjacent joint points can reflect the degree of limb extension; calculating the angle between joint points can reflect the bending state of the limb. Temporal analysis can also be performed on the sequence of skeleton key point coordinates to extract features such as motion speed and acceleration to describe the dynamic changes of the human body posture. Further, a recurrent neural network (RNN) and its variants are used to process the sequence information of the skeleton key point coordinates to obtain the extraction of the skeleton key point coordinates.
[0078] For denoising voice signals, multiple acoustic features are extracted, including Mel Frequency Cepstral Coefficients (MFCC), Linear Prediction Cepstral Coefficients (LPCC), short-time energy, and zero-crossing rate. Among them, MFCC can simulate the frequency response characteristics of the human auditory system, convert the voice signal into the Mel frequency domain, and then perform cepstral analysis to obtain a set of coefficients, representing the spectral characteristics of the sound. Short-time energy reflects the energy change of the voice signal in a short period, and the zero-crossing rate represents the number of times the signal crosses the zero level per unit time. These features can be used to distinguish different speech states to determine whether the student is speaking.
[0079] Furthermore, the features extracted from different modal data are fused to obtain a comprehensive representation of the behavioral state. Schematically, the fusion can be performed during the feature extraction process, that is, combining some of the features extracted from different modal data and then continuing with the subsequent feature extraction operations. Exemplarily, the facial image features extracted by CNN and the skeleton key point sequence features extracted by RNN are concatenated at the intermediate layer and then further feature learning and integration are performed through the subsequent network layers; the fusion can also be performed after the feature extraction is completed, that is, performing operations such as concatenation or weighted summation on the final feature vectors extracted from different modal data to obtain a comprehensive representation of the behavioral state. Exemplarily, the CNN feature vector of the facial image, the RNN feature vector of the skeleton key points, and the MFCC feature vector of the voice signal are concatenated to form a higher-dimensional feature vector.
[0080] The behavioral state classification unit is used to obtain the behavioral classification result based on the representation of the behavioral state features.
[0081] The behavioral state classification unit classifies the classroom behavior of the student according to the representation of the behavioral state features to determine the behavior type, behavior duration, and intervention level label. Among them, the behavior types include momentary distraction, persistent distraction, brief speech, persistent disruptive speech, fidgeting, etc. Momentary distraction means that the student has a dull or wandering look in the eyes and no physical movements for a short time; persistent distraction is closing the eyes or nodding frequently without any physical movements; brief speech is the mouth opening and closing for a short time and sound features are detected; persistent disruptive speech is continuously detecting the mouth opening and closing, having physical movements, and sound features; fidgeting is detecting physical movements for a long time.
[0082] The behavior duration is divided according to the characteristics of different behavior types and the degree of impact on the classroom. Exemplarily, brief behaviors are usually within 1 - 3 minutes, including momentary distraction, brief speech, etc.; medium-duration behaviors are within 3 - 10 minutes, including longer periods of distraction or fidgeting; behaviors with a long duration exceed 10 minutes, such as persistent disruptive speech, long periods of dozing off, etc.
[0083] The intervention level labels can be divided into level 0 (no intervention required), level 1 (system automatic reminder), and level 2 (prompt teacher intervention). Level 0 is applicable to occasional short-term minor abnormal behaviors that do not affect learning and classroom order; level 1 is for short-term distractions and short-term fidgeting, etc., with the system automatically reminding; level 2 is for behaviors such as continuous disturbing speech, long-term distraction, or dozing off, etc., which require teachers to intervene in a timely manner.
[0084] In one of the embodiments, feature extraction is performed on the standardized multi-modal data, and the extracted features are fused to obtain a behavioral state feature representation, including:
[0085] S21. Perform temporal variation extraction and facial detail extraction on the facial region image matrix to obtain facial features.
[0086] The facial region image matrix is a series of images that change over time. Using the method of time series analysis, pay attention to the changes in facial expressions and movements between consecutive frames. Exemplarily, use the optical flow method to calculate the motion vectors of facial pixels between adjacent frames, and judge the blinking frequency of the eyes, the opening and closing speed of the mouth, etc. Further, identify the closed and open detail states of the eyes and mouth in the convolutional neural network.
[0087] S22. Extract the time domain and frequency domain features of the denoised sound signal to obtain sound features.
[0088] The time domain features reflect the changes of the sound signal on the time axis. Among them, the short-time energy can reflect the intensity change of the sound. By dividing the sound signal into multiple short time periods and calculating the energy of the signal in each period, it can be judged whether the sound is loud or weak to distinguish between loud speech and whispering. The zero-crossing rate represents the number of times the sound signal crosses the zero level within a unit time, and it can be used to distinguish between voiceless and voiced sounds, as well as judge the presence or absence of sound. The frequency domain features of the sound signal reveal the distribution of the signal in different frequency components. Among them, the Mel frequency cepstral coefficients can simulate the perceptual differences of the human ear to sounds of different frequencies, so as to judge its interference to the classroom.
[0089] S23. Extract the limb movement amplitude, frequency, and change trend based on the skeleton key point coordinates to obtain skeleton motion features.
[0090] The coordinates of the skeleton key points contain the position information of each joint of the human body. By calculating the distance changes between adjacent key points, the amplitude of limb movements can be obtained. Exemplarily, calculate the distance between the wrist and elbow key points. When the arm is extended, this distance will increase; when the arm is bent, the distance will decrease. By calculating and analyzing the distances between multiple adjacent key points, the amplitude of limb movements can be comprehensively described. The overall amplitude of limb movements can also be measured by calculating the relative position changes between different limb parts, such as the position change of the head relative to the body, to determine whether the student is listening attentively facing the podium.
[0091] Perform time series analysis on the sequence of skeleton key point coordinates, count the number of occurrences of specific limb movements within a certain period of time, and obtain the movement frequency. Exemplarily, count the number of head nodding movements within one minute, or count the frequency of hand waving movements within a certain period of time. Determine whether a certain movement occurs by setting a threshold for the movement, that is, when the displacement of the head in the vertical direction exceeds a certain threshold and repeatedly appears within a short period of time, it is considered a nodding movement. Use the sliding window method to count the number of movements within different time windows, so as to obtain the movement frequencies at different time scales.
[0092] S24. Perform multi-modal feature fusion on the facial features, voice features, and skeleton movement features to obtain a representation of the behavior state features.
[0093] When performing multi-modal feature fusion, in order to better reflect the importance of different modal features for behavior state judgment, weights can be assigned to the features of each modality. The weights can be learned through a data-driven method. Exemplarily, use machine learning algorithms to optimize on the training data so that the fused features can be better used for behavior state classification. Further, use an adaptive weighted fusion algorithm to dynamically adjust the weights according to the performance of different modal features on different samples to improve the fusion effect. The obtained representation of the behavior state features is a feature vector that combines facial, voice, and limb movement information.
[0094] In one embodiment, performing multi-modal feature fusion on the facial features, voice features, and skeleton movement features to obtain a representation of the behavior state features includes:
[0095] S31. Use a sliding time window to segment the facial features, voice features, and skeleton movement features, and calculate feature vectors for the facial features, voice features, and skeleton movement features within each sliding time window to obtain multi-modal raw features.
[0096] Facial features, voice features, and skeletal motion features are sequential data that change over time. Exemplarily, a fixed 3-second time window is set, and then this window is slid on the time axis with a step size of 0.5 seconds. Each time it slides, the window covers a new time interval, dividing the continuous feature sequence into multiple overlapping segments. This enables observing and analyzing the changes in features at different time scales.
[0097] For the facial features within each sliding time window, since the facial features may be multi-dimensional features extracted by models such as CNN, the statistics of the multi-dimensional features within the window can be calculated, such as the mean and standard deviation, or pooling operations can be used to compress the multiple facial feature values within the window into a feature vector of a fixed length to highlight the main features. In terms of voice features, for frequency-domain features such as MFCC or time-domain features such as short-time energy, the statistics within the window are also calculated to obtain a more representative feature vector that reflects the frequency distribution and energy changes of the voice during that time period. For skeletal motion features, based on the changes in the coordinates of the skeletal key points within the window, derivative features such as the speed and acceleration of limb movements are calculated, and then combined with the original features such as the movement amplitude and frequency to construct a feature vector that can comprehensively describe the limb movement state within the window.
[0098] S32. Concatenate multiple multi-modal raw features in the order of the sliding time window and perform preliminary fusion through a fully connected layer to obtain multi-modal preliminary fusion features.
[0099] Concatenate multiple multi-modal raw features in the order of the sliding time window, that is, connect the facial, voice, and skeletal motion feature vectors obtained within different time windows in sequence. The formed long sequence vector contains the changes in multi-modal features over a period of time, retaining the order information of the features changing over time, which is convenient for subsequent processing to capture the time series features of the behavior state. Input the concatenated multi-modal raw feature vector into the fully connected layer, use the weight matrix of the fully connected layer to perform a linear transformation on the input features, add the bias term, and then perform a non-linear transformation through the activation function to learn the preliminary correlation relationship between features of different modalities and preliminarily fuse the features from different modalities in a unified space. By adjusting the weights and biases of the fully connected layer, the fused features can better reflect the comprehensive information of the behavior state.
[0100] S33. Perform weighted fusion on the multi-modal preliminary fusion according to the preset weight strategy to obtain the feature representation of the behavior state.
[0101] The preset weight strategy is a way of weight allocation preset according to the importance of different modal features for judging the behavior state. These weights can be determined based on prior knowledge and experimental analysis. Exemplarily, in a discussion class, the weight of the voice feature can be appropriately reduced. According to the preset weights, the multi-modal preliminary fusion features are weighted and fused to highlight the influence of the modal features that are more important for judging the behavior state and suppress the relatively unimportant features, so as to obtain a feature vector that can more accurately represent the student's behavior state, that is, the behavior state feature representation.
[0102] In one embodiment, the behavior state classification model is obtained in the following manner:
[0103] S41. Obtain a multi-modal data set, which includes multiple modal data, as well as the behavior types and intervention level labels corresponding to the multi-modal data one by one; the behavior types include normal listening, short-term distraction, continuous distraction, dozing off, short-term speech, and continuous disturbing speech; the intervention level labels include no intervention, system automatic reminder, and prompt teacher intervention.
[0104] S42. Extract and fuse features from the multi-modal data to obtain a time series feature vector.
[0105] S43. Divide the time series feature vector according to the preset time window length to obtain a multi-modal feature vector sequence; the multi-modal feature vector sequence includes time windows and their corresponding multi-modal feature vectors.
[0106] S44. Mark the behavior types and intervention levels within the time window in the multi-modal feature vector sequence to obtain a training set.
[0107] S45. Based on the self-attention mechanism, iteratively learn the mapping relationship between the multi-modal feature vectors and the behavior classification results in the training set to obtain model parameters.
[0108] S46. Construct a state classification model according to the model parameters.
[0109] Collect multi-modal data covering facial expression images, sound signals, body movement images, etc., and label corresponding behavior types and intervention level tags for them to form a multi-modal data set. Using technologies such as convolutional neural networks, recurrent neural networks and their variants, extract features such as facial expression changes, sound frequency and intensity, body movement amplitude and frequency from different modal data respectively, and then adopt a fusion strategy to integrate these features into a time-series feature vector to comprehensively present the changes in students' behaviors over time. According to the preset time window length, divide the time-series feature vector into multiple multi-modal feature vector sequences, and each sequence contains a specific time window and its corresponding comprehensive feature vector. Mark the behavior type and intervention level within the time window for the feature vector sequence to construct a training set and provide data samples with clear classification labels for model training. Based on the self-attention mechanism, the model automatically learns the associations and importance weights between different time steps and different modal features during the training process. By calculating the attention scores between feature vectors, the model can focus on key features and ignore secondary information. In each iteration, the model makes a behavior classification prediction for the multi-modal feature vectors in the training set according to the current parameters, calculates the loss between the prediction result and the true label, and uses the backpropagation algorithm to calculate the gradient of the loss with respect to the model parameters, and adjusts the parameters according to the gradient to make the model prediction result gradually approach the true label and gradually optimize the model's mapping ability for behavior classification results. After multiple rounds of iterative training, when the performance of the model on the training set and the validation set reaches the expectation, save the model parameters at this time. Construct a behavior state classification model based on these optimized parameters. In practical applications, the model receives new multi-modal data, and after feature extraction, fusion and division, inputs it into the model for behavior state classification prediction, and outputs the corresponding behavior type and intervention level tags to assist teachers in classroom management and teaching intervention decision-making.
[0110] In one embodiment, based on the intervention decision logic, generate teaching intervention instructions using the behavior classification result, including:
[0111] S51. Calculate the abnormal behavior index according to the behavior type, behavior duration and intervention level tag.
[0112] The impact degrees of behavior type, behavior duration, and intervention level label on the severity of abnormal behavior are different. Therefore, it is necessary to assign weights to them. The weights can be determined based on the experience of education experts, statistical analysis of a large amount of classroom observation data, or optimized through machine learning algorithms. Exemplarily, assume that through analysis, it is found that the behavior type has the greatest impact on the severity of abnormal behavior, followed by behavior duration, and then the intervention level label. Therefore, weights of 0.5, 0.3, and 0.2 are assigned to the behavior type, behavior duration, and intervention level label respectively. Further, different behavior types are quantified. A short distraction is quantified as 1, a continuous distraction is quantified as 2, dozing off is quantified as 3, a short speech is quantified as 2, a continuous disturbing speech is quantified as 4, etc. Quantify according to the length of the behavior duration. Set a time benchmark. If the behavior duration is within 1 minute, it is quantified as 1; 1 - 5 minutes is quantified as 2; 5 - 10 minutes is quantified as 3; more than 10 minutes is quantified as 4, etc. The system automatically reminding is quantified as 1, and prompting the teacher to intervene is quantified as 2. According to the quantified values and weights, use the weighted summation method to calculate the abnormal behavior index.
[0113] S52. When the abnormal behavior index reaches the first intervention range, generate a system intervention instruction; the system intervention instruction is used to instruct the system to send a reminder to the student terminal corresponding to the abnormal behavior.
[0114] According to the actual teaching needs and experience, determine the first intervention range. For abnormal behaviors within this range, it is possible for students to adjust their behaviors in time and return to the normal learning state through the system's automatic reminder. Therefore, when the calculated abnormal behavior index falls within the first intervention range, the system will generate a system intervention instruction. The system intervention instruction contains clear operation instructions, that is, to send a reminder to the student terminal corresponding to the abnormal behavior. The reminder method can be a pop-up reminder, a vibration reminder, etc., to attract the student's attention, make them aware of their abnormal behavior, and then correct it.
[0115] S53. When the abnormal behavior index reaches the second intervention range, generate a teacher intervention reminder instruction; the teacher intervention reminder instruction is used to instruct the system to send the student information corresponding to the abnormal behavior and the behavior type to the corresponding teacher terminal.
[0116] Similarly, according to the actual teaching situation, determine the second intervention range, which represents that the abnormal behavior is relatively serious and interferes with the normal progress of teaching. The system will generate a teacher intervention reminder instruction. In addition to instructing the system to send a reminder message to the corresponding teacher terminal, the teacher intervention reminder instruction will also contain student information such as the name and seat number of the student corresponding to the abnormal behavior and the specific behavior type. This helps the teacher quickly understand which student has what kind of abnormal behavior, so as to take corresponding intervention measures in time, maintain classroom order, and ensure the normal progress of teaching activities.
[0117] In one embodiment, the system further includes a feedback collection and model adaptive adjustment layer. The feedback collection and model adaptive adjustment layer is used to collect teacher intervention feedback and adjust the weight factors in the weight strategy and the model parameters of the behavior state classification model based on a learning algorithm; the teacher intervention feedback includes accepting the intervention reminder and ignoring the intervention reminder.
[0118] After the teacher uses the teaching intervention instruction, the system provides a feedback entry for the teacher to record the processing result of the intervention reminder. When the teacher accepts the intervention reminder and executes the corresponding intervention measure, the system records this information, indicating that the intervention reminder issued by the system is recognized by the teacher and considered necessary; if the teacher ignores the intervention reminder, the system also records this feedback, which means that the teacher believes that this intervention reminder may be inaccurate or unnecessary. The system associates these feedback information with the corresponding behavior classification results, student information, and the multi-modal data at that time to form structured feedback data, providing a basis for subsequent model adjustment. When the system collects the teacher's intervention feedback, it adjusts the weight factor through a learning algorithm. Exemplarily, if the teacher ignores the intervention reminder multiple times and a certain modal feature in the corresponding multi-modal data has a high weight, it may mean that this modal feature has an excessive impact on behavior classification under the current weight allocation, resulting in misjudgment by the system. At this time, the learning algorithm will appropriately reduce the weight factor of this modal feature to make the multi-modal feature fusion more reasonable. On the contrary, if the teacher frequently accepts the intervention reminder and the behavior classification is accurate, the learning algorithm may appropriately increase the weight factor of the modal feature that plays a key role in the correct classification to enhance the system's recognition ability for such behaviors.
[0119] The accuracy of the behavior state classification model is crucial for the performance of the system. The model parameters are adjusted using the collected teacher intervention feedback in combination with a learning algorithm. When the teacher accepts the intervention reminder and the intervention is effective, it indicates that the model has made a correct judgment in this classification. The learning algorithm will fine-tune the model parameters in the direction of strengthening this judgment, enabling the model to classify more accurately when encountering similar situations. If the teacher ignores the intervention reminder, it indicates that the model's judgment may not match the actual situation. The learning algorithm will analyze the reason for the deviation of the model in this classification, calculate the error gradient through the backpropagation algorithm, and adjust the parameters such as the weights and biases of the model according to the gradient, enabling the model to more accurately identify the behavior state in subsequent classifications. By continuously collecting teacher intervention feedback and adjusting the model parameters, the behavior state classification model can gradually adapt to different classroom scenarios and student behavior patterns, improving the overall classification performance.
[0120] Feedback collection and model adaptive adjustment are an ongoing process. As the system is used in different classroom environments, a large amount of feedback data on teacher interventions will be continuously accumulated. The system continuously optimizes the weight strategy and behavior state classification model based on this data. At the same time, due to factors such as classroom environment and student population that may change, the system's dynamic adjustment ability enables it to adapt to these changes in a timely manner and maintain a high detection accuracy and intervention effectiveness. For example, the overall behavior habits of students in the new semester may be different. Through feedback collection and model adaptive adjustment, the system can automatically adjust to better adapt to the new situation.
[0121] In one of the embodiments, the system further includes an information security and encrypted transmission layer. The information security and encrypted transmission layer is used to encrypt the collection, transmission, and storage of multimodal data.
[0122] The information security and encrypted transmission layer plays a key role in ensuring the security and data privacy of the intelligent detection of classroom abnormal behaviors and teaching intervention system. It prevents data from being stolen, tampered with, and leaked through the encryption process of multimodal data in the collection, transmission, and storage links. Schematically, in the data collection stage, multimodal data is obtained from the collection device. To prevent data from being stolen at the source, a symmetric encryption algorithm or an asymmetric encryption algorithm is used to encrypt the original data. Exemplarily, when the camera captures images, the built-in encryption module uses the AES (Advanced Encryption Standard Technology) algorithm to encrypt the image data in real time, converting the plaintext image into ciphertext. Even if the data is intercepted before being transmitted to the system, it is difficult for attackers to obtain the original image information.
[0123] Data faces the risk of network attacks during the transmission process. The information security and encrypted transmission layer uses the transport layer security protocol for encrypted transmission. Exemplarily, at the data sending end, the encrypted data is encapsulated with the TLS / SSL protocol header and transmitted through the network; after receiving the data, the receiving end first decrypts and verifies the integrity according to the TLS / SSL protocol header to ensure that the data has not been tampered with and the source is reliable. When the system transmits the collected multimodal data from the classroom local device to the remote server, the TLS / SSL protocol ensures the secure transmission of data in the network and prevents the data from being eavesdropped, stolen, or tampered with during the transmission.
[0124] When data is stored in a server or storage device, the information security and encrypted transmission layer also takes encryption measures. The stored data is encrypted and stored in a hierarchical manner. For example, sensitive data such as student identity information is encrypted using strong encryption algorithms such as the Advanced Encryption Standard, and ordinary data can use relatively lightweight encryption algorithms. To ensure data traceability and immutability, digital signatures are introduced. The data owner signs the data using a private key, and the receiver or verifier uses the public key to verify the authenticity of the signature and the integrity of the data. Exemplarily, when the server stores student multimodal data, the student's personal sensitive information is encrypted and stored using AES, and at the same time, digital signature technology is used to sign the key data to ensure the security and integrity of the data.
[0125] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown in the direction of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0126] Based on the same inventive concept, the embodiments of the present application also provide a method for intelligent detection and teaching intervention of classroom abnormal behaviors for implementing the classroom abnormal behavior intelligent detection and teaching intervention system involved above. The implementation solutions provided by this method to solve problems are similar to the implementation solutions recorded in the above system. Therefore, the specific limitations in one or more embodiments of the method for intelligent detection and teaching intervention of classroom abnormal behaviors provided below can refer to the limitations on the classroom abnormal behavior intelligent detection and teaching intervention system in the above text, and will not be repeated here.
[0127] In an exemplary embodiment, as Figure 2 shown, a method for intelligent detection and teaching intervention of classroom abnormal behaviors is provided, including:
[0128] S201. Obtain multimodal data of students in the classroom.
[0129] S202. Preprocess and synchronize the time of the multimodal data to obtain standardized multimodal data.
[0130] S203. Input the standardized multimodal data into the behavior state classification model to obtain a behavior classification result; the behavior classification result includes a behavior type, a behavior duration, and an intervention level label.
[0131] S204. Generate teaching intervention instructions based on the intervention decision logic using the behavior classification results.
[0132] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0133] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0134] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0135] The above embodiments only represent several implementation manners of the embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.
Claims
1. An intelligent detection and teaching intervention system for classroom abnormal behaviors, characterized in that, Including: A data acquisition layer for obtaining multi-modal data in a student classroom; the multi-modal data includes a set of student classroom performance images and sound signals corresponding to the time. A data standardization layer for preprocessing and time synchronization of the multi-modal data to obtain standardized multi-modal data. An abnormal behavior detection layer for inputting the standardized multi-modal data into a behavior state classification model to obtain a behavior classification result; the behavior classification result includes a behavior type, a behavior duration, and an intervention level label. An intervention decision layer for generating a teaching intervention instruction based on an intervention decision logic using the behavior classification result; the teaching intervention instruction includes a system prompt instruction and a teacher intervention reminder instruction.
2. The system according to claim 1, characterized in that, The preprocessing and time synchronization of the multi-modal data to obtain standardized multi-modal data includes: Using a face detection algorithm to locate the facial region in the set of student classroom performance images, and segmenting the located facial region image to obtain a facial region image matrix. Using a pose estimation algorithm to extract skeleton key points from the set of student classroom performance images to obtain skeleton key point coordinates. Filtering and frame segmentation of the sound signal to obtain a denoised sound signal. Performing time alignment of cross-modal data on the facial region image matrix, the skeleton key point coordinates, and the denoised sound signal based on a timestamp to obtain the standardized multi-modal data; the standardized multi-modal data includes multiple sets of the facial region image matrix, the skeleton key point coordinates, and the denoised sound signal that are matched within the same time window.
3. The system according to claim 2, wherein: The abnormal behavior detection layer includes a feature extraction and fusion unit and a behavior state classification unit. The feature extraction and fusion unit is used to extract features from the standardized multi-modal data and fuse the extracted features to obtain a behavior state feature representation. The behavior state classification unit is used to obtain a behavior classification result based on the behavior state feature representation.
4. The system according to claim 3, characterized in that, The extraction of features from the standardized multi-modal data and the fusion of the extracted features to obtain a behavior state feature representation includes: Performing temporal change extraction and facial detail extraction on the facial region image matrix to obtain facial features. Extracting the time domain and frequency domain features of the denoised sound signal to obtain sound features. Extracting the limb movement amplitude, frequency, and change trend based on the skeleton key point coordinates to obtain skeleton motion features. Performing multi-modal feature fusion on the facial features, sound features, and skeleton motion features to obtain a behavior state feature representation.
5. The system according to claim 4, wherein The performing multi-modal feature fusion on the facial features, sound features, and skeleton motion features to obtain a behavior state feature representation includes: Using a sliding time window to segment the facial features, sound features, and skeleton motion features, and calculating feature vectors for the facial features, sound features, and skeleton motion features within each sliding time window to obtain multi-modal raw features. Splice the multiple multimodal original features in the order of the sliding time window, and perform preliminary fusion through a fully connected layer to obtain multimodal preliminary fusion features; Perform weighted fusion on the multimodal preliminary fusion according to a preset weight strategy to obtain a behavioral state feature representation.
6. The system according to claim 1, characterized in that, The behavioral state classification model is obtained through the following method: Obtain a multimodal dataset, where the multimodal dataset includes multiple modal data, as well as the behavior types and intervention level labels corresponding to the multimodal data one by one; the behavior types include normal class listening, short-term distraction, continuous distraction, dozing off, short-term speaking, and continuous disruptive speaking; the intervention level labels include no intervention, system automatic reminder, and reminder to the teacher for intervention; Extract features and perform feature fusion on the multimodal data to obtain a time series feature vector; Divide the time series feature vector according to a preset time window length to obtain a multimodal feature vector sequence; the multimodal feature vector sequence includes time windows and their corresponding multimodal feature vectors; Mark the behavior types and intervention levels within the time window in the multimodal feature vector sequence to obtain a training set; Iteratively learn the mapping relationship between the multimodal feature vectors and the behavior classification results in the training set based on the self-attention mechanism to obtain model parameters; Construct the state classification model according to the model parameters.
7. The system according to claim 1, characterized in that The generating of teaching intervention instructions using the behavior classification result based on the intervention decision logic includes: Calculate an abnormal behavior index according to the behavior type, the behavior duration, and the intervention level label; When the abnormal behavior index reaches the first intervention range, generate a system intervention instruction; the system intervention instruction is used to instruct the system to send a reminder to the student terminal corresponding to the abnormal behavior; When the abnormal behavior index reaches the second intervention range, generate a teacher intervention reminder instruction; the teacher intervention reminder instruction is used to instruct the system to send the student information corresponding to the abnormal behavior and the behavior type to the corresponding teacher terminal.
8. The system according to any one of claims 1 to 7, characterized in that: The system further includes a feedback collection and model adaptive adjustment layer; The feedback collection and model adaptive adjustment layer is used to collect teacher intervention feedback, and based on a learning algorithm, use the teacher intervention feedback to adjust the weight factors in the weight strategy and the model parameters of the behavioral state classification model; the teacher intervention feedback includes accepting the intervention reminder and ignoring the intervention reminder.
9. The system according to claim 8, characterized in that: The system further includes an information security and encrypted transmission layer; The information security and encrypted transmission layer is used to perform data encryption on the acquisition, transmission, and storage of the multimodal data.
10. An intelligent detection and teaching intervention method for classroom abnormal behaviors, characterized in that, The method includes: Obtain multimodal data in the student classroom; Perform preprocessing and time synchronization on the multimodal data to obtain standardized multimodal data; Input the standardized multimodal data into a behavioral state classification model to obtain a behavior classification result; The behavior classification result includes a behavior type, a behavior duration, and an intervention level label; Generate teaching intervention instructions using the behavior classification result based on the intervention decision logic.
Citation Information
Cited By
Classroom interaction reasoning auxiliary method, system and equipment based on AI drive
CN121684335A