Teaching behavior intelligent identification method and system based on low-computing-power equipment

A low-computational system using lightweight neural networks and ROI-based processing for classroom behavior and emotion recognition addresses high computational demands, enabling efficient and accurate student monitoring with scalable integration.

CN120318753APending Publication Date: 2025-07-15HANGZHOU LEZHIXING EDUCATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510372448.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing behavior and emotion recognition technology has a large amount of computing power on low-computing devices, making it difficult to effectively apply, and the recognition accuracy is insufficient when multiple people speak or the human body is blocked.

Method used

The lightweight MobileNet and ResNet neural network models are adopted, combined with the ROI rotation algorithm, and the teaching behavior video stream is collected and processed in real time, user behavior and emotions are identified, and data is monitored and stored in real time through the console to generate abnormal alarm signals.

Benefits of technology

Accurate teaching behavior and emotion recognition is achieved on low-computing equipment, reducing system load, supporting reporting of recognition results, and providing reference for classroom teaching effectiveness evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318753A_ABST
    Figure CN120318753A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent teaching behavior recognition method and system based on low-computing-power equipment, and aims at recognizing behaviors and emotions of students by analyzing the body and face shapes of the students and the interactive relationship between the students and facilities and equipment such as tables and chairs, teaching aids and the like in a classroom, and recording related data. According to the system, targeted optimization is carried out on a low computing power scene, the operation pressure and the system load are reduced, and through task scheduling, parallel processing and ROI setting based on time rotation, the processing speed of the system is improved on the premise that a target area can be completely covered. The system supports the reporting of an identification result to an upstream service, thereby providing a reference for classroom teaching effect evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent management technologies, and in particular, to an intelligent teaching behavior recognition system, an intelligent teaching behavior recognition method, and an electronic device based on low-computing-power devices. Background Art

[0002] With the continuous progress of technologies, the importance of information technologies in campus teaching has become increasingly prominent, and among them, behavior and emotion recognition technologies have been widely applied. Technically, behavior and emotion recognition mainly includes two categories: based on vision and speech recognition. In the vision-based behavior recognition method, one category explicitly obtains body posture features through human body detection and key point extraction, and then conducts behavior classification; the other category takes video segments as input and directly obtains recognition results. Both categories of methods are implemented using a series of neural networks with different structures. Vision-based emotion recognition is similar to behavior recognition in principle, but mainly focuses on the facial area. On the other hand, speech recognition can also be used to judge personnel behavior or assist the visual recognition results. By recognizing the intonation of speech, the emotional state of personnel can also be judged.

[0003] Behavior and emotion recognition technologies can monitor the attention and emotional state of students in real time, thereby helping teachers comprehensively understand the learning status and needs of students. Based on these data, teachers can adjust teaching strategies in a timely manner, promote the implementation of personalized teaching, and further enhance the learning experience and participation of students.

[0004] In existing behavior and emotion recognition methods, for example, Patent No. CN106851216B discloses a classroom behavior monitoring system and method based on face and speech recognition. Video information of students and teachers in the classroom is collected through a camera; voice information of students and teachers in the classroom is collected through a recording device; the main control processor preprocesses the received video information of students and teachers, extracts facial expression features and behavior features of students and teachers; the main control processor processes the received voice information of students, extracts student voice features; the main control processor processes the received voice data information of teachers, extracts teacher voice features, calculates the score of teacher teaching effect, makes an evaluation of teacher teaching according to the score, and provides guiding suggestions. These existing technologies based on video segment solutions have a large amount of computation and are generally not applicable to low-computing-power scenarios. The key point-based solutions usually require the main part of the human body to be clearly visible in the image. When the human body is blocked, key points may not be normally extracted. Moreover, for emotion recognition, the key point-based solutions also have relatively high requirements for resolution.

[0005] In addition, in existing speech-based solutions, when multiple people are speaking simultaneously in the classroom, if the recognition results need to be accurate to each person, the voices of each person need to be distinguished, which will significantly increase the system complexity and computing power requirements at this time. Summary of the Invention

[0006] On the one hand, the present application proposes an intelligent teaching behavior recognition system based on low-computing power devices, including a data acquisition module, a processing module, a console, and a storage module, wherein:

[0007] The data acquisition module is used to collect the teaching behavior video stream in the classroom environment in real time, and obtain the static teaching behavior images from the continuous video stream by means of regular or intermittent frame extraction, and transmit them to the processing module;

[0008] The processing module is used to detect the user behavior and face in the static teaching behavior image in parallel, identify the user behavior, mark the interaction relationship between the user and the devices around the classroom environment, and identify the user emotion according to the face recognition, and transmit the recognition result to the console;

[0009] The console is used to monitor and schedule the data acquisition module, the processing module, and the storage module in real time, provide algorithm support and system visualization operation services, and report the system data to the upstream system;

[0010] The storage module is used for storing system data;

[0011] The data acquisition module, the processing module, the storage module, and the console are respectively communicatively connected.

[0012] As an optional implementation of the present application, optionally, when the processing module detects the user behavior, it includes:

[0013] Adopt a lightweight MobileNet neural network model to identify and extract the behavior features of the user in the static teaching behavior image and the key point features of the devices around the classroom environment, and label the interaction relationship between the user and the devices around the classroom environment for the behavior features of the user according to the identified key point features of the devices around the classroom environment;

[0014] Among them, when the MobileNet neural network model runs, it performs model pruning according to the device performance, and for each block in the model, iteratively reduces the number of convolutional layers and channels therein to reduce the number of parameters.

[0015] As an optional implementation of the present application, optionally, when the processing module detects the user emotion, it includes:

[0016] Adopt a lightweight MobileNet neural network model to identify and extract the face image of the user in the static teaching behavior image;

[0017] Use a lightweight ResNet neural network model to recognize the emotional features in the face image and output the corresponding user emotions;

[0018] Among them, when the MobileNet neural network model and the ResNet neural network model are running, they are respectively pruned according to the device performance. For each block in the model, the number of convolutional layers and channels in it are iteratively reduced to reduce the number of parameters.

[0019] As an alternative implementation of this application, optionally, the model pruning includes the following steps:

[0020] Step 1: Set the target number of parameters and MAP;

[0021] Step 2: Set the value of variable i to 1;

[0022] Step 3: For the i-th block in the model, remove the convolutional layer in the middle position and adjust the input and output of the front and back layers accordingly;

[0023] Step 4: Retrain the model. If MAP is higher than the target, execute Step 5, otherwise execute Step 12;

[0024] Step 5: If the current number of parameters is lower than the target, execute Step 12, otherwise execute Step 6;

[0025] Step 6: Reduce the number of channels of all convolutional layers in the i-th block and adjust the input and output of the front and back layers accordingly;

[0026] Step 7: Retrain the model. If MAP is higher than the target, execute Step 8, otherwise execute Step 12;

[0027] Step 8: If the current number of parameters is lower than the target, execute Step 12, otherwise execute Step 9;

[0028] Step 9: If i < total number of blocks, execute Step 10, otherwise execute Step 11;

[0029] Step 10: i = i + 1, execute Step 3;

[0030] Step 11: If the current number of model parameters is lower than the target, execute Step 12, otherwise execute Step 2;

[0031] Step 12: End.

[0032] As an alternative implementation of this application, optionally, when using a lightweight MobileNet neural network model to recognize and extract the face image of the user in the static image of the teaching behavior, it includes:

[0033] Statistically analyze the data throughput of the currently input static images of teaching behaviors. When the data throughput is lower than the preset target, execute the ROI rotation algorithm according to time, perform block recognition based on the ROI region on the input static images of teaching behaviors, and adopt corresponding ROI configurations in different time periods to perform time-division recognition on each region in the classroom environment.

[0034] As an alternative implementation of this application, optionally, the ROI rotation algorithm includes the following steps:

[0035] Step 1: Input the ROI rotation configuration, including the position, shape, and quantity of the ROI;

[0036] Step 2: Input the static images of teaching behaviors;

[0037] Step 3: Process the static images of teaching behaviors based on the preset target ROI region;

[0038] Step 4: If the processing speed in Step 3 reaches the expectation, execute Step 6; otherwise, execute Step 5;

[0039] Step 5: Further divide the ROI, re-divide the target ROI region, and update the relevant configuration, then execute Step 8;

[0040] Step 6: If the currently used ROI reaches the rotation time, execute Step 7; otherwise, execute Step 8;

[0041] Step 7: Switch to the next ROI;

[0042] Step 8: Output the image recognition results of each divided ROI;

[0043] Step 9: Repeat Steps 2 to 8 until the end.

[0044] As an alternative implementation of this application, optionally, the console is further configured to:

[0045] Control the data acquisition module to collect the teaching behavior video stream according to the preset sampling frequency;

[0046] Real-time record the user's behavior and emotion, and bind the user's recognition result with the pre-calibrated user seat ID according to the interaction relationship;

[0047] According to the preset behavior anomaly monitoring strategy and emotion anomaly monitoring strategy, respectively determine whether the user's behavior and emotion are abnormal:

[0048] If so, generate the corresponding anomaly warning signal and report it to the upstream system;

[0049] Otherwise, continue to monitor.

[0050] On the other hand, this application proposes an intelligent recognition method for teaching behaviors based on low-computing-power devices, which is implemented based on the above-mentioned system and includes the following steps:

[0051] The console initializes the system, enters the teaching behavior monitoring program, and schedules the work of each module;

[0052] The data acquisition module continuously captures the teaching behavior video stream in the classroom environment, and obtains static images of teaching behaviors from the continuous video stream by means of regular or intermittent frame extraction, and transmits them to the processing module;

[0053] The processing module parallelly detects the user behaviors and faces in the static images of the teaching behaviors, identifies the user behaviors, marks the interaction relationship between the user and the devices around the classroom environment, and identifies the user emotions based on face recognition, and transmits the recognition results to the console;

[0054] The console records the user behaviors and user emotions in real time, and binds the recognition results of the user to the pre-calibrated user seat ID according to the interaction relationship; according to the preset behavior anomaly monitoring strategy and emotion anomaly monitoring strategy, it respectively judges whether the user behaviors and the user emotions of the user are abnormal: if so, it generates corresponding anomaly warning signals and reports them to the upstream system; otherwise, it continues to monitor;

[0055] The storage module stores system data.

[0056] On the other hand, this application also proposes an electronic device, including:

[0057] A processor;

[0058] A memory for storing instructions executable by the processor;

[0059] Wherein, the processor is configured to implement the above-mentioned method when executing the executable instructions.

[0060] The technical effects of the present invention:

[0061] This application captures videos and images during classroom teaching through a monocular camera, analyzes students' body and facial features as well as their interaction relationships with facilities and equipment such as desks, chairs, and teaching aids in the classroom to identify students' behaviors and emotions, and records relevant data. The system is specifically optimized for low-computing-power scenarios to reduce the operation pressure and system load. Through task scheduling, parallel processing, and time-rotated ROI setting, it ensures that the system processing speed can be improved on the premise of completely covering the target area. The system supports reporting the recognition results to the upstream service, thereby providing a reference for classroom teaching effect evaluation.

[0062] The application and implementation of the present invention have the following technical advantages:

[0063] Optimized for low-computing-power devices, accurate recognition results are provided while maintaining a small amount of computation;

[0064] Optimized for the classroom environment, using the classroom environment, facilities, and items to assist in behavior and emotion recognition;

[0065] The system has good scalability and can enable algorithms with different complexities according to the computing power conditions of the device;

[0066] Supports transmitting the recognition results to external systems such as the educational administration system, thus providing a reference for classroom teaching.

[0067] With reference to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The drawings included in and constituting a part of this specification, together with the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.

[0069] Figure 1 Shows a schematic diagram of the system composition of the present invention;

[0070] Figure 2 Shows a flowchart of the image processing of the present invention;

[0071] Figure 3 Shows a flowchart of the model pruning of the present invention;

[0072] Figure 4 Shows a flowchart of the ROI rotation algorithm of the present invention;

[0073] Figure 5 Shows a schematic diagram of the system deployment architecture of the present invention;

[0074] Figure 6 Shows a schematic diagram of the camera acquiring an image of the present invention;

[0075] Figure 7 Shows a schematic diagram of the ROI division of the present invention;

[0076] Figure 8 Shows a schematic diagram of the secondary ROI division of the present invention;

[0077] Figure 9 Shows a schematic diagram of the recognition result of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0078] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Identical reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0079] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.

[0080] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, well-known means, elements, and circuits have not been described in detail so as to highlight the gist of the present disclosure.

[0081] Embodiment 1

[0082] As Figure 1 shown, in one aspect of the present application, a teaching behavior intelligent recognition system based on low-computing power devices is proposed, including a data acquisition module, a processing module, a console, and a storage module, where:

[0083] The data acquisition module is configured to collect the teaching behavior video stream in the classroom environment in real time, and obtain the teaching behavior static images from the continuous video stream by means of timed or intermittent frame extraction and transmit them to the processing module;

[0084] The processing module is configured to detect the user behavior and face in the teaching behavior static images in parallel, identify the user behavior, mark the interaction relationship between it and the devices around the classroom environment, and recognize the user emotion according to face recognition, and transmit the recognition result to the console;

[0085] The console is configured to monitor and schedule the data acquisition module, the processing module, and the storage module in real time, provide algorithm support and system visualization operation services, and report the system data to the upstream system;

[0086] The storage module is used for system data storage;

[0087] The data acquisition module, the processing module, the storage module, and the console are respectively communicatively connected.

[0088] This system mainly monitors the behavior and emotion of students in the classroom environment, such as a classroom. If it is other environments, such as outdoor learning, it can also be applied based on the principle of this embodiment.

[0089] The hardware devices of each module are not limited, and are purchased and deployed by the user according to the deployment environment.

[0090] The implementation scheme of the present invention will be described in detail below.

[0091] The data acquisition module includes one or more cameras, which perform real-time video acquisition on the classroom environment, and obtain static images from the continuous video stream by means of timed or intermittent frame extraction, and directly transmit them to the processing module or forward them to the processing module via NVR for analysis and recognition. When the same camera corresponds to multiple independent downstream processing flows, the sampling frequencies of each downstream processing flow are aligned, and the output image of the camera is only decoded once and sent to each downstream processing flow to reduce the resource consumption caused by image decoding. When the algorithm requires continuous time-series data, the camera stops frame extraction and outputs a coherent video stream, and each frame of the image is still only decoded once.

[0092] The camera also supports real-time adjustment of parameters such as exposure time, brightness, contrast, and gamma, so as to maintain high imaging quality under changing environmental light conditions, and thus ensure the subsequent processing effect.

[0093] The processing module receives the images from the camera and calls the neural network for processing to obtain a preliminary recognition result. Each task has multiple processing flows in series according to business requirements, and the processing module matches the camera output image with the corresponding flow according to the task ID.

[0094] As shown in the appendix Figure 2 In order to reduce the computing power requirement, the module adopts a lightweight neural network to extract the overall appearance of the students and the features of facilities and items such as desks, chairs, and teaching aids around them for behavior recognition. At the same time, the facial image block of the student is extracted and the whole is input into the neural network for classification to obtain the facial expression features to determine the student's emotion. When processing, the two tasks are carried out in parallel, and the steps are as follows:

[0095] Step 1: Input the image;

[0096] Step 2: Input the image into the behavior recognition network to recognize the behavior of the person in the image;

[0097] Step 3: Input the image into the face detection network to detect the face in the image;

[0098] Step 4: Input the block containing the face in the image into the emotion recognition network to recognize the emotion of the person;

[0099] Step 5: Output the result.

[0100] The behavior recognition network, the face detection network, and the emotion recognition network are respectively trained based on the corresponding lightweight neural network models, which will be described separately below.

[0101] As an alternative implementation of the present application, optionally, when detecting the user's behavior, the processing module includes:

[0102] Adopt a lightweight MobileNet neural network model to identify and extract the behavior features of the user in the static image of the teaching behavior and the key point features of the devices around the classroom environment. According to the identified key point features of the devices around the classroom environment, label the interaction relationship between the user's behavior features and the devices around the classroom environment for the user;

[0103] Among them, when the MobileNet neural network model is running, model pruning is performed according to the device performance. For each block in the model, the number of convolutional layers and channels in it is iteratively reduced to reduce the number of parameters.

[0104] Behavior recognition is implemented using a single model. The direct detection of the human form and surrounding objects (devices around the classroom environment) is used to replace the commonly used key point extraction or temporal feature extraction. In the model design, MobileNet is used as the basic network model, and the model is further pruned according to the device performance. For each block in the model, the number of convolutional layers and channels in it is iteratively reduced to reduce the number of parameters.

[0105] The MobileNet neural network model can be pre-trained based on a large number of classroom (or other environment) behavior images of students. In order to implement the use of the MobileNet neural network model to identify and extract the behavior features of the user in the static image of the teaching behavior and the key point features of the devices around the classroom environment, and further label the interaction relationship between the user and the device, the following technical implementation solutions can be carried out:

[0106] 1. Model training objectives

[0107] Identify and extract the classroom behavior features of each student in the static image.

[0108] Identify and extract the key point features of the devices around the classroom environment.

[0109] According to the device key point features and behavior features, label the interaction relationship between the user and the device.

[0110] 2. Technical solutions

[0111] 2.1. Data preparation

[0112] Dataset collection: Collect a static image dataset containing various teaching scenarios and student behaviors to ensure the diversity and representativeness of the data.

[0113] Data annotation: Manually annotate the behaviors of students in the image (such as raising hands, reading, writing, etc.) and the key points of the devices (such as the positions and contours of the blackboard, computer, projector, etc.) to form the labeled data required for training.

[0114] 2. Model Selection and Training

[0115] MobileNet Model: Select lightweight neural network models such as MobileNetV3 as the basis because of their efficient computing performance and small model size.

[0116] Multi-task Learning Framework: Construct a multi-task learning framework containing two sub-tasks, one for behavioral feature classification and the other for device key point detection.

[0117] Loss Function Design: The cross-entropy loss function is used for the behavior classification task, and the heatmap regression loss function or a combination of L1 / L2 losses is used for the key point detection task.

[0118] Model Training and Optimization: Use the labeled dataset for model training and optimize the model performance by adjusting strategies such as the learning rate, batch size, and regularization.

[0119] 2.3 Feature Extraction and Matching

[0120] Behavioral Feature Extraction: Use the trained behavior classification network to extract the behavioral feature vectors of students from the image.

[0121] Device Key Point Detection: Use the trained key point detection network to detect and extract the key point coordinates and contour information of the device from the image.

[0122] Feature Matching and Association: Match and associate the student behavior with the nearest device key point according to the spatial position relationship between the device key point coordinates and the behavioral feature vectors.

[0123] 2.4 Interaction Relationship Annotation

[0124] Semantic Rule Formulation: According to the common sense and logic in the teaching scenario, formulate a series of semantic rules for judging the interaction relationship between students and devices (such as "the student is using the computer", "the student is watching the projector", etc.).

[0125] Interaction Relationship Judgment: Combine the feature matching results and semantic rules to automatically judge and annotate the interaction relationship between students and devices.

[0126] Result Output and Visualization: Output the annotated interaction relationship in the form of text or image annotation for subsequent analysis and application.

[0127] 3. Model Training and Verification

[0128] Use the labeled dataset for model training and evaluate the model's performance through methods such as cross - validation. It can be implemented in combination with the training steps of existing models. For example, divide the training set and the validation set, and verify the model performance based on F1, AUC values, etc. If the verification passes, deploy the model and put it into application.

[0129] Feature extraction and matching experiment: Conduct feature extraction and matching experiments on the test set to verify the accuracy and stability of the model.

[0130] Interactive relationship annotation experiment: Combine semantic rules to conduct interactive relationship annotation experiments and evaluate the accuracy and consistency of the annotation results.

[0131] System optimization and integration: Optimize and improve the system according to the experimental results, and finally integrate the model into the teaching behavior analysis system.

[0132] This technology is based on the MobileNet neural network model, which realizes the recognition and extraction of user behavior features in static images of teaching behaviors and key point features of surrounding devices in the classroom environment, and further annotates the interactive relationship between users and devices.

[0133] The model can be loaded on the platform where the processing module is located or deployed on the console (at this time, the model is called through the task request method) for self - deployment.

[0134] As an optional implementation of this application, optionally, when the processing module detects the user's emotion, it includes:

[0135] Adopt a lightweight MobileNet neural network model (face detection model) to identify and extract the face image of the user in the static image of the teaching behavior;

[0136] Adopt a lightweight ResNet neural network model (emotion classification model) to identify the emotion features in the face image and output the corresponding user emotion;

[0137] Among them, when the MobileNet neural network model and the ResNet neural network model are running, they respectively perform model pruning according to the device performance. For each block in the model, the number of convolutional layers and channels in it are iteratively reduced to reduce the number of parameters.

[0138] Similarly, face detection and emotion recognition are also carried out through pre - trained models here to achieve intelligent emotion monitoring. The training and application steps of the two models will be given below. Please specifically understand them in combination with the training and application principles of the corresponding basic models.

[0139] (1) Training and Application Steps of the Lightweight MobileNet Neural Network Model (Face Detection Model)

[0140] Training Steps:

[0141] Data Preparation:

[0142] Collect a large number of static image datasets of teaching behaviors containing faces.

[0143] Annotate the faces in the images, marking the bounding boxes of the faces.

[0144] Divide the dataset into a training set, a validation set, and a test set.

[0145] Data Preprocessing:

[0146] Normalize the images so that their pixel values are between 0 and 1.

[0147] Perform data augmentation on the images, such as rotation, flipping, scaling, etc., to increase the diversity of the data.

[0148] Model Selection and Modification:

[0149] Select the lightweight MobileNet neural network model as the basic architecture.

[0150] According to the requirements of the face detection task, modify the output layer of the model so that it can output the coordinates and confidence levels of the face bounding boxes.

[0151] Model Training:

[0152] Use the training set data to train the model, and adopt a suitable loss function (such as Intersection over Union Loss - IOU Loss) to optimize the model performance.

[0153] During the training process, use the validation set data to monitor the model performance, and adjust hyperparameters such as the learning rate and batch size as needed.

[0154] Model Evaluation:

[0155] Use the test set data to evaluate the trained model, and calculate metrics such as accuracy and recall rate in the face detection task.

[0156] Fine-tune the model according to the evaluation results to improve its performance.

[0157] Application Steps:

[0158] Load the Model:

[0159] In actual applications, first load the trained MobileNet face detection model.

[0160] Image input:

[0161] Input the static image of the teaching behavior to be detected into the model.

[0162] Face detection:

[0163] Use the model to perform face detection on the image, and output the coordinates and confidence of the face bounding box.

[0164] Face extraction:

[0165] Crop the face image from the original image according to the output face bounding box coordinates.

[0166] (2) Training and application steps of the lightweight ResNet neural network model (emotion classification model)

[0167] Training steps:

[0168] Data preparation:

[0169] Collect a face image dataset containing different emotions (such as happy, sad, angry, surprised, fearful, disgusted, neutral, etc.).

[0170] Manually annotate the emotions in the images to form emotion labels.

[0171] Divide the dataset into a training set, a validation set, and a test set.

[0172] Data preprocessing:

[0173] Normalize the images.

[0174] Perform data augmentation on the images to increase data diversity.

[0175] Model selection and modification:

[0176] Select the lightweight ResNet neural network model as the basic architecture.

[0177] According to the requirements of the emotion classification task, modify the output layer of the model so that it can output the probability distribution of different emotions.

[0178] Model training:

[0179] Use the training set data to train the model, and adopt the cross-entropy loss function to optimize the model performance.

[0180] During the training process, use the validation set data to monitor the model performance, and adjust hyperparameters such as the learning rate and batch size as needed.

[0181] Model evaluation:

[0182] Evaluate the trained model using the test set data and calculate metrics such as its accuracy and F1 score on the emotion classification task.

[0183] Fine-tune the model based on the evaluation results to improve its performance.

[0184] Application steps:

[0185] Load the model:

[0186] In practical applications, first load the trained ResNet emotion classification model.

[0187] Input of face images (the face of the student recognized and output by the above face detection model):

[0188] Input the face images extracted from the MobileNet face detection model into the ResNet emotion classification model.

[0189] Emotion recognition:

[0190] Use the model to perform emotion recognition on the face images and output the probability distribution of different emotions.

[0191] Emotion output:

[0192] According to the output probability distribution, select the emotion with the highest probability as the emotion output of the user.

[0193] Through the above steps, a lightweight MobileNet face detection model and a ResNet emotion classification model can be trained respectively and applied to the face detection and emotion recognition tasks in static images of teaching behaviors. In practical applications, these two models are used in combination. First, use the MobileNet model to detect faces in the image, and then use the ResNet model to perform emotion recognition on the detected faces to achieve automatic recognition and analysis of the emotions of users in static images of teaching behaviors.

[0194] When the MobileNet neural network model is running, in order to reduce system pressure and save computing power, the present invention performs model pruning according to the device performance. For each block in the model, the number of convolutional layers and channels is iteratively reduced to reduce the number of parameters. Specifically:

[0195] As Figure 3 shown, as an optional implementation of this application, optionally, the model pruning includes the following steps:

[0196] Step 1, set the target number of parameters and MAP;

[0197] Step 2, set the value of variable i to 1;

[0198] Step 3: For the i-th block in the model, remove the convolutional layer in the middle position, and correspondingly adjust the inputs and outputs of the front and back layers.

[0199] Step 4: Retrain the model. If the MAP is higher than the target, execute Step 5; otherwise, execute Step 12.

[0200] Step 5: If the current number of parameters is lower than the target, execute Step 12; otherwise, execute Step 6.

[0201] Step 6: Reduce the number of channels of all convolutional layers in the i-th block, and correspondingly adjust the inputs and outputs of the front and back layers.

[0202] Step 7: Retrain the model. If the MAP is higher than the target, execute Step 8; otherwise, execute Step 12.

[0203] Step 8: If the current number of parameters is lower than the target, execute Step 12; otherwise, execute Step 9.

[0204] Step 9: If i < the total number of blocks, execute Step 10; otherwise, execute Step 11.

[0205] Step 10: i = i + 1, and execute Step 3.

[0206] Step 11: If the current number of parameters of the model is lower than the target, execute Step 12; otherwise, execute Step 2.

[0207] Step 12: End.

[0208] When the model is deployed, additional quantization needs to be performed according to the characteristics of the hardware platform. The quantization bits depend on the bits supported by the platform, the inference speed of the model, and the target MAP. By performing lightweight processing on the MobileNet model, such as pruning, quantization, etc., the model size and computational complexity are reduced, the computing power requirements of the system device are lowered, and the application breadth of the system is improved.

[0209] Emotion recognition includes two models: face detection and emotion classification. The basic network models are MobileNet and ResNet respectively. Similar to the behavior recognition model, the models are also pruned and quantized according to the device performance.

[0210] As an optional implementation of the present application, optionally, when using the lightweight MobileNet neural network model to identify and extract the face image of the user in the static image of the teaching behavior, it includes:

[0211] Statistically calculate the data throughput of the currently input static images of teaching behaviors. When the data throughput is lower than the preset target, execute the ROI rotation algorithm according to time, perform block recognition based on the ROI region on the input static images of teaching behaviors, and adopt corresponding ROI configurations in different time periods to perform time-division recognition on each region in the classroom environment.

[0212] Here, considering that there are usually multiple faces in the same image, the data of multiple image blocks are additionally combined into a batch and input into the model to further improve the processing efficiency. During the processing, the processing module can perform block division or filtering on the obtained images based on the ROI region to achieve key monitoring or information extraction within the ROI (Region of Interest), thereby effectively reducing the influence of irrelevant backgrounds or noises during the processing and improving the processing efficiency and recognition accuracy.

[0213] As Figure 4 shown, as an optional implementation scheme of the present application, optionally, the ROI rotation algorithm includes the following steps:

[0214] Step 1: Input the ROI rotation configuration, including the ROI position, shape, and quantity;

[0215] Step 2: Input the static images of teaching behaviors;

[0216] Step 3: Process the static images of teaching behaviors based on the preset target ROI region;

[0217] Step 4: If the processing speed in Step 3 reaches the expectation, execute Step 6; otherwise, execute Step 5;

[0218] Step 5: Further divide the ROI, divide the target ROI region again, and update the relevant configuration, then execute Step 8;

[0219] Step 6: If the currently used ROI reaches the rotation time, execute Step 7; otherwise, execute Step 8;

[0220] Step 7: Switch to the next ROI;

[0221] Step 8: Output the image recognition results of each divided ROI;

[0222] Step 9: Repeat Steps 2 to 8 until the end.

[0223] The module can statistically calculate the current processing speed of the algorithm in real time. When the data throughput is lower than the preset target, it executes the ROI rotation algorithm according to time (the system automatically enables ROI rotation, divides the screen into 4 regions, and only processes 1 region at a time and rotates them in sequence). At different time periods of a day, corresponding ROI configurations are adopted to perform segmented recognition on each region of the classroom.

[0224] Through the ROI rotation algorithm, it is possible to avoid missed detections caused by ROI segmentation, thereby improving the segmentation and recognition accuracy of classroom students.

[0225] In terms of hardware, the processing module can include one or more embedded processing platforms. For each platform, it supports simultaneously invoking multiple processing cores of the platform, and setting the binding relationship according to the computational complexity of the processing flow and the computing power of the processing cores to process data in parallel to improve the system throughput.

[0226] When a preset behavior or emotion is recognized, use a rectangular box to mark the student and match it with the pre-calibrated seat.

[0227] As an optional implementation of this application, optionally, the console is further used for:

[0228] Controlling the data acquisition module to acquire the teaching behavior video stream according to a preset sampling frequency;

[0229] Real-time recording of the user's behavior and emotion, and binding the recognition result of the user with the pre-calibrated user seat ID according to the interaction relationship;

[0230] According to the preset behavior anomaly monitoring strategy and emotion anomaly monitoring strategy, respectively judge whether the user's behavior and emotion are abnormal:

[0231] If so, generate a corresponding anomaly warning signal and report it to the upstream system;

[0232] Otherwise, continue to monitor.

[0233] The console is the working core of the system, used for hardware access and management, configuring algorithms, and summarizing and reporting recognition information. Through direct communication with various external devices, the console can monitor and schedule other modules in real time. When accessing hardware, the console can perform permission verification and function initialization according to a preset strategy; after successful access, it can configure the subordinate relationship between hardware as needed, and deploy corresponding algorithm functions and recognition logics. When receiving the processing result, the console integrates the processing result hierarchically and reports it to the upstream system.

[0234] The console supports the monitoring of abnormal states. When software and hardware anomalies occur, it can restart the subsystem or switch to a standby device and send an alarm to the corresponding upstream system.

[0235] Deploy high-definition cameras inside the teaching area to ensure that the faces and behaviors of each user can be clearly captured. See the previous description for details.

[0236] Seat ID calibration: Assign a unique seat ID to each user and record the seat position information for subsequent binding. The administrator can perform the calibration in advance.

[0237] Face detection and recognition:

[0238] Use the lightweight MobileNet model to perform real-time face detection on the images captured by the camera.

[0239] Extract the detected face images and try to match them with the preset user face database to identify the user's identity (if the system supports the user identification function).

[0240] If the user's identity cannot be directly recognized, continue the process but record the user as an anonymous user.

[0241] Emotion recognition:

[0242] Input the detected face images into the lightweight ResNet emotion classification model to obtain the user's emotion classification results.

[0243] Seat ID binding:

[0244] Determine the user's seat according to the detected face position information (or combined with other sensor data, such as infrared sensing, etc.).

[0245] Bind and record the recognized user behaviors and emotions with the corresponding seat IDs.

[0246] Behavior anomaly monitoring strategy:

[0247] Strategy example: Set specific behavior patterns as anomalies, such as constantly looking down at the mobile phone for a long time, frequently leaving the seat, etc. Analyze the user behavior data in real time and compare it with the preset abnormal behavior patterns.

[0248] Emotion anomaly monitoring strategy:

[0249] Strategy example: Set the duration of a specific emotion state as an anomaly, such as anger lasting for more than 5 minutes. Analyze the user emotion data in real time and judge whether there is an abnormal emotion state in combination with the time dimension.

[0250] Abnormal alarm:

[0251] If it is detected that the user's behavior or emotion is abnormal, immediately generate an abnormal alarm signal. The alarm signal includes information such as the user's seat ID, abnormal type (behavior / emotion), abnormal description, and occurrence time.

[0252] Alarm reporting:

[0253] The generated abnormal alarm signals are reported to the upstream management system (such as the teaching management system, security monitoring system, etc.) through a preset interface. The upstream system can take corresponding measures according to the received alarm information, such as sending notifications to teachers, activating security plans, etc. Regularly or in real-time, synchronize user behavior, emotion, and seat ID binding data to the upstream system for subsequent analysis or report generation.

[0254] By integrating a lightweight neural network model and a user seat ID management system, real-time recording, binding, and abnormal monitoring of user behavior and emotion are achieved.

[0255] The console also provides a visual human-machine interaction interface for operators to intuitively and quickly operate various functions.

[0256] Storage module

[0257] The storage module is used to save the original data collected by the camera, the recognition results obtained by the processing module, as well as the configuration information and data related to the system operation, mainly including: original image data, image recognition results, algorithm model files, software images, and system operation configurations.

[0258] The storage module can include a complete NVR to store the original videos captured by the camera for distribution to multiple embedded processing platforms or provided to the upstream business system.

[0259] Through targeted data structure design, the storage module can orderly archive and manage a large number of continuous image frames and their corresponding metadata. In this process, the storage module can not only maintain a complete time series index for the original data but also synchronously retain the corresponding processing results.

[0260] The storage module supports adopting differential encryption and backup strategies for data with different importance levels or sensitivity levels. By integrating high reliability and scalability design at the system architecture level, this storage module can adapt to a distributed storage environment and also cooperate with cloud storage services to achieve remote backup or data mirroring.

[0261] The storage module provides external access interfaces, supporting fast retrieval by keywords, timestamps, or other custom tags for being called by other upstream systems.

[0262] In one embodiment, please understand the application principle of the present invention in combination with the following implementation cases.

[0263] Such as Figure 5For the system deployment architecture shown, the system proposed by the present invention is used to identify the behaviors and emotions of students in the classroom. The system uses an embedded platform as the computing core, and the console is deployed on the server. The server also simultaneously hosts a storage service and an upstream business platform. The embedded platform and the server are installed in the campus computer room. A monocular camera is used in the classroom to collect images, and the camera, the embedded platform, and the server are all connected through a local area network.

[0264] In the classroom, the camera is installed at the front, shooting diagonally downward to capture the entire classroom. The collected images are as Figure 6 shown.

[0265] The behaviors recognized by the system include four categories: reading and writing, raising hands, standing, and lying on the table. The emotion recognition includes five categories: calm, happy, angry, sad, and afraid. When the system is running, two threads for behavior recognition and emotion recognition run in parallel and are simultaneously allocated to different cores of the processing chip of the embedded platform.

[0266] The system supports setting the image resolution input to the neural network. Using a higher resolution results in higher accuracy but slower processing speed. The system will monitor the image processing speed in real time. When the set resolution is high and the processing speed is lower than 5 fps, the system automatically enables ROI rotation, divides the screen into four regions, and only processes one of the regions at a time and rotates them in sequence. To avoid missed detections caused by ROI segmentation, there is a certain overlap between ROIs. The position and shape of the ROIs are set according to the seat arrangement in the image. As Figure 7 shown.

[0267] If the processing speed is still lower than 5 fps, the ROIs are further divided. For example, ROI 1 in Figure 7 can be further divided into four sub-regions, as Figure 8 shown.

[0268] When a preset behavior or emotion is recognized, the student is marked with a rectangular box and corresponding to the pre-calibrated seat map to determine the student's identity. Subsequently, the original image and the annotated image are uploaded to the storage service and then reported to the upstream business platform. An example of the reported image is as Figure 9 shown. For example, when it is recognized that "Zhang San. Reading, calm", it means that student Zhang San is reading and is in a calm mood.

[0269] Obviously, those skilled in the art should understand that to implement all or part of the processes in the above embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above control embodiments. Those skilled in the art can understand that to implement all or part of the processes in the above embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above control embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (abbreviation: HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.

[0270] Embodiment 2

[0271] Based on the implementation principle of Embodiment 1, on the other hand, this application proposes an intelligent recognition method for teaching behaviors based on low-computing-power devices, which is implemented based on the above-mentioned system, and includes the following steps:

[0272] The console initializes the system, enters the teaching behavior monitoring program, and schedules the work of each module;

[0273] The data acquisition module continuously acquires the teaching behavior video stream in the classroom environment, and obtains the static teaching behavior images from the continuous video stream by means of regular or intermittent frame extraction, and transmits them to the processing module;

[0274] The processing module parallelly detects the user behaviors and faces in the static teaching behavior images, identifies the user behaviors, marks the interaction relationship between them and the devices around the classroom environment, and identifies the user emotions according to the face recognition, and transmits the recognition results to the console;

[0275] The console records the user behaviors and user emotions of the user in real time, and binds the recognition results of the user with the pre-calibrated user seat ID according to the interaction relationship; according to the preset behavior anomaly monitoring strategy and emotion anomaly monitoring strategy, respectively judge whether the user behaviors and the user emotions of the user are abnormal: if so, generate corresponding anomaly warning signals and report them to the upstream system; otherwise, continue to monitor;

[0276] The storage module stores the system data.

[0277] Please understand the above method steps in combination with the modules and their interactions described in Embodiment 1, and they will not be elaborated here.

[0278] Each module or step of the present invention described above can be implemented by a general computing system. They can be concentrated on a single computing system or distributed on a network composed of multiple computing systems. Optionally, they can be implemented by program codes executable by the computing system. Thus, they can be stored in a storage system for execution by the computing system, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0279] Embodiment 3

[0280] Furthermore, on the other hand, the present application also proposes an intelligent recognition method and system electronic device for teaching behaviors based on low-computing-power devices, including:

[0281] A processor;

[0282] A memory for storing instructions executable by the processor;

[0283] Wherein, when the processor is configured to execute the executable instructions, it implements an intelligent recognition method and system for teaching behaviors based on low-computing-power devices described in Embodiment 2.

[0284] The electronic device of the embodiments of the present disclosure includes a processor and a memory for storing instructions executable by the processor. Wherein, when the processor is configured to execute the executable instructions, it implements an intelligent recognition method and system for teaching behaviors based on low-computing-power devices described in Embodiment 2 above.

[0285] Here, it should be noted that the number of processors can be one or more. At the same time, in the electronic device of the embodiments of the present disclosure, an input system and an output system can also be included. Among them, the processor, the memory, the input system and the output system can be connected through a bus or in other ways, and specific limitations are not made here.

[0286] As a computer-readable storage medium, the memory can be used to store software programs, computer-executable programs and various modules, such as: the programs or modules corresponding to an intelligent recognition method and system for teaching behaviors based on low-computing-power devices of the embodiments of the present disclosure. The processor executes various functional applications and data processing of the electronic device by running the software programs or modules stored in the memory.

[0287] The input system can be used to receive input numbers or signals. Among them, the signal can be a key signal related to the user settings and function control of the device / terminal / server. The output system can include a display device such as a display screen.

[0288] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. An intelligent recognition system for teaching behaviors based on low-computing-power devices, characterized in that It includes a data acquisition module, a processing module, a console, and a storage module, where: The data acquisition module is used to collect the teaching behavior video stream in the classroom environment in real time, and obtain the static teaching behavior images from the continuous video stream by means of timed or interval frame extraction, and transmit them to the processing module; The processing module is used to detect the user behavior and face in the static teaching behavior images in parallel, identify the user behavior, label the interaction relationship between the user and the surrounding devices in the classroom environment, and recognize the user's emotion based on face recognition, and transmit the recognition results to the console; The console is used to monitor and schedule the data acquisition module, the processing module, and the storage module in real time, provide algorithm support and system visualization operation services, and report the system data to the upstream system; The storage module is used for system data storage; The data acquisition module, the processing module, the storage module, and the console are respectively communicatively connected.

2. The intelligent recognition system for teaching behaviors based on low-computing-power devices according to claim 1, wherein When the processing module detects the user behavior, it includes: Adopt a lightweight MobileNet neural network model to identify and extract the behavior features of the user in the static teaching behavior image and the key point features of the surrounding devices in the classroom environment, and label the interaction relationship between the user and the surrounding devices in the classroom environment for the behavior features of the user according to the identified key point features of the surrounding devices in the classroom environment; Among them, when the MobileNet neural network model is running, it performs model pruning according to the device performance, and for each block in the model, iteratively reduces the number of convolutional layers and channels in it to reduce the number of parameters.

3. The intelligent recognition system for teaching behaviors based on low-computing-power devices according to claim 2, wherein When the processing module detects the user emotion, it includes: Adopt a lightweight MobileNet neural network model to identify and extract the face image of the user in the static teaching behavior image; Adopt a lightweight ResNet neural network model to identify the emotion features in the face image and output the corresponding user emotion; Among them, when the MobileNet neural network model and the ResNet neural network model are running, they respectively perform model pruning according to the device performance, and for each block in the model, iteratively reduce the number of convolutional layers and channels in it to reduce the number of parameters.

4. The intelligent recognition system for teaching behaviors based on low-computing-power devices according to claim 3, wherein The model pruning includes the following steps: Step 1: Set the target number of parameters and MAP; Step 2: Set the value of variable i to 1; Step 3: For the i-th block in the model, remove the convolutional layer in the middle position, and adjust the input and output of the front and back layers accordingly; Step 4: Retrain the model. If MAP is higher than the target, execute Step 5, otherwise execute Step 12; Step 5: If the current number of parameters is lower than the target, execute Step 12, otherwise execute Step 6; Step 6: Reduce the number of channels of all convolutional layers in the i-th block, and adjust the input and output of the front and back layers accordingly; Step 7: Retrain the model. If MAP is higher than the target, execute Step 8, otherwise execute Step 12; Step 8: If the current number of parameters is lower than the target, execute Step 12, otherwise execute Step 9; Step 9: If i < total number of blocks, execute Step 10; otherwise, execute Step 11; Step 10: i = i + 1, and execute Step 3; Step 11: If the current model parameter quantity is lower than the target, execute Step 12; otherwise, execute Step 2; Step 12: End.

5. The intelligent recognition system for teaching behaviors based on low-computing-power devices according to claim 3, wherein When using the lightweight MobileNet neural network model to identify and extract the face image of the user in the static image of the teaching behavior, it includes: Statistical data throughput of the currently input static image of the teaching behavior. When the data throughput is lower than the preset target, execute the ROI rotation algorithm according to time, perform block recognition based on the ROI region on the input static image of the teaching behavior, and use corresponding ROI configurations in different time periods to perform time-division recognition on each region in the classroom environment.

6. The intelligent recognition system for teaching behaviors based on low-computing-power devices according to claim 5, wherein, The ROI rotation algorithm includes the following steps: Step 1: Input the ROI rotation configuration, including ROI position, shape, and quantity; Step 2: Input the static image of the teaching behavior; Step 3: Process the static image of the teaching behavior based on the preset target ROI region; Step 4: If the processing speed in Step 3 reaches the expectation, execute Step 6; otherwise, execute Step 5; Step 5: Further divide the ROI, re-divide the target ROI region, and update the relevant configuration, then execute Step 8; Step 6: If the currently used ROI reaches the rotation time, execute Step 7; otherwise, execute Step 8; Step 7: Switch to the next ROI; Step 8: Output the image recognition results of each divided ROI; Step 9: Repeat Steps 2 to 8 until the end.

7. The intelligent recognition system for teaching behaviors based on low-computing-power devices according to claim 1, wherein, The console is also used for: Controlling the data acquisition module to collect the teaching behavior video stream according to the preset sampling frequency; Real-time recording of the user's behavior and user emotion, and binding the recognition result of the user with the pre-calibrated user seat ID according to the interaction relationship; According to the preset behavior anomaly monitoring strategy and emotion anomaly monitoring strategy, respectively judge whether the user's behavior and user emotion are abnormal: If so, generate the corresponding anomaly warning signal and report it to the upstream system; Otherwise, continue to monitor.

8. An intelligent recognition method for teaching behaviors based on low-computing-power devices, which is implemented based on the system described in any one of claims 1-7, characterized in that, It includes the following steps: The console performs system initialization, enters the teaching behavior monitoring program, and schedules the work of each module; The data acquisition module real-time collects the teaching behavior video stream in the classroom environment, and obtains the static image of the teaching behavior from the continuous video stream by means of regular or intermittent frame extraction and transmits it to the processing module; The processing module parallelly detects the user behavior and face in the static image of the teaching behavior, identifies the user behavior, marks the interaction relationship between it and the surrounding devices in the classroom environment, and recognizes the user emotion according to face recognition, and transmits the recognition result to the console; The console real-time records the user's behavior and user emotion, and binds the recognition result of the user with the pre-calibrated user seat ID according to the interaction relationship; According to the preset abnormal behavior monitoring strategy and abnormal emotion monitoring strategy, respectively determine whether the user's behavior and emotion are abnormal: if so, generate corresponding abnormal alarm signals and report them to the upstream system; otherwise, continue to monitor; The storage module stores system data.

9. An electronic device, characterized in that, It includes: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in claim 8 when executing the executable instructions.

Citation Information

Patent Citations

  • A Classroom Behavior Monitoring System and Method Based on Face and Voice Recognition

    CN106851216B