A webcam and alarm system for online teaching

By using cameras and alarm systems in the online teaching system, combined with image processing and deep learning technology, students' behavior monitoring and analysis are realized, solving the problem that the existing system cannot accurately identify students' emotions and behavior patterns, and improving teaching effectiveness and student participation.

CN118553017BActive Publication Date: 2025-05-27WUYUEFENG (SHENZHEN) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410930473.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2025-05-27
Estimated Expiration
2044-07-11

AI Technical Summary

Technical Problem

The existing online teaching monitoring system cannot accurately identify and analyze students' emotional changes and behavior patterns, and it is difficult to effectively evaluate students' participation and teaching activity, and it lacks comprehensive analysis tools and personalized teaching support.

Method used

The camera and alarm system for online teaching are adopted, combined with image processing and deep learning technology to realize student behavior monitoring and analysis. Through the video deployment module, emotional behavior recognition module, comprehensive multi-dimensional module and group behavior analysis module, we ensure that each student’s facial expressions and upper body movements are fully captured, emotions and behavior patterns are identified, and participation and teaching effects are evaluated.

Benefits of technology

It realizes accurate monitoring and analysis of student behavior, improves the quality of teaching videos, provides a reliable data foundation, provides support for behavior analysis and participation assessment, optimizes teaching content and methods, and improves student participation and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118553017B_ABST
    Figure CN118553017B_ABST
Patent Text Reader

Abstract

The present invention discloses a webcam and alarm system for online teaching, specifically in the field of online teaching, which is used to solve the problem of online teaching quality analysis. Through image processing and deep learning technologies, the facial expressions and upper body movements of each student can be completely captured in the teaching video, improving the image processing efficiency and the accuracy of feature recognition, ensuring data integrity, and providing a reliable basis for behavior analysis and engagement assessment. Combining Haar cascade classifier, VGG-Face and SVM classification technologies, it accurately identifies and analyzes the emotions and behaviors of students, optimizes teaching content and methods, and enhances students' engagement and satisfaction. Using clustering algorithms to analyze the student participation time series, identify teaching peaks and troughs, evaluate the teaching effect through the teaching quality index, guide teaching adjustments, enhance the pertinence and effectiveness of curriculum design, improve teaching transparency and adjustability, and better respond to students' needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of online teaching, and more specifically, to a webcam and alarm system for online teaching. Background Art

[0002] With the popularization and development of online education, using technical means to monitor and improve teaching quality has become an important issue. Currently, there are various systems in the market that monitor classroom activities through webcams, and these systems can capture videos in real time and analyze the interactions between teachers and students. However, these existing systems mainly focus on teachers' teaching behaviors rather than students' reactions and engagement levels, and most systems lack the ability to deeply analyze students' emotions and behavior patterns. Although existing monitoring systems provide basic video monitoring functions, they have obvious deficiencies in limited emotional and behavioral analysis capabilities, lack of comprehensive analysis tools, insufficient ability to respond and adjust teaching strategies, and insufficient personalized teaching support. These deficiencies indicate that the existing technology cannot accurately identify and analyze students' emotional changes and behavior patterns, is difficult to effectively evaluate students' engagement and teaching activity levels, and cannot provide a comprehensive analysis of the performance of student groups during various periods of the teaching process.

[0003] To solve the above problems, a technical solution is provided now. Summary of the Invention

[0004] To overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a webcam and alarm system for online teaching. Through image processing and deep learning technologies, it realizes the monitoring and analysis of students' behaviors, ensuring that the facial expressions and upper body movements of each student are completely captured in the teaching video. By using real-time video analysis, image preprocessing, edge detection, and face recognition technologies, it improves the efficiency of image processing and the accuracy of feature recognition, ensures data integrity, and provides a reliable basis for behavior analysis and engagement assessment. Combining Haar cascade classifier, VGG-Face, and SVM classification technologies, it accurately identifies and analyzes students' emotions and behaviors, optimizes teaching content and methods, improves students' engagement and satisfaction. Using clustering algorithms to analyze students' participation time series, identify teaching peaks and valleys, evaluate teaching effects through teaching quality indices, guide teaching adjustments, enhance the pertinence and effectiveness of curriculum design, improve teaching transparency and adjustability, and better respond to students' needs, so as to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solutions: including: a video deployment module, an emotion and behavior recognition module, a comprehensive multi-dimensional module, and a group behavior analysis module;

[0006] The video deployment module installs and configures cameras in the learning space, captures the video stream of each student in real time, controls the cameras to focus on the corresponding positions of the students, ensures that the videos clearly record the facial expressions and upper body movements of each student, and sends the captured video stream to the emotion and behavior recognition module;

[0007] The emotion and behavior recognition module uses video capture and multi-level machine learning techniques to analyze the facial expressions and body movements of students in online teaching. By identifying the emotional states and behavior patterns of students, it understands and measures the engagement and activity levels, and sends the calculated result values of engagement and activity to the comprehensive multi-dimensional module;

[0008] The comprehensive multi-dimensional module quantifies the activity and engagement of students, calculates the teaching highlights per unit time, and sends the measured values of the teaching highlights per unit time to the group behavior analysis module;

[0009] The group behavior analysis module applies clustering algorithms to the time series of students' highlights, analyzes and identifies the performance of student groups in different time periods of the teaching video, determines the peak and trough periods of teaching by comparing density estimates with corresponding thresholds, analyzes whether the teaching needs to be adjusted, and gives an alarm prompt.

[0010] In a preferred embodiment, the operation process of the video deployment module includes the following:

[0011] Preprocess each frame of the collected images based on the supervision cameras provided by each student;

[0012] Use edge detection algorithms to identify the edges in the images;

[0013] Apply face recognition technology to locate the faces in the images;

[0014] Use the trained upper body detection model to identify the upper body parts in the images;

[0015] Analyze the positions and sizes of the detected faces and upper bodies, and determine whether they are fully displayed within the predefined camera capture area;

[0016] If it is found that the face or upper body of a student is not fully captured, an automatic reminder signal will be sent, and the position of the camera or the student needs to be adjusted.

[0017] In a preferred embodiment, the operation process of the emotion and behavior recognition module includes the following:

[0018] Capture video streams from the cameras of each student;

[0019] First, use the Haar cascade classifier to detect faces in the videos;

[0020] Extract key features of the face using a pre-trained convolutional neural network;

[0021] Input the extracted face features into a classification model to recognize different emotional states;

[0022] Use a pose estimation tool to extract the positions and motion information of body key points from video data;

[0023] Combine the face and body features into a single feature vector through feature concatenation, as follows:

[0024] Let the face feature vector be , where represents the th feature recognized from the face, is the dimension of the face features;

[0025] Let the body feature vector be , where represents the th feature obtained from body language analysis, is the dimension of the body features;

[0026] For each feature combination of the face and body , define an interaction function ;

[0027] Construct a matrix , where each element , where the matrix will contain the interaction information of all face feature and body feature pairs;

[0028] Connect the head and tail of each row or column of the matrix, then the matrix is expanded into a vector, and the length of the vector is , and each element represents a specific feature interaction;

[0029] For each student, there is a series of feature vectors at a series of time points, that is, the feature vector at each time point, which is marked as , where represents the time step;

[0030] To capture the change trend of features, calculate the first-order difference of the feature vector at each time point at each time point ;

[0031] Obtain the vector norm of each time step by calculating the Euclidean norm of the vector: ;

[0032] Detect peaks in the time series of the vector norm at each time step, representing significant moments of feature change.

[0033] In a preferred embodiment, the process of obtaining the activity index is as follows:

[0034] Calculate the activity index based on the number, amplitude, and frequency of the peaks , and the calculation formula is as follows: ;

[0035] Where:

[0036] is the number of detected peaks;

[0037] is the height of each peak;

[0038] is the average interval between peaks.

[0039] In a preferred embodiment, the process of obtaining the participation index is as follows:

[0040] Analyze the consistency of students' participation over a period of time by calculating the standard deviation of the vector norm at each time step: ;

[0041] Where:

[0042] is the mean of the vector norm at each time step;

[0043] provides the average intensity of students' behavior changes, and a high mean may indicate frequent participation or activity;

[0044] is the variance of the vector norm at each time step;

[0045] By penalizing high variance, students with stable behavior are preferred because stability is usually associated with deep participation;

[0046] is the maximum autocorrelation coefficient of the vector norm at each time step at a given maximum lag .

[0047] In a preferred embodiment, the operation of the comprehensive multi-dimensional module includes the following:

[0048] Normalize and sum the activity index and the participation index with weights to obtain the wonderfulness coefficient, which is used to comprehensively evaluate the wonderfulness of the teaching video per unit time.

[0049] In a preferred embodiment, the operation process of the group behavior analysis module includes the following:

[0050] Apply the K-means clustering algorithm to the time series of the wonderful coefficient of all the students in the class to identify the group performance of the wonderful coefficient of the students in each time period;

[0051] First, collect the wonderful coefficient of all the students in the class at each time point;

[0052] Organize the wonderful coefficient according to the time series to form a two-dimensional data matrix, where each row represents a student and each column corresponds to the wonderful coefficient at a time point;

[0053] Apply the K-means clustering algorithm to the standardized data set using the selected k value;

[0054] Obtain the clustering result of the wonderful coefficient at each time point, and each cluster represents a group performance pattern;

[0055] Use kernel density estimation to perform density estimation on each cluster.

[0056] In a preferred embodiment, for each cluster, calculate the ratio of its density estimation value to the wonderful threshold, which reflects the performance of the teaching activity relative to the expected standard in each time period: ;

[0057] Wherein, represents the ratio of the density estimation value of the th cluster to the wonderful threshold, represents the density estimation value of the th cluster, represents the wonderful threshold;

[0058] Combine the relative density scores of all time periods (i.e., all clusters), and calculate the overall teaching quality index using the following formula: ;

[0059] Wherein, represents the overall teaching index, is the total number of clusters, is used to enhance the dynamic range of the cluster score;

[0060] If the overall teaching index is close to or greater than 1, it indicates that the overall teaching effect is positive relative to the set wonderful threshold, and the teaching method can effectively stimulate students' interest and participation in most time periods;

[0061] If the overall teaching index is less than 1, it indicates that the teaching effect is generally lower than the expected level, and adjustment or improvement is needed, and an alarm prompt is issued.

[0062] Technical effects and advantages of a webcam and alarm system for online teaching according to the present invention:

[0063] 1. Through image processing and deep learning technologies, the present invention realizes the monitoring and analysis of students' behaviors, ensuring that the facial expressions and upper body movements of each student can be completely captured in the teaching video. By performing real-time analysis on video frames and combining preprocessing steps such as grayscale conversion, noise removal, and image enhancement, the processing efficiency of images and the accuracy of feature recognition are improved. Using edge detection and face and upper body recognition technologies, the behaviors of students can be accurately located and evaluated, so that when the key parts of students are not completely captured, a reminder signal can be automatically sent out to guide the necessary camera or position adjustment. This process not only improves the quality of teaching videos, but also provides a reliable basis for subsequent behavior analysis and engagement assessment by ensuring data integrity, thus helping to optimize the teaching process and enhance the learning experience of students.

[0064] 2. By capturing facial and body languages through the video head and comprehensively using Haar cascade classifiers, VGG-Face deep learning models, and SVM classification, the present invention realizes the accurate recognition and analysis of students' emotional states and behaviors, helping to monitor and evaluate the emotional responses and participation dynamics of students in the online learning environment in real time, so as to deeply understand the learning experience and interaction of students. It is convenient for teachers to adjust teaching methods and course content more effectively, ensuring that teaching strategies can better meet the needs of students, enhancing the participation and student satisfaction in the learning process, and ultimately improving the teaching effect and learning outcomes. This not only helps to identify and support students who need extra attention emotionally, but also can optimize teaching design to make it more attractive and educational.

[0065] 3. By using clustering algorithms to analyze the time series of the wonderful coefficients of all students in the class, the present invention identifies the peaks and valleys of students' participation at different time periods in the teaching video, providing extremely valuable insights for teaching. It enables teachers to accurately understand which teaching contents or methods are most effective in stimulating students' interest and participation, and at the same time identify those periods that fail to effectively attract students' attention. By comparing the density estimation value of each cluster with the set wonderful threshold and calculating the overall teaching quality index, teachers can quantify the overall effect of teaching activities and judge whether they have achieved the expected teaching goals. If the overall teaching index is lower than 1, it indicates that the teaching content or method needs to be adjusted and optimized, thus promoting teachers to re-evaluate and improve teaching strategies to improve students' learning effects and overall satisfaction. Furthermore, it not only enhances the pertinence and effectiveness of curriculum design, but also improves the transparency and adjustability of the teaching process, enabling educators to better respond to the needs and preferences of students. Description of the Drawings

[0066] Figure 1Schematic flow diagram of a camera and alarm system for online teaching according to the present invention;

[0067] Figure 2 Schematic diagram of an example step of a video deployment module of a camera and alarm system for online teaching according to the present invention. Detailed implementation manners

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0069] Embodiment 1

[0070] Figure 1 A camera and alarm system for online teaching according to the present invention is provided, including: a video deployment module, an emotion and behavior recognition module, a comprehensive multi-dimensional module, and a group behavior analysis module;

[0071] The video deployment module installs and configures a camera in the learning space, captures the video stream of each student in real time, controls the camera to focus on the corresponding position of the student, ensures that the video clearly records the facial expressions and upper body movements of each student, and sends the captured video stream to the emotion and behavior recognition module;

[0072] The emotion and behavior recognition module uses video capture and multi-level machine learning techniques to analyze the facial expressions and body movements of students in online teaching, understands and measures the participation and activity levels by identifying the emotional states and behavior patterns of students, and sends the calculated result values of participation and activity to the comprehensive multi-dimensional module;

[0073] The comprehensive multi-dimensional module quantifies the activity and participation of students, calculates the teaching wonderfulness in each unit time, and sends the measured value of the teaching wonderfulness in each unit time to the group behavior analysis module;

[0074] The group behavior analysis module applies a clustering algorithm to the time series of the wonderfulness of students, analyzes and identifies the performance of student groups in each time period in the teaching video, determines the peak and trough periods of teaching by comparing density estimation with corresponding thresholds, analyzes whether the teaching needs to be adjusted, and gives an alarm prompt.

[0075] The operation process of the video deployment module includes the following:

[0076] According to the cameras for supervision provided by each student independently, control the capture of the desktops and facial expressions of each student, and identify whether the captured picture contains the facial expressions and upper body movements of the students;

[0077] Through an image processing algorithm, analyze video frames in real time to detect whether the student's face and upper body are completely within the camera's field of view.

[0078] Preprocess each frame of the captured image to improve the efficiency and accuracy of subsequent processing. The preprocessing steps include:

[0079] Gray-scale conversion: Convert the color image into a grayscale image to reduce computational complexity.

[0080] Noise removal: Use techniques such as Gaussian blur or median filtering to reduce image noise.

[0081] Image enhancement: Improve the image quality through methods such as contrast enhancement to make features more obvious.

[0082] Use an edge detection algorithm (such as the Canny edge detector) to identify the edges in the image. This helps in subsequent object recognition steps and can highlight the student's silhouette and important parts.

[0083] Apply face recognition technology (such as the Haar cascade classifier) to locate the face in the image. This is to ensure that the student's face can be accurately captured.

[0084] Utilize a trained upper body detection model (which can be a deep learning-based convolutional neural network) to identify the upper body part in the image. These models can usually identify specific parts of the human body in the image.

[0085] Analyze the position and size of the detected face and upper body to determine whether they are completely displayed within the predefined camera capture area.

[0086] If it is found that the student's face or upper body is not fully captured, an automatic reminder signal will be issued, and the position of the camera or the student needs to be adjusted.

[0087] For example, as Figure 2 shown, the following example can be combined to understand the determination of whether it is completely displayed within the predefined camera capture area:

[0088] Camera perspective: Set the camera's perspective coverage frame to a video stream of 1920x1080 pixels.

[0089] Capture area: The predefined capture area is 80% of the video center area, that is, the center area is 1536x864 pixels.

[0090] Face location: Assume that the face recognition algorithm has identified the face position and given the bounding box of the face. For example, (x = 800, y = 400, width = 200, height = 200).

[0091] Upper body positioning: The upper body detection also gives a bounding box. For example, (x = 700, y = 300, width = 400, height = 600).

[0092] Face detection: Calculate the overlapping part between the face bounding box and the predefined capture area. For example, if the entire face is within the central area of 1536x864, then the coverage rate is 100%.

[0093] Upper body detection: Similarly, calculate the coverage percentage of the upper body bounding box in the central area.

[0094] Feedback: If the coverage percentage of any detection is lower than 90%, an automatic reminder signal will be sent, and the position of the camera or the student needs to be adjusted.

[0095] Through the above steps, it can effectively ensure that the key parts of each student are fully captured, thus providing high-quality video data for subsequent behavior analysis and engagement assessment.

[0096] The present invention realizes efficient student behavior monitoring and analysis through advanced image processing and deep learning technologies, ensuring that the facial expressions and upper body movements of each student can be completely captured in the teaching video. By performing real-time analysis on video frames and combining preprocessing steps such as grayscale conversion, noise removal, and image enhancement, the processing efficiency of the image and the accuracy of feature recognition are improved. Using edge detection and face and upper body recognition technologies, it is possible to accurately locate and evaluate the behavior of students, so that when the key parts of students are not fully captured, an automatic reminder signal will be sent to guide the necessary camera or position adjustment. This process not only improves the quality of the teaching video, but also provides a reliable basis for subsequent behavior analysis and engagement assessment by ensuring the integrity of the data, thus helping to optimize the teaching process and enhance the learning experience of students.

[0097] The operation process of the emotion behavior recognition module includes the following:

[0098] Capture high-resolution video streams from the cameras of each student to ensure clear images for better recognition of expressions and actions.

[0099] First, use the Haar cascade classifier to detect faces in the video;

[0100] The Haar cascade is a machine learning-based method trained using a large number of positive and negative instances (facial and non-facial images). This method is fast and effective and can accurately detect faces in various environments.

[0101] Use a pre-trained convolutional neural network, such as VGG-Face, to extract the key features of the face.

[0102] VGG-Face is a deep learning model trained based on the VGG-16 architecture, specialized for face recognition, which can identify and extract detailed facial features such as eyebrows, eyes, mouth, etc.

[0103] The extracted facial features are input into a classification model such as SVM, which has been trained to recognize different emotional states (such as happy, sad, angry, etc.).

[0104] The following is a possible Python code implementation that shows the basic framework of the above process (assuming the relevant models and libraries have been installed and configured):

[0105] import cv2

[0106] import numpy as np

[0107] from keras.models import load_model

[0108] from sklearn.svm import SVC

[0109] # Load the Haar cascade face detector

[0110] face_cascade = cv2.CascadeClassifier('haarcascade_frontalface_default.xml')

[0111] # Load the pre-trained VGG-Face model

[0112] vgg_face = load_model('vgg_face.h5')

[0113] # Load the trained SVM model

[0114] svm_classifier = joblib.load('emotion_classifier.pkl')

[0115] def process_frame(frame):

[0116] # Convert to grayscale image

[0117] gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)

[0118] # Detect faces

[0119] faces = face_cascade.detectMultiScale(gray, 1.3, 5)

[0120] for (x, y, w, h) in faces:

[0121] # Extract the face region

[0122] face_region = frame[y:y+h, x:x+w]

[0123] # Preprocess to match VGG-Face input

[0124] face_resized = cv2.resize(face_region, (224, 224))

[0125] # Feature extraction

[0126] features = vgg_face.predict(face_resized[np.newaxis, :, :, :])

[0127] # Emotion classification

[0128] emotion = svm_classifier.predict(features)

[0129] In the provided code example, first, OpenCV is used to load the Haar cascade classifier to detect faces from video frames. Each detected face region is cropped and resized to match the input requirements of the pre-trained VGG-Face model, which is used to extract key features of the face. These features are then fed into a trained support vector machine (SVM) model, which classifies these features to identify the current emotional state, such as happy, sad, or angry. This process provides real-time facial expression recognition and emotion classification for each frame of the video, which can be used to monitor and analyze students' emotional responses.

[0130] Fuse the features extracted from facial expressions and body language to create a comprehensive feature set for more comprehensive emotion and behavior analysis.

[0131] Use facial recognition technology (such as using the VGG-Face model) to extract key features of the face. These features may include high-level descriptors extracted from various facial regions (eyes, mouth, nose, etc.).

[0132] Use a pose estimation tool (such as OpenPose) to extract the position and movement information of body key points from video data. This includes position data of the head, shoulders, arms, etc., as well as possible movement patterns such as crossed arms, nodding, etc.

[0133] To make data from different sources compatible, all features need to be normalized to ensure they are on the same scale. This may include scaling features to the same range, such as 0 to 1 or a standard normal distribution.

[0134] By concatenating features, combine the facial and body features into a single feature vector. For example, concatenate the facial feature vector and the body feature vector end-to-end.

[0135] Let the facial feature vector be , where represents the th feature obtained from facial recognition, and is the dimension of the facial features.

[0136] Let the body feature vector be , where represents the th feature obtained from body language analysis, and is the dimension of the body features.

[0137] For each feature combination of the face and body , define an interaction function . The interaction function can be non-linear. For example, . Non-linearity can reveal more complex relationships between features.

[0138] Construct a matrix , where each element , and the matrix will contain the interaction information of all face and body feature pairs.

[0139] Connect the head and tail of each row (or column) of the matrix, then the matrix is expanded into a vector, and the length of the vector is , and each element represents a specific feature interaction.

[0140] The final feature vector will be provided as input data to subsequent machine learning models for further analysis or prediction of emotional states or behavior patterns.

[0141] For each student, there is a series of feature vectors at a series of time points, that is, the feature vector at each time point, which is labeled as , where represents the time step.

[0142] To capture the changing trend of features, calculate the first-order difference of the feature vectors at each time point ;

[0143] Obtain the vector norm of each time step by calculating the Euclidean norm of the vectors: ;

[0144] Detect peaks in the time series of the vector norms at each time step, which represent significant moments of feature changes and may be related to high participation or high activity.

[0145] The process of obtaining the activity index is as follows:

[0146] Calculate the activity index based on the number, amplitude (peak height), and frequency (interval between peak appearances) of the peaks , and the calculation formula is as follows: ;

[0147] where:

[0148] is the number of detected peaks;

[0149] is the height of each peak;

[0150] is the average interval between peaks.

[0151] The activity index aims to quantify the level of activity that occurs per unit time for students. When measuring, it not only considers the frequency and intensity of students' behaviors but also combines the dynamic characteristics of behavior changes. By calculating the norm of the first-order difference of the feature vectors, the instantaneous changes in students' behaviors can be captured, and these changes reflect the reaction speed and activity of students towards teaching content. Further, by performing peak analysis on these norms (i.e., finding significant behavior changes), significant events of activity are evaluated, such as frequent gestures or strong facial expression changes. The calculation formula utilizes the number of peaks, the height of each peak, and the average interval between peaks, and these factors together reflect the frequency, intensity, and duration of activity. The purpose is to comprehensively evaluate the behavioral dynamics of students in order to more fully understand their level of participation and activity during the learning process.

[0152] The activity index is used to reflect the dynamic participation of students in online classes. When the activity index is large, it indicates that students are more active and involved in interactions, discussions, or other class activities, and may frequently answer questions, participate in discussions, or show a high level of physical activity (such as frequent use of gestures). On the contrary, a smaller activity index indicates less student participation, and they may be quieter and more passive in class, or show less physical activity. Therefore, by monitoring and analyzing the activity index, teachers can better understand the participation of each student and the classroom dynamics, and thus adjust teaching strategies to increase student participation and the effectiveness of classroom interactions.

[0153] The process of obtaining the participation index is as follows:

[0154] Analyze the consistency of student participation over a period of time by calculating the standard deviation of the vector norm for each time step:

[0155] A low standard deviation indicates that student participation is relatively consistent, which may imply continuous participation.

[0156] A high standard deviation indicates large fluctuations in participation, which may imply intermittent participation.

[0157] Use a formula similar to the activity one, incorporating the influence of the standard deviation: ;

[0158] Where:

[0159] is the mean of the vector norm for each time step;

[0160] provides the average intensity of changes in student behavior, and a high mean may indicate frequent participation or activity;

[0161] is the variance of the vector norm for each time step;

[0162] By penalizing high variance, students with stable behavior are preferred because stability is usually associated with deep participation;

[0163] is the maximum autocorrelation coefficient of the vector norm for each time step at a given maximum lag .

[0164] Considering that autocorrelation provides a measure of behavior persistence, by emphasize long-term consistency and the repeatability of behavior patterns.

[0165] The engagement index reflects the degree of students' continuous attention and active participation in classroom content. When the engagement index is large, it indicates that students not only participate in classroom activities frequently, but also their participation is continuous and consistent, showing a deep understanding and high interest in the learning content. On the contrary, a smaller engagement index indicates that students may participate occasionally but are generally passive or have unstable participation, showing a lack of interest and inattention to the course content. Therefore, by analyzing the engagement index, teachers can more accurately evaluate students' learning attitudes and behavior patterns, and then provide personalized support and intervention to improve the overall learning effect of students.

[0166] The operation of the comprehensive multi-dimensional module includes the following:

[0167] The wonderfulness coefficient is obtained by normalizing and weighted summing the activity index and the engagement index, and is used to comprehensively evaluate the wonderfulness degree of teaching videos per unit time. For example, it can be calculated through the following calculation formula: ;

[0168] Among them, represents the wonderfulness coefficient; and are the normalized activity index and engagement index respectively; and are the preset proportional coefficients of the normalized activity index and engagement index respectively, and are both greater than zero.

[0169] The wonderfulness coefficient is obtained by normalizing and weighted summing the activity index and the engagement index, and is used to represent and reflect the overall dynamics and participation depth of classroom activities per unit time. Specifically, the larger the wonderfulness coefficient, the more active students are not only physically (such as through gestures, expressions and actions) but also psychologically (such as sustained attention, emotional participation and cognitive processing) during this time period. This usually means that the classroom content can effectively attract and maintain students' interest, the teaching interaction design is appropriate, and can stimulate students' enthusiasm. On the contrary, a lower wonderfulness coefficient indicates that the classroom may lack sufficient interaction and stimulation during this time period, students' participation is not active enough or the teaching content fails to arouse students' sufficient interest, resulting in a decline in students' activity and participation. Therefore, the wonderfulness coefficient provides an important quantitative indicator for teachers to help them evaluate the attractiveness of teaching content and the effectiveness of teaching methods, so as to adjust teaching strategies in a timely manner to optimize the learning process.

[0170] The method of the present invention for comprehensively analyzing activity and engagement is an effective approach in teaching evaluation, mainly because these two indicators reflect students' performance in class from different perspectives. The activity index mainly measures the physical activity of students, such as gestures, facial expression changes, and other body movements, which helps observe students' immediate reactions and external manifestations to teaching content. The engagement index delves into students' mental activities and cognitive investment, measuring their attention, understanding, and emotional responses to course content, thereby revealing the internal acceptance and processing of learning materials by students. By combining these two indicators, teachers can obtain a more comprehensive perspective, not only seeing students' external manifestations but also gaining in-depth understanding of their learning attitudes and the quality of classroom behavior. This comprehensive analysis helps teachers identify the successful parts and areas for improvement in teaching activities, thus timely adjusting teaching strategies and methods to improve teaching effectiveness, ensuring that course content can continuously attract and motivate all students, and ultimately achieving the goal of enhancing teaching quality and learning outcomes.

[0171] The present invention captures facial and body language through a video camera, and comprehensively uses Haar cascade classifiers, VGG-Face deep learning models, and SVM classification to achieve precise identification and analysis of students' emotional states and behaviors, helping to monitor and evaluate students' emotional responses and participation dynamics in an online learning environment in real time, thereby deeply understanding students' learning experiences and interaction situations. It is convenient for instructors to more effectively adjust teaching methods and course content, ensuring that teaching strategies can better meet students' needs, enhancing engagement and student satisfaction during the learning process, and ultimately improving teaching effectiveness and learning outcomes. This not only helps identify and support students who need extra emotional attention but also can optimize teaching design to make it more attractive and educational.

[0172] The operation process of the group behavior analysis module includes the following:

[0173] Use the K-means clustering algorithm to analyze the wonderfulness coefficients of students at different time periods in the course.

[0174] Apply the clustering algorithm to the time series of the wonderfulness coefficients of all students in the class to identify the group performance of students' wonderfulness coefficients in each time period. Identify the high-value group (peak) and low-value group (trough) of the wonderfulness coefficient in each time period.

[0175] First, collect the wonderfulness coefficients of all students in the class at each time point.

[0176] Organize the wonderfulness coefficients in a time series to form a two-dimensional data matrix, where each row represents a student and each column corresponds to the wonderfulness coefficient at a time point.

[0177] Using the selected value of k, apply the K-means clustering algorithm to the standardized dataset, which can be implemented using sklearn.cluster.KMeans in Python.

[0178] Obtain the clustering results of the wonderfulness coefficients at each time point, where each cluster represents a group performance pattern.

[0179] Among them, the elbow method can be used to determine the value of k (i.e., the number of clusters) in K-means clustering.

[0180] Use kernel density estimation to estimate the density of each cluster. Compare the density estimation value of each cluster with the wonderfulness threshold;

[0181] When the density estimation value of a certain cluster is greater than or equal to the set wonderfulness threshold, it indicates that the students' participation degree is very high during the time period represented by this cluster, or in other words, the classroom activities are very attractive to students during this time period. This high density indicates that there are a large number of student data points at the wonderfulness coefficient level of this cluster, indicating that many students have shown similar highly active or deeply involved behaviors during this time period. This situation can be regarded as a "successful moment" of the classroom content or activities, perhaps because the teaching method is particularly effective, or the teaching content is particularly interesting to students;

[0182] If the density estimation value of the cluster is lower than the wonderfulness threshold, this means that the students' participation degree is not high enough during the time period represented by this cluster, or the classroom activities fail to fully attract students. The low density may indicate that only a few students show a high wonderfulness coefficient, while most students show low activity and participation. This situation may prompt teachers to adjust teaching strategies or content to increase students' participation and interest, especially during these low periods.

[0183] For each cluster, calculate the ratio of its density estimation value to the wonderfulness threshold. It reflects the performance of teaching activities in each time period relative to the expected standard: ;

[0184] Among them, represents the ratio of the density estimation value of the th cluster to the wonderfulness threshold, represents the density estimation value of the th cluster, represents the wonderfulness threshold.

[0185] Combine the relative density scores of all time periods (i.e., all clusters), and calculate the overall teaching quality index using the following formula: ;

[0186] Among them, represents the overall teaching index, is the total number of clusters, is used to enhance the dynamic range of the clustering score.

[0187] If the overall teaching index is close to or greater than 1, it indicates that the overall teaching effect is positive relative to the set excellent threshold, and the teaching method can effectively stimulate students' interest and participation during most periods;

[0188] If the overall teaching index is less than 1, it indicates that the teaching effect is generally lower than the expected level, and adjustment or improvement is needed, and an alarm prompt is issued.

[0189] The present invention analyzes the time series of the excellent coefficients of all students in the class by using a clustering algorithm, identifies the peaks and troughs of students' participation at different time periods in the teaching video, and provides extremely valuable insights for teaching. It enables teachers to accurately understand which teaching contents or methods are most effective in stimulating students' interest and participation, and at the same time identify those periods that fail to effectively attract students' attention. By comparing the density estimation value of each cluster with the set excellent threshold and calculating the overall teaching quality index, teachers can quantify the overall effect of teaching activities and judge whether they have achieved the expected teaching goals. If the overall teaching index is less than 1, it indicates that the teaching contents or methods need to be adjusted and optimized, thereby promoting teachers to re-evaluate and improve teaching strategies to improve students' learning effects and overall satisfaction. Furthermore, it not only enhances the pertinence and effectiveness of curriculum design, but also improves the transparency and adjustability of the teaching process, enabling educators to better respond to students' needs and preferences.

[0190] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.

[0191] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0192] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0193] In the several embodiments provided in this application, it should be understood that the disclosed systems and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0194] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0195] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, can also exist physically separately for each unit, or two or more units can be integrated in one unit.

[0196] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps described in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0197] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

[0198] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0199] In the implementation of the present invention, the acquisition of all user face images is based on the informed consent of the user, and the collected face data strictly complies with relevant laws, regulations, and requirements for data privacy protection.

Claims

1. A camera and alarm system for online teaching, characterized in that: include: Video deployment module, emotional behavior recognition module, comprehensive multi-dimensional module and group behavior analysis module; The video deployment module installs and configures cameras in the learning space, captures the video stream of each student in real time, controls the camera to focus on the corresponding position of the student, ensures that the video clearly records the facial expression and upper body movement of each student, and sends the captured video stream to the emotion behavior recognition module; The emotional behavior recognition module uses video capture and multi-level machine learning technology to analyze students' facial expressions and body movements during online teaching. By identifying students' emotional states and behavioral patterns, it understands and measures their participation and activity, and sends the calculated results of participation and activity to the comprehensive multidimensional module. The comprehensive multi-dimensional module quantifies the activity and participation of students, calculates the wonderful teaching situation in each unit time, and sends the measurement value of the wonderful teaching situation in each unit time to the group behavior analysis module; The group behavior analysis module applies clustering algorithms to the time series of students' outstanding performances, analyzes and identifies the group performance of students in each time period in the teaching video, determines the peak and trough periods of teaching by comparing density estimation with corresponding thresholds, analyzes whether the teaching needs to be adjusted, and gives alarm prompts; The operation process of the video deployment module includes the following: Based on the monitoring camera provided by each student, each frame of the collected image is pre-processed; Use edge detection algorithms to identify edges in images; Apply facial recognition technology to locate faces in images; Use the trained upper body detection model to identify the upper body part in the image; Analyze the position and size of the detected face and upper body to determine whether they are fully displayed within the predefined camera capture area; If it is found that the student's face or upper body is not fully captured, an automatic reminder signal will be issued, requiring the camera or student's position to be adjusted; The operation process of the emotional behavior recognition module includes the following: Capture video streams from each student’s webcam; First, use the Haar cascade classifier to detect faces in the video; Use pre-trained convolutional neural networks to extract key facial features; The extracted facial features are fed into a classification model to identify different emotional states; Use pose estimation tools to extract the position and motion information of body key points from video data; Through feature concatenation, the features of the face and body are combined into a single feature vector, as follows: Assume the facial feature vector is ,in Represents the first Features, It is the dimension of facial features; Let the body feature vector be ,in Indicates the first Features, It is a dimension of physical characteristics; For each combination of face and body features , define an interaction function ; Build a The matrix , where each element , where the matrix will contain the interaction information of all facial feature and body feature pairs; Connect each row or column of the matrix end to end, and the matrix will be expanded into a vector with a length of , each element represents a specific feature interaction; For each student, there is a series of feature vectors at each time point, that is, the feature vector at each time point, which is labeled ,in represents the time step; In order to capture the changing trend of features, the first-order difference of the feature vector at each time point is calculated. ; The vector norm at each time step is obtained by calculating the Euclidean norm of the vector: ; Detect peaks in the time series of the vector modulus length at each time step, representing significant moments of feature change; The process of obtaining the activity index is as follows: Calculate activity index based on number, amplitude and frequency of peaks , the calculation formula is as follows: ; in: is the number of peaks detected; is the height of each peak; is the average interval between peaks; The process of obtaining the participation index is as follows: Analyze the consistency of student engagement over time by calculating the standard deviation of the vector magnitude at each time step: ; in: is the mean of the vector modulus at each time step; is the variance of the vector magnitude at each time step; is the maximum hysteresis at a given The maximum autocorrelation coefficient of the vector modulus at each time step; The operation of the integrated multidimensional module includes the following: The activity index and participation index are normalized and weighted to obtain the wonderful coefficient, which is used to comprehensively evaluate the wonderfulness of the teaching video in a unit of time; The operation process of the group behavior analysis module includes the following: Apply K-means clustering algorithm to the time series of the wonderful coefficient of all students in the class to identify the group performance of the wonderful coefficient of students in each time period; First, collect the wonderful coefficients of all students in the class at each time point; Organize the wonderful coefficients in time series to form a two-dimensional data matrix, where each row represents a student and each column corresponds to the wonderful coefficient at a time point; Using the selected k value, apply the K-means clustering algorithm to the standardized dataset; Get the clustering results of the wonderful coefficient at each time point, and each cluster represents a group performance mode; Density estimation is performed for each cluster using kernel density estimation; For each cluster, the ratio of its density estimate to the excellence threshold is calculated, reflecting the performance of teaching activities in each time period relative to the expected standard: ; in, Indicates The ratio of the density estimate of the cluster to the wonderful threshold, Indicates The density estimate of the clusters, Indicates the wonderful threshold; The relative density scores of all time periods are combined to calculate the overall teaching quality index using the following formula: ; in, represents the overall teaching index, is the total number of clusters, is the dynamic range used to enhance the clustering score; If the overall teaching index is greater than 1, it indicates that the overall teaching effect is positive relative to the set wonderful threshold, and the teaching method can effectively stimulate students' interest and participation in most of the time; If the overall teaching index is less than 1, it indicates that the teaching effect is generally lower than the expected level and needs to be adjusted or improved, and an alarm will be issued.

Citation Information

Patent Citations

  • Man-machine interaction method and system for online education based on artificial intelligence

    CN107958433A

  • Online teaching system and online teaching method

    CN108281052A