A method and system for identifying classroom cognitive input based on multimodal data
By using multimodal data fusion and deep learning models, the problem of real-time, non-intrusive assessment of students' cognitive engagement in the classroom has been solved, enabling multi-dimensional, fine-grained identification of cognitive engagement and supporting personalized teaching.
Patent Information
- Application Number
- CN202310856502.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-12
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-07-12
AI Technical Summary
Existing technologies struggle to assess students’ cognitive engagement in the classroom in a real-time, non-intrusive manner, particularly in extracting multidimensional, fine-grained, and implicit dynamic engagement features.
A multimodal data fusion approach is adopted, which uses body posture, head posture, eye movement changes, facial expressions, facial units, classroom audio and speech text, combined with deep learning models (such as YOLOv8, EfficientNet, and TextCNN) to identify cognitive input, adaptively adjust the weights of modal data, and integrate learner feedback to achieve accurate perception of classroom cognitive input.
It enables multi-dimensional, fine-grained, and implicit dynamic identification of cognitive input in the classroom, meeting the real-time and accurate perception needs of cognitive input in classroom teaching, and supporting automatic perception of multi-granular classroom cognitive input and personalized teaching.
Smart Images

Figure CN117237766B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of image recognition, image classification, text classification, and text recognition technology. Specifically, it relates to a classroom cognitive input recognition method based on multimodal data. The aim is to infer classroom cognitive input by integrating implicit and psychological information contained in multimodal data, providing technical support for educational applications such as student learning in natural states and teacher classroom intervention, and helping education to develop towards precision, personalization, and intelligence. Background Technology
[0002] The deep integration of emerging information technologies such as artificial intelligence and big data with education and teaching has propelled the vigorous development of smart education. The classroom, as the main arena for education and teaching, supports diverse teaching activities and accommodates students with varying abilities, serving as a crucial place for students to acquire knowledge and master skills. However, students often exhibit inattentiveness, lack of focus, and uneven engagement in the classroom, leading to insufficient student participation. Teachers, especially novice teachers, cannot monitor each student's engagement in real time and intervene effectively. Therefore, monitoring student engagement is crucial for providing teachers with a basis for precise interventions in the classroom. Cognitive engagement is a fundamental dimension of learning engagement, but its highly implicit nature makes it difficult to directly model and measure. Existing research often uses traditional methods such as self-report for assessment, which fails to capture the dynamic development of cognitive engagement and reflect the holistic implicit mechanisms of cognitive states. In the process of classroom teaching, innovating the assessment methods of classroom cognitive engagement and constructing a scientific assessment framework to guide the non-invasive collection of multimodal learning data and the comprehensive assessment of students' cognitive engagement status are key to overcoming the problems of the current research on classroom cognitive engagement assessment, such as its one-sided and superficial content and its highly intrusive assessment methods.
[0003] Currently, common methods for measuring cognitive engagement include manual observation, video recording, self-report, interviews, teacher rating, experience sampling, physiological measurement, and text coding. Due to the time-consuming and labor-intensive nature of classroom observation and interviews, existing research typically uses them as supplementary methods to assess cognitive states. Considering the psychological characteristics of cognitive engagement, researchers often use self-report methods for assessment, with common scales including the JES scale and the SCCEI scale. These methods have relatively low requirements for the learning environment and are widely applicable to different classroom settings, often used in conjunction with other methods such as experience sampling and teacher rating. Physiological measurement methods are common in laboratory settings, but their invasiveness and high equipment costs make them difficult to meet the needs of cognitive engagement assessment in classroom settings. Video recording provides convenience for collecting cognitive engagement data in classroom settings; visual cue-based representation methods are widely popular, often using cameras in the corner of the classroom to directly record students' faces, upper bodies, and classroom audio. This type of method differs from data collected in online contexts. Its explicit form is richer than information such as text, clickstream, and timestamps. It can capture the characteristics of students' cognitive temporal changes and also places higher demands on complex interaction coding involving individuals, teachers and students, and students and students. It requires capturing more cognitive information from a broader perspective.
[0004] In conclusion, automatic perception of classroom cognitive engagement is an important direction for the development of smart education. Although some studies have explored the perception of classroom cognitive engagement using visual cues such as facial expressions and body language, difficulties remain in areas such as multi-dimensional fine-grained representation of cognitive engagement, extraction of implicit dynamic engagement features, and multi-granularity engagement recognition.
[0005] Therefore, based on the research content, this invention designs a classroom cognitive input identification method based on multimodal data to realize the automatic perception of classroom cognitive input, and provides technical support for accurate identification and perception of classroom cognitive input. Summary of the Invention
[0006] This invention addresses the current challenges of multidimensional fine-grained representation of classroom cognitive engagement, extraction of implicit dynamic engagement features, and multi-granularity engagement recognition. Starting with multimodal data, it designs an intelligent method for recognizing classroom cognitive engagement based on multimodal data to assess learners' cognitive engagement status. This invention provides a method for recognizing classroom cognitive engagement based on multimodal data, supporting non-contact, non-intrusive automatic perception of classroom cognitive engagement.
[0007] This invention provides a method for identifying classroom cognitive engagement based on multimodal data, comprising the following steps:
[0008] Step 1: From the perspective of multimodal data, construct a classroom cognitive input perception database based on multimodal data;
[0009] Step 2: Extract multimodal data from the natural classroom, conduct multimodal data analysis of classroom cognitive engagement perception, and construct a multidimensional summary model of classroom cognitive engagement representation based on the multimodal data;
[0010] Step 3: Based on learners' multimodal data and multidimensional representations of classroom cognitive engagement, deep learning methods are used to identify multidimensional cognitive engagement based on multimodal data, and finally, the cognitive engagement identification results for different modal data are output.
[0011] Step 4: Integrate the cognitive engagement results of each modality obtained in Step 3, and then adaptively adjust the weights of the recognition results of different modalities. Train the classroom cognitive engagement weight parameters based on the learners' cognitive engagement questionnaire feedback to perceive the learners' overall classroom cognitive engagement level.
[0012] Furthermore, the multimodal data includes body posture, head posture, eye movement changes, facial expressions, facial units, classroom audio, and spoken text.
[0013] Furthermore, in step 2, based on multimodal data, the characteristics of each modality are mined, thereby linking the multimodal data with multidimensional classroom cognitive engagement representation, and thus determining the performance of classroom cognitive engagement in a certain modality.
[0014] Furthermore, for the representation of body posture, head posture, eye movement change modal data and cognitive behavior dimension in step 3, the YOLOv8 model is used to mine the change features in learners' body posture, head posture, and eye movement changes, and to determine the cognitive behavior information mapped by learners' body posture, etc.
[0015] (1) Data preprocessing
[0016] Align the input body pose, head pose, and eye movement modal image sizes to 640*640, set them to RGB images, and set the channels to CHW arrangement format, etc.
[0017] (2) Backbone layer
[0018] Feature extraction was performed on body pose, head pose, and eye movement modal data. First, two consecutive 3×3 convolutions were used to reduce the resolution by a factor of 4, resulting in feature maps with 64 and 128 channels, respectively. Then, a c2f module was employed to enrich the model's gradient flow through branched cross-layer connections.
[0019] (3) Neck layer and Head layer
[0020] The features output from different stages of the backbone layer are directly fed into the upsampling operation. The feature maps are combined using a decoupled head and an anchor-free mechanism, and convolution calculations are performed on the bounding boxes.
[0021] (4) Target detection loss calculation
[0022] For target detection tasks using body pose, head pose, and eye movement modal data, the loss calculation process includes a positive / negative sample allocation strategy and loss calculation. Considering the superiority of the dynamic allocation strategy, a task-aligned strategy is adopted, selecting positive samples t based on a weighted average of classification and regression scores, calculated as follows:
[0023] t = s α ×u β (1)
[0024] s α It is the predicted value with α parameter corresponding to the labeled category, u β The calculation method is as follows:
[0025]
[0026] Where Y represents the actual labeled boxes for all students' behavioral information. The bounding box represents the predicted behavioral information of all students. Loss calculation includes classification loss and regression loss. The classification loss is calculated using the BCE Loss method, and the regression loss is calculated using the Distribution Focal Loss method and the CIoU Loss method. Finally, the above three loss calculations are weighted according to a certain weight ratio to obtain the final loss function.
[0027] 1) The classification loss CLS value is calculated as follows:
[0028]
[0029] Where M represents the number of students in the class, Y i This is the actual frame containing the behavioral information of the i-th student. It is the prediction box for the behavior information of the i-th student.
[0030] 2) The regression loss DFL and CIL values are calculated as follows:
[0031] DFL(S i ,S i+1 )=-((Y i+1 -Y)log(S i )+(YY i )log(S i+1 (4)
[0032] CIL = 1 - (u β -(loss(length)+loss(width))) (5)
[0033] Among them, S i This represents the calculation of the Softmax activation function on the limb movement features of the i-th student, transforming the new limb movement feature values into a probability distribution ranging from [0,1] to 1. `loss(length)` represents the predicted bounding boxes for the behavior information of all students. The loss value over length is the sum of the actual bounding box Y and the predicted bounding boxes for all students' behavior information. The loss value in width compared to the actual bounding box Y.
[0034] 3) Define the final loss value as L. The fusion method of the three types of losses, CLS, DFL, and CIL, is as follows:
[0035] L=λ1·CLS+λ2·DFL+λ3·CIL (6)
[0036] Wherein, λ1, λ2 and λ3 are the fusion weight parameters, and their values range from [0,1].
[0037] Furthermore, for facial expression and facial unit modal data, the EfficientNet model is used to mine the variation features in facial expressions and facial units to determine the cognitive and emotional information mapped by the learner's face. This process consists of nine computational stages, as shown below;
[0038] (1) Stage 1: The learner’s facial expression shallow feature values are obtained by ordinary convolution with a kernel size of 3×3 and a stride of 2.
[0039] (2) Stages 2 to 8: Deep facial expression feature information of learners is output through repeated stacked MBConv structures. The MBConv structure mainly improves the dimensionality of facial expression modality data through a 1×1 ordinary convolution. The number of convolution kernels is p times the number of channels of the input feature matrix, where p∈{1,6}. Then, a q×q Depthwise Conv convolution (where q=3 or 5) and an SE module are used to extract key facial features. Then, a 1×1 ordinary convolution is used to reduce the dimensionality of the dataset with key facial features. Finally, a Dropout layer is used to prevent overfitting and generate a new facial expression feature map with deep feature information.
[0040] (3) Stage 9: It consists of one ordinary convolutional layer, one max pooling layer, and one fully connected layer, and finally outputs the cognitive and emotional information reflected by the learner's facial expressions.
[0041] Furthermore, using the TextCNN model, we can mine the changing features of learners' classroom audio and speech text modal data to determine the cognitive speech information mapped by learners' utterances.
[0042] (1) The first layer is the input layer. The input layer is an n×k matrix input, where n is the number of words in a sentence and k is the dimension of the word vector corresponding to each word. That is, each row of the input layer is a k-dimensional word vector corresponding to a word.
[0043] (2) The second layer is a convolutional layer. The convolution kernel is set to [value] on the input matrix. The output of the convolutional layer is c, calculated using the following formula:
[0044]
[0045] in, This represents a convolution operation, where c = [c1, c2, ..., c...]. n-h+1 ] represents the new word feature vector extracted by the convolutional layer, c1, c2, ..., c n-h+1 Let c represent the feature vector of each student's first sentence (s1), second sentence (s2), and so on, up to the last sentence spoken in the lesson. i The calculation method is as follows:
[0046] c i =f(w·x) i:i+h-1 +b) (8)
[0047] Where, x i:i+h-1 This represents a window of size h×k consisting of rows i to i+h-1 of the input matrix, defined by x. i x i+1 ... x i+h-1 It is composed of two parts, where b is the bias parameter and f is the nonlinear activation function.
[0048] (3) The third layer is the pooling layer. Max pooling, K-Max pooling, or average pooling can be used to further filter the new text feature vectors output by the convolutional layer. Max pooling selects the largest feature from the feature vectors generated by each sliding window and then concatenates these features to form a vector representation. K-Max pooling selects the K largest features from each feature vector, and average pooling averages each dimension of the feature vector. All of these methods achieve the same effect: obtaining a fixed-length vector representation from sentences of different lengths through pooling.
[0049] (4) The fourth layer is a fully connected layer and a text classification output. The goal of the TextCNN model is to classify learner speech data, so a fully connected layer is concatenated and the Softmax activation function is used to output the probability of each class of speech data.
[0050] Furthermore, a summary model of classroom cognitive engagement representation is constructed from three dimensions: learners' cognitive behavior, cognitive emotion, and cognitive language. The specific construction steps are as follows;
[0051] (1) The cognitive behavior dimension in classroom cognitive input is comprehensively represented by body posture, head posture and eye movement change modal data. For the classroom video frame at time f, the image corresponding to that time is vectorized, and each pixel in the image is represented by a number in [0,9] as the representation result A of body posture, head posture and eye movement change modal data.
[0052] (2) The cognitive-emotional dimension of classroom cognitive input is represented by facial expressions and facial unit modal data. For the classroom video frame at time f, the face is first automatically extracted using the OpenCV library. The extracted face image is used as the basis for the cognitive-emotional representation at that time. Then, each pixel in the color image is represented by the number [0,9] to form the final representation result B.
[0053] (3) The cognitive verbal dimension in classroom cognitive input is quantitatively represented by classroom audio and speech text modal data. The cognitive verbal dimension is represented by two methods: pre-trained word vectors and parameterized word vectors. The representation result is C.
[0054] Furthermore, in step 1, a classroom cognitive input perception database based on multimodal data is constructed, and the specific implementation method is as follows;
[0055] (1) In the classroom environment, teachers conduct classroom teaching in a natural state, which includes several classroom learners who participate in classroom activities and knowledge construction under the guidance of teachers. Teachers are also allowed to integrate advanced technological tools and different teaching models to carry out rich classroom activities to meet the needs of different learning stages.
[0056] (2) Capture learners' cognitive state in a non-invasive and non-perceptive manner. Install several high-definition cameras at the front and back of the classroom. Turn them on before class to record students' classroom learning in real time. Turn them off after class and export the classroom video data from the terminal system as a data source for classroom cognitive input perception.
[0057] (3) During the data annotation process, learners’ real cognitive engagement feelings were obtained through post-class questionnaires. The Likert five-point scoring method was used as the annotation standard for learners’ overall engagement status. Video image data was used as the annotation source for learners’ cognitive behavior. Automatic face extraction was performed on the video image data. Learners’ cognitive emotions were annotated through facial image data. Classroom audio data was used as the annotation source for learners’ cognitive speech.
[0058] (4) For each data label, a portion of classroom video data is selected first, and multiple coders are used to label it simultaneously. Any inconsistencies are discussed before large-scale classroom cognitive input data labeling is carried out.
[0059] (5) In order to obtain the cognitive engagement status of learners at different granularities, classroom video frames can be extracted as needed. If the classroom video frame rate is 25fps, the frame extraction rate can be selected as 25 frames / time, 50 frames / time, ..., 25*f frames / time, etc. (f is an integer) to train the summary model of classroom cognitive engagement representation.
[0060] Furthermore, in step 4, the recognition of cognitive input across the three dimensions of cognitive behavior, cognitive emotion, and cognitive speech is achieved through the following methods;
[0061] (1) Assume that the three input vectors of the i-th learner perceived at time j are respectively and in For cognitive-behavioral input, For cognitive and emotional engagement, Representing cognitive verbal input, n1, n2, and n3 represent the dimensions of the three input feature vectors, respectively;
[0062] (2) Given a learning activity, assuming there are F real-time input recognitions during the entire learning activity, train three cognitive input state recognition networks respectively, and calculate and infer the cognitive behavior input results through a deep learning model. Cognitive-emotional engagement results and cognitive verbal input results
[0063] (3) Based on the three cognitive input results automatically identified, and according to the overall level of joint perceptual cognitive input feedback from the learner, the overall cognitive input level of learner j at time i is then determined. j The calculation is as follows:
[0064]
[0065] Here, β1, β2, and β3 are the network parameters to be learned.
[0066] This invention also provides a classroom cognitive engagement recognition system based on multimodal data, comprising the following modules:
[0067] The database construction module is used to build a classroom cognitive input perception database based on multimodal data from the perspective of multimodal data.
[0068] The multi-dimensional representation module is used to extract multimodal data from the natural classroom, conduct multimodal data analysis of classroom cognitive engagement perception, construct a multi-dimensional summary model of classroom cognitive engagement representation based on the multimodal data, and obtain a multi-dimensional representation of classroom cognitive engagement.
[0069] The multi-dimensional recognition module is used to identify multi-dimensional cognitive input based on learners' multimodal data and classroom cognitive input. It uses deep learning methods to identify multi-dimensional cognitive input based on multimodal data and finally outputs the cognitive input recognition results for different modal data.
[0070] The results fusion module is used to fuse the cognitive engagement results of each modality obtained by the multi-dimensional recognition module, and then adaptively adjust the weights of the recognition results of different modalities. Based on the learners' cognitive engagement questionnaire feedback, the classroom cognitive engagement weight parameters are trained to perceive the learners' overall classroom cognitive engagement level.
[0071] Compared with existing research and technology, this invention has the following advantages:
[0072] 1. This invention combines educational and psychological theories with deep learning methods to establish a multi-modal data-driven, multi-dimensional, fine-grained classroom cognitive engagement representation computational model. It analyzes the internal mechanism of classroom cognitive engagement from three aspects: cognitive behavior, cognitive emotion, and cognitive language, meeting the real-time and accurate perception needs of cognitive engagement during classroom learning and laying the foundation for multi-granular classroom cognitive engagement perception.
[0073] 2. This invention addresses the problem of recognizing cognitive input in different modalities by using three types of deep learning models for automatic recognition. It proposes a Yolov8 model based on body posture, an EfficientNet model based on facial expressions, and a TextCNN model based on speech and text, thereby improving the recognition model's ability to learn implicit dynamic input features and facilitating its application in actual classrooms.
[0074] 3. This invention constructs a fine-grained classroom cognitive engagement identification method that integrates multimodal data, and on this basis, integrates learners' self-reported feedback to design a cognitive engagement identification method based on real-world perception data. This adaptively adjusts the contribution of different modal data to the final classroom cognitive engagement, meeting the multi-level and multi-stage cognitive engagement perception needs in actual classroom applications. Attached Figure Description
[0075] Figure 1 A diagram illustrating a multimodal data-driven model for representing classroom cognitive engagement.
[0076] Figure 2 A diagram illustrating the observation indicator system for data input in classroom cognition;
[0077] Figure 3 This is a structural diagram of the Yolov8 model based on body posture, etc.
[0078] Figure 4 The structure diagram of the EfficientNet model based on facial expressions, etc.
[0079] Figure 5 This is a diagram of the TextCNN model structure based on speech and text.
[0080] Figure 6 A flowchart for identifying cognitive input in a natural classroom setting;
[0081] Figure 7 The image shows the training results of the Yolov8 model based on body posture, etc.
[0082] Figure 8 The image shows the training results of the EfficientNet model based on facial expressions, etc.
[0083] Figure 9 This is a diagram showing the training results of the TextCNN model based on speech and text. Detailed Implementation
[0084] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.
[0085] This invention provides an intelligent recognition method for classroom cognitive input based on multimodal data, comprising the following steps:
[0086] Step 1: From the perspective of multimodal data, construct a classroom cognitive input perception database based on multimodal data;
[0087] Step 2: Extract multimodal data from the natural classroom, conduct multimodal data analysis of classroom cognitive engagement perception, and construct a multidimensional summary model of classroom cognitive engagement representation based on the multimodal data;
[0088] Step 3: Based on learners' multimodal data and multidimensional representations of classroom cognitive engagement, deep learning methods are used to identify multidimensional cognitive engagement based on multimodal data, and finally, the cognitive engagement identification results for different modal data are output.
[0089] Step 4: Integrate the cognitive engagement results of each modality obtained in Step 3, and then adaptively adjust the weights of the recognition results of different modalities. Train the classroom cognitive engagement weight parameters based on the learners' cognitive engagement questionnaire feedback to perceive the learners' overall classroom cognitive engagement level.
[0090] To achieve the above objectives, a classroom cognitive engagement perception database based on multimodal data is constructed from the perspective of multimodal data. Specific steps include:
[0091] (1) In the classroom environment, teachers conduct classroom teaching in a natural state, which includes several classroom learners who participate in classroom activities and knowledge construction under the guidance of teachers. Teachers are also allowed to integrate advanced technological tools and different teaching models to carry out rich classroom activities to meet the needs of different learning stages.
[0092] (2) Capture learners' cognitive state in a non-invasive and non-perceptive manner. Install several high-definition cameras at the front and back of the classroom. Turn them on before class to record students' classroom learning in real time. Turn them off after class and export the classroom video data from the terminal system as a data source for classroom cognitive input perception.
[0093] (3) During the data annotation process, learners' actual cognitive engagement was obtained through post-class questionnaires. The Likert five-point scoring method was used as the annotation standard for learners' overall engagement. Video image data was used as the annotation source for learners' cognitive behavior. Automatic face extraction was performed on the video image data, and learners' cognitive emotions were annotated using facial image data. Classroom audio data was used as the annotation source for learners' cognitive speech. Data annotation observation indicators were proposed, such as… Figure 2 As shown.
[0094] (4) For each data label, a portion of classroom video data is selected first, and multiple coders are used to label it simultaneously. Any inconsistencies are discussed before large-scale classroom cognitive input data labeling is carried out.
[0095] (5) In order to obtain the cognitive engagement status of learners at different granularities, classroom video frames can be extracted as needed. If the classroom video frame rate is 25fps, the frame extraction rate can be selected as 25 frames / time, 50 frames / time, ..., 25*f frames / time, etc. (f is an integer and ≤ the total number of classroom video frames / 25) to train the summary model of classroom cognitive engagement representation.
[0096] Furthermore, this invention extracts multimodal data on perceived cognitive engagement in the classroom and constructs a summary model representing this engagement. The multimodal data involved includes body posture, head posture, eye movement changes, facial expressions, facial units, classroom audio, and spoken text. The cognitive engagement dimensions covered include cognitive behavior, cognitive emotion, and cognitive speech. Figure 1 As shown. The specific construction steps of the summary model of classroom cognitive engagement representation are as follows;
[0097] (1) The cognitive behavior dimension of classroom cognitive engagement is comprehensively represented by body posture, head posture, and eye movement modal data. For the classroom video frame at time f, the image corresponding to that time is vectorized, and each pixel in the image is represented by a number in the range [0,9]. Color images have three channels: R, G, and B, so pixels in all three channels are represented as the representation results of body posture, head posture, and eye movement modal data. Representation method A is as follows:
[0098]
[0099] Among them, a fi Let F represent the cognitive behavior feature matrix of the i-th student at time f, where i ranges from [1, M], M represents the number of students in the class, and F represents the total number of frames extracted from the class video.
[0100] (2) The cognitive-emotional dimension of classroom cognitive engagement is represented by facial expressions and facial unit modal data. For each classroom video frame at any given moment, faces are first automatically extracted using the OpenCV library. The extracted face images are used as the basis for cognitive-emotional representation at that moment. Then, each pixel in the color image is represented by a number in the range [0,9], forming the final representation result B as follows:
[0101]
[0102] Among them, b fi Let b represent the cognitive-emotional feature matrix of the i-th student at time f, and b fi For a fi A subset of.
[0103] (3) The cognitive verbal dimension in classroom cognitive input is quantitatively represented by classroom audio and speech text modal data. The cognitive verbal dimension is represented by two methods: pre-trained word vectors and parameterized word vectors. The representation method is as follows:
[0104]
[0105] in, Let f represent the cognitive speech feature word vector of the i-th student at time f. Let f represent the cognitive speech feature word vector with μ parameters for the i-th student at time f.
[0106] Furthermore, fine-grained recognition of multi-dimensional cognitive input is carried out for multi-modal data, including body posture, head posture, eye movement changes, facial expressions, facial units, classroom audio and speech text;
[0107] Furthermore, the methods for extracting and calculating modal data such as body posture, head posture, and eye movement changes are as follows;
[0108] The input cognitive-behavioral dimension representation A mainly adopts the Yolov8 model, such as... Figure 3 As shown, we mine the changing features in learners' body movements to determine the cognitive and behavioral information mapped from learners' body postures. To further verify the effectiveness of the model, we have conducted relevant experiments on a self-built dataset, and the training results are as follows. Figure 7 As shown, passive behavior was found to be the most effective at being identified.
[0109] (1) Data preprocessing
[0110] Since the cognitive behavior dimension representation A is key behavioral information extracted from images of students learning in the classroom, further settings are needed for representation A to match the input format of the Yolov8 model. Therefore, the input body pose, head pose, and eye movement modal images are aligned to 640*640, set to RGB images, and the channels are set to CHW arrangement format, etc.
[0111] (2) Backbone layer
[0112] Feature extraction was performed on body pose, head pose, and eye movement modal data. First, two consecutive 3×3 convolutions were used to reduce the resolution by a factor of 4, resulting in feature maps with 64 and 128 channels, respectively. Then, a c2f module was employed to enrich the model's gradient flow through branched cross-layer connections.
[0113] (3) Neck layer and Head layer
[0114] The features output from different stages of the backbone layer are directly fed into the upsampling operation. The feature maps are combined using a decoupled head and an anchor-free mechanism, and convolution calculations are performed on the bounding boxes.
[0115] (4) Target detection loss calculation
[0116] For target detection tasks using body pose, head pose, and eye movement modal data, the loss calculation process includes a positive / negative sample allocation strategy and loss calculation. Considering the superiority of the dynamic allocation strategy, a task-aligned strategy is adopted, selecting positive samples t based on a weighted average of classification and regression scores, calculated as follows:
[0117] t = s α ×u β (9)
[0118] s α It is the predicted value with α parameter corresponding to the labeled category, u β The calculation method is as follows:
[0119]
[0120] Where Y represents the actual labeled boxes for all students' behavioral information. The bounding box represents the predicted behavioral information of all students. Loss calculation includes classification loss and regression loss. The classification loss is calculated using the BCE Loss method, and the regression loss is calculated using the Distribution Focal Loss method and the CIoU Loss method. Finally, the above three loss calculations are weighted according to a certain weight ratio to obtain the final loss function.
[0121] 1) The classification loss CLS value is calculated as follows:
[0122]
[0123] Where M represents the number of students in the class, Y i This is the actual frame containing the behavioral information of the i-th student. It is the prediction box for the behavior information of the i-th student.
[0124] 2) The regression loss DFL and CIL values are calculated as follows:
[0125] DFL(S i ,S i+1 )=-((Y i+1 -Y)log(S i )+(YY i )log(S i+1 (12)
[0126] CIL = 1 - (u β -(loss(length)+loss(width))) (13)
[0127] Among them, S iThis represents the calculation of the Softmax activation function on the limb movement features of the i-th student, transforming the new limb movement feature values into a probability distribution ranging from [0,1] to 1. `loss(length)` represents the predicted bounding boxes for the behavior information of all students. The loss value over length is the sum of the actual bounding box Y and the predicted bounding boxes for all students' behavior information. The loss value in width compared to the actual bounding box Y.
[0128] 3) Define the final loss value as L. The fusion method of the three types of losses, CLS, DFL, and CIL, is as follows:
[0129] L=λ1·CLS+λ2·DFL+λ3·CIL (14)
[0130] Wherein, λ1, λ2 and λ3 are the fusion weight parameters, and their values range from [0,1].
[0131] Furthermore, the methods for extracting and calculating modal data such as facial expressions and facial units are as follows;
[0132] The input representation B of the cognitive-emotional dimension primarily employs the EfficientNet model, such as... Figure 4 As shown, we mine the variation features in facial expressions and facial units to determine the cognitive and emotional information mapped by the learner's face. To further verify the effectiveness of the model, we have conducted relevant experiments on our self-built dataset, obtaining results by setting different hyperparameters. Figure 8 The experimental results shown indicate that the final accuracy of emotion recognition reached 91%. Since the cognitive emotion dimension representation B is extracted from key cognitive emotion information in images of students' classroom learning, further settings are needed for representation B to match the input format of the EfficientNet model. These settings include aligning the image size to 224*224, setting it to RGB image, and setting the channels to CHW arrangement format. Then, nine computational stages are performed, as shown below:
[0133] (1) Stage 1: Perform ordinary convolution calculation with convolution kernel w' (size 3×3, stride 2) and the representation input B of the cognitive emotion dimension. Obtain learner's superficial facial feature values (FACES):
[0134]
[0135] (2) Stages 2 to 8: Corresponding Figure 3 Blocks 1 to 7 are the core modules for feature computation. They mainly output the learner's deep facial feature map FACES' through repeated stacking of MBConv structures.
[0136] FACES'=MBConv⊙FACES (16)
[0137] Where ⊙ represents the feature calculation method of the MBConv structure, the MBConv structure mainly uses a 1×1 ordinary convolution to increase the dimensionality of modal data such as facial expressions (corresponding to...). Figure 3 The Module 1 module has a convolution kernel number that is p times the number of channels in the input feature matrix, where p ∈ {1, 6}. It then uses a q × q Depthwise Conv convolution (where q = 3 or 5) and an SE module to further extract key facial features (corresponding to...). Figure 4 The Module 2 module in the text then uses a 1×1 ordinary convolution to reduce the dimensionality of the dataset containing key facial features (corresponding to...). Figure 4 The module (Module 3) is used to generate a new facial feature map FACES' with deep feature information, which is then prevented from overfitting by using a Dropout layer.
[0138] (3) Stage 9: Consists of one ordinary convolutional layer, one max pooling layer, and one fully connected layer (corresponding to...) Figure 4 The Final Layers module outputs cognitive and emotional information mapped from the learner's facial expressions at time j.
[0139]
[0140] Where pool() represents pooling computation, fc() represents fully connected computation, and b represents the real bias to be trained.
[0141] Furthermore, the methods for extracting and calculating classroom audio and speech text modal data are as follows;
[0142] The input cognitive verbal dimension representation C mainly employs the TextCNN model, such as... Figure 5 As shown, we mine the variation features in learners' classroom audio and speech text to determine the cognitive verbal information mapped by the learners' classroom audio and speech text modalities. To further verify the effectiveness of the model, we have conducted multiple rounds of hyperparameter tuning experiments on our self-built dataset. The training results with the hyperparameter epoch set to 50 are shown below. Figure 9 As shown.
[0143] (1) The first layer is the input layer (corresponding to) Figure 5The encoding module (of the input layer) is an n×k matrix, where n is the number of words in a sentence and k is the dimension of the word vector for each word. That is, each row of the input layer is a k-dimensional word vector corresponding to a word. Each word vector is the cognitive speech dimension representation C of the aforementioned input, which can be pre-trained on other corpora or obtained by training the network as unknown parameters. Here, a dual-channel approach is used, meaning the cognitive speech dimension representation C has two input matrices: a pre-trained word vector and a parameterized word vector.
[0144] (2) The second layer is a convolutional layer (corresponding to) Figure 5 (The Convolution layer module). On the input matrix, set the convolution kernel to... The output of the convolutional layer is c, calculated using the following formula:
[0145]
[0146] in, This represents a convolution operation, where c = [c1, c2, ..., c...]. n-h+1 ] represents the new word feature vector extracted by the convolutional layer, c1, c2, ..., c n-h+1 Let c represent the feature vector of each student's first sentence (s1), second sentence (s2), and so on, up to the last sentence spoken in the lesson. i The calculation method is as follows:
[0147] c i =f(w·x) i:i+h-1 +b) (19)
[0148] Where, x i:i+h-1 This represents a window of size h×k consisting of rows i to i+h-1 of the input matrix, defined by x. i x i+1 ... x i+h-1 It is composed of two parts, where b is the bias parameter and f is the nonlinear activation function.
[0149] (3) The third layer is the pooling layer (corresponding to) Figure 5(The Pooling layer module). You can choose max pooling, K-Max pooling, or average pooling to further filter the new feature vectors (Vec) output by the convolutional layer. Max pooling selects the largest feature from the feature vectors generated by each sliding window and then concatenates these features to form a vector representation. K-Max pooling selects the K largest features from each feature vector, and average pooling averages each dimension of the feature vector. All of these methods achieve the same effect: obtaining a fixed-length vector representation from sentences of different lengths through pooling.
[0150] Vec=pool(c) (20)
[0151] (4) The fourth layer is a fully connected layer and text classification output (corresponding to) Figure 5 The TextCNN model aims to classify learner speech data. Therefore, it uses fully connected layers and a Softmax activation function to output the probability of each class of speech data perceived at time j, denoted as .
[0152]
[0153] like Figure 6 As shown, based on the learners' classroom cognitive engagement questionnaire feedback results, the decision-making process integrates cognitive behavior recognition results, cognitive emotion recognition results, and cognitive verbal recognition results. Specific steps include:
[0154] (1) The three input vectors of the i-th learner perceived at time j are respectively and in For cognitive-behavioral input, For cognitive and emotional engagement, Representing cognitive verbal input, n1, n2, and n3 represent the dimensions of the three input feature vectors, respectively;
[0155] (2) Given a learning activity, assuming there are F real-time input recognitions during the entire learning activity, train three cognitive input state recognition networks respectively, and calculate and infer the cognitive behavior input results through a deep learning model. (Passive behavior, active behavior, constructive behavior, interactive behavior, behavioral disengagement), cognitive and emotional engagement results (Positive emotions, negative emotions, emotional detachment) and cognitive-verbal engagement results (low-level speech, high-level speech, speech disengagement);
[0156] (3) Based on the three cognitive input results automatically identified, and according to the overall level of joint perceptual cognitive input (low input, medium input, high input) reported by the learner, the overall cognitive input level of learner j at time i is then determined. j The calculation is as follows:
[0157]
[0158] Among them, Engagement j ∈{0,1,2}, where β1, β2 and β3 are the network parameters to be learned, ranging from [0,1].
[0159] Similarly, we can perceive higher-level or more nuanced levels of engagement in a progressively coarser manner, which helps in identifying online learning engagement at different levels and stages. Based on this, we collected classroom cognitive engagement data in primary schools and applied our representation model and data annotation observation index system to this dataset. We obtained experimental results for the classroom cognitive engagement identification model under optimal hyperparameter conditions, as shown in the table below. Our identification method achieved good results across all evaluation indicators, including P, R, and F1, further validating the effectiveness of the aforementioned method. Going forward, we will further expand the scale of the educational dataset and integrate different category weights to improve the model's generalization ability. This has high application value in the classroom environment, helping teachers understand students' learning status and thus optimize classroom teaching.
[0160] Table 1. Results of Classroom Cognitive Engagement Identification
[0161]
[0162] This invention also provides a classroom cognitive engagement recognition system based on multimodal data, comprising the following modules:
[0163] The database construction module is used to build a classroom cognitive input perception database based on multimodal data from the perspective of multimodal data.
[0164] The multi-dimensional representation module is used to extract multimodal data from the natural classroom, conduct multimodal data analysis of classroom cognitive engagement perception, construct a multi-dimensional summary model of classroom cognitive engagement representation based on the multimodal data, and obtain a multi-dimensional representation of classroom cognitive engagement.
[0165] The multi-dimensional recognition module is used to identify multi-dimensional cognitive input based on learners' multimodal data and classroom cognitive input. It uses deep learning methods to identify multi-dimensional cognitive input based on multimodal data and finally outputs the cognitive input recognition results for different modal data.
[0166] The results fusion module is used to fuse the cognitive engagement results of each modality obtained by the multi-dimensional recognition module, and then adaptively adjust the weights of the recognition results of different modalities. Based on the learners' cognitive engagement questionnaire feedback, the classroom cognitive engagement weight parameters are trained to perceive the learners' overall classroom cognitive engagement level.
[0167] The specific implementation methods of each module are the same as those of each step, and will not be described in this invention.
[0168] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
Claims
1. A method for identifying classroom cognitive engagement based on multimodal data, characterized in that, Includes the following steps: Step 1: From the perspective of multimodal data, construct a classroom cognitive input perception database based on multimodal data; Step 2: Extract multimodal data from the natural classroom, conduct multimodal data analysis of classroom cognitive engagement perception, and construct a multidimensional summary model of classroom cognitive engagement representation based on the multimodal data to obtain a multidimensional representation of classroom cognitive engagement. The multimodal data in step 2 includes body posture, head posture, eye movement changes, facial expressions, facial units, classroom audio, and spoken text; In step 2, a summary model of classroom cognitive engagement representation is constructed from three dimensions: learners' cognitive behavior, cognitive emotion, and cognitive language. The specific steps for constructing the summary model of classroom cognitive engagement representation are as follows; (1) The cognitive behavior dimension in classroom cognitive input is comprehensively represented by body posture, head posture and eye movement change modal data. For the classroom video frame at time f, the image corresponding to that time is vectorized and each pixel in the image is represented by a number in [0,9] as the representation result A of body posture, head posture and eye movement change modal data. (2) The cognitive-emotional dimension of classroom cognitive input is represented by facial expressions and facial unit modal data. For each classroom video frame at each moment, the face is first automatically extracted using the OpenCV library. The extracted face image is used as the basis for the cognitive-emotional representation at that moment. Then, each pixel in the color image is represented by the number [0,9] to form the final representation result B. (3) The cognitive verbal dimension in classroom cognitive input is quantitatively represented by classroom audio and speech text modal data. The cognitive verbal dimension is represented by two methods: pre-trained word vectors and parameterized word vectors. The representation result is C. Step 3: Based on learners' multimodal data and multidimensional representations of classroom cognitive engagement, deep learning methods are used to identify multidimensional cognitive engagement based on multimodal data, and finally, the cognitive engagement identification results for different modal data are output. Step 4: Integrate the cognitive engagement results of each modality obtained in Step 3, and then adaptively adjust the weights of the recognition results of different modalities. Train the classroom cognitive engagement weight parameters based on the learners' cognitive engagement questionnaire feedback to perceive the learners' overall classroom cognitive engagement level.
2. The classroom cognitive engagement recognition method based on multimodal data as described in claim 1, characterized in that: Using the Yolov8 model to analyze body posture, head posture, and eye movement modal data, we can mine the changing features in learners' limb movements and determine the cognitive and behavioral information mapped by the learners' body posture, head posture, and eye movement modalities. Specifically, this includes the following steps: (1) Data preprocessing Align the input body pose, head pose, and eye movement modal images to 640*640 pixels, set them to RGB images, and set the channels to CHW arrangement format. (2) Backbone layer Feature extraction is performed on body pose, head pose, and eye movement modal data. First, two consecutive 3×3 convolutions are used to reduce the resolution by a factor of 4, resulting in feature maps with 64 and 128 channels, respectively. Then, a c2f module is used to enrich the gradient flow of the model through branch cross-layer connections. (3) Neck layer and Head layer The features output from different stages of the Backbone layer are directly fed into the upsampling operation. The feature maps are combined using a decoupling head and an anchor-free mechanism, and convolution calculations are performed on the bounding boxes. (4) Target detection loss calculation For target detection tasks using body pose, head pose, and eye movement modal data, the loss calculation process includes a positive / negative sample allocation strategy and loss calculation. Considering the superiority of the dynamic allocation strategy, a task-aligned strategy is adopted, selecting positive samples t based on a weighted average of classification and regression scores, as calculated below: t=s α ×u β (1) Among them, s α It is the predicted value with α parameter corresponding to the labeled category, u β The actual labeled box Y and the predicted box for behavioral information represent the behavior information of all students. The loss value with the β parameter, Loss calculation includes classification loss and regression loss. The classification loss is calculated using the BCE Loss method, and the regression loss is calculated using the Distribution Focal Loss method and the CIoU Loss method. Finally, the above three loss calculations are weighted by a certain weight ratio to obtain the final loss function. 1) The classification loss CLS value is calculated as follows: Where M represents the number of students in the classroom, Y i This is the actual frame containing the behavioral information of the i-th student. It is the prediction box for the behavior information of the i-th student; 2) The regression loss DFL and CIL values are calculated as follows: DFL(S i ,S i+1 )=-((Y i+1 -Y)log(S i )+(Y-Y i )log(S i+1 ))(3) CIL=1-(u β -(loss(length)+loss(width)))(4) Among them, S i This represents the calculation of the Softmax activation function on the limb movement features of the i-th student, transforming the new limb movement feature values into a probability distribution ranging from [0,1] to 1. The loss(length) represents the predicted bounding boxes for the behavior information of all students. The loss value over length is the sum of the actual bounding box Y and the predicted bounding boxes for all students' behavior information. The loss value in width relative to the actual bounding box Y; Finally, the three types of losses, CLS, DFL, and CIL, are fused and weighted to obtain the final target loss function.
3. The classroom cognitive input identification method based on multimodal data as described in claim 1, characterized in that: For facial expression and facial unit modal data, the EfficientNet model is used to mine the variation features in facial expressions and facial units, and to determine the cognitive and emotional information mapped by the learner's facial expressions. The specific processing procedure is as follows: (1) Stage 1: Obtain the shallow facial feature values of the learner by performing ordinary convolution with a kernel size of 3×3 and a stride of 2. (2) Stages 2 to 8: Deep facial feature information of the learner is output through repeated stacked MBConv structures; where the MBConv structure increases the dimensionality of the facial modality data through a 1×1 ordinary convolution, the number of convolution kernels is p times the number of channels of the input feature matrix, p∈{1,6}, and then extracts key facial features through a q×q Depthwise Conv convolution and a SE module. Then, a 1×1 ordinary convolution is used to reduce the dimensionality of the dataset with key facial features. Finally, a Dropout layer is used to prevent overfitting and generate a new facial feature map with deep feature information; where q=3 or 5; (3) Stage 9: It consists of one ordinary convolutional layer, one max pooling layer, and one fully connected layer, and finally outputs the emotional information reflected by the learner's facial expressions.
4. The classroom cognitive engagement identification method based on multimodal data as described in claim 1, characterized in that: For classroom audio and speech text modal data, the TextCNN model is used to mine the variation features in learners' classroom audio and speech text, and to determine the cognitive speech information mapped by learners' utterances; the specific processing procedure is as follows; (1) The first layer is the input layer: The input layer is an n×k matrix input, where n is the number of words in a sentence and k is the word vector dimension of each word. That is to say, each row of the input layer is a k-dimensional word vector corresponding to a word. (2) The second layer is a convolutional layer: On the input matrix, the convolution kernel is set to... The output of the convolutional layer is the feature vector c of all sentences, where c is the feature vector of each sentence. i The calculation method is as follows: c i =f(w·x i:i+h-1 +b) (5) Where, x i:i+h-1 This represents a window of size h×k consisting of rows i to i+h-1 of the input matrix, defined by x. i x i+1 ... x i+h-1 It is composed of multiple parts, where b is the bias parameter and f is the nonlinear activation function; (3) The third layer is the pooling layer: select one of max pooling, K-Max pooling or average pooling operations to further filter the new text feature vectors output by the convolutional layer; (4) The fourth layer is a fully connected layer and text classification output: The goal of the TextCNN model is to classify learner speech data. Therefore, a fully connected layer is concatenated and the Softmax activation function is used to output the probability of each class of speech data.
5. The classroom cognitive input identification method based on multimodal data as described in claim 1, characterized in that: Step 1 involves constructing a classroom cognitive engagement perception database based on multimodal data. The specific implementation method is as follows: (1) In the classroom environment, the teacher conducts classroom teaching in a natural state, which includes several classroom learners who participate in classroom activities and knowledge construction under the guidance of the teacher; (2) Capture learners’ cognitive state in a non-invasive and non-perceptive manner; install several high-definition cameras at the front and back of the classroom, turn them on before class to record students’ classroom learning in real time, turn them off after class and export classroom video data from the terminal system as a data source for classroom cognitive input perception. (3) During the data annotation process, learners’ real cognitive engagement feelings were obtained through post-class questionnaires. The Likert five-point scoring method was used as the annotation standard for learners’ overall engagement status. Video image data was used as the annotation source for learners’ cognitive behavior. Automatic face extraction was performed on the video image data. Learners’ cognitive emotions were annotated through facial image data. Classroom audio data was used as the annotation source for learners’ cognitive speech. (4) For each data label, a portion of classroom video data is selected first, and multiple coders are used to label it simultaneously. Any inconsistencies are negotiated before large-scale classroom cognitive input data labeling is carried out. (5) In order to obtain the learner’s cognitive engagement status at different granularities, classroom video frames are extracted as needed. If the classroom video frame rate is 25fps, the frame extraction rate is selected as 25*f frames / time, where f is an integer and ≤ the total number of classroom video frames / 25, which is used to train the summary model of classroom cognitive engagement representation.
6. The classroom cognitive input identification method based on multimodal data as described in claim 1, characterized in that: In step 4, the following methods are used to identify cognitive input across the three dimensions of cognitive behavior, cognitive emotion, and cognitive speech; (1) Assume that the three input vectors of the i-th learner perceived at time j are respectively and in For cognitive-behavioral input, For cognitive and emotional engagement, Representing cognitive verbal input, n1, n2, and n3 represent the dimensions of the three input feature vectors, respectively; (2) Given a learning activity, assuming there are F real-time input recognitions during the entire learning activity, train three cognitive input state recognition networks respectively, and calculate and infer the cognitive behavior input results through a deep learning model. Cognitive-emotional engagement results and cognitive verbal input results (3) Based on the three cognitive input results automatically identified, and according to the overall level of joint perceptual cognitive input feedback from the learner, the overall cognitive input level of learner j at time i is then determined. j The calculation is as follows: Here, β1, β2, and β3 are the network parameters to be learned.
7. A classroom cognitive engagement recognition system based on multimodal data, characterized in that, Includes the following modules: The database construction module is used to build a classroom cognitive input perception database based on multimodal data from the perspective of multimodal data. The multi-dimensional representation module is used to extract multimodal data from the natural classroom, conduct multimodal data analysis of classroom cognitive engagement perception, construct a multi-dimensional summary model of classroom cognitive engagement representation based on the multimodal data, and obtain a multi-dimensional representation of classroom cognitive engagement. Multimodal data includes body posture, head posture, eye movement changes, facial expressions, facial units, classroom audio, and spoken text; A summary model of classroom cognitive engagement representation is constructed from three dimensions: learners’ cognitive behavior, cognitive emotion, and cognitive language. The specific steps for constructing the summary model of classroom cognitive engagement representation are as follows; (1) The cognitive behavior dimension in classroom cognitive input is comprehensively represented by body posture, head posture and eye movement change modal data. For the classroom video frame at time f, the image corresponding to that time is vectorized and each pixel in the image is represented by a number in [0,9] as the representation result A of body posture, head posture and eye movement change modal data. (2) The cognitive-emotional dimension of classroom cognitive input is represented by facial expressions and facial unit modal data. For each classroom video frame at each moment, the face is first automatically extracted using the OpenCV library. The extracted face image is used as the basis for cognitive-emotional representation at that moment. Then, each pixel in the color image is represented by the number [0,9] to form the final representation result B. (3) The cognitive verbal dimension in classroom cognitive input is quantitatively represented by classroom audio and speech text modal data. The cognitive verbal dimension is represented by two methods: pre-trained word vectors and parameterized word vectors. The representation result is C. The multi-dimensional recognition module is used to identify multi-dimensional cognitive input based on learners' multimodal data and classroom cognitive input. It uses deep learning methods to identify multi-dimensional cognitive input based on multimodal data and finally outputs the cognitive input recognition results for different modal data. The results fusion module is used to fuse the cognitive engagement results of each modality obtained by the multi-dimensional recognition module, and then adaptively adjust the weights of the recognition results of different modalities. Based on the learners' cognitive engagement questionnaire feedback, the classroom cognitive engagement weight parameters are trained to perceive the learners' overall classroom cognitive engagement level.