A learner confusion degree classification method, a teaching effect evaluation method, and a device
Patent Information
- Application Number
- CN202410358519.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-27
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-03-27
AI Technical Summary
[0007]针对现有技术的缺陷,本申请的目的在于提供一种学习者困惑度分类方法、教学效果评价方法及设备,旨在解决现有技术不能对困惑表情强度进行良好的估计,导致无法对教学效果进行有效评估的问题
[0057]本申请提供一种学习者困惑度分类方法、教学效果评价方法及设备,提出了一种面向视频的基于多层级特征融合的困惑表情监测方法与教学评价应用,该方法能在无监督的条件下自动提取图片序列中的代表帧,并利用图片的动作单元信息,融合图片的表情关系特征;考虑了从局部到全局的重要视觉特征,实现了困惑表情的细粒度分类;并通过构建的积极困惑与消极困惑模型,最终对教学效果做出量化评价。
Smart Images

Figure CN118155265B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence, and more specifically, relates to a learner confusion classification method, a teaching effectiveness evaluation method, and a device. Background Technology
[0002] Traditional teaching evaluation methods suffer from low efficiency and high labor costs, leading many researchers to propose intelligent methods for evaluating teacher effectiveness. Existing research mainly falls into two categories: process-based evaluation and outcome-based evaluation. Process-based evaluation typically considers classroom behavior data generated during learning, quantifying indicators such as learner participation and concentration to reflect teacher effectiveness. Outcome-based evaluation usually analyzes students' post-lesson test or exam scores to reflect teacher effectiveness.
[0003] However, both of these methods have certain drawbacks. The former typically assesses teaching effectiveness based on learner-generated process data, such as characterizing various classroom behaviors by detecting student actions to measure teaching effectiveness. This approach provides real-time feedback on the classroom situation, but commonly used classroom behaviors (such as raising hands, standing, and looking around) do not directly reflect the learner's level of knowledge acquisition and lack a strong correlation with teaching effectiveness. While outcome-based methods can directly reflect the teacher's teaching effectiveness, the outcome data is post-hoc and lacks real-time relevance, resulting in a shortcoming of not being able to provide real-time feedback during the learning process. Therefore, existing research suffers from drawbacks such as indirect and non-real-time evaluation, and cannot effectively solve the problem of evaluating classroom teaching effectiveness.
[0004] A teacher's teaching effectiveness is directly reflected in students' learning behavior, and the occurrence of learning behavior is directly related to the confusion and distress experienced during the learning process. In fact, learning confusion, as a frequent learning emotion, has a direct and significant impact on learning activities. Therefore, evaluating teaching effectiveness through the identification of confusion is a direct and real-time method. However, because learning confusion differs from other learning emotions, and compared to other emotion measurement methods, confusion is more implicit and more difficult to measure, and several common methods for measuring learning confusion also have different limitations in specific scenarios.
[0005] Current research largely focuses on identifying the occurrence of confused emotions, with limited attention paid to the fine-grained intensity monitoring of confused expressions themselves. Furthermore, existing expression intensity estimation methods are not well-suited for tasks involving the estimation of confused expression intensity in learning scenarios. This is because, in real-world learning environments, confused emotions are often intermingled with other emotions, resulting in significant class gaps between confused expressions of varying intensities.
[0006] Existing intensity estimation methods based on six basic facial expressions often model the intensity of expressions within the same category, thus failing to provide accurate estimations of the intensity of confused expressions. Specifically, some existing methods perform intensity identification within a single category, first determining the expression's category and then classifying its intensity with fine granularity. However, there are numerous expression transformation relationships between emotions; for example, confused emotions can easily transform into positive or negative emotions. When the level of confusion is high or low, it may exceed the original expression category. Therefore, existing intensity monitoring models based on single-category expressions are not suitable for confusion level monitoring tasks. Summary of the Invention
[0007] In view of the shortcomings of the prior art, the purpose of this application is to provide a learner confusion classification method, a teaching effectiveness evaluation method and device, which aims to solve the problem that the prior art cannot accurately estimate the intensity of confused expressions, thus making it impossible to effectively evaluate the teaching effectiveness.
[0008] To achieve the above objectives, firstly, this application provides a learner confusion classification method, comprising the following steps:
[0009] Acquire video sequences containing learner facial images in a teaching setting;
[0010] The video sequence is divided into multiple video sub-sequences, and the peak frames in each video sub-sequence are located based on the facial expression contour features of each frame image.
[0011] Extract global features of the learner's face from each peak frame;
[0012] Local features of the learner's face and activation patterns of facial action units (AUs) corresponding to each peak frame are extracted from each peak frame. Based on the similarity of the activation patterns of facial action units between any two peak frames, all peak frames are combined to construct an unweighted graph convolutional neural network. The global and local features of each peak frame are fused as the features of the corresponding nodes in the graph convolutional neural network.
[0013] The output features of the graph convolutional neural network are used as the final features of the video sequence and input into the model classification layer to predict the perplexity, thereby determining the learner's perplexity reflected in each peak frame.
[0014] Based on the distribution of confusion intensity of each peak frame within a preset time period, the learner's confusion state within that time period is classified as either positive or negative confusion.
[0015] In one possible implementation, the peak frames in each video subsequence are located based on the facial expression contour features of each frame image, including:
[0016] The differences in facial expression contour features between any two frames in this video subsequence are compared using cosine similarity.
[0017] Compare the distances in time between the two frames with the largest differences from the center of the video subsequence, and take the frame with the smallest distance as the peak frame.
[0018] In one possible implementation, the facial expression contour features are local binary encoded histograms.
[0019] In one possible implementation, the peak frame (peak) is determined by the following formula:
[0020]
[0021] Among them, F i and F j Let F be any two image frames in a video subsequence F, where LBPH() represents the LBPH feature of the image frames and cos_sim() represents the cosine similarity. The time at the center of the sequence is indicated by n, and n represents the total number of image frames in the video subsequence. This represents the time between two image frames when the cosine similarity is minimum.
[0022] In one possible implementation, global features of the face in the peak frame are extracted using a well-trained global face feature extraction model;
[0023] Local facial features and facial action unit activation patterns in peak frames are extracted using a trained facial action unit intensity detector.
[0024] The unweighted graph neural network is specifically defined as follows: when the cosine similarity between the features of two peak frames is greater than or equal to a preset similarity threshold, there is an edge between the nodes corresponding to the two peak frames; otherwise, there is no edge.
[0025] In one possible implementation, the learner's perplexity within a preset time period is classified based on the distribution of perplexity intensity of each peak frame, including:
[0026] When the duration of confusion exceeding a preset intensity within a preset time period is less than the preset time, the intensity of confusion changes periodically, and the duration of high-intensity confusion is shorter than the duration of low-intensity confusion, it is classified as positive confusion.
[0027] A situation is classified as negative confusion when at least one of the following conditions is met: the duration of confusion exceeding a preset intensity within a preset time period is greater than the preset time; the intensity of confusion increases linearly; or the duration of high-intensity confusion is longer than the duration of low-intensity confusion.
[0028] Secondly, this application provides a method for evaluating teaching effectiveness, including the following steps:
[0029] The perplexity classification method, based on the first aspect or any possible implementation of the first aspect, determines the occurrence time of positive perplexity and the occurrence time of negative perplexity in the video sequence;
[0030] The occurrence rate and continuity of positive confusion in the teaching scenario are determined based on the frequency and timing of the occurrence of positive and negative confusion.
[0031] The teaching effectiveness in the teaching scenario is evaluated based on the product of the proportion of positive confusion and the continuity of confusion. The larger the value of the product, the better the teaching effectiveness.
[0032] In one possible implementation, when there is only one learner in the teaching scenario, the quantified value TE of the teaching effectiveness is:
[0033]
[0034] Among them, the larger the TE, the higher the teaching quality; P positive C represents the proportion of positive confusion; C represents the continuity of confusion.
[0035] P positve = count(po) / ALLstate
[0036]
[0037] The `count` function is used to calculate the number of times a corresponding state occurs. `count(po)` represents the number of times positive confusion occurs, `ALLstate` represents the total number of states with both positive and negative confusion, and `s`... i This represents the state of confusion within the i-th time window, with a value of -1 or 1, representing negative or positive confusion respectively.
[0038] In one possible implementation, when there are multiple learners in the teaching scenario, the teaching effectiveness is determined through the following steps:
[0039] The average perplexity distribution of multiple learners is determined based on the perplexity of each learner in each peak frame;
[0040] Based on the average confusion level of multiple learners in each time window, using the confusion level classification method described in the first aspect or any possible implementation of the first aspect, determine the occurrence time of average positive confusion and the occurrence time of average negative confusion in the video sequence.
[0041] The occurrence rate and duration of average positive confusion and average negative confusion in the teaching scenario are determined based on the frequency and timing of their occurrence.
[0042] The teaching effectiveness in the teaching scenario is evaluated based on the product of the average positive confusion occurrence rate and the average confusion continuity.
[0043] In one possible implementation, when there is only one learner in the teaching scenario, the average quantitative value ATE of the teaching effectiveness is:
[0044] ATE = AC * AP positive
[0045] Among them, the larger the ATE value, the higher the teaching quality; AC represents the average confusion continuity, and AP... positive This indicates the average percentage of positive confusion.
[0046] The AC and AP positive The following steps were used to calculate the result:
[0047] The average perplexity distribution ACD of n learners is determined based on the perplexity of each learner in each peak frame: ACD = {ac1, ac2, ..., ac3}. tk},in, c in The degree of confusion of the nth student within the i-th time window, ac i This represents the average perplexity of the i-th time window;
[0048] The average occurrence time of positive confusion and the average occurrence time of negative confusion in the video sequence are determined based on the average confusion level of n learners in each time window.
[0049] AP positve = count(po) / ALLstate
[0050]
[0051] The `count` function is used to calculate the average number of occurrences of the corresponding state. `count(po)` represents the average number of occurrences of the positive confusion state, `ALLstate` represents the sum of the number of occurrences of both positive and negative confusion states, and `s`... i This represents the state of confusion within the i-th time window, with a value of -1 or 1, representing negative or positive confusion respectively.
[0052] Thirdly, this application provides an electronic device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method described in the first aspect or any possible implementation thereof, or the second aspect or any possible implementation thereof.
[0053] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed on a processor, causes the processor to perform the methods described in the first aspect or any possible implementation of the first aspect, or the second aspect or any possible implementation of the second aspect.
[0054] Fifthly, this application provides a computer program product that, when run on a processor, causes the processor to execute the method described in the first aspect or any possible implementation of the first aspect, or the second aspect or any possible implementation of the second aspect.
[0055] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.
[0056] Overall, the technical solutions conceived in this application have the following beneficial effects compared with the prior art:
[0057] This application provides a learner confusion classification method, a teaching effectiveness evaluation method, and a device. It proposes a video-based method for monitoring confused expressions based on multi-level feature fusion and its application in teaching evaluation. This method can automatically extract representative frames from image sequences under unsupervised conditions and utilize the action unit information of the images to fuse the facial expression relationship features. It considers important visual features from local to global perspectives, achieving fine-grained classification of confused expressions. Finally, through the construction of positive and negative confusion models, it makes a quantitative evaluation of teaching effectiveness. Attached Figure Description
[0058] Figure 1 A flowchart of the learner confusion classification method provided in the embodiments of this application.
[0059] Figure 2 A flowchart of a teaching effectiveness evaluation method provided in an embodiment of this application;
[0060] Figure 3 A diagram of the peak frame localization algorithm based on maximizing feature difference provided in the embodiments of this application.
[0061] Figure 4A flowchart of a perplexity detection method based on the fusion of AU intensity cues and visual features provided in an embodiment of this application.
[0062] Figure 5 A comparison diagram of the confusion state transformation model provided in this application embodiment and an existing model.
[0063] Figure 6(a) is a flowchart of the teaching effect evaluation application in the "one-on-one" scenario provided in the embodiment of this application.
[0064] Figure 6(b) is a flowchart of the teaching effect evaluation application in the "one-to-many" scenario provided in the embodiments of this application.
[0065] Figure 7 The overall flowchart of the image-based student confusion classification method and teaching evaluation application provided in the embodiments of this application is shown.
[0066] Figure 8 This is an architectural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0068] In this application, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A existing alone, A and B existing simultaneously, and B existing alone. In this application, the symbol " / " indicates that the related objects are in an "or" relationship, for example, A / B means A or B.
[0069] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0070] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, for example, multiple learners means two or more learners, etc.
[0071] First, the technical terms involved in the embodiments of this application will be introduced.
[0072] (1) Confusion
[0073] Confusion refers to feeling doubtful and unsure of what to do.
[0074] (2) Confusion
[0075] Confusion level refers to the degree of confusion felt.
[0076] Next, the technical solutions provided in the embodiments of this application will be described.
[0077] To address the shortcomings of existing methods for identifying confusion levels and their applications in teaching evaluation, this application proposes a teaching effectiveness evaluation method based on confusion level analysis. This method monitors students' confusion levels, assesses their confusion states, and then analyzes the effectiveness of classroom teaching. Specifically, this application constructs a multi-feature fusion graph convolutional network using facial action unit relationships to perform fine-grained segmentation of confusion emotions. Subsequently, based on the changing patterns of confusion levels, a confusion state model is constructed, and a time window-based confusion state determination method is designed for real-world scenarios. Finally, a classroom teaching effectiveness evaluation application is built in common "one-on-one" and "one-to-many" teaching scenarios. Practice has proven that the proposed method can effectively monitor learners' confusion emotions at a fine-grained level, assess students' confusion states accordingly, and ultimately reflect the teacher's teaching effectiveness. Compared with existing solutions, this application has advantages in terms of directness, efficiency, and real-time performance.
[0078] Currently, there are two main approaches to recognizing facial expression intensity: static expression-based recognition and dynamic expression sequence-based recognition. Static expression-based methods can predict the intensity of an expression in a single image, but their effectiveness is generally limited by the amount of information in that image. Video-based recognition can completely preserve the process of an expression's occurrence, thus potentially allowing for effective judgment of the expression intensity in each frame of a sequence; however, this sacrifice of efficiency does not result in a significant improvement in accuracy. In fact, due to limited datasets, researchers often focus on static image-based expression recognition. However, this approach requires extensive manual annotation and ignores the inherent intensity differences between different expressions and the possibility of expression transformation. Therefore, it is not suitable for learning scenarios.
[0079] Meanwhile, most existing work focuses on detecting confusion levels without establishing a reliable mapping between confusion and learning outcomes. In other words, current work lacks confusion-based teaching evaluation methods. Specifically, using confusion levels alone is insufficient for evaluating teaching effectiveness. In a classroom setting, determining the positive or negative nature of current confusion is beneficial for more efficient classroom assessment, but existing work does not effectively distinguish between positive and negative confusion.
[0080] Based on this, this application proposes a video-based method for detecting confused expressions and its application in teaching evaluation. This method can automatically extract representative frames from image sequences under unsupervised conditions and fuse the facial expression relationship features of the images using the AU information of the images. It considers important visual features from local to global perspectives, achieving fine-grained classification of confused expressions. Finally, through the construction of positive and negative confusion models, it makes a quantitative evaluation of teaching effectiveness.
[0081] Figure 1 A flowchart of the learner confusion classification method provided in the embodiments of this application is shown below. Figure 1 As shown, it includes the following steps:
[0082] S101, Obtain a video sequence containing learner facial images in a teaching scenario;
[0083] S102, the video sequence is divided into multiple video sub-sequences, and the peak frame in each video sub-sequence is located based on the facial expression contour features of each frame image.
[0084] S103, extract global features of the learner's face in each peak frame;
[0085] S104, extract the local features of the learner's face in each peak frame and the activation pattern of the facial action unit corresponding to each peak frame, and combine all peak frames based on the similarity of the activation patterns of the facial action unit between any two peak frames to construct an unweighted graph convolutional neural network, and fuse the global and local features of each peak frame as the features of the corresponding node in the graph convolutional neural network.
[0086] S105, the output features of the graph convolutional neural network are used as the final features of the video sequence and input into the model classification layer to predict the perplexity, thereby determining the learner's perplexity reflected in each peak frame;
[0087] S106, classify the learner's confusion state within the preset time period according to the distribution of confusion intensity of each peak frame within the preset time period, and determine whether it is positive confusion or negative confusion.
[0088] Further, see Figure 2 The diagram shown is a flowchart of a teaching effectiveness evaluation method provided in an embodiment of this application, including the following steps:
[0089] based on Figure 1 The provided perplexity classification method determines the occurrence time of positive perplexity and the occurrence time of negative perplexity in the video sequence;
[0090] The occurrence rate and continuity of positive confusion in the teaching scenario are determined based on the frequency and timing of the occurrence of positive and negative confusion.
[0091] The teaching effectiveness in the teaching scenario is evaluated based on the product of the proportion of positive confusion and the continuity of confusion. The larger the value of the product, the better the teaching effectiveness.
[0092] In a specific embodiment, in order to obtain representative peak frames in an image sequence under unsupervised information conditions, this application proposes a feature difference maximization peak frame localization algorithm (MFD-PFL) based on the LBP method, which reduces video information to image information.
[0093] In emotion quantification tasks within learning scenarios, facial expressions exhibit inconsistent distributions, such as large intra-class spacing and small inter-class spacing. Therefore, general expression intensity recognition networks do not perform well in learning scenarios. This application proposes a perplexity level detection method based on the fusion of AU intensity cues and visual features (IV-CDM). This method utilizes the similarity between action unit patterns, incorporating various expression transformation relationships present in learned emotions, and then fuses relational features with multi-level visual features to effectively determine the student's perplexity level.
[0094] To address the issue of complete isolation between positive and negative emotions in existing research, this application provides definitions and characteristics of positive and negative confusion. Based on the given positive and negative confusion, quantitative state models of positive and negative confusion are proposed, using the time-domain and frequency-domain characteristics of confusion degree: Positive Confusion State Model (PCSM) and Negative Confusion State Model (NCSM).
[0095] To determine positive or negative confusion during extended classroom sessions, this application proposes a confusion assessment method based on time windows. This method monitors the confusion distribution of students within each window using a sliding time window approach. Based on the previously proposed PCSM and NCSM models, it calculates the confusion state corresponding to each time window. In short, this application determines the student's confusion state within a time window based on the sequence of image confusion within that time window.
[0096] Based on changes in students' confusion levels, this application proposes a confusion state determination method to identify students' confusion states. To further evaluate the teacher's teaching effectiveness, this section introduces a method for evaluating teaching effectiveness based on confusion states. Considering diverse learning scenarios, subsequent analysis and modeling were conducted under the common "one-on-one" and "one-to-many" learning models.
[0097] Step 1 is a prerequisite for subsequent steps 2 and 4. This step, through research on traditional unsupervised peak frame localization algorithms and analysis based on the needs of actual teaching scenarios, proposes a peak frame annotation model for scenarios with small facial expression changes: the Feature Difference Maximization Peak Frame Localization Algorithm (MFD-PFL) based on the LBP method. The overall model flow framework is as follows: Figure 3 As shown.
[0098] Step 2, based on the representative frames obtained in Step 1, constructs a graph neural network using the similarity between action unit patterns across multiple frames. While incorporating various facial expression transformation relationships, it uses information fused from local and global visual features as graph nodes, ultimately effectively determining the student's level of confusion. Figure 4 .
[0099] See Figure 4 As shown, on the one hand, this application uses a training set (X) train ,Y train The intensity label Y of X,Y is used to fine-tune the pre-trained face model. This fine-tuning process enables the model to adapt to the expression recognition problem in the current problem domain. The fine-tuned model is then used to extract features from the overall data X to obtain global features. On the other hand, this application utilizes the open-source facial action unit detector OpenFace for facial action unit recognition. The identified facial action unit activation patterns are used to determine the edge relationships in the graph convolutional network, while the extracted facial action unit intensity features serve as local features of the face. Based on the above design, this application integrates local and global features and constructs a graph convolutional neural network based on facial action unit activation patterns to ultimately achieve perplexity classification.
[0100] Step 3 analyzes the shortcomings of existing definitions of confusion and, based on existing educational viewpoints, constructs a reasonable model of "positive confusion" and "negative confusion." It provides definitions and salient features of positive and negative confusion, ultimately determining the student's state of confusion. Existing confusion models are as follows... Figure 5 As shown in the middle left figure, the confusion state transition model proposed in this application is as follows: Figure 5As shown in the right-middle figure, based on this confusion state transformation model, a method for determining confusion states in real-world scenarios is presented. First, the concept of a time window is adopted, and the confusion degree distribution over a period of time is obtained based on step 2; then, the proposed confusion state model is used to determine the confusion state of students within this time window.
[0101] Step 4 presents a teaching evaluation scheme for a learning scenario. Specifically, based on the proposed confusion level identification method and teaching effectiveness evaluation method, teaching effectiveness evaluation applications for "one-to-one" and "one-to-many" scenarios are constructed respectively. The business process diagrams of the applications are shown in Figure 6(a) and Figure 6(b).
[0102] Compared with existing technologies, this application has at least one of the following advantages:
[0103] The perplexity detection method proposed in this application, which is based on the fusion of AU intensity cues and visual features, constructs a similarity relationship between different expressions and intensities by using the action unit activation patterns in different facial images, enabling the model to adapt to intensity monitoring tasks when there are expression transitions.
[0104] This paper improves the definition of confusion state in learning scenarios, overcomes the challenge that existing works do not distinguish well between positive and negative confusion, and proposes a clear Positive Confusion State Model (PCSM) and Negative Confusion State Model (NCSM).
[0105] This application constructs a reasonable "classroom confusion-teaching effect" mapping model and provides a technical solution for evaluating teaching effectiveness in "one-on-one" and "one-to-many" teaching scenarios.
[0106] like Figure 7 This application presents the overall process architecture for teaching evaluation based on perplexity identification. The method can be divided into four steps: selection of representative frames in the image sequence, classification of perplexity in static images, determination of perplexity state, and quantification of teaching effectiveness.
[0107] Step 1: Selection of representative frames from the image sequence
[0108] To extract peak frames from an image sequence without supervision, this application proposes a Feature Difference Maximization Peak Frame Localization (MFD-PFL) algorithm based on the LBP method. Its formal expression is: From the input image sequence F = {f0, f1, ... f...} n-1 Extracting peak frames f peak =MFDPFL(F). This algorithm consists of two parts: facial expression feature extraction and feature selection.
[0109] The feature extraction part employs the traditional Local Binary Pattern (LBP) algorithm to extract texture features near facial expressions. The feature extraction module extracts the Local Binary Pattern Histogram (LBPH) features for each face in the image sequence. Let the number of images in the sequence be k. The obtained features can be represented by a matrix L_h = {A_1,…,A_k}^T, where A_i has dimensions (m,n), corresponding to the size of the original image.
[0110] After the feature extraction module completes its work, the feature selection module needs to obtain the peak frame from all the acquired features. This application proposes that, in a complete expression sequence, the following two assumptions are satisfied:
[0111] Assumption 1: Among the two frames with the greatest difference in the same expression sequence, one of them must be the peak frame F_peak.
[0112] Hypothesis 2: F_peak is most likely to appear in the middle of the sequence.
[0113] Based on the above assumptions, this application presents the design of the MFD-PFL algorithm. First, LBPH features are extracted from all images in the sequence F = {F_0,…F_(n-1)}. Cosine similarity is used to compare the differences in features between any two images, and the frame with the largest difference is selected. Then, the temporal distance between the two frames with the largest differences and the sequence center is compared; the frame with the smallest distance is the peak frame. Formalized as:
[0114]
[0115] Step 2: Perplexity Classification of Static Images
[0116] Building upon step 1, a perplexity detection method based on the fusion of facial action unit intensity cues and visual features (IV-CDM) is used to detect facial perplexity in the sample images. Specifically, the process of constructing the IV-CDM model can be formally described as follows: Given an input expression image X = {x_i ∈ R^(C*H*W)}, labeled perplexity Y = [y_0, y_1, y_2, y_3], the model F(f(EVIT(x), AU, GraphER(X))) = Y is trained using a pre-trained downstream expression intensity model EVIT, action unit intensity features AU = {au_0, au_1, ..., au_n}, and the expression relation extraction method GraphER based on GCN. Here, C, H, and W represent the number of channels, pixel height, and pixel width of the image itself, respectively; y refers to the perplexity level labeled in the dataset.
[0117] Specifically, in constructing the graph network, this application proposes a GCN method based on facial expression relationships. It uses the cosine similarity of action unit patterns between images to establish undirected edge features E. This approach considers that there is greater similarity in action unit patterns between expressions of the same intensity. That is:
[0118]
[0119] Since different identities and environments can lead to changes in appearance in real-world scenarios, directly measuring the similarity of AUs (Authors and Entities) between two images using the aforementioned similarity metrics could make the constructed relationships more susceptible to the influence of non-expression factors. Therefore, this application employs unweighted undirected edge graph construction and adopts... Define the relationships between nodes. Here, th represents a manually set similarity threshold for action units.
[0120] Step 3: Determine the state of confusion
[0121] Building upon step 2, this step presents a method for determining a state of confusion.
[0122] a) Positive Confusion State Model
[0123] Since positive confusion ultimately leads to positive learning outcomes, students generally move on to the next knowledge point after resolving the confusion; therefore, positive confusion exhibits a cyclical pattern. Conversely, prolonged unresolved confusion transforms into negative confusion, thus positive confusion typically lasts shorter, with higher intensity periods shorter than lower intensity periods. In other words, positive confusion manifests as a cyclical increase in confusion level, lasting a short period at its peak before rapidly declining. Specific characteristics are given in Table 1. Based on these characteristics, this application presents the following model of positive confusion states.
[0124] Table 1
[0125]
[0126] Positive Confusion State Model (PCSM): For a sequence of images within a time window of duration T and frame rate K, given the perplexity distribution:
[0127] CD={c_1,…,c_k}(c_k∈{c_1,…c_M},M≤th*K)
[0128] Where th represents the confusion time period threshold, and C_M can be represented as the union of four sets of images with different confusion levels: C M =C0∪C1∪C2∪C3, and satisfy |C3|<|C2|; at the same time, if the perplexity can be expressed as a Fourier expansion of a periodic function over time:
[0129]
[0130] The C_M phase is then called the active confusion phase. Where, a i ,b i Let be a constant coefficient in the Fourier expansion, and t be used to determine the possible time period, which is limited by a threshold th. This formula attempts to determine whether there exists a t in the interval [0, th] such that the entire perplexity distribution undergoes a periodic transformation.
[0131] b) Negative Confusion State Model (NCSM): For a sequence of images within a time window of duration T and frame rate K, given the confusion distribution:
[0132] CD = {c1,…,c} k}(c k ∈{c1,…c M},M≤th*K)
[0133] g can be expressed as a function of time:
[0134]
[0135] α is a hyperparameter that represents the degree of confusion development that varies from person to person; or when C_M can be represented as the union of four levels: C M When C0∪C1∪C2∪C3, the following conditions are met: Both of the above situations can be considered as C M The stage is the negative confusion stage. This formula attempts to determine, within the interval [0, th], the existence of a t such that the overall confusion distribution increases.
[0136] It is worth noting that in actual implementation, it is generally not possible to find t such that the perplexity distribution exactly fits the two distributions mentioned above. Therefore, this application needs to make approximations during implementation.
[0137] Step 4: Quantifying Teaching Effectiveness
[0138] Building upon step 3, this section introduces a method for evaluating teaching effectiveness based on the state of confusion in order to further evaluate the teacher's teaching effectiveness.
[0139] In one-on-one teaching scenarios, fewer factors need to be considered. This application uses the time-window-based confusion state determination method proposed earlier to evaluate the teaching effectiveness in one-on-one scenarios. Factors that can be considered for teaching effectiveness in single-student scenarios include: the proportion of positive confusion (P). positive The proportion of negative confusion P negativeThe continuity of confusion, C, is defined as follows:
[0140] P negative = count(ne) / ALLstate
[0141] P positve = count(po) / ALLstate
[0142]
[0143] The `count` function is used to calculate the number of times a corresponding state occurs, `ALLstate` represents the total number of states, and `s`... i This represents the state of confusion within the i-th time window, with a value of -1 or 1, representing negative or positive confusion respectively.
[0144] The above method, based on a simple state of confusion, constructs three factors, P, to measure teaching effectiveness. negative ,P negative With C. Where P positive With P negative The duration of students' active or passive confusion during a lesson can be quantified, and C measures the stability of students' confusion (the closer the value is to 1, the more stable the confusion). Overall, this application reflects the teacher's teaching effectiveness (TE) by considering the state and quality of learning confusion.
[0145]
[0146] The closer the value is to 1, the higher the teaching quality; 0 is the lowest.
[0147] For one-to-many teaching environments, further definitions are needed based on the one-to-one scenario due to the involvement of monitoring multiple students. Considering the potential instability in multi-student learning situations (such as students leaving midway or inaccurate image capture), this section presents a method for calculating the average teaching quality (ATE) based on the PCSM and NCSM models proposed in this application.
[0148] An extended definition of the perplexity distribution in PCSM and NCSM is given based on the number of visible students n within the current time window: Average perplexity distribution ACD = {ac1, ac2, ..., ac...} tk}, where ac is the average based on the number of students: c in This refers to the confusion level of the nth student within the i-th SW. Other calculations are the same as for a single student. Therefore:
[0149] ATE = AC * AP positive
[0150] Here, AC represents the average continuity of confusion, and AP_positive represents the average proportion of positive confusion. Similarly, the closer the ATE value is to 1, the higher the teaching quality, with 0 being the lowest.
[0151] Based on the methods in the above embodiments, this application provides an electronic device, see [link to relevant documentation]. Figure 8 The system includes a processor 810, a communications interface 820, a memory 830, and a communication bus 840. The processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions stored in the memory 830 to execute the methods described in the above embodiments.
[0152] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0153] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0154] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.
[0155] It is understood that the processor in the embodiments of this application can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.
[0156] The method steps in this application embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.
[0157] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0158] It is understood that the various numerical designations used in the embodiments of this application are merely for the convenience of description and are not intended to limit the scope of the embodiments of this application.
[0159] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for classifying learner confusion, characterized in that, Includes the following steps: Acquire video sequences containing learner facial images in a teaching setting; The video sequence is divided into multiple video sub-sequences, and the peak frames in each video sub-sequence are located based on the facial expression contour features of each frame image. Extract global features of the learner's face from each peak frame; Local features of the learner's face and the activation patterns of facial action units corresponding to each peak frame are extracted from each peak frame. Based on the similarity of the activation patterns of facial action units between any two peak frames, all peak frames are combined to construct an unweighted graph convolutional neural network. The global and local features of each peak frame are fused as the features of the corresponding node in the graph convolutional neural network. The output features of the graph convolutional neural network are used as the final features of the video sequence and input into the model classification layer to predict the perplexity, thereby determining the learner's perplexity reflected in each peak frame. Based on the distribution of confusion intensity of each peak frame within a preset time period, the learner's confusion state within that time period is classified as either positive or negative confusion. Based on the facial expression contour features of each image frame, the peak frames in each video subsequence are located, including: The differences in facial expression contour features between any two frames in this video subsequence are compared using cosine similarity. Compare the distances in time between the two frames with the largest differences from the center of the video subsequence, and take the frame with the smallest distance as the peak frame.
2. The method according to claim 1, characterized in that, The facial expression contour features are local binary encoded histograms.
3. The method according to claim 2, characterized in that, The peak frame Determined by the following formula: in, and Let F be any two image frames in a video subsequence F. ( ) represents the LBPH feature of the image frame. Represents cosine similarity. Indicates the time at the center of the sequence. This represents the total number of image frames in a video subsequence. This represents the time between two image frames when the cosine similarity is minimum.
4. The method according to claim 1, characterized in that, Global facial features are extracted from peak frames using a well-trained global facial feature extraction model. Local facial features and facial action unit activation patterns in peak frames are extracted using a trained facial action unit intensity detector. The unweighted graph neural network is defined as follows: when the cosine similarity between the features of two peak frames is greater than or equal to a preset similarity threshold, there is an edge between the nodes corresponding to the two peak frames; otherwise, there is no edge.
5. The method according to claim 1, characterized in that, Based on the distribution of perplexity intensity in each peak frame within a preset time period, the learner's perplexity within that time period is classified, including: When the duration of confusion exceeding a preset intensity within a preset time period is less than the preset time, the intensity of confusion changes periodically, and the duration of high-intensity confusion is shorter than the duration of low-intensity confusion, it is classified as positive confusion. A situation is classified as negative confusion when at least one of the following conditions is met: the duration of confusion exceeding a preset intensity within a preset time period is greater than the preset time; the intensity of confusion increases linearly; or the duration of high-intensity confusion is longer than the duration of low-intensity confusion.
6. A method for evaluating teaching effectiveness, characterized in that, Includes the following steps: The occurrence time of positive confusion and the occurrence time of negative confusion in the video sequence are determined based on the confusion classification method according to any one of claims 1 to 5. The occurrence rate and continuity of positive confusion in the teaching scenario are determined based on the frequency and timing of the occurrence of positive and negative confusion. The teaching effectiveness in the teaching scenario is evaluated based on the product of the proportion of positive confusion and the continuity of confusion. The larger the value of the product, the better the teaching effectiveness.
7. The method according to claim 6, characterized in that, When there is only one learner in the teaching scenario, the quantitative value of the teaching effect is... for: in, The larger the value, the higher the teaching quality; The proportion of positive confusion indicates the presence of such a person. Indicates the continuity of confusion; in, The function is used to calculate the number of times a corresponding state occurs. Indicates the number of times positive confusion occurs. This represents the total number of states where positive and negative confusion occur. This represents the state of confusion within the i-th time window, with a value of -1 or 1, representing negative or positive confusion respectively.
8. The method according to claim 6, characterized in that, When there are multiple learners in the teaching scenario, the teaching effectiveness is determined through the following steps: The average perplexity distribution of multiple learners is determined based on the perplexity of each learner in each peak frame; Based on the average confusion level of multiple learners in each time window, the confusion level classification method described in any one of claims 1-6 is used to determine the occurrence time of average positive confusion and average negative confusion in the video sequence. The occurrence rate and duration of average positive confusion and average negative confusion in the teaching scenario are determined based on the frequency and timing of their occurrence. The teaching effectiveness in the teaching scenario is evaluated based on the product of the average positive confusion occurrence rate and the average confusion continuity.
9. An electronic device, characterized in that, include: At least one memory for storing computer programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-5 or 6-8.
Citation Information
Patent Citations
Face emotion recognition method based on graph convolutional neural network
CN111339847A
Emotion classification method and system based on graph convolutional neural network
CN113569997A