Visual analysis method and system for correlation between teacher multi-modal sentiment and student behavior
By constructing a visual analysis method for the correlation between multimodal emotions and student behavior, and using visualization technology to analyze the dynamic changes in teachers' emotions and students' learning behaviors, this method solves the problem of difficulty in analyzing the correlation between emotions and behaviors in online classrooms, and enables in-depth exploration and understanding of the relationship between teachers' emotional expression and students' learning behaviors.
Patent Information
- Application Number
- CN202211334285.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2042-10-28
AI Technical Summary
Existing technologies struggle to effectively analyze the correlation between teachers' multimodal emotions and students' behavior, especially in massive open online classrooms. Traditional methods cannot comprehensively consider multiple factors and cannot intuitively analyze the dynamic evolution process, leading to biases in the understanding of students' learning behavior and the assessment of their emotional expressions.
A visual analysis method for the correlation between teachers' multimodal emotions and students' behavior is adopted. By acquiring teachers' lecture videos and students' learning records, multimodal features of faces, audio, and text, as well as features of students' faces, heads, and mouse click behaviors, are extracted to construct summary views, learning behavior pattern views, correlation views, and detailed views. Visualization techniques such as stacked barcode maps, river stacked maps, Sankey diagrams, and projection maps are used for analysis.
It enables in-depth exploration of the correlation between teachers' emotional expression and students' behavior from multiple perspectives, helps teachers understand students' learning behavior, provides personalized intervention and support, and intuitively and effectively analyzes large-scale dynamic data, thereby improving the effectiveness of online teaching.
Smart Images

Figure CN115641537B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to visualization technology, and in particular to a visual analysis method and system for the correlation between teachers' multimodal emotions and students' behavior. Background Technology
[0002] Emotions play an important role in human communication and teaching. Studies have shown that appropriate emotional expression can improve audience engagement and lead to successful knowledge transfer [1]. Therefore, exploring the correlation between teachers’ multimodal emotions and students’ behavior is of great value for teachers to understand students’ learning performance in massive open online classrooms (MOOCs) and improve their teaching skills. However, the openness of MOOCs makes it difficult for teachers to understand students’ learning behavior and choose appropriate emotional expression. The separation of time and space in the context of MOOCs also hinders the transmission of emotional information [2]. Compared with traditional face-to-face teacher-student communication, online classroom teachers cannot directly observe the reactions of participating students. In addition, the complex structure of multimodal emotion and student behavior data increases the difficulty of correlation analysis. Previous studies have divided emotions into two groups, namely positive emotions and negative emotions, and studied their impact on students’ dropout behavior in MOOCs videos [3,4]. Since humans express emotions through a variety of behaviors, such as changes in facial and vocal expressions, it is too simplistic to describe them as positive and negative, resulting in a large bias in the assessment of the impact of specific emotions on student behavior.
[0003] Some studies based on traditional mathematical statistical methods rely heavily on student ratings of teacher emotional outcomes and behaviors when investigating the correlation between teacher emotions and student behaviors [9,10,11,12]. This can be problematic due to comprehension biases inherent in conventional methods or the same raters
[13] . Consequently, the true link between teacher emotions and student behaviors may be overestimated, as both are reported by students through questionnaires. To avoid these common methodological biases, some researchers [14,15,16] have used different perspectives (i.e., observational, student, and teacher perspectives). However, this can introduce another problem, as the overlap between different perspectives appears to be only weak to moderate, at least for teacher emotions and behaviors [14,15]. Furthermore, including all three perspectives is often impractical, and it is difficult to know whether and how the choice of a particular perspective affects the degree of correlation with the results and conclusions. Traditional statistical methods can only solve specific problems. However, analyzing the correlation between teachers' emotions and students' learning behaviors in a dynamic and evolving scenario like an online classroom requires comprehensive consideration of multiple explicit or implicit features. Visual analytics is a method that can comprehensively consider multiple factors and intuitively analyze dynamic evolution processes. Therefore, an effective visual analytics technique is needed to systematically explore and interpret the relationship between teachers' emotions and students' learning behaviors in order to gain deeper insights.
[0004] According to our investigation, many methods [3,4,5,6] have been proposed to analyze how emotions affect student dropout / repetition rates in MOOCs courses, but few studies have investigated how dynamic changes in emotions affect student academic performance. Furthermore, some studies [7,8,17,18,19] have attempted to use visualization techniques to visually analyze the multimodal emotions of speakers in lecture videos or the learning behaviors of students in online courses. However, these techniques only study the emotional expression of the speaker or the learning behavior of the participants from a single perspective, without conducting a holistic analysis from both the speaker's and participant's perspectives. Moreover, these techniques do not consider the potential relationship between the speaker's emotional expression and the participant's behavior. Furthermore, techniques for analyzing the multimodal emotional expression of speakers in lecture videos cannot be directly applied to the analysis of teacher emotions in massive open online teaching videos. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide a visual analysis method and system for the correlation between teachers' multimodal emotions and students' behavior, addressing the shortcomings of existing technologies. It analyzes the changes in teachers' emotional states and students' learning behaviors from the perspectives of the presenter and participants, and explores the correlation between the changes in teachers' emotional states and students' learning behaviors.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a visual analysis method for the correlation between teachers' multimodal emotions and students' behavior, comprising the following steps:
[0007] S1. Obtain the teacher's lecture videos and the students' learning record videos and mouse click data in the online classroom;
[0008] S2. Preprocess the data obtained in step S1 to extract multimodal features of the teacher's face, audio, and text, as well as features of the student's face, head, and mouse click behavior.
[0009] S3. Using the multimodal features of teachers and the behavioral features of students extracted in step S2, construct a summary view; the summary view is used to present the emotional changes of teachers and the corresponding changes in students' learning behaviors at the video level.
[0010] S4. Construct a student learning behavior pattern view using the summary view; the behavior pattern view is used to present the overall behavior pattern of students during the learning process.
[0011] S5. Using the summary view and the learning behavior pattern view, construct a relevance view; the relevance view is used to present the teacher's multimodal emotions and the changes in students' learning behaviors at the sentence level.
[0012] S6. Based on the relevance view, construct a detailed view; the detailed view is used to further present the teacher's multimodal emotions and the changes in students' learning behaviors at the sentence level.
[0013] After step S6, the following also includes:
[0014] S7. Construct a playback view that provides original online course videos and keyframe images of students' learning process for interactive analysis of teachers' emotional state and students' learning behavior.
[0015] In step S3, the specific implementation process of constructing the summary view using the multimodal features of teachers and the behavioral features of students extracted in step S2 includes:
[0016] Stacked barcode images and embedded donut charts are used to summarize the teacher's emotional state and student's learning behavior information corresponding to each online course video. The barcode images and donut charts are correlated with the video duration. The horizontal axis of the barcode image represents the video length, and the vertical axis, from top to bottom, represents the emotional state of the facial, audio, and text channels, with colors encoding the emotional type. The donut chart corresponding to the barcode image represents student behavior information within 90 seconds. The outer ring represents six click events, with colors encoding the event type and arc length encoding the proportion of click events. The inner ring represents the student's emotional type, with color representing the emotional type, arc length representing the proportion of different emotional types, and arc height representing the accuracy rate of emotional type recognition. The six click events are "video play," "video pause," "video fast forward," "video rewind," "video speed change," and "video volume change."
[0017] Using the keywords {emotional diversity, learning behavior ratio}, the teacher and student feature information obtained in step S2 is filtered and sorted based on the number of emotional types of teachers in the three channels of facial, audio and text at each time point in the course video, or the proportion of common behaviors of students at each time point in the learning process. The teacher's multimodal emotional information and student's learning behavior information are then visualized and summarized according to the order of the sorting results.
[0018] In step S4, the specific implementation process of constructing a student learning behavior pattern view using the summary view includes:
[0019] A river stacking diagram is used to visualize the click flow of all students during the course video learning process; the horizontal axis of the river stacking diagram represents the video duration, and the vertical axis represents the number of six click events, with different colored rivers encoding the changes of different click events; the six click events are "video play", "video pause", "video fast forward", "video rewind", "video rate change" and "video volume change".
[0020] A ring chart is used to visualize the head posture and facial expressions of all students during video viewing. The middle part of the ring chart represents the overall head yaw angle of all students, and the scale in the middle represents the yaw value. The histogram on the outer layer of the ring chart encodes the students' facial expressions, with color representing emotion type and height representing the accuracy of identifying a certain emotion type. The top of the ring chart indicates the start and end positions of the video, and the time gradually increases in a clockwise direction.
[0021] In step S5, the relevance view includes a Sankey diagram and a projection diagram. Each node in the Sankey diagram represents an emotion type, encoded with different colors. The nodes on the left represent facial expressions, the nodes in the middle represent text emotions, and the nodes on the right represent the overall facial emotions of students corresponding to the teacher's emotions. The links between different nodes represent the set of sentences corresponding to the emotion type. The right side of the relevance view is a projection diagram of all students at the sentence level. The outer arc of the projection diagram represents the student's facial emotions, with color encoding type and length encoding algorithm confidence. The internal pointer design represents the overall head deflection of the student in the sentence, with numbers representing the angle deflection value. The discrete arcs outside the pointers represent click events generated by students in a certain time period.
[0022] The specific implementation process of constructing a relevance view includes:
[0023] Based on the summary view showing the teacher's multimodal emotional changes corresponding to the video and the corresponding student learning behavior information, Sankey diagrams are used to visualize the teacher's facial, audio, and textual emotions at the sentence level.
[0024] Based on student head posture, facial expressions, and clickstream information, the following indicators {sentence number, yaw angle, emotion, mouse click event} are projected to observe the temporal changes in student behavior during the learning process. These indicators represent the sentence number, yaw angle, head posture, facial emotion, and click event corresponding to different timestamps, respectively. These indicators form an n-dimensional feature vector {x1,…,…} n}, where n represents the number of selected metrics; the clickstream information refers to mouse click behavior;
[0025] Based on the overall behavioral patterns of students in the learning process presented in the behavioral pattern view, the clusters corresponding to specific behavioral patterns in the projection map are located, and the students' learning behaviors and the temporal changes of learning behaviors are analyzed.
[0026] By analyzing the changes in teachers' emotions at the sentence level in the Sankey diagram and the changes in students' learning behaviors at the sentence level in the projected view, we can analyze the correlation between teachers' multimodal emotions and students' learning behaviors.
[0027] The horizontal axis of the multiple graphs in the detailed view represents video time. These graphs include a barcode graph for visualizing teacher emotions in online course videos, a stacked histogram for visualizing student head posture at yaw and pitch angles, and a pie chart and a thread river graph for visualizing the encoding of detailed student click events. The vertical axis of the barcode graph represents the emotional state of the facial, audio, and text channels from top to bottom, with colors encoding the emotion type. The vertical axis of the stacked histogram represents the student's deflection values at two angles. The pie chart represents a single click event, and the thread river graph represents consecutive click events.
[0028] The present invention also provides a visual analysis system for the correlation between teachers' multimodal emotions and students' behavior, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the steps of the method described above.
[0029] The present invention also provides a computer-readable storage medium having a computer program / instructions stored thereon; when the computer program / instructions are executed by a processor, they implement the steps of the method described above.
[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0031] (1) This invention helps users to deeply explore the correlation between teachers' emotional expression and students' behavior. A persuasive presentation is crucial to the final outcome; facial expressions, tone of voice, and word choice can influence knowledge delivery, meaning that teachers' emotional expression is related to participants' learning behavior. Therefore, this invention takes a dual perspective—from the knowledge presenter and the participant—and conducts multi-granular visual analysis of teachers' multimodal emotions and students' learning behavior at three levels: video, sentence, and frame. This helps users deeply explore teachers' emotional expression and its potential relationship with students' learning behavior, thereby helping teachers choose appropriate emotional expressions to improve their teaching skills.
[0032] (2) This invention helps users better understand students' complex learning behaviors and provide timely personalized intervention and support. Compared to traditional offline teaching where teachers and students interact face-to-face, online teaching scenarios do not allow teachers to directly observe students' behavioral responses, thus posing a significant challenge to understanding students' learning behaviors. To better grasp students' behavioral performance during the learning process, this invention provides visual analysis of various student learning behaviors, including facial expressions, head postures, and mouse click behavior. It not only analyzes their dynamic evolution but also explores and analyzes their overall behavioral patterns, thereby helping users understand changes in students' learning behaviors more quickly and effectively.
[0033] (3) This invention can intuitively and effectively analyze large-scale dynamic data. The emotions of teachers and the learning behaviors of students in MOOCs videos are large-scale, high-dimensional time-series data. The multimodal and multigranular nature of teachers' emotional behaviors, as well as the randomness and dynamism of large-scale student learning behaviors, pose significant challenges to analyzing the correlation between emotional expression and student learning behaviors. Therefore, this invention employs visual analytics technology and proposes various novel visualization designs to process and present this large-scale time-series data, in order to better analyze the correlation between teachers' emotional expressions and student learning behaviors. Attached Figure Description
[0034] Figure 1 This is a flowchart of the visual analysis method according to an embodiment of the present invention;
[0035] Figure 2 This is a schematic diagram of the angles used to extract head posture in the method of this embodiment of the invention;
[0036] Figure 3 This is a summary view of the teacher's multimodal emotions and student learning behaviors at the video level in the method of this embodiment of the invention;
[0037] Figure 4 This is a view of the overall learning behavior pattern of students at the video level in the method of this embodiment of the invention;
[0038] Figure 5 This is a summary view of the correlation between teachers' emotional states and students' learning behaviors at the sentence level in the method of this embodiment of the invention;
[0039] Figure 6 This is a projection diagram of student learning behavior in the method of this embodiment of the invention;
[0040] Figure 7 This is a detailed view of the correlation between teachers' multimodal emotions and students' learning behaviors in a sentence, as described in the embodiments of the present invention.
[0041] Figure 8 This is a detailed view of student learning behavior in the method of this embodiment of the invention;
[0042] Figure 9 This is a playback view in the method of the embodiments of the present invention that can help users further verify changes in the teacher's emotional state and the student's learning behavior;
[0043] Figure 10 This is a flowchart illustrating the overall analysis of cases in the method of this invention embodiment;
[0044] Figure 11 This is a diagram illustrating the overall learning behavior pattern of students at the video level in the case analysis of the real-time example method of this invention.
[0045] Figure 12 This is a Sankey diagram of teacher multimodal sentiment at the sentence level in the case analysis of the real-time example method of this invention;
[0046] Figure 13 This is a projection diagram of students at the sentence level in the case analysis of the real-time example method of this invention;
[0047] Figure 14 This is a graph showing the changes in pitch information of teachers' multimodal emotions in case analysis during the real-time example method of this invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0049] Example 1
[0050] Figure 1 This is a schematic flowchart of the method in Embodiment 1 of the present invention. This embodiment provides a visual analysis method for analyzing the correlation between teachers' multimodal emotions and students' learning behaviors in large-scale online open classrooms, specifically including the following steps:
[0051] S1. Obtain the teacher's lecture videos and students' learning record videos and mouse click data in the online classroom, specifically including the following steps:
[0052] (1) Collect lecture videos from teachers on open online course platforms;
[0053] (2) Use the front-facing camera of the laptop to capture students' learning records at 30 frames per second (FPS) with a resolution of 640×480;
[0054] (3) A JavaScript plugin was developed to collect mouse click data from students in online classes. This plugin is based on the DOM structure of HTML files and can collect mouse clicks on video web pages. Six types of mouse click events were collected: {play,pause,seek forward,seek back,rate change,volume change}, representing "video playback", "video pause", "video fast forward", "video rewind", "video rate change", and "video volume change", respectively. Each click event contains four attributes: {[classID],[studentID],[event type],[current time / duration]}, representing "course name", "student ID", "click type", and "click timestamp / video length", respectively.
[0055] S2. Preprocess the collected teacher lecture videos, student classroom learning record videos, and mouse click data to extract multimodal features of the teacher's face, audio, and text, as well as features of the student's face, head, and mouse click behavior. This includes the following steps:
[0056] (1) Extracting the teacher's facial features from online course videos, mainly including facial landmark detection and facial analysis processing;
[0057] (2) Extracting the teacher's text features from online course videos, mainly including the extraction of voice data from the video, the conversion of voice data into text data, and text sentiment analysis;
[0058] (3) Extracting the teacher's audio features from online course videos, mainly including speech data segmentation and audio sentiment analysis;
[0059] (4) Preprocess the student learning record videos collected in S1 and extract the students’ facial behavioral features;
[0060] (5) Preprocess the student learning record videos collected in S1 and extract the students’ head posture and behavioral features;
[0061] (6) Preprocess the student mouse click data collected in S1 and extract the student click flow behavior features;
[0062] (7) The teacher's facial emotions, text emotions and audio emotions extracted from (1)(2)(3) are integrated.
[0063] In step (1), the facial features of the teacher are extracted from the online course video, which specifically includes the following steps:
[0064] A. Facial landmark detection:
[0065] First, the feature extraction framework Pliers is used to extract image frames of the teacher from the collected course videos. After processing each video, a series of images V = {I1, I2, ..., I...} are obtained. n}, where I i Let represent the i-th frame of the image, and n represent the total number of frames in the video. Then, Face++ is used to identify the teacher's face in the video frames and extract the coordinates of the facial landmarks {x}. i ,y i The coordinate unit is pixels, and the objects detected include facial contours, eyes, eyebrows, lips, and nose contours.
[0066] B. Face analysis and processing:
[0067] The extracted facial landmark information of teachers is used as input, and a CNN convolutional neural network model is used to identify teachers' facial emotions in real time. The identified facial emotions include {anger, disgust, fear, happiness, neutral, sadness, surprise}, which represent anger, disgust, fear, happiness, neutrality, sadness, and surprise, respectively.
[0068] In step (2), the textual features of the teacher are extracted from the online course videos, specifically including the following steps:
[0069] First, the DeepSeech speech recognition model was used to extract the audio data of the teacher's lecture from the course video. Then, the speech data was converted into corresponding text using the speech-to-text conversion API provided by Google Cloud. The converted text contains a series of words, and the timestamp information for each word is timestamps = {onsets, offsets, duration}, where onsets represents the appearance time of the word in the video, offsets represents the end time, and duration represents the duration of the word. The text also includes the corresponding punctuation marks. After obtaining this text, a CNN-LSTM model was used to analyze the sentiment of the text, identifying seven emotions: anger, disgust, fear, happiness, sadness, surprise, and neutral.
[0070] In step (3), the audio features of the teacher are extracted from the online course video, which specifically includes the following steps:
[0071] First, the audio data extracted in (2) is segmented into words and sentences. Then, openSMILE is used to extract audio features such as pitch variations and Mel Frequency Cepstral Coefficient for audio emotion recognition. These features are then fed into a neural network model. In the RAVDESS data test, the voice emotion recognition accuracy reached 96%. Finally, there are the same seven emotions as in (2).
[0072] In step (4), facial behavioral features of students are extracted from the collected student learning record videos, mainly focusing on extracting students' facial emotions. The processing steps are the same as in (1).
[0073] In step (5), the head posture and behavioral features of students are extracted from the collected student learning record videos, which specifically includes the following steps:
[0074] A. Head position detection:
[0075] Before head pose estimation analysis, head position detection is performed. The facial positions of students in the collected learning recording videos are extracted to depict head position, such as... Figure 2 The location shown is marked as a rectangular bounding box that appropriately surrounds the student's face. This rectangle has four attributes: {top, left, width, height}, representing the ordinate of the top-left corner pixel, the x-coordinate of the top-left corner pixel, the width of the rectangle, and the height of the rectangle, respectively.
[0076] B. Head pose estimation:
[0077] First, the student learning recording video is processed using the method in step (1) to extract the student's image frames. Then, a deep learning model is used to estimate the student's head pose in each image frame. Figure 2 As shown, the head pose extracted from each image frame is a three-dimensional vector. Where s represents student s, i represents the i-th frame of student s's head image, and the three dimensions are the student's pitch angle, yaw angle, and roll angle, representing the angles of head rotation around the X, Y, and Z axes, respectively. Among these three angles, pitch and yaw are important for analyzing student behavior because they represent the student's vertical and horizontal viewing positions on the screen, respectively. The roll angle is not very meaningful in this embodiment because head rotation along the Z-axis does not significantly affect the student's viewing position. Therefore, this angle is not considered in this embodiment. Since different students may have different device settings during learning, resulting in different ranges of head position and head posture, max-min normalization is used to normalize each student's head posture to (-1, 1).
[0078] In step (6), the student mouse click data collected in S1 is preprocessed to extract the student click flow behavior features, including the following steps:
[0079] First, the collected student mouse click data is filtered to remove missing or erroneous values. Then, the time attributes [current timestamp / video duration] of individual click events such as play, pause, speed change, and volume change, and continuous click events such as fast forward and replay are split. The final output is in the following two formats: {[Course Name],[Student ID],[Click Event Type],[Current Timestamp],[Video Duration]} and {[Course Name],[Student ID],[Click Event Type],[Event Start Timestamp],[Event End Timestamp]}, because continuous click events have a start and end time.
[0080] In this process, the teacher's facial emotions, text emotions, and audio emotions extracted in (1), (2), and (3) are fused together, specifically including the following steps:
[0081] A. Multimodal fusion:
[0082] For multimodal fusion, we use the emotion categories common to each channel emotion recognition model, and finally fuse them to produce a total of seven emotions: anger, disgust, fear, happiness, neutrality, sadness, and surprise.
[0083] B. Multi-level integration:
[0084] For multi-level fusion, we consider three temporal granularities: sentence level, word level, and frame level. In steps (2) and (3), text and audio sentiment have already been aligned at the sentence level. Facial sentiment is extracted frame by frame. To perform sentence fusion, the most common facial emotion in each sentence is used to represent the facial sentiment at the sentence level. For word-level fusion, the text information extracted in step (2) includes the timestamp information of the words. Based on the detected time, facial, text, and audio sentiment can be mapped to each word, thereby achieving word-level fusion.
[0085] S3. Based on the teacher's multimodal emotional characteristics and students' various behavioral characteristics obtained in step S2, a summary view is constructed for visualization. A barcode image and nested circular design are used to present the teacher's emotional changes and the corresponding changes in students' learning behaviors at the video level. This mainly includes the following steps:
[0086] (1) As Figure 3 As shown, stacked barcode images and embedded ring designs are used to summarize the teacher's emotional state and the student's learning behavior information corresponding to each online course video. The barcode images and ring images are correlated with the video time.
[0087] (2) The view information corresponding to the online course video is filtered and sorted by keywords such as {emotional diversity, learning behavior ratio} to better present the information to the user.
[0088] In the summary view above, the horizontal axis of the barcode chart represents the video duration, and the vertical axis, from top to bottom, represents the emotional state of the facial, audio, and text channels, with colors encoding the emotional type. Each corresponding donut represents student behavior information within 90 seconds. The outer donut represents six clickstream events, with colors encoding the event type and arc length encoding the proportion of each click event. The inner donut represents the student's emotional type; color indicates the emotional type, arc length indicates the proportion of different emotional types, and the arc height indicates the accuracy of the algorithm in identifying the emotional type. This view helps users quickly gain an overview of the overall information about teachers and students by summarizing teacher emotions and student learning behaviors at the video level.
[0089] S4. Based on the summary view obtained in step S3, such as Figure 4 As shown, a stacked river map and an enhanced circular design are used to construct a view of student learning behavior patterns, which is used to present the overall behavior patterns of students in the learning process. Users can quickly locate specific behavior patterns and perform further analysis, specifically including the following steps:
[0090] (1) Visualize the click flow of all students during the course video learning process using a river stacking diagram;
[0091] (2) A novel ring diagram is used to visualize the head posture and facial expressions of all students during the video viewing process.
[0092] The aforementioned river-stacked graph of student learning behavior views uses the horizontal axis to represent video duration and the vertical axis to represent the number of six types of click events. Different colored rivers encode variations in different click events. The central section of the novel ring plot represents the overall head yaw angle of all students, with the central scale indicating the yaw value. This angle represents the left and right head yaw of students, a relatively common behavior, and is therefore chosen for overall visualization. The histogram of the outer ring encodes students' facial expressions, with color representing emotion type and height representing the accuracy of the algorithm in identifying that emotion type. Furthermore, the top of the ring plot indicates the start and end positions of the video, with time gradually increasing in a clockwise direction.
[0093] S5. Based on the summary view obtained in step S3 and the learning behavior pattern view obtained in step S4, such as Figure 5As shown, an enhanced Sankey diagram and a novel projected view design are used to construct a relevance view to help users further explore changes in teachers' multimodal emotions and students' learning behaviors at the sentence level, thereby analyzing the correlation between teachers' emotions and students' learning behaviors. Specifically, the following steps are included:
[0094] (1) Based on the teacher and student overview information corresponding to the video presented in S2, Sankey diagrams are used to further visualize the teacher's facial, audio, and textual emotions at the sentence level.
[0095] (2) Figure 6 As shown, further analysis of students' learning behavior information in S3 is conducted. Based on students' head posture, facial expressions, and click flow information, the following indicators {sentence number, yaw angle, emotion, mouse click event} are projected to observe the temporal changes in students' behavior during the learning process. These indicators represent the sentence to which the student belongs at that time, the head posture at the yaw angle, facial emotion, and click event, respectively. These measures form an n-dimensional feature vector {x1,…,x...} n}, where n represents the number of indicators selected.
[0096] (3) Figure 6 As shown in the black dashed box, based on the student learning behavior patterns presented in S4, the clusters corresponding to the specific behavior patterns in the projection diagram can be quickly located, thereby enabling further exploration and analysis of students' learning behaviors and their temporal changes.
[0097] (4) Visually explore the correlation between teachers' multimodal emotions and students' learning behaviors by observing the changes in teachers' emotions at the sentence level in the Sankey diagram and the changes in students' learning behaviors at the sentence level in the projected view.
[0098] In the enhanced Sankey diagram within the aforementioned relevance view, each node represents an emotion type, encoded with different colors. Nodes on the left represent facial expressions, those in the middle represent text emotions, and those on the right represent the overall facial emotions of students corresponding to the teacher's emotions. Links between different nodes represent sets of sentences corresponding to the emotion type. The right side of the view shows a projection of all students at the sentence level, with each point represented using a novel design. In this design, the outer arc represents the student's facial emotion, with color-coded type and length-coded confidence level. Internally, pointers represent the student's overall head rotation within the sentence, with numbers representing the angle of rotation. Discrete arcs outside the pointers represent click events generated by students during that time period.
[0099] S6. Based on the correlation view obtained in step S5, such as Figure 7As shown, a detailed view is constructed to explore the changes in teachers' multimodal emotions and students' learning behaviors at the sentence level, in order to further analyze and summarize the correlation between teachers' emotions and students' learning behaviors. The specific steps include the following:
[0100] (1) Use the same barcode image as in step S2 to visualize the teacher's emotions in the online course video in a more granular way, and use a line graph below it to present the teacher's pitch information during the teaching process to better understand the changes in the teacher's emotions.
[0101] (2) Figure 8 As shown, a stacked histogram was designed to visualize students' detailed head postures at pitch and yaw angles. A stacked histogram of students' corresponding facial emotions was added near the yaw angle to better understand students' learning behavior.
[0102] (3) Figure 8 As shown, a circle and thread river design are used to encode and visualize the detailed click events of students.
[0103] In the detailed views above, the horizontal axis of each graph represents video time. In step (1), the vertical axis from top to bottom represents the emotional state of the three channels: face, audio, and text, with color coding for emotion type. In step (2), the vertical axis represents the student's deflection value at two angles, which can be positive or negative. The stacked histogram near the yaw line represents the student's facial emotion at that time, with color coding for emotion type as well. In step (3), the circle represents a single click event (e.g., pause, play, etc.). However, for continuous click events (e.g., fast forward, rewind), using multiple circles would cause visual confusion. Therefore, a thread river is used to bind them together to represent continuous click events, with color coding for the type of click event.
[0104] S7. Construct the playback view, such as Figure 9 As shown, original online course videos and keyframe images of students' learning processes are provided for interactive analysis of teachers' emotional states and students' learning behaviors to further explore the correlation between the two. Specifically, the following steps are included:
[0105] (1) Provide a function to play the original online course video, allowing users to combine the original video to interactively explore and analyze the teacher's multimodal emotions in different views.
[0106] (2) Provide keyframe images of students to help users better analyze students’ learning behavior and better understand the correlation between teachers’ emotions and students’ learning behavior.
[0107] The following examples illustrate this:
[0108] Research question: Does a teacher's extensive emotional expression during the teaching process correspond to a variety of learning behaviors in students?
[0109] Teachers pay close attention to students' learning behaviors during the teaching process. "Whether a teacher's extensive emotional expression corresponds to rich student learning behaviors" is a topic of great interest to teachers. However, manually browsing large-scale video sets and identifying representative segments is tedious and time-consuming. Therefore, the visual analysis method provided in this invention can be used to analyze the complex relationship between teachers' multimodal emotions and student learning behaviors. The overall analysis process is as follows: Figure 10 As shown, the specific steps include the following:
[0110] S1. Obtain the teacher's lecture videos and the students' learning record videos and mouse click data in the online classroom;
[0111] S2. Import the data and perform preprocessing operations;
[0112] S3. The summary view presents changes in the teacher's emotional state and information on students' learning behaviors at the video level, and filters and sorts the overview information corresponding to the video set according to user needs to help users quickly grasp the overall information of the teacher and students. For example... Figure 3 As shown, the middle section of the view contains three columns: the first column represents the video course title, the second the course type, and the third the teacher's facial, audio, and text emotion changes, as well as students' click stream and facial emotional behavior changes. The graph shows that the teacher's emotions in the first video in the list encompass multiple types and change frequently over time, indicating that the teacher in this course video expressed a wide range of emotions during the lesson. The nested donut chart also shows multiple colors and significant differences in arc length and height, suggesting substantial variations in students' emotional expression and click events during their learning of this video. This initially suggests a correlation between the teacher's rich emotional expression and the students' rich learning behaviors, prompting further exploration by clicking on this video.
[0113] S4. After quickly summarizing the changes in teacher emotional states and overall student learning behaviors corresponding to the video set in S3, in order to further explore the overall learning behavior patterns of students and locate specific behavior patterns, the overall learning behavior patterns and changes of all students at the video level are presented in the student learning behavior view. For example... Figure 11As shown in the click flow graph, numerous "pause" and "replay" click events can be observed around the 520-second mark of the video. Furthermore, the student head posture during this time period is significantly altered in the right-hand annular graph, indicating the emergence of specific learning behavior patterns during this segment. The click flow, head posture, and facial expression information in the graph also reveal multiple learning behavior patterns exhibited by students during the video learning process.
[0114] S5. To help users further analyze teachers' multimodal emotions and students' temporal changes in learning behavior, and to explore the correlation between the two, the correlation view presents information on teachers' emotions and students' learning behaviors at the sentence level. For example... Figure 12 As shown, a large proportion of the nodes on the left side of the Sankey diagram are gray-coded, indicating that the teacher's primary emotional expression in the lesson is neutral. Furthermore, the teacher's keyframe images and the emotional type of the text word cloud in the middle also show that the teacher indeed conveyed a rich range of emotions in the video. The links between the Sankey diagrams reveal complex relationships between the teacher's different modalities of emotion, demonstrating a variety of emotional combinations during knowledge transmission. Additionally, as... Figure 13 As shown by the links on the right side of the Sankey diagram, there are multiple clusters of different colors in the student learning behavior projection map, and these clusters are connected to teacher sentiment in the Sankey diagram. This indicates a correlation between student clusters with different learning behaviors and teacher multimodal sentiment, which also explains why students participating in this course exhibit rich facial expressions and diverse behavioral patterns during the learning process. This further illustrates the correlation between teacher multimodal sentiment and student learning behavior. Furthermore, to quickly locate and analyze students' specific learning behavior patterns and their temporal changes, S4 is used to quickly locate specific behavioral patterns in the student learning process, such as... Figure 11 The black dashed box in the middle indicates the time period around 520 seconds described in S4. Next, based on the behavioral pattern presented in S4, the cluster corresponding to that behavioral pattern in the projection map is located, such as... Figure 13 As shown in the dotted box below, the students' heads turned significantly during this period, and their facial expressions were one of disgust and confusion. They also made many pause and replay clicks during the learning process, which indicates that they encountered difficulties during this period.
[0115] S6 and S7. To further explore the teacher's emotional state and student learning behavior information corresponding to the different sentences presented in S5, the detailed view presents more specific changes in the teacher's emotional state, as well as detailed head posture, facial expressions, and clickstream information of the students. For example... Figure 7As shown in the figure, it is immediately apparent that the teacher's emotions and the students' head posture and facial expressions changed frequently over time. This further illustrates that the teacher engaged in a great deal of emotional expression during the lesson, and the students also exhibited a variety of learning behaviors, indicating a correlation between the two. In the sentence corresponding to 500-600 seconds, it can be observed that the teacher primarily used "angry" facial and audio emotions to explain the content of the sentence (see...). Figure 14 During this period, the students' head posture changed significantly, and their facial expressions were mainly "surprise" and "disgust," resulting in a large number of "replay" mouse click events. Figure 8 (As shown). To further verify this result, by reviewing the original video provided in step S7, it can be seen that the teacher used the emotion of "anger" to emphasize the importance of the content in the sentence with a stern expression and tone, thus leaving a deep impression on the students. Meanwhile, as... Figure 9 As shown, the student keyframe images provided in S6 also show that the teacher's emotional expression successfully attracted the students' attention, that is, the students turned their heads to focus on the screen position where the teacher was emphasizing the content.
[0116] Furthermore, the significant color changes in the barcode images before and after the current sentence indicate a substantial shift in the teacher's emotions across these three consecutive sentences, primarily from "anger-neutral-happiness." The students' head postures also change from "positive angle-0-negative angle," and their facial expressions undergo significant shifts. Correspondingly, the student clickstream shifts from "replay-fast forward." Additionally, a large distance can be observed between the clusters corresponding to these three sentences in the projection graph of S5 (see...). Figure 13 This further illustrates the significant changes in student behavior across these three consecutive sentences. To verify this finding from visual analysis, interactive analysis of the original video in S6 revealed that the teacher initially explained the knowledge in a gentle manner, then began to emphasize key points in a stern tone, and finally summarized the content in a relaxed manner. This further validates the feasibility of the method in this embodiment. To verify whether changes in the teacher's emotions affect changes in student behavior, keyframe images of students provided in S7 showed that students' head posture and facial expressions did indeed change with the teacher's attitude and tone, demonstrating a correlation between changes in the teacher's emotions and student learning behavior.
[0117] Example 2
[0118] Embodiment 2 of the present invention provides a visualization analysis system corresponding to Embodiment 1 above, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program in the memory to implement the steps of the method in Embodiment 1 above.
[0119] In some implementations, the memory may be high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk storage device.
[0120] In other implementations, the processor can be any type of general-purpose processor, such as a central processing unit (CPU) or a digital signal processor (DSP), and there is no limitation here.
[0121] Example 3
[0122] Embodiment 3 of the present invention provides a computer-readable storage medium corresponding to Embodiment 1 above, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, they implement the steps of the method of Embodiment 1 above.
[0123] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination thereof.
[0124] References
[0125] [1]Gallo,Carmine.Talk like TED:the 9 public speaking secrets of the world's top minds.Pan Macmillan,2014.
[0126] [2] Liu, Zhi, et al. "Exploring students' engagement patterns in SPOCforums and their association with course performance." EURASIA Journal of Mathematics, Science and Technology Education 14.7 (2018): 3143-3158.
[0127] [3]Dmoshinskaia,Nataliia.Dropout prediction in MOOCs:using sentimentanalysis of users'comments to predict engagement.MS thesis.University ofTwente,2016.
[0128] [4]Dillon,John,et al."Student Emotion,Co-Occurrence,and Dropout in aMOOC Context."International Educational Data Mining Society(2016).
[0129] [5]Wen,Miaomiao,Diyi Yang,and Carolyn Rose."Sentiment Analysis inMOOC Discussion Forums:What does it tell us?."Educational data mining2014.2014.
[0130] [6]Yang,Diyi,et al."Exploring the effect of confusion in discussionforums of massive open online courses."Proceedings of the second(2015)ACMconference on learning@scale.2015.
[0131] [7]Zeng,Haipeng,et al."EmoCo:Visual analysis of emotion coherence inpresentation videos."IEEE transactions on visualization and computer graphics26.1(2019):927-937.
[0132] [8]Aslan,Sinem,et al."Investigating the impact of a real-time,multimodal student engagement analytics technology in authentic classrooms."Proceedings of the 2019 CHI conference on human factors in computingsystems.2019.
[0133] [9]Becker,Eva Susann,et al."The importance of teachers'emotions andinstructional behavior for their students'emotions–An experience samplinganalysis."Teaching and Teacher Education 43(2014):15-26.
[0134]
[10] Irena."The role of social factors in shaping students’testemotions:A mediation analysis of cognitive appraisals."Social Psychology ofEducation 18.4(2015):785-809.
[0135]
[11] Irena."The role of social factors in shaping students’testemotions:A mediation analysis of cognitive appraisals."Social Psychology ofEducation 18.4(2015):785-809.
[0136]
[12] Mazer,Joseph P.,et al."The dark side of emotion in the classroom:Emotional processes as mediators of teacher communication behaviors and student negative emotions."Communication Education 63.3(2014):149-168.
[0137]
[13] Spector,Paul E.,et al."A new perspective on method variance: Ameasure-centric approach."Journal of Management 45.3(2019):855-880.
[0138]
[14] Dobbelaer,Marjoleine Jolie."The quality and qualities ofclassroom observation systems."(2019).
[0139]
[15] Fauth, Benjamin, et al." Primary school teaching from the perspective of students, teachers and observers. and prediction of learning success."Journal for Psychology 28.3(2014):127-137.
[0140]
[16] Scherzinger,Marion,and Alexander Wettstein."Classroomdisruptions,the teacher–student relationship and classroom management fromthe perspective of teachers,students and external observers:A multimethodapproach."Learning Environments Research 22.1(2019):101-116.
[0141]
[17] Chen,Yuanzhe,et al."DropoutSeer:Visualizing learning patterns inMassive Open Online Courses for dropout reasoning and prediction."2016 IEEEConference on Visual Analytics Science and Technology(VAST).IEEE,2016.
[0142]
[18] Shi,Conglei,et al."VisMOOC:Visualizing video clickstream datafrom massive open online courses."2015 IEEE Pacific visualization symposium(PacificVis).IEEE,2015.
[0143]
[19] Maher,Kevin,et al."E-ffective:A Visual Analytic System forExploring the Emotion and Effectiveness of Inspirational Speeches."IEEETransactions on Visualization and Computer Graphics 28.1(2021):508-517.
Claims
1. A visual analysis method for the correlation between teachers' multimodal emotions and students' behavior, characterized in that, Includes the following steps: S1. Obtain the teacher's lecture videos and the students' learning record videos and mouse click data in the online classroom; S2. Preprocess the data obtained in step S1 to extract multimodal features of the teacher's face, audio, and text, as well as features of the student's face, head, and mouse click behavior. S3. Using the multimodal features of teachers and the behavioral features of students extracted in step S2, construct a summary view; the summary view is used to present the emotional changes of teachers and the corresponding changes in students' learning behaviors at the video level. S4. Construct a student learning behavior pattern view using the summary view; The behavior pattern view is used to present the overall behavior pattern of students during the learning process; S5. Construct a relevance view using the summary view and the learning behavior pattern view; The relevance view is used to present changes in teachers' multimodal emotions and students' learning behaviors at the sentence level; S6. Based on the relevance view, construct a detailed view; the detailed view is used to further present the teacher's multimodal emotions and the changes in students' learning behaviors at the sentence level.
2. The visual analysis method for the correlation between teachers' multimodal emotions and students' behavior according to claim 1, characterized in that, After step S6, the following also includes: S7. Construct a playback view that provides original online course videos and keyframe images of students' learning process for interactive analysis of teachers' emotional state and students' learning behavior.
3. The visual analysis method for the correlation between teachers' multimodal emotions and students' behavior according to claim 1, characterized in that, In step S3, the specific implementation process of constructing the summary view using the multimodal features of teachers and the behavioral features of students extracted in step S2 includes: Stacked barcode images and embedded donut charts are used to summarize the teacher's emotional state and student's learning behavior information corresponding to each online course video. The barcode images and donut charts are correlated with the video duration. The horizontal axis of the barcode image represents the video length, and the vertical axis, from top to bottom, represents the emotional state of the facial, audio, and text channels, with colors encoding the emotional type. The donut chart corresponding to the barcode image represents student behavior information within 90 seconds. The outer ring represents six click events, with colors encoding the event type and arc length encoding the proportion of click events. The inner ring represents the student's emotional type, with color indicating the emotional type and arc length indicating the proportion of different emotional types. The proportion and the height of the arc represent the accuracy of identifying emotion types; the six click events are "video playback", "video pause", "video fast forward", "video rewind", "video speed change", and "video volume change"; using the keywords {emotional diversity, learning behavior ratio}, based on the number of emotion types in the teacher's facial, audio, and text channels corresponding to each time point in the course video, or the proportion of behaviors shared by students at each time point during the learning process, the teacher and student feature information obtained in step S2 is filtered and sorted, and the teacher's multimodal emotional information and student learning behavior information are visualized and summarized according to the order of the sorting results.
4. The visual analysis method for the correlation between teachers' multimodal emotions and students' behavior according to claim 1, characterized in that, In step S4, the specific implementation process of constructing a student learning behavior pattern view using the summary view includes: A river stacking graph is used to visualize the click flow of all students during the course video learning process. The horizontal axis of the river stacking graph represents the video duration, and the vertical axis represents the number of six click events. Different colored rivers encode changes in different click events. The six click events are "video play," "video pause," "video fast forward," "video rewind," "video rate change," and "video volume change." A ring graph is used to visualize the head posture and facial expressions of all students during video viewing. The middle part of the ring graph represents the overall head yaw angle of all students, with the yaw angle being the yaw angle. The scale in the middle represents the yaw value. The histogram on the outer layer of the ring graph encodes the students' facial expressions, with color representing emotion type and height representing the accuracy of identifying a certain emotion type. The top of the ring graph represents the start and end positions of the video, with time gradually increasing in a clockwise direction.
5. The visual analysis method for the correlation between teachers' multimodal emotions and students' behavior according to claim 1, characterized in that, In step S5, the relevance view includes a Sankey diagram and a projection diagram. Each node in the Sankey diagram represents an emotion type, encoded with different colors. The nodes on the left represent facial expressions, the nodes in the middle represent text emotions, and the nodes on the right represent the overall facial emotions of students corresponding to the teacher's emotions. The links between different nodes represent the set of sentences corresponding to the emotion type. The right side of the relevance view is a projection diagram of all students at the sentence level. The outer arc of the projection diagram represents the student's facial emotions, with color encoding type and length encoding algorithm confidence. The internal pointer design represents the overall head deflection of the student in the sentence, with numbers representing the angle deflection value. The discrete arcs outside the pointers represent click events generated by students in a certain time period.
6. The visual analysis method for the correlation between teachers' multimodal emotions and students' behavior according to claim 5, characterized in that, The specific implementation process of constructing a relevance view includes: Based on the summary view showing the teacher's multimodal emotional changes corresponding to the video and the corresponding student learning behavior information, Sankey diagrams are used to visualize the teacher's facial, audio, and textual emotions at the sentence level. Based on students' head posture, facial expressions, and clickstream information, the following indicators {sentence number, yaw angle, emotion, mouse click event} are projected to observe the temporal changes in students' behavior during the learning process. These indicators represent the sentence number, yaw angle, head posture, facial emotion, and click event corresponding to different timestamps, respectively. These indicators form an n-dimensional feature vector {x1,…,x}. n }, where n represents the number of selected metrics; the clickstream information refers to mouse click behavior; Based on the overall behavioral patterns of students in the learning process presented in the behavioral pattern view, the clusters corresponding to specific behavioral patterns in the projection map are located, and the students' learning behaviors and the temporal changes of learning behaviors are analyzed. By analyzing the changes in teachers' emotions at the sentence level in the Sankey diagram and the changes in students' learning behaviors at the sentence level in the projected view, we can analyze the correlation between teachers' multimodal emotions and students' learning behaviors.
7. The visual analysis method for the correlation between teachers' multimodal emotions and students' behavior according to claim 1, characterized in that, The horizontal axis of the multiple graphs in the detailed view represents video time. These graphs include a barcode graph for visualizing teacher emotions in online course videos, a stacked histogram for visualizing student head postures at yaw and pitch angles, and a pie chart and a thread river graph for visualizing the encoding of detailed student click events. The vertical axis of the barcode graph represents the emotional state of the facial, audio, and text channels from top to bottom, with colors encoding the emotion type. The vertical axis of the stacked histogram represents the student's deflection values at two angles; the pie chart represents a single click event, and the thread river graph represents consecutive click events.
8. A visual analysis system for the correlation between teachers' multimodal emotions and students' behavior, comprising a memory, a processor, and a computer program stored in the memory; characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program / instructions stored thereon; characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 7.