An online learning engagement recognition method based on multi-vision cue fusion

This online learning input recognition method, which integrates multiple visual cues, solves the challenges of multi-dimensional fine-grained representation and implicit dynamic feature extraction of online learning input. It achieves non-contact, accurate perception and multi-granularity recognition, making it suitable for large-scale personalized online learning.

CN115424336BActive Publication Date: 2026-03-31HUAZHONG NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-05
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Online learning input perception methods face difficulties in multidimensional fine-grained representation, implicit dynamic feature extraction, and multi-granularity recognition. Traditional methods are time-consuming, labor-intensive, and unsuitable for large-scale personalized applications.

Method used

This paper proposes an online learning-based input recognition method based on multi-visual cue fusion. By constructing a multi-visual cue perception database, extracting multi-visual cue data, fusing features using deep learning methods, and performing fine-grained recognition using deep convolutional networks and the Grad-CAM method, combined with a graph network model for multi-granular input recognition.

Benefits of technology

It enables contactless and non-intrusive automatic sensing of online learning input, meets the needs of real-time and accurate sensing, and is suitable for large-scale personalized online learning applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424336B_ABST
    Figure CN115424336B_ABST
Patent Text Reader

Abstract

The application discloses an online learning input recognition method based on multi-vision clue fusion. Firstly, the application faces the demand of large-scale online learning input perception, starts from the multi-vision clue angle, excavates the associated vision clues of the online learning input, and constructs a multi-dimensional fine-grained representation model of the online input. Secondly, the feature learning problem of the time sequence is converted into a graph-based feature learning problem, a graph network model based on mutual information regularization is proposed, meanwhile, training support is provided for the machine learning method used by the application, and a learning input perception database based on multi-vision clues is constructed. Finally, a fine-grained learning input recognition method fusing multi-vision clues is constructed, and on this basis, a coarse-grained learning input recognition method based on the input graph is designed to integrate the fine-grained variable-length learning input sequence, so that multi-granularity online learning input recognition is finally realized, and the multi-level and multi-stage learning input perception demand in actual application is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image recognition and image classification technology, specifically involving an online learning engagement recognition method based on multi-visual cue fusion. The aim is to infer online learning engagement by fusing the psychological and physiological information implied in multiple visual cues, providing technical support for educational applications such as personalized learning and adaptive intervention, and helping education develop towards precision, personalization, and intelligence. Background Technology

[0002] The deep integration of emerging information technologies such as artificial intelligence and big data with education and teaching has fueled the rapid and vigorous development of online education, driving its large-scale and normalized application. Online education has become an indispensable and important component of the education ecosystem.

[0003] While online education provides learners with cross-temporal and spatial support and resource sharing guarantees, it faces an increasingly prominent "quality crisis," mainly manifested in "high dropout rates" and "low completion rates." Research indicates that one of the main reasons for this "quality crisis" is the lack of accurate perception of learning engagement in online learning systems. The implicit, dynamic, and complex nature of learning engagement poses a significant challenge to its perception. Therefore, the perception of online learning engagement has become a focus of attention in the field of online education. Traditional perception methods based on self-reporting and manual observation are time-consuming and labor-intensive, and can no longer meet the needs of large-scale online learning. Therefore, the field of online learning urgently needs learning engagement perception methods suitable for large-scale applications.

[0004] Currently, there are two main approaches to perceiving engagement in online learning: manual perception and automatic perception. Manual perception methods are time-consuming and labor-intensive, making them unsuitable for large-scale personalized online learning applications. Therefore, researchers have gradually shifted their focus to automatic perception methods, with log data, wearable devices, and computer vision methods being three representative categories. However, log data primarily records learning behavior, emphasizing behavioral engagement and having limitations in representing emotional and cognitive engagement. Wearable device-based methods are often constrained by the application environment, lacking ease of use and cost-effectiveness, and can cause inconvenience and anxiety for learners, making them unsuitable for large-scale online learning applications. Automatic perception methods based on computer vision have become an important research direction in the field of online learning engagement perception due to their advantages such as non-contact, non-invasiveness, low cost, and good real-time performance. Computer vision-based methods typically utilize various visual cues such as facial expressions, eye movements, body posture, and physiological signals to perceive online learning engagement. In recent years, feature learning methods based on deep learning have gradually gained attention. Although current deep learning methods can improve the recognition of learning engagement, these methods lack sufficient attention to multi-dimensional learning engagement perception. Meanwhile, online learning input recognition based on multiple visual cues also faces challenges such as difficulty in extracting implicit input features and difficulties in multi-granularity recognition.

[0005] In conclusion, automatic perception of learning engagement is an important direction for the development of online education. Although some research has utilized visual cues such as facial expressions and body language to perceive learning engagement, challenges remain in areas such as multidimensional fine-grained representation of learning engagement, extraction of implicit dynamic engagement features, and multi-granularity engagement recognition.

[0006] Therefore, based on the research content, this invention designs a recognition method based on multi-visual cue fusion to realize the perception of online learning input, providing technical support for accurate identification and perception of online learning input. Summary of the Invention

[0007] This invention addresses the current challenges of multi-dimensional fine-grained representation of online learning engagement, extraction of implicit dynamic engagement features, and multi-granularity engagement recognition. Starting with visual cues, it designs an intelligent perception method for online learning engagement based on multi-visual-cue fusion to assess learners' engagement status. This invention provides an online learning engagement recognition method based on multi-visual-cue fusion, supporting non-contact, non-intrusive automatic perception of online learning engagement.

[0008] This invention provides an online learning input recognition method based on multi-visual cue fusion, comprising the following steps:

[0009] Step 1: From the perspective of multiple visual cues, construct a learning input perception database based on multiple visual cues;

[0010] Step 2: Extract multi-visual cue data, perform visual cue analysis on learning engagement perception, construct an online learning engagement representation summary model from different dimensions based on multi-visual cues, and extract multi-dimensional engagement features based on multi-visual cues.

[0011] Step 3: The input features obtained in Step 2 are fused using deep learning methods. The fused features are then input into a deep convolutional network for cognitive input recognition. Furthermore, the Grad-CAM method is used to perceive the learner's fine-grained online learning input level in different dimensions.

[0012] Furthermore, multiple visual cues include facial expressions, body posture, head posture, rPPG signals, and eye movement signals.

[0013] Furthermore, in step 2, based on multi-visual cue data, deep correlation analysis is used to mine the correlation of visual cues, thereby determining the visual cues to be used in a certain dimension.

[0014] The objective function of deep correlation analysis is to maximize the correlation of the network output, i.e.

[0015]

[0016] Where X i and X j These represent the features of two different visual cues. This indicates the network parameters that need to be optimized. f represents the optimal network parameters. i ,f j This indicates the final output;

[0017] Furthermore, after network mapping, the non-linear correlation between the two visual cues is determined. This determination can be evaluated by calculating the correlation of the output, as shown in the following formula:

[0018]

[0019] Where ρ i,j This indicates the correlation between two visual cues. They represent f respectively i ,f j The standard deviation of the evaluation results will be used as the basis for determining which visual cues to represent each dimension.

[0020] Furthermore, a summary model of online learning engagement representation is constructed from three dimensions: behavior, emotion, and cognition.

[0021] Furthermore, the specific steps for constructing the online learning input representation summary model are as follows;

[0022] A. Graph Construction

[0023] Given a dynamic input sequence, we first construct an undirected graph G = (V, E) to obtain the learning input features, where V is the number of nodes, E is the set of edges between all nodes, and the adjacency matrix A ∈ R. M×M First, it represents the adjacency relationships of nodes in an undirected graph G; second, in order to obtain dynamic information through the graph structure, segments or frames in the sequence are transformed into nodes in the graph, using... This means that each node v i Both can be represented by a single eigenvector n i ∈R f Relatedly, f represents the feature dimension, and the weights between nodes in the adjacency matrix A are obtained through learning.

[0024] B. Learnable Graph Networks

[0025] The graph network model used specifically includes the following four parts;

[0026] (1) Nonlinear graph convolution

[0027] First, we define the graph convolution operation as follows:

[0028] G * (H k )=σ(MLP k (ReLu(A)H k (3)

[0029] Where k represents the number of layers, k = 0, ..., K), H k For the input of the (k+1)th layer, MLP k Let G represent a multilayer perceptron at layer k, where σ is a non-linear activation function, A is the adjacency matrix to be learned, and G... * This represents a graph convolutional layer, where ReLU is a non-linear activation function.

[0030] (2) Diagram Inception

[0031] Given input H k Then the Inception diagram can be represented as:

[0032]

[0033] in and These are two different graph convolution operations, maxpool(H) k () is the pooling operation, and the output of this layer is composed of two graph convolutional layers and one pooling layer.

[0034] (3) Learnable pooling

[0035] A learnable pooling method is used for H. K The layer-based pooling method uses the following specific calculation formula:

[0036] h G =[maxpool(H K )|H K P|meanpool(H K (5)

[0037] Where P is the pooling vector, h is the learnable pooling parameter, and ρ is the pooling vector. G It is the output of the pooling layer;

[0038] (4) Objective function

[0039] Graph classification loss L GC The cross-entropy loss, commonly used in classification problems, can be defined using the following formula:

[0040]

[0041] in It is the result of the input measurement on the nth input sequence, i.e., the output of the model;

[0042] Graph structure learning loss L GL The design is to facilitate learning the pooling vector P and the adjacency matrix A, which is defined as follows:

[0043]

[0044] Where ⊙ represents element-wise product, e is a vector with all elements equal to 1, λ1, λ2, and λ3 control the weights of each part, and the structure matrix A... d Defined as:

[0045] (A d ) ij =(ij) 2 (8)

[0046] Where i and j represent node numbers, A d This can force nodes that are temporally adjacent to each other to have a stronger correlation;

[0047] To obtain a compact representation, the graph representation loss is defined as:

[0048] L GR =λ4I α (N;H K )=λ4(E α (N)+E α (H K )-E α (N,H K (9)

[0049] Where N represents the input of the model, λ4 is the weight coefficient, and I α It is mutual information, E α Let the α-order matrix represent the Renyi entropy, therefore the overall optimization objective function is:

[0050]

[0051] Here, θ represents other parameters to be learned.

[0052] Furthermore, in step 1, a learning input perception database based on multiple visual cues is constructed, and the specific implementation method is as follows;

[0053] (1) A number of undergraduate and graduate students who use cloud platforms to learn were recruited to conduct self-directed learning activities online, covering different learning activities and learning stages;

[0054] (2) Use the front-facing camera to record students’ online learning videos. During the recording process, the experience sampling method is used to obtain students’ instantaneous learning engagement. The data obtained through wearable devices is used as the source of annotation for learners’ physiological signals. Eye-tracking devices are used to capture learners’ gaze data. After each learning activity, a questionnaire is used to obtain their engagement in the learning activity. The engagement status of the learners is corrected manually.

[0055] (3) During the video recording process, the learning platform records learning behavior and learning outcome data to provide a reference for data annotation;

[0056] (4) Cut the video into video segments as needed and perform multi-dimensional and multi-granular data annotation to train the online learning input representation summary model.

[0057] Furthermore, in step 3, cognitive input identification across the three dimensions of behavior, emotion, and cognition is achieved through the following methods;

[0058] (1) Assume that the three input feature vectors obtained at time i are respectively and in For emotional investment, For behavioral input, Let n1, n2, and n3 represent the dimensions of the feature vectors for the three types of inputs, which are obtained through a deep convolutional network.

[0059] (2) Given a learning activity, assume that there are a total of [number] activities during the entire learning activity. In the next real-time input recognition, all feature vectors can form three learning input maps, namely the emotional input map. Behavioral input diagram Cognitive Input Map

[0060] (3) Learners’ input in learning activities can be obtained from three input maps. Each map can correspond to a type of input, and the three maps together can be used to perceive the overall input.

[0061] Furthermore, the extraction of rPPG signals includes three steps as follows;

[0062] a) Region of Interest (ROI) Selection: First, face images are obtained from video frames using a face detection and tracking algorithm, and an Active Appearance Model (AAM) is fitted. Second, the ROI is selected based on the key points determined by the AAM model. Here, the ROI is the upper half of the face excluding the eyes. Finally, the mean of each color channel of the ROI is taken, so the video sequence can form three one-dimensional signals, corresponding to the R, G, and B channels of the ROI.

[0063] b) Signal extraction: First, select a window and extract head motion signals and color signals from the video. The head motion signals include tracking signals corresponding to the three angles of pitch, roll and yaw, and the color signals include tracking signals corresponding to the three channels of R, G and B. Second, extract the original rPPG signals from the three color channel signals of R, G and B using the POS method.

[0064] c) Signal filtering: First, both the head motion signal and the rPPG signal are transformed to the frequency domain using FFT; second, the average of the head motion spectrum is subtracted from the rPPG spectrum to obtain a new spectrum; then, the passband range of the filter is determined based on the maximum value of the new spectrum; finally, the filtered rPPG signal is obtained by bandpass filtering, and the final rPPG signal is obtained through post-processing.

[0065] Compared with existing research and technology, this invention has the following advantages:

[0066] 1. This invention combines educational theory and deep learning methods to establish a multi-dimensional fine-grained learning engagement representation model driven by multiple visual cues, analyzes the intrinsic mechanism of learning engagement, meets the real-time and accurate perception requirements of online learning engagement, and lays the foundation for multi-granularity online learning engagement perception.

[0067] 2. This invention transforms the feature learning problem of time series data into a graph-based feature learning problem, proposing a graph network model based on mutual information regularization to improve the model's ability to learn implicit dynamic input features. This network model has a small parameter size, making it convenient for practical applications.

[0068] 3. This invention constructs a fine-grained learning input recognition method that integrates multiple visual cues, and on this basis, designs a coarse-grained learning input recognition method based on input graphs to integrate fine-grained variable-length learning input sequences, ultimately realizing multi-granularity online learning input recognition, meeting the multi-level and multi-stage learning input perception needs in practical applications. Attached Figure Description

[0069] Figure 1 Input representation model graphs for multi-visual cue-driven online learning;

[0070] Figure 2 This is a diagram of the basic network structure used for head pose estimation;

[0071] Figure 3 This is a diagram of the basic network structure used for line-of-sight estimation;

[0072] Figure 4 Here is a flowchart of the rPPG signal extraction process;

[0073] Figure 5 A diagram illustrating the deep correlation analysis method for visual cues;

[0074] Figure 6 A diagram of a network model based on mutual information regularization;

[0075] Figure 7 A flowchart illustrating fine-grained learning input identification, using cognitive input as an example;

[0076] Figure 8 A flowchart for identifying learning inputs for learning activities. Detailed Implementation

[0077] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.

[0078] This invention provides an online learning input recognition method based on multi-visual cue fusion, comprising the following steps:

[0079] Step 1: From the perspective of multiple visual cues, construct a learning input perception database based on multiple visual cues, including facial expressions, body posture, head posture, rPPG signals, and eye movement signals.

[0080] Step 2: Extract multi-visual cue data, perform visual cue analysis of learning engagement perception, further combine manual features and learning features, construct an online learning engagement representation summary model from different dimensions based on multi-visual cues, extract multi-dimensional engagement features based on multi-visual cues, and perform learning engagement recognition.

[0081] Step 3: Use an interpretable deep learning model to perform feature fusion for fine-grained learning input identification, and propose a graph-based deep learning method to perceive the input level of coarse-grained learning activities.

[0082] To achieve the above objectives, according to a first aspect of the present invention, a learning input perception database based on multiple visual cues is constructed from the perspective of multiple visual cues, specifically including the following steps:

[0083] 1. Recruit approximately 100 undergraduate and graduate students to participate in online learning activities using the cloud platform, covering different learning activities and stages;

[0084] 2. Use the front-facing camera to record students' online learning videos. During the recording process, use the experience sampling method to obtain students' instantaneous learning engagement. Use the data obtained through wearable devices as the source of annotation for learners' physiological signals. Use eye-tracking devices to capture learners' gaze data. After each learning activity, use a questionnaire to obtain their engagement in the learning activity. Use manual methods to correct the engagement status of the learners.

[0085] 3. During video recording, learning behavior and learning outcome data are recorded through the learning platform to provide a reference for data annotation;

[0086] 4. Cut the video into video segments as needed and perform multi-dimensional, multi-granular data annotation to train key algorithms.

[0087] Furthermore, according to a second aspect of the present invention, multi-visual cue data extraction is performed for input perception in online learning. The visual cues involved in the present invention include facial expressions, body posture, head posture, rPPG signals, and eye-tracking signals, such as... Figure 1 As shown;

[0088] Furthermore, the method for extracting facial expression cue data is as follows;

[0089] 1) Detect faces in each frame of the image using the YOLOv4 face detection algorithm;

[0090] 2) Form a sequence of face images from each frame;

[0091] 3) Extract the low-level features of each frame to form a sequence of facial expression cues.

[0092] Furthermore, visual cues of body posture will be used to extract learners' skeletal features, providing support for capturing dynamic behavioral inputs.

[0093] Furthermore, regarding head pose cues, this invention employs a deep learning-based method to estimate the head's pitch and yaw angles, primarily comprising three parts: head detection, preprocessing, and CNN. Figure 2 This is a diagram of the basic network structure used for head pose estimation.

[0094] Figure 3 The network structure diagram is shown below. This method estimates the learner's gaze point on the screen using the coordinates of the left and right eyes and the corner of the eye, forming a gaze tracking sequence.

[0095] Furthermore, such as Figure 4 As shown, the extraction of rPPG signals mainly includes three steps as follows;

[0096] a) Region of Interest (ROI) Selection: First, face images are acquired from video frames using a face detection and tracking algorithm, and an Active Appearance Model (AAM) is fitted. Second, the ROI is selected based on the keypoints determined by the AAM model; here, the ROI can be the upper half of the face excluding the eyes. Finally, the mean value of each color channel of the ROI is taken, resulting in three one-dimensional signals in the video sequence, corresponding to the R, G, and B channels of the ROI.

[0097] b) Signal Extraction: First, select a window and extract head motion signals and color signals from the video. The head motion signals include tracking signals corresponding to the pitch, roll, and yaw angles, while the color signals include tracking signals corresponding to the R, G, and B channels. Second, extract the raw rPPG signals from the R, G, and B color channel signals using the POS method.

[0098] c) Signal Filtering: First, both the head motion signal and the rPPG signal are transformed to the frequency domain using FFT. Second, the average of the head motion spectrum is subtracted from the rPPG spectrum to obtain a new spectrum. Then, the passband range of the filter is determined based on the maximum value of the new spectrum. Finally, the filtered rPPG signal is obtained through bandpass filtering, and the final rPPG signal is obtained through post-processing.

[0099] Furthermore, such as Figure 5 As shown, based on visual cues such as facial expressions, rPPG, body posture, and eye movements, deep correlation analysis is used to explore the correlation between visual cues. The correlation between cues is analyzed in combination with different learning activities to determine which cues can be used in different learning scenarios, and further reveal the internal mechanism of learning engagement.

[0100] The objective function of deep correlation analysis is to maximize the correlation of the network output, i.e.

[0101]

[0102] Where X i and X j These represent the features of two different visual cues. This indicates the network parameters that need to be optimized. f represents the optimal network parameters. i ,f j This indicates the final output.

[0103] Furthermore, after network mapping, we can determine the non-linear correlation between the two visual cues. This determination can be evaluated by calculating the correlation of the output, as shown in the following formula:

[0104]

[0105] Where ρ i,j This indicates the correlation between two visual cues. They represent f respectively i ,f j The standard deviation of the evaluation results was used as the basis for determining which visual cues to use to represent each dimension. Further analysis was conducted in conjunction with different learning activities and learning scenarios to finally determine a multi-dimensional fine-grained representation model of online learning input based on multiple visual cues, laying the foundation for subsequent learning input perception.

[0106] Furthermore, based on multi-visual cue data extraction, this invention utilizes online learning based on multi-visual cue to input a multi-dimensional fine-grained representation model for feature extraction in each dimension.

[0107] like Figure 6 As shown, this invention provides a graph network model based on mutual information regularization, which improves the model's ability to learn implicit dynamic input features. This network model has a small parameter size, making it convenient for practical applications. Specific steps include:

[0108] 1. Graph Construction

[0109] Given a dynamic input sequence (such as an expression sequence), we first construct an undirected graph G = (V, E) to obtain the learning input features, where V is a set of nodes (assuming there are M nodes), E is the set of edges between all nodes, and the adjacency matrix A ∈ R. M×M This represents the adjacency relationships of nodes in an undirected graph G. Secondly, in order to obtain dynamic information through the graph structure, segments (or frames) in the sequence are transformed into nodes in the graph (using...). (represented), each node v i Both can be represented by a single eigenvector n i ∈R fRelatedly (f represents the feature dimension), the weights between nodes in the adjacency matrix A can be obtained through learning.

[0110] 2. Learnable Graph Networks

[0111] The graph network model used in this invention mainly includes the following four parts;

[0112] 1) Nonlinear graph convolution:

[0113] First, we define the graph convolution operation as follows:

[0114] G * (H k )=σ(MLP k (ReLu(A)H k )), Formula 3

[0115] Where k represents the number of layers (k = 0, ..., K), H k For the input of the (k+1)th layer, MLP k Let G represent a multilayer perceptron at layer k, where σ is a non-linear activation function, A is the adjacency matrix to be learned, and G... * This represents a graph convolutional layer, where ReLU is a non-linear activation function.

[0116] 2) Diagram Inception:

[0117] Given input H k Then the graph Inception network can be represented as:

[0118]

[0119] in and These are two different graph convolution operations, maxpool(H) k ) is the pooling operation, and the output of this layer is composed of two graph convolutional layers and a pooling layer.

[0120] 3) Learnable pooling:

[0121] A learnable pooling method is used for H. K The layer-based pooling method uses the following specific calculation formula:

[0122] h G =[maxpool(H K )|H K P|meanpool(H K )] Formula 5

[0123] Where P is the pooling vector, h is the learnable pooling parameter, and ρ is the pooling vector. G It is the output of the pooling layer.

[0124] 4) Objective function:

[0125] Graph classification loss L GC The cross-entropy loss, commonly used in classification problems, can be defined using the following formula:

[0126]

[0127] in It is the result of the input measurement of the nth input sequence, i.e., the output of the model.

[0128] Graph structure learning loss L GL The design is to facilitate learning the pooling vector P and the adjacency matrix A, which is defined as follows:

[0129]

[0130] Where ⊙ represents element-wise product, e is a vector with all elements equal to 1, λ1, λ2, and λ3 control the weights of each part, and the structure matrix A... d Defined as:

[0131] (A d ) ij =(ij) 2 , Formula 8 Where i and j represent node numbers, A d This can force nodes that are adjacent in time to have a stronger correlation.

[0132] To obtain a compact representation, the graph representation loss is defined as:

[0133] L GR =λ4I α (N;H K )=λ4(E α (N)+E α (H K )-E α (N,H K )) Formula 9

[0134] Where N represents the input of the model, λ4 is the weight coefficient, and I α It is mutual information, E α Let α be the matrix representing the Renyi entropy. Therefore, the overall optimization objective function is:

[0135]

[0136] Here, θ represents other parameters to be learned.

[0137] like Figure 7As shown, according to the third part of this invention, a deep learning model with interpretability is used to perform fine-grained learning and feature fusion for input recognition. Specific steps include:

[0138] 1. Combine the extracted facial expression (appearance features, FACS features), body posture, head posture, rPPG signal, and eye movement signal features;

[0139] 2. Input it into a 1D deep convolutional network for cognitive input recognition;

[0140] 3. Based on the Grad-CAM (Gradient-weighted Class Activation Mapping) method, activation maps for different levels of engagement are generated. Histograms are used to statistically interpret learners' engagement levels across different dimensions, thereby explaining the distribution of features affecting cognitive engagement and providing a reference for understanding the cognitive engagement recognition results.

[0141] like Figure 8 As shown, this invention designs a coarse-grained learning input recognition method based on input graphs, integrating fine-grained variable-length learning input sequences to ultimately achieve multi-granularity online learning input recognition. Specific steps include:

[0142] 1. Suppose that the three input feature vectors (taken from the fully connected layer in the deep convolutional network) perceived at time i are respectively and in For emotional investment, For behavioral input, Let n1, n2, and n3 represent the dimensions of the feature vectors for the three types of inputs, respectively.

[0143] 2. Given a learning activity, assume there are a total of [number] activities during the entire learning activity. In the next real-time input recognition, all feature vectors can form three learning input maps, namely the emotional input map. Behavioral input diagram Cognitive Input Map

[0144] 3. This invention obtains the learner's input in the learning activity from three input graphs. Each graph can correspond to a type of input, and the three graphs can be combined to perceive the overall input.

[0145] Similarly, we can perceive higher levels of input in a way that progresses from fine to coarse, which helps to identify online learning input at different levels and stages.

[0146] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.

Claims

1. An online learning engagement recognition method based on multi-vision cue fusion, characterized in that, Comprising the following steps: Step 1, from the perspective of multi-vision clues, a learning input perception database based on multi-vision clues is constructed; Step 2, multi-vision clue data is extracted, vision clue analysis of learning input perception is carried out, online learning input representation summary model is constructed based on multi-vision clues from different dimensions, and multi-dimensional input feature extraction based on multi-vision clues is carried out; The specific construction steps of the online learning input representation summary model are as follows: A. Graph construction Given a dynamic input sequence, first construct an undirected graph. To obtain the learning input characteristics, among which It is a node. It is the set of edges between all nodes, the adjacency matrix. Representing an undirected graph The adjacency relationships of nodes are defined; secondly, in order to obtain dynamic information through the graph structure, segments or frames in the sequence are transformed into nodes in the graph, using... This indicates that each node Both can be achieved using a single feature vector. Related to this, Represents the feature dimension, adjacency matrix The weights between nodes are obtained through learning. B. Learnable graph network The graph network model used specifically includes the following four parts; (1) Nonlinear graph convolution Firstly, the graph convolution operation is defined, which is defined as: (3) wherein denotes the number of layers, , is the input of the layer, denotes the multi-layer perceptron on the layer, is a non-linear activation function, is the adjacency matrix to be learned, denotes the graph convolution layer, ReLu is a non-linear activation function; (2) Graph inception Given input Then the graph Inception can be represented as: (4) wherein and are two different graph convolution operations, is a pooling operation, and the graph inception is composed of two graph convolution layers and one pooling layer. (3) Learnable pooling A learnable pooling method is adopted to The pooling method is designed for the layer, and the specific calculation formula is as follows: (5) wherein is a pooling vector, and h G is a pooling layer output; (4) Objective function graph classification loss The graph classification loss can be defined using the cross-entropy loss common in classification problems, which is computed as follows: (6) wherein is a true input measurement corresponding to the th input sequence, is an input measurement prediction output by the model for the th input sequence, i.e., the output of the model; graph structure learning loss The design is to facilitate learning of the pooling vectors and the adjacency matrix which is defined as: (7) where represents the element product, is a vector with all elements equal to 1, , and control the weight of each part, respectively, the structure matrix is defined as: (8) wherein, i, j denotes a node number, one can force nodes that are adjacent in time to have stronger associations; In order to obtain a compact representation, the graph representation loss is defined as: (9) wherein N represents an input of the model, is a weight coefficient, is mutual information, is a matrix of order N represents the Renyi entropy, and thus the overall optimization objective function is: (10) wherein, represents other parameters to be learned; Step 3, the input features obtained in step 2 are fused by using the deep learning method, then the fused features are input into the deep convolution network for cognitive input recognition, and the fine-grained online learning input level of the learner in different dimensions is further perceived by the Grad-CAM method.

2. The online learning engagement recognition method based on multi-vision cue fusion of claim 1, wherein: The multi-vision clues include facial expressions, body postures, head postures, rPPG signals and eye movement signals.

3. The method of claim 1, wherein the method is based on fusion of multiple visual cues for online learning engagement recognition. In step 2, based on multi-vision clue data, the correlation of vision clues is mined by using deep correlation analysis method, and then the vision clues used in a certain dimension are determined; The objective function of the deep correlation analysis method is to maximize the correlation of the network output, that is, (1) wherein and represent features of two different visual cues, represents a network parameter to be optimized, represents an optimal network parameter, represents a final output; Further, after network mapping, the nonlinear correlation of the two kinds of vision clues is judged, which can be evaluated by calculating the correlation of the output, and the calculation formula is as follows: (2) wherein represents the correlation between two visual cues, respectively represent the standard deviation of the evaluation results as a basis for determining which visual cues represent which dimensions.

4. The method of claim 1, wherein the method is based on fusion of multiple visual cues for online learning engagement recognition. The online learning input representation summary model is constructed from the three dimensions of behavior, emotion and cognition.

5. The method of claim 1, wherein: In step 1, the learning input perception database based on multi-vision clues is constructed, and the specific implementation manner is as follows: (1) A number of subjects of undergraduate and graduate students using cloud platform learning are recruited, so that they carry out online self-learning activities, covering different learning activities and learning stages; (2) A front camera is used to record the online learning video of the students, and the experience sampling method is used to obtain the instantaneous learning input of the students during the recording process, the data obtained by the wearable device is used as the labeling source of the physiological signals of the learners, and the eye movement device is used to capture the gaze data of the learners, and after each learning activity is completed, the input condition of the learning activity is obtained through a questionnaire, and the input state of the learners is corrected through manual method; (3) During the video recording process, the learning behavior and learning result data are recorded through the learning platform to provide reference for data labeling; (4) The video is cut into video clips according to needs, and multi-dimensional and multi-granularity data labeling is carried out, which is used to train the online learning input representation summary model.

6. The method of claim 1, wherein: In step 3, the cognitive input recognition in the three dimensions of behavior, emotion and cognition is realized by the following methods: (1) Assume that at time The three perceived input feature vectors are , and , where is the emotional input, is the behavioral input, represents the cognitive input, respectively represents the dimension of the feature vector of the three inputs, and the three input feature vectors are obtained through a deep convolutional network; (2) Given a learning activity, assume that there are three real-time inputs identified throughout the learning activity, then all feature vectors can form three learning engagement graphs, respectively, the emotional engagement graph , the behavioral engagement graph and the cognitive engagement graph ; (3) The input of the learners in the learning activity is obtained from the three input graphs, each graph can correspond to one input, and the overall input can be perceived by combining the three graphs.

7. The method of claim 2, wherein the method is based on fusion of multiple visual cues for online learning engagement recognition. The extraction of rPPG signal includes the following three links; a) Region of interest selection: Firstly, the face detection and tracking algorithm is used to obtain the face image from the video frame, and the active appearance model AAM is fitted; secondly, the region of interest is selected according to the key points determined by the AMM model, and the region of interest is the part of the upper half of the face excluding the eyes; finally, the mean value of each color channel of the region of interest is taken, then the video sequence can form three one-dimensional signals, respectively corresponding to the R, G and B three channels of the video region of interest; b) Signal extraction: Firstly, a window is selected, and the head motion signal and the color signal are obtained from the video, wherein the head motion signal includes the tracking signals corresponding to the three angles of pitch, roll and yaw, and the color signal includes the tracking signals corresponding to the three channels of R, G and B; secondly, the original rPPG signal is extracted from the R, G and B three color channel signals by the POS method; c) Signal filtering: Firstly, the head motion signal and the rPPG signal are transformed into the frequency domain by FFT; secondly, the average of the head motion spectrum is subtracted from the rPPG spectrum to obtain a new spectrum; then, the passband range of the filter is determined according to the maximum value of the new spectrum; finally, the filtered rPPG signal is obtained by band-pass filtering, and the final rPPG signal is obtained by post-processing.