An online learning status recognition method and related device based on multitasks

By constructing a multi-task recognition model of gaze point and facial expression, combined with adaptive alternating joint scheduling and personalized calibration, the problem of insufficient model accuracy and generalization ability in online learning state recognition is solved, and better learning effect and training efficiency are achieved.

CN116758623BActive Publication Date: 2025-07-11XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310698144.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2025-07-11
Estimated Expiration
2043-06-13

AI Technical Summary

Technical Problem

In online learning state recognition, the existing technology has problems such as low model accuracy and generalization ability, difficulty in model training and fine-tuning, and a single scheduling method leads to excessive bias in the result or wasted data samples.

Method used

Using a multi-task-based online learning state recognition method, the gaze point estimation sub-task model and the facial expression recognition sub-task model are constructed, combining adaptive alternating joint scheduling and a two-stage gaze point estimation model that integrates Landmarks and personalized calibration, an alternating joint training algorithm for joint loss function and task adaptation is designed to balance the deviations caused by the data set and task scheduling method.

Benefits of technology

It improves the accuracy and generalization ability of learning state recognition, reduces the complexity of the model, solves the training difficulties in online education scenarios, and achieves better learning results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116758623B_ABST
    Figure CN116758623B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-task based online learning state recognition method and related devices, belonging to the field of learning state prediction. The present invention provides a multi-task learning framework that combines a neural network with a fixation point estimation model that fuses Landmarks and personalized calibration to predict two learning state indicators, namely expression and fixation point; adopts a two-stage fixation point estimation algorithm that fuses Landmarks and personalized calibration, which solves the problems of insufficient accuracy and generalization performance of general models, difficulty in systematic implementation and large-scale application, etc.; adopts task-adaptive alternating joint scheduling for multi-task training, effectively improving the learning performance of each task and achieving better results for the main task. The present invention solves the deficiencies in accuracy and generalization ability of general fixation point estimation models, and at the same time overcomes the difficulties in systematic implementation and large-scale application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of learning state prediction, and specifically relates to an online learning state recognition method and related device based on multi-tasks. Background Art

[0002] Online education provides learners with a rich and flexible learning method, and learners can access various excellent courses anytime and anywhere. However, due to the separation of teachers and students, it is difficult to guarantee the teaching effect of online education. The learning state is composed of basic indicators describing the psychological and physiological states of learners. Related research has qualitatively shown that the learning state can directly affect the learning effect. In order to improve the teaching effect, the state detection of online classroom learners has become one of the key issues in related research.

[0003] The learning state involves various states such as the cognition and emotion of learners, and can be further subdivided into specific indicators such as eye movement state, behavior actions, postures, facial expressions, environmental interactions, and biometric features. Among them, eye features are a very important type of feature. The fixation behavior of the eyes indicates the real-time attention position of the learner and is an important clue to explain the learning intention of the learner. By sorting and analyzing the fixation behavior, the deep internal cognitive state of the learner can be mined. There are many correlation relationships among the above indicators, such as the eye movement state and facial expression, and the posture and environmental interaction. Multi-Task Learning (MTL) can explore the commonalities between sub-tasks, greatly improve the utilization of the correlation relationships and common information among tasks, and thus improve the learning effect and performance of all sub-tasks. In multi-task learning based on hard parameter sharing, there are two sub-task training methods: Alternate training and Joint training. Alternate training has lower requirements for the data set and can use a larger range of data sets. However, due to the different sub-task scheduling orders and weights, it is easy to cause the model inference result to be biased towards one or more tasks and cannot obtain a stable learning effect. Joint training has very high requirements for the data set. In actual research, a large number of data set labels cannot meet the training needs of all sub-tasks, so it will cause a great waste of the data set, thereby greatly reducing the shared information contained in the shared feature layer.

[0004] Generally speaking, the fixation point detection task currently has problems such as low model accuracy and generalization ability, and difficulties in model training and fine-tuning in the online education scenario. In addition, multi-task learning using a single scheduling method inevitably has problems such as too strong result bias or a large waste of data samples. How to perform gaze estimation, how to make full use of the data set, and balance the deviation caused by different data sets or task scheduling methods on the basis of ensuring shared feature extraction are important issues that need to be further explored in online learning state recognition. Summary of the Invention

[0005] The object of the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a multi-task-based online learning state recognition method and related device.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A multi-task-based online learning state recognition method includes the following steps:

[0008] (1) Construct a multi-task recognition model for the gaze points and facial expressions of online learners, which consists of a gaze point estimation sub-task model and a facial expression recognition sub-task model;

[0009] The gaze point estimation sub-task model is used to merge the left eye feature, right eye feature and Landmarks Grid feature and output the gaze landing point; the facial expression recognition sub-task model is used to merge the Landmarks Grid and the face branch feature and output the facial expression;

[0010] (2) Use adaptive alternating joint scheduling for multi-task training

[0011] During training, first alternately train the gaze point estimation sub-task model or the facial expression recognition sub-task model according to the feature labels of the samples in the training dataset; after completing the alternating training, increase the task joint level to level 2, and use the sample data containing both gaze point and facial expression labels to perform combined training on the multi-task recognition model.

[0012] Further, in step (1), the gaze point estimation sub-task model is a two-stage gaze point estimation model that combines a general gaze point estimation model that fuses Landmarks and a gaze point personalized calibration model;

[0013] When running the gaze point estimation sub-task model, first use the general gaze point estimation model that fuses Landmarks to estimate the gaze point of the learner in the online learning scenario, so as to obtain the general gaze point coordinates; then, use the gaze point personalized calibration model to estimate the deviation of the general gaze point coordinates, and combine the general gaze point coordinates and the deviation of the general gaze point coordinates to finally obtain the real gaze point coordinates.

[0014] Further, in step (1), running the gaze point estimation sub-task model specifically includes:

[0015] Input the learner's image into the general gaze point estimation model that fuses Landmarks, and use VGG to extract the left-eye and right-eye image features; use the FaceAlignment algorithm to extract the facial key points of the learner, and at the same time intercept the eye images to construct a Landmarks Grid, and use Micro CNN to extract the Landmarks Grid features; then splice the left-eye features, right-eye features and Landmarks Grid features in the fully connected layer, and finally output the general gaze landing point of the learner;

[0016] At the same time, use the learner's facial image as the input of the gaze point personalized calibration model, extract features through a convolutional neural network, and finally output the difference between the actual gaze landing point and the estimated landing point of the general model;

[0017] Estimate the learner's final actual gaze landing point by combining the learner's general gaze landing point and the difference between the actual gaze landing point and the estimated landing point of the general model.

[0018] Further, in step (2), adaptive alternating joint scheduling is used for multi-task training. The specific process is as follows:

[0019] (201) Perform sub-task encoding

[0020] Multi-task learning contains n sub-tasks, and perform Onehot encoding on the sub-tasks;

[0021] Construct a loss function Li for each sub-task, then the multi-task loss function vector Considering the sample S in the dataset D, if the sample S has the sub-task T a , T b , …, T k required labels, then the sample S for the sub-task T a , T b , …, T k joint loss function L ab...k is:

[0022]

[0023] where, S ab...k is the Onehot encoding value that enables the task T a , T b , …, T k , that is, the binary encoding with the a, b, …, k bits being 1;

[0024] According to the labels of the tasks contained in the sample S, divide the original dataset D into alternating training datasets {D1, D2, …, D n} and joint training datasets {D12 , D 123 , …D ij...k , …, D 12...n};

[0025] Alternating training dataset D i only contains the label data required for task T i , and the joint training dataset D ij...k contains the label data required for tasks T i , T j , …, T k;

[0026] (202) Design the loss function

[0027] During the alternating training schedule, for subtask T i the weight coefficient is λ i . Let the weight coefficient λ i be inversely proportional to the ratio of the size of its own dataset D i to the size of the overall dataset D total :

[0028]

[0029] Then normalize λ i , and the weight coefficient λ i of the subtask T i is:

[0030]

[0031] where λ i represents the weight coefficient of subtask T i , N represents the total number of subtasks, and mount(D i ) represents the size of the dataset of subtask T i ;

[0032] During the joint training schedule, divide the training dataset into the main task set T m and the auxiliary task set T n . When conducting joint training, the weight coefficients of each task are defined as follows:

[0033]

[0034] where λ j represents the weight coefficient of the main task T j , L j represents the loss function of the main task T j , and T j belongs to the main task set T m ; λ k represents the auxiliary task Tk The weight coefficient, L k represents the auxiliary task T k The loss function of, T k belongs to the auxiliary task set T n ;

[0035] On the premise that the importance levels of the main tasks are the same and the importance levels of the auxiliary tasks are the same, the task weights are updated by Equation (8);

[0036] Organize the weight coefficients of the tasks according to Equation (9):

[0037]

[0038] Then for the k - th level of joint training scheduling combination, the joint loss function is as follows:

[0039]

[0040] Furthermore, in step (2), adaptive alternating joint scheduling is used for multi - task training, specifically:

[0041] Define the task scheduling queue Q list List, traverse the data set, and construct the data subset D according to the sub - task set supported by the data sample taskset ;

[0042] Insert the data subset D taskset into the corresponding scheduling queue list Q according to the number of tasks it contains i During training, extract the alternating training or joint training combinations at each level from the task scheduling queue in order, construct the joint loss function according to Equation (10), and update the model parameters through backpropagation.

[0043] Furthermore, step (2) uses the ELearner Multistate data set for training.

[0044] An online learning state recognition device for multi - tasks, including a multi - task recognition model construction module for the online learner's gaze point and facial expression and an adaptive alternating joint scheduling module;

[0045] The multi - task recognition model construction module for the online learner's gaze point and facial expression is used to construct a multi - task recognition model for the online learner's gaze point and facial expression, which consists of a gaze point estimation sub - task model and a facial expression recognition sub - task model;

[0046] The fixation point estimation sub-task model is used to merge the left-eye features, right-eye features, and Landmarks Grid features and output the fixation landing point; the facial expression recognition sub-task model is used to merge the Landmarks Grid and the face branch features and output the facial expression.

[0047] The adaptive alternating joint scheduling module is used to alternately train the fixation point estimation sub-task model or the facial expression recognition sub-task model according to the feature labels; after completing the alternating training, the task joint level is increased to level 2, and the multi-task recognition model is combinedly trained using the sample data containing both fixation point and facial expression labels.

[0048] Furthermore, the fixation point estimation sub-task model is composed of a general fixation point estimation model that fuses Landmarks and a fixation point personalized calibration model.

[0049] The general fixation point estimation model that fuses Landmarks is used to estimate the fixation point of the learner in the online learning scenario, so as to obtain the general fixation point coordinates.

[0050] The fixation point personalized calibration model is used to estimate the deviation of the general fixation point coordinates, and combine the general fixation point coordinates and the deviation of the general fixation point coordinates to finally obtain the true fixation point coordinates.

[0051] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the multi-task-based online learning state recognition method of the present invention are implemented.

[0052] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the multi-task-based online learning state recognition method of the present invention are implemented.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] The present invention provides a multi-task-based online learning state recognition method, which uses a multi-task learning framework combining a neural network and a fixation point estimation model that fuses Landmarks and personalized calibration to predict two learning state indicators of expression and fixation point; adopts task-adaptive alternating joint scheduling multi-task training to effectively improve the learning performance of each task and make the main task achieve better results. The present invention solves the deficiencies of the general fixation point estimation model in terms of accuracy and generalization ability, and at the same time overcomes the difficulties of systematic implementation and large-scale application; compared with single-task and general scheduling multi-task training, the method of the present invention has achieved better results in both the facial expression recognition and fixation point estimation tasks.

[0055] Furthermore, a two-stage fixation point estimation algorithm that combines Landmarks and personalized calibration is adopted, which solves problems such as insufficient accuracy and generalization performance of general models, difficulty in systematic implementation, and large-scale application.

[0056] The present invention provides an online learning status recognition device based on multi-tasks, including specific modules for completing the above-mentioned working methods;

[0057] A computer device and a storage medium for an online learning status recognition method based on multi-tasks are used to implement the specific steps of the above-mentioned working methods. Description of the Drawings

[0058] Figure 1 It is a diagram of a two-stage personalized fixation point estimation fusion model;

[0059] Figure 2 It is a general fixation point estimation model that combines Landmarks;

[0060] Figure 3 It is a personalized fixation point calibration model architecture;

[0061] Figure 4 It is an extensible multi-task learning framework;

[0062] Figure 5 It is a multi-task learning adaptive scheduling training algorithm framework;

[0063] Figure 6 It is a sub-task Onehot encoding;

[0064] Figure 7 It is a multi-task framework for online learner fixation point and facial expression recognition;

[0065] Figure 8 It is an example of the Multi-Mnist dataset;

[0066] Figure 9 It is the expression distribution of the ELearner MultiState dataset. Detailed Embodiments

[0067] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0068] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0069] Aiming at the problems of strong result bias and large waste of data samples in the single scheduling method for multi-task learning, a task-adaptive alternating joint scheduling multi-task training algorithm is proposed, and a joint loss function is designed. On the basis of ensuring the full use of the data set and shared feature extraction, the deviation caused by different data sets or task scheduling methods is balanced, and the learning effect of the main task is improved. A model for estimating the fixation points of an online learner that fuses Landmarks and personalized calibration is proposed. While reducing the complexity of the model, this model ensures the detection effect of the model and solves the problem of fixation point personalization to a certain extent.

[0070] The present invention will be further described in detail below with reference to the accompanying drawings:

[0071] A method for identifying the online learning state based on multi-tasks includes two steps of constructing a model and training scheduling, which are specifically as follows:

[0072] 1) A two-stage fixation point estimation model that fuses Landmarks and personalized calibration

[0073] Aiming at the problems of low accuracy and generalization ability of the fixation point detection task model and the difficulties in model training and fine-tuning in the online education scenario, the present invention fuses the fixation point estimation model that fuses Landmarks Grid and the fixation point personalization calibration model based on deviation estimation, and proposes an online learner fixation point estimation model that fuses Landmarks and personalized calibration, as Figure 1 shown.

[0074] Facial image features usually contain information about a person's head pose, and this information is very helpful for gaze estimation. The present invention integrates the face key point information into the Face Grid and proposes a general fixation point estimation CNN model that fuses Landmarks, as Figure 2 shown.

[0075] The model input consists of three parts: the left-eye image, the right-eye image, and the Landmarks Grid feature that fuses Landmarks and the facial Grid. The VGG network is used to perform convolution operations on the input images in the convolutional layer, and the output results are concatenated together through the fully connected layer. Finally, the position of the final fixation point is estimated. Compared with the traditional model, this model can not only track the fixation characteristics of learners more accurately, but also significantly reduce the complexity of the model and the training time.

[0076] Fixation Point Personalized Calibration Model Based on Deviation Estimation

[0077] Affected by various factors, the fixation points calculated by the general model are not completely accurate. For example, there are differences in camera equipment, camera position / angle, and human eye line-of-sight deviation among different learners. Therefore, the present invention proposes a fixation point personalized calibration model based on deviation estimation, and the model architecture is as Figure 3 shown.

[0078] The input of the model is the learner image and its Landmarks Grid. Features are extracted through a small convolutional neural network (Micro CNN), and the output is the difference between the actual fixation point and the estimated value of the general model. The loss function is defined as:

[0079]

[0080] where, g p is the actual fixation point, g gt is the benchmark fixation point estimated by the general model, The structure of the small convolutional neural network is shown in Table 1.

[0081] Table 1 Network Structure of Micro CNN Model

[0082]

[0083] 2) Multi-Task Learning Framework with Adaptive Training Scheduling

[0084] There are usually two frameworks for multi-task learning methods based on deep learning: one is the multi-task deep learning method based on hard constraints; the other is the multi-task deep learning method based on soft constraints. Among them, the method based on hard parameter constraints can achieve better results when the sub-tasks have a strong correlation. To facilitate the introduction of new task branches, a multi-task extension framework is designed as Figure 4 shown.

[0085] Inspired by the methods of hybrid sample data augmentation and adaptive task relationship context, etc., the present invention proposes a task-adaptive alternating joint scheduling multi-task training framework, as Figure 5 shown. The steps are as follows:

[0086] (1) Perform subtask encoding

[0087] Perform Onehot encoding on the subtasks. Assume that multi-task learning contains a total of n subtasks, and the k-th task T k is encoded as Figure 6 as shown.

[0088] Construct a loss function L for each subtask i , then the multi-task loss function vector Consider a sample S in the dataset D. Assume that the sample S has the subtasks T a , T b , …, T k required labels. Then the sample has a joint loss function L for the subtasks T a , T b , …, T k as follows: ab...k as follows:

[0089]

[0090] where S ab...k is the Onehot encoding value enabling the tasks T a , T b , …, T k , that is, the binary encoding with the a, b, …, k bits being 1. According to the labels of the tasks contained in the sample, the original dataset D is divided into alternating training datasets {D1, D2, …, D n} and joint training datasets {D 12 , D 123 , … D ij...k , …, D 12...n}. The alternating training dataset D i only contains the label data required for the task T i , and the joint training dataset D ij...k contains the label data required for the tasks T i , T j , …, T k .

[0091] (2) Design the loss function

[0092] For the case where there are main tasks and auxiliary tasks, the present invention designs the loss functions for alternating training and joint training at all levels as follows.

[0093] For the scheduling of alternating training, let λ i represent the task T during alternating training scheduling iThe weight coefficient is inversely proportional to the ratio of its own dataset size to the overall dataset size:

[0094]

[0095] Then, normalize λ i For task T i , its weight coefficient is:

[0096]

[0097] For the scheduling of joint training, the present invention considers the case where there is a primary task T m , and an auxiliary task T n . In this case, improving the performance of the primary task T m is the primary goal, and a certain degree of reduction in the inference effect of the auxiliary task T n is allowed. Therefore, during joint training, the weight coefficients of each task are defined as follows:

[0098]

[0099] Where λ j represents the weight coefficient of task T j , L j represents the loss function of task T j , and T j belongs to the primary task set T m . According to experience, assuming that the importance levels of the primary tasks are the same and the importance levels of the auxiliary tasks are the same, the task weights are updated by the above formula. When the loss function of task T i becomes smaller relative to the loss functions of all primary tasks, its weight coefficient λ i becomes larger; conversely, its weight coefficient λ i becomes smaller. Thus, the loss functions of all tasks are aligned with the primary tasks, achieving the purpose of primarily considering the primary tasks. Organize the weight coefficients of the tasks in the following manner:

[0100]

[0101] Then, for the k-level joint training scheduling combination, there is the following loss function:

[0102]

[0103] This formula can calculate the joint loss function for any combination of subtasks. For all dataset samples, regardless of the information and labels required to support multiple tasks, they can be fully utilized and learned, thereby training a better multi-task model.

[0104] (3) Design of the joint training adaptive algorithm

[0105] The specific joint training adaptive scheduling algorithm is shown in Table 2.

[0106] Table 2 Joint Training Adaptive Scheduling Algorithm

[0107]

[0108] First, define the task scheduling queue Q list list. Traverse the dataset, and construct data subsets D for data samples according to the subtask sets they support taskset . Insert D taskset into the corresponding scheduling queue list Q according to the number of tasks it contains i . During training, extract alternating training or multi-level joint training combinations from the task scheduling queue in sequence, construct a joint loss function according to formula (10), and update the model parameters through backpropagation.

[0109] Based on the above work, on the basis of the general model for fixation point detection, add facial expression recognition to construct a multi-task architecture for two tasks, as Figure 7 shown. The framework has a total of four branches, namely the left eye, right eye, LandmarksGrid, and face branches. Among them, the Landmarks Grid branch is shared by the fixation point estimation task and the facial expression recognition task. After extracting the features of each path respectively, the fixation point estimation task combines the features of the left eye, right eye, and Landmarks Grid branches as the input for subsequent steps; while the expression recognition task combines the features of the Landmarks Grid and face branches as the input for subsequent steps. Finally, set a joint loss function for the fixation landing point and facial expression tasks. Then splice the features of the four branches and use them as the unified shared layer for the fixation point estimation and facial expression recognition subtasks. Among them, the fixation point estimation task combines the features of the left eye, right eye, and Landmarks Grid branches as the input for subsequent steps; while the expression recognition task combines the features of the Landmarks Grid and face branches as the input for subsequent steps. During training, first alternately select one task from the two subtasks at a fixed or random ratio for training in the current batch. After completing the alternating training, set a joint loss function for the fixation landing point and facial expression tasks, and further perform secondary joint training. Add the loss functions of the two subtasks, and use the optimization function to optimize the joint loss of the two subtasks to achieve the overall optimum.

[0110] The present invention also provides a multi-task online learning state recognition device, including a multi-task recognition model construction module for the gaze point and facial expression of an online learner and an adaptive alternating joint scheduling module; the multi-task recognition model construction module for the gaze point and facial expression of the online learner is used to construct a multi-task recognition model for the gaze point and facial expression of the online learner, which is composed of a gaze point estimation sub-task model and a facial expression recognition sub-task model; the gaze point estimation sub-task model is used to merge the left eye feature, the right eye feature and the Landmarks Grid feature and output the gaze landing point; the facial expression recognition sub-task model is used to merge the Landmarks Grid and the face branch feature and output the facial expression; the adaptive alternating joint scheduling module is used to alternately train the gaze point estimation sub-task model or the facial expression recognition sub-task model according to the feature label; after the alternating training is completed, the task joint level is increased to level 2, and the multi-task recognition model is combined and trained with the sample data containing both the gaze point and the facial expression label.

[0111] In one embodiment, a computer device is provided, and the computer device may be a server. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it realizes the recognition method of the multi-task online learning state recognition device.

[0112] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are realized: (1) Construct a multi-task recognition model for the gaze point and facial expression of an online learner, which is composed of a gaze point estimation sub-task model and a facial expression recognition sub-task model;

[0113] The gaze point estimation sub-task model is used to merge the left eye feature, the right eye feature and the Landmarks Grid feature and output the gaze landing point; the facial expression recognition sub-task model is used to merge the Landmarks Grid and the face branch feature and output the facial expression;

[0114] (2) Perform multi-task training using adaptive alternating joint scheduling

[0115] During training, first perform alternating training on the fixation point estimation sub-task model or the facial expression recognition sub-task model according to the feature labels; after completing the alternating training, raise the task joint level to level 2, and use sample data containing both fixation point and facial expression labels to perform combined training on the multi-task recognition model.

[0116] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: (1) Construct a multi-task recognition model of the fixation point and facial expression of an online learner composed of a fixation point estimation sub-task model and a facial expression recognition sub-task model;

[0117] The fixation point estimation sub-task model is used to merge the left eye feature, the right eye feature, and the Landmarks Grid feature and output the fixation landing point; the facial expression recognition sub-task model is used to merge the Landmarks Grid and the face branch feature and output the facial expression.

[0118] (2) Perform multi-task training using adaptive alternating joint scheduling

[0119] During training, first perform alternating training on the fixation point estimation sub-task model or the facial expression recognition sub-task model according to the feature labels; after completing the alternating training, raise the task joint level to level 2, and use sample data containing both fixation point and facial expression labels to perform combined training on the multi-task recognition model.

[0120] Embodiment

[0121] The fixation point estimation model is respectively verified for effectiveness using the ELearner Multistate dataset and the Gaze Capture dataset. The multi-task learning model is respectively verified for effectiveness using the ELearner Multistate and the Multi-Mnist datasets.

[0122] The ELearner Multistate dataset is as follows:

[0123] Relying on the online learning scenario video resources collected in real time by the education big data online education platform, in the single-person online education scenario of the present invention, the videos collected by the front screen camera and the logs collected by the Tobii eye tracker are used as data sources to construct an online learner fixation point estimation dataset. It includes: the learner facial RGB image dataset face-images, the learning interface screen capture dataset screen-images, and the eye movement log record dataset eye-record. The present invention names the self-constructed online education scenario learner fixation point and facial expression dataset as the ELearner MultiState dataset.

[0124] The Gaze Capture dataset is as follows:

[0125] The Gaze Capture dataset is a large-scale gaze estimation dataset collected by Krafka, covering 1474 volunteers and containing 2,445,504 samples. Among them, 1249 volunteers used iPhones and 225 volunteers used iPads to provide relevant data.

[0126] The Multi-Mnist dataset is as follows:

[0127] To convert the digit classification task into a multi-task problem, Sabour et al. merged multiple Mnist handwritten digit images. The present invention constructs a three-task Multi-Mnist dataset in a similar way, as Figure 8 shown.

[0128] The test results are as follows:

[0129] To verify the effects of the two-stage gaze point estimation model that fuses Landmarks and personalized calibration and the multi-task adaptive training scheduling algorithm adopted by the present invention, experimental comparisons are carried out respectively. Among them,

[0130] (1) The two-stage gaze point estimation model that fuses Landmarks and personalized calibration

[0131] Based on the self-constructed online learner state dataset ELearner MultiState mentioned above, experiments are carried out respectively using Itracker, AFFNet, and the two-stage personalized fusion model proposed by the present invention. Among them, the general gaze point estimation model is named GGL (General Gaze Locating), and the gaze point estimation model that fuses personalized calibration is named GLIA (Gaze Locating with Individual Adaptation). The experimental results are shown in Table 3.

[0132] Table 3 Experimental results of ELearner MultiState

[0133]

[0134] As can be seen from the chart, the method proposed by the present invention has achieved the best results on the ELearner MultiState dataset. The average error of the fixation point of the general gaze estimation model GGL is 5.30 cm. After integrating the personalized calibration model, the average error of the fixation point is reduced to 4.69 cm. The notebook screen used in the ELearner MultiState dataset is 14 inches, that is, the diagonal length of the screen is about 46.6 cm. That is, the average errors of the fixation point positions estimated by the general estimation model and the personalized calibration model are about 11.37% and 10.06% of the entire screen respectively.

[0135] The present invention further verifies the model on the Gaze Capture dataset, and the experimental results are shown in Table 4.

[0136] Table 4 Gaze Capture Experimental Results

[0137]

[0138] Since the device screen sizes used in the Gaze Capture dataset are 9.7 inches and 4.0 inches respectively, that is, the diagonal lengths of the device screens are 24.64 cm and 10.16 cm respectively, the experimental average error is relatively small numerically, but the error percentage is about 10%-20%. Due to the small size of the iPhone screen, the error percentage even exceeds 20% of the screen. The proposed gaze point model GLIA integrating personalized calibration of the present invention can achieve good results on all devices involved in the experiment. The general gaze point detection model can achieve better results on iPads and laptops. However, the estimation effect on the iPhone is reduced compared to Itracker. The reason for this problem is that when using the iPhone, there is a high probability that the target face is not completely included in the image.

[0139] (2) Multi-task Adaptive Training Scheduling Algorithm

[0140] The statistical situation of the facial expression part of the ELearner MultiState dataset is as Figure 9 shown.

[0141] The present invention trains the two subtasks of gaze point and facial expression based on the ELearner MultiState dataset, and compares the effects of single-task training, alternating training, and alternating + joint training methods respectively. The experimental results are shown in Table 5.

[0142] Table 5 ELearner MultiState Dataset Experimental Results

[0143]

[0144] The proposed multi-task algorithm for adaptive training of the present invention is verified by constructing a Multi-Mnist dataset.

[0145] Each Multi-Mnist multi-task digital classification image sample is divided into three parts: left, middle, and right. A handwritten digit image from the Mnist dataset is randomly selected for each part. The final three tasks are: identifying the digit in the left picture (Task-L), identifying the digit in the middle picture (Task-M), and identifying the digit in the right picture (Task-R). The experimental results are shown in Table 6. In the experiment, both +2-level joint training and +3-level joint training use Task-M as the main task.

[0146] Table 6 Experimental Results of Multi-Mnist Dataset

[0147]

[0148] In single-task learning, Task-R has better results. When using alternating training, the model performance decreases to a certain extent. The main reason is that alternating training causes the model to no longer fully learn in the direction of Task-R during training, and the shared information obtained from Task-L and Task-M is not sufficient to compensate for the decrease in the performance of Task-R caused by the scattered learning direction. When using alternating training for Task-L and Task-M, sufficient shared information is obtained from other tasks, so the learning performance of both increases to a certain extent. On the basis of alternating training, 2-level joint training is further introduced, that is, the subtasks are combined in pairs in different ways and jointly trained: Task-L and Task-M are combined, Task-M and Task-R are combined, and Task-L and Task-R are combined. +2-level joint training sets Task-M as the main task, and the task weight coefficients are all initialized to 0.3. The experimental results show that the results of Task-L, Task-M, and Task-R are all improved to a certain extent compared with single-task training and alternating training, and the main task-M has the largest improvement, about 0.53%, and Task-L and Task-R are improved by about 0.24% and 0.16% respectively. Finally, on the basis of alternating training + 2-level joint training, 3-level joint training is further introduced, that is, Task-L, Task-M, and Task-R are combined together for joint training. It can be found from the results that the model performance of the main task-M has been further improved, while the performance of the auxiliary tasks Task-L and Task-R has hardly improved and even decreased slightly.

[0149] The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention fall within the protection scope of the claims of the present invention.

Claims

1. An online learning state recognition method based on multi-tasks, characterized in that, Including the following steps: (1) Construct a multi-task recognition model for the gaze point and facial expression of an online learner, which consists of a gaze point estimation sub-task model and a facial expression recognition sub-task model; The gaze point estimation sub-task model is used to merge the left eye feature, the right eye feature, and the Landmarks Grid feature, and output the gaze landing point; the facial expression recognition sub-task model is used to merge the Landmarks Grid and the face branch feature, and output the facial expression; (2) Perform multi-task training using adaptive alternating joint scheduling During training, first perform alternating training on the gaze point estimation sub-task model or the facial expression recognition sub-task model according to the feature labels of the samples in the training dataset; after completing the alternating training, increase the task joint level to level 2, and use the sample data containing both gaze point and facial expression labels to perform combined training on the multi-task recognition model; in step (1), the gaze point estimation sub-task model is a two-stage gaze point estimation model that combines a general gaze point estimation model integrating Landmarks and a gaze point personalized calibration model; When running the gaze point estimation sub-task model, first use the general gaze point estimation model integrating Landmarks to estimate the gaze point of the learner in the online learning scenario, so as to obtain the general gaze point coordinates; then, use the gaze point personalized calibration model to estimate the deviation of the general gaze point coordinates, and combine the general gaze point coordinates and the deviation of the general gaze point coordinates to finally obtain the true gaze point coordinates; In step (2), perform multi-task training using adaptive alternating joint scheduling. The specific process is as follows: (201) Perform sub-task encoding The multi-task learning contains a total of n sub-tasks, and perform Onehot encoding on the sub-tasks; Construct the loss function \(L\) for each subtask i , then the multi-task loss function vector Consider the sample \(S\) in the dataset \(D\). If the sample \(S\) has the subtasks \(T\) a , \(T\) b , …, \(T\) k required labels, then the joint loss function \(L\) of the sample \(S\) for the subtasks \(T\) a , \(T\) b , …, \(T\) k is as follows: ab...k is: Among them, S ab...k is the Onehot encoding value enabling tasks T a , T b , …, T k , that is, the binary encoding with the a, b, …, k bits being 1; According to the labels of the tasks contained in the sample S, the original dataset D is divided into alternating training datasets {D1, D2, …, D n} and joint training datasets {D 12 , D 123 , … D ij...k , …, D 12...n}; Alternating training dataset D i Contains only task T i The required labeled data, combined training dataset D ij...k Contains task T i ,T j ,…,T k The required labeled data; (202) Design the loss function During alternating training scheduling, sub-task T i has a weight coefficient of λ i , and let the weight coefficient λ i be inversely proportional to the ratio of the size of its own dataset D i to the size of the overall dataset D total : Then normalize λ i and the weight coefficient λ i of the subtask T i is as follows: Among them, λ i represents the weight coefficient of subtask T i , N represents the total number of subtasks, mount(D i ) represents the dataset size of subtask T i ; When scheduling joint training, the training dataset is divided into a main task set T m and an auxiliary task set T n . When conducting joint training, the weight coefficients of each task are defined as follows: Among them, λ j represents the weight coefficient of the main task T j , L j represents the loss function of the main task T j , T j belongs to the main task set T m ; λ k represents the weight coefficient of the auxiliary task T k , L k represents the loss function of the auxiliary task T k , T k belongs to the auxiliary task set T n ; On the premise that the importance levels of the main tasks are the same and the importance levels of the auxiliary tasks are the same, update the task weights through formula (8); Organize the weight coefficients of the tasks according to formula (9): Then for the k-level joint training scheduling combination, the joint loss function is as follows: In step (2), perform multi-task training using adaptive alternating joint scheduling, specifically: Define the task scheduling queue Q list A list is used to traverse the data set, and data subsets D are constructed for data samples according to the set of subtasks they support taskset ; Insert the data subset D taskset into the corresponding scheduling queue list Q according to the number of tasks it contains i During training, alternately extract the training or combined training at all levels from the task scheduling queue in sequence, construct the combined loss function according to Equation (10), and update the model parameters through backpropagation.

2. The multi-task-based online learning status recognition method according to claim 1, wherein In step (1), when running the gaze point estimation sub-task model, specifically: Input the learner's image into the general gaze point estimation model integrating Landmarks, and use VGG to extract the left eye and right eye image features; use the Face Alignment algorithm to extract the facial key points of the learner, and at the same time intercept the eye images to construct the Landmarks Grid, and use Micro CNN to extract the Landmarks Grid features; then splice the left eye feature, the right eye feature, and the Landmarks Grid features in the fully connected layer, and finally output the general gaze landing point of the learner; At the same time, use the learner's facial image as the input of the gaze point personalized calibration model, extract features through the convolutional neural network, and finally output the difference between the actual gaze landing point and the estimated landing point of the general model; Estimate the learner's final actual fixation point by combining the general fixation point of the learner and the difference between the actual fixation point and the estimated fixation point of the general model.

3. An online learning status recognition device for multi-tasks, characterized in that, It includes a multi-task recognition model construction module for the fixation point and facial expression of online learners and an adaptive alternating joint scheduling module; The multi-task recognition model construction module for the fixation point and facial expression of online learners is used to construct a multi-task recognition model for the fixation point and facial expression of online learners, which consists of a fixation point estimation sub-task model and a facial expression recognition sub-task model; The fixation point estimation sub-task model is used to merge the left eye feature, the right eye feature, and the Landmarks Grid feature and output the fixation point; the facial expression recognition sub-task model is used to merge the Landmarks Grid and the face branch feature and output the facial expression; The adaptive alternating joint scheduling module is used to alternately train the fixation point estimation sub-task model or the facial expression recognition sub-task model according to the feature label; after completing the alternating training, the task joint level is increased to level 2, and the multi-task recognition model is combined and trained with the sample data containing both the fixation point and the facial expression label; The fixation point estimation sub-task model is a two-stage fixation point estimation model that combines a general fixation point estimation model integrating Landmarks and a fixation point personalized calibration model; When running the fixation point estimation sub-task model, first use the general fixation point estimation model integrating Landmarks to estimate the fixation point of the learner in the online learning scenario, so as to obtain the general fixation point coordinates; then, use the fixation point personalized calibration model to estimate the deviation of the general fixation point coordinates, and combine the general fixation point coordinates and the deviation of the general fixation point coordinates to finally obtain the true fixation point coordinates; Adaptive alternating joint scheduling is used for multi-task training, and the specific process is as follows: (201) Perform sub-task encoding Multi-task learning includes a total of n sub-tasks, and Onehot encoding is performed on the sub-tasks; Construct the loss function \(L\) for each subtask i , then the multi-task loss function vector Considering the sample \(S\) in the dataset \(D\), if the sample \(S\) has the subtasks \(T\) a , \(T\) b , …, \(T\) k with the required labels, then the joint loss function \(L\) of the sample \(S\) for the subtasks \(T\) a , \(T\) b , …, \(T\) k is as follows: ab...k For: Among them, S ab...k is the Onehot encoding value that enables tasks T a , T b , …, T k , that is, the binary encoding with the a, b, …, k bits being 1; According to the labels of the tasks contained in the sample S, the original dataset D is divided into alternating training datasets {D1, D2, …, D n} and joint training datasets {D 12 , D 123 , … D ij...k , …, D 12...n}; Alternating training dataset D i Contains only the task T i Required labeled data, combined training dataset D ij...k Contains the task T i , T j , …, T k Required labeled data; (202) Design the loss function During alternating training scheduling, subtask T i has a weight coefficient of λ i , and let the weight coefficient λ i be inversely proportional to the ratio of the size of its own dataset D i to the size of the overall dataset D total : Then, normalize λ i and the weight coefficient λ i of the sub-task T i is as follows: Among them, λ i represents the weight coefficient of subtask T i , N represents the total number of subtasks, mount(D i ) represents the dataset scale of subtask T i ; When scheduling joint training, the training dataset is divided into a main task set T m and an auxiliary task set T n . When conducting joint training, the weight coefficients of each task are defined as follows: Among them, λ j represents the weight coefficient of the main task T j , L j represents the loss function of the main task T j , T j belongs to the main task set T m ; λ k represents the weight coefficient of the auxiliary task T k , L k represents the loss function of the auxiliary task T k , T k belongs to the auxiliary task set T n ; On the premise that the importance levels of the main tasks are the same and the importance levels of the auxiliary tasks are the same, the task weights are updated by formula (8); The weight coefficients of the tasks are organized according to formula (9): Then for the k-level joint training scheduling combination, the joint loss function is as follows: In step (2), adaptive alternating joint scheduling is used for multi-task training, specifically: Define the task scheduling queue Q list A list is used to traverse the data set, and data subsets D are constructed from data samples according to the set of subtasks they support taskset ; Insert the data subset D taskset into the corresponding scheduling queue list Q according to the number of tasks it contains i During training, alternately extract training or combined training at all levels from the task scheduling queue in sequence, construct a combined loss function according to Equation (10), and update the model parameters through backpropagation.

4. The multi-task online learning status recognition device according to claim 3, characterized in that The fixation point estimation sub-task model consists of a general fixation point estimation model integrating Landmarks and a fixation point personalized calibration model; The general fixation point estimation model integrating Landmarks is used to estimate the fixation point of the learner in the online learning scenario, so as to obtain the general fixation point coordinates; The fixation point personalized calibration model is used to estimate the deviation of the general fixation point coordinates, and combine the general fixation point coordinates and the deviation of the general fixation point coordinates to finally obtain the true fixation point coordinates.

5. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the multi-task-based online learning state recognition method according to any one of claims 1-2.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the multi-task-based online learning status recognition method according to any one of claims 1-2.