A multi-modal network training system student learning state evaluation method and system
By constructing a three-dimensional emotion model and combining it with EEG signals and facial images, the problem of accuracy in assessing learners' learning status in online training systems has been solved, enabling a comprehensive and accurate assessment of learners' learning status and analysis of course effectiveness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-20
- Publication Date
- 2026-04-14
AI Technical Summary
Existing online training systems struggle to accurately assess learners' learning status, especially in situations where online learning involves separation of time and space. Current methods fail to fully reflect learners' attention and learning progress, resulting in low accuracy of assessment results.
A multimodal online training system is used to construct a three-dimensional emotion model by collecting trainees' EEG signals and facial images. Statistical indicators and hierarchical analysis are used to comprehensively evaluate trainees' learning status, including three dimensions: arousal, interest, and satisfaction. The system makes accurate judgments by combining EEG signals and facial expressions.
It enables a comprehensive and accurate assessment of trainees' learning status, analyzes the effectiveness of different training courses, optimizes evaluation methods, and ensures the rationality and accuracy of evaluation results.
Smart Images

Figure CN117315420B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote training, and more specifically, to a method and system for evaluating the learning status of trainees in a multimodal network training system related to power grid dispatchers. Background Technology
[0002] With profound changes in power systems across sources, grids, and loads, the characteristics of power grid operation are exhibiting new features. This necessitates that power grid dispatching agencies at all levels improve their perception, understanding, and prediction of these new power system operating characteristics, and enhance their ability to manage new AC / DC hybrid power grids. By using a Dispatcher Training Simulation System (DTS) to simulate various power grid operating states, dispatchers can be provided with personalized and intelligent online knowledge learning, skills practice, and joint dispatching exercises. This establishes a new training platform integrating learning, training, evaluation, and guidance, becoming the safest, most economical, and most effective means of dispatching business training. This comprehensively improves the efficiency and effectiveness of dispatcher training, providing intellectual support for the safe, stable, and economical operation of large power grids.
[0003] A stable learning state is crucial for achieving good learning outcomes. However, the spatial and temporal separation inherent in online learning makes it difficult to monitor learners' focus in a timely manner. Therefore, exploring feasible methods for accurately identifying online learning status is essential. In online training environments, analyzing changes in learners' facial expressions can provide a better understanding of their learning status.
[0004] In 1971, American scholars Ekman et al., through numerous facial expression recognition experiments, first categorized human faces into six basic expressions: happiness, surprise, fear, sadness, disgust, and anger. Recognizing these basic expressions allows for the classification of learners' emotions into positive, negative, and neutral emotions. However, basic expressions cannot fully and accurately reflect learners' attention and learning status. Further exploration of monitoring systems for online education and training has led to detection methods that combine factors such as head posture, facial expression, eye state, mouth state, and gaze. Attention is categorized into a limited number of levels, such as focused, normal, and fatigued. This limited categorization results in low accuracy of feedback information, hindering the accurate assessment of learners' focus and compromising practical application effectiveness. While some studies yield relatively high-reliability evaluations, their final judgments only consider whether learners are distracted, which is too simplistic and fails to comprehensively reflect the significant differences in learning status among different learners. Therefore, it is necessary to provide a comprehensive and accurate method and system for evaluating learners' learning status.
[0005] On the other hand, regardless of whether the brain is in a normal state of function or a state of brain disease, as long as our brain is still operating, there is neural activity within it. Neural activity generates electrical signals that propagate to the surface of the head. If electrodes are inserted into the head, we can observe an electroencephalogram (EEG). In 1924, Hans Berger discovered human brain activity. He first named alpha waves and beta waves. EEG provides two basic pieces of information: spatial distribution and temporal progression. Neural activity can generate topographic maps and temporal progression. With these topographic maps, we can infer where neural activity exists in the brain, thus achieving spatial localization. These EEG results can help us understand cognitive processes or brain diseases, allowing us to implement regulatory measures and interventions. After regulation, changes in EEG can lead to treatment of diseases or improvement of cognitive abilities. Furthermore, with authorization, EEG signals can also be used to perceive an individual's emotional indicators of their external environment, such as level of attention and responsiveness.
[0006] However, the aforementioned advanced technologies have not been applied effectively to assess the learning status of learners in online training systems. Therefore, there is an urgent need for a multimodal online training system learner learning status evaluation method, system, terminal, and computer-readable storage medium. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this invention provides a method, system, terminal, and computer-readable storage medium for evaluating the learning status of trainees in a multimodal network training system. The method obtains the trigger count of each dimension and level in a three-dimensional emotion model by acquiring facial images from electroencephalogram (EEG) signals, and then scores the learning status of trainees.
[0008] The present invention adopts the following technical solution.
[0009] The first aspect of this invention relates to a method for evaluating the learning status of trainees in a multimodal network training system. The method includes the following steps: Step 1, collecting EEG signals and facial images of trainees within a preset time period, and preprocessing the EEG signals and facial images to extract multiple EEG signals and multiple facial images to be identified; Step 2, defining the dimensions and levels of a three-dimensional emotion model, and using statistical indicators to calculate the dimensions and levels of each EEG signal and facial image to be identified in the three-dimensional emotion model; Step 3, using the number of recognitions of each level in each dimension of the three-dimensional emotion model as input, constructing an analytic hierarchy process using the three-dimensional emotion model to score the learning status of trainees.
[0010] Preferably, the process of collecting EEG signals from trainees within a preset time period includes: placing multiple EEG acquisition electrodes at preset positions on the trainees' heads, with the EEG acquisition electrodes being divided into a first electrode, a second electrode, and a third electrode according to the preset head position.
[0011] Preferably, the EEG signal is preprocessed, including: using a low-pass filtering method to filter out high-frequency interference above 70Hz and power frequency interference from the EEG lead signals acquired by the first, second, and third electrodes to obtain a noise-filtered signal; using independent component analysis to decompose the noise-filtered signal into independent components, removing the electrooculogram artifacts, and then performing db4 wavelet decomposition and reconstruction to filter out the DC component and low-frequency drift less than 0.1Hz, thereby obtaining the real EEG signal; and truncating the real EEG signal into a fixed-length time dimension to obtain multiple EEG signals to be identified.
[0012] Preferably, the acquisition of facial images of trainees within a preset time period includes: acquiring facial depth images of the trainees using a depth camera and generating a three-dimensional point cloud, and acquiring RGB two-dimensional images using an RGB area array camera.
[0013] Preferably, the facial images are preprocessed, including: defining facial feature points for the trainees, with the facial feature points being T1 group feature points, T2 group feature points, and T3 group feature points; extracting depth data of the T1 group feature points from the facial depth image, and extracting the two-dimensional coordinates of the T1 group feature points from the RGB two-dimensional image; converting the depth data and two-dimensional coordinates into three-dimensional coordinates of the T1 group feature points; directly extracting the coordinates of the T2 group feature points from the three-dimensional point cloud, and calculating the three-dimensional coordinates of the T3 group feature points using the three-dimensional coordinates of the T2 group feature points and the three-dimensional coordinates of the T1 group feature points; calculating the Euclidean distance between preset feature points to obtain multiple feature vectors, performing principal component analysis on the multiple feature vectors, and extracting principal component features.
[0014] Preferably, the dimensions and levels of the three-dimensional emotion model are defined, including: the dimensions of the three-dimensional emotion model include arousal, interest and satisfaction, and each dimension is defined with dimensional indicators based on EEG signals and facial images; each dimensional indicator includes the same number of levels, and each level is assigned a level weight.
[0015] Preferably, statistical indicators are used to calculate the dimensions and levels of each EEG signal to be identified in the three-dimensional emotion model. This includes: calculating the mean frequency area, the mean frequency corresponding to the peak power, and the mean median frequency of each real EEG signal collected by the first electrode in multiple preset frequency ranges to form a spectral feature vector for each real EEG signal; calculating the autoregressive coefficient of each real EEG signal collected by the first electrode; standardizing the spectral feature vector and the autoregressive coefficient, using them as input data for EEG arousal, and using the nearest neighbor regression algorithm to identify the arousal level of each real EEG signal.
[0016] Preferably, statistical indicators are used to calculate the dimensions and levels of each EEG signal to be identified in the three-dimensional emotion model. This includes: performing db4 wavelet transform on each real EEG signal collected by the second electrode and extracting the alpha and beta rhythm signals; calculating the average and variance of the power spectral density of each rhythm signal; using the average and variance of the power spectral density of each rhythm signal as features of each real EEG signal; and using the nearest neighbor regression algorithm to identify the level of interest of each real EEG signal.
[0017] Preferably, statistical indicators are used to calculate the dimensions and levels of each EEG signal to be identified in the three-dimensional emotion model, including: the area and peak power of the power spectrum of each real EEG signal collected by the third electrode within a set frequency range; using the area and peak power as input data for EEG satisfaction, the nearest neighbor regression algorithm is used to identify the level of satisfaction of each real EEG signal.
[0018] Preferably, statistical indicators are used to calculate the dimension and level of each facial image to be identified in the three-dimensional emotion model, including: using a support vector machine model to classify the principal component features.
[0019] Preferably, the hierarchical analysis method is constructed using a three-dimensional emotion model, including: the hierarchical analysis method is a two-layer model, the first layer model is an evaluation matrix of arousal, interest and satisfaction, and the second layer model is an evaluation matrix of EEG and facial expression for each dimension.
[0020] Preferably, the construction process of the first-layer model is as follows: compare the importance of arousal, interest and satisfaction in pairs, obtain the evaluation value based on the comparison results, and use the evaluation value as the value of each item in the evaluation matrix.
[0021] Preferably, the construction process of the second-layer model is as follows: Expert scores based on arousal level are obtained for EEG signals and facial images; the average of these expert scores is calculated to generate an arousal level dataset; the arousal level of EEG signals and facial images is obtained through the EEG and facial expression evaluation matrices under arousal level; the matching degree between the arousal level dataset and the arousal level is used to calculate the EEG signal recognition accuracy 'a' and the facial image recognition accuracy 'b', respectively, thereby obtaining the weight coefficients of EEG signals and facial expressions in the second-layer model, which are respectively...
[0022] Preferably, the number of recognitions for each level in each dimension of the three-dimensional emotion model is input into the first-layer model and the second-layer model respectively, thereby obtaining the learning status score of the trainees.
[0023] A second aspect of this invention relates to a learning status evaluation device for a multimodal network training system. The device includes a data acquisition module, a judgment module, and a scoring module. The data acquisition module is used to acquire EEG signals and facial images of trainees within a preset time period, and preprocess the EEG signals and facial images to extract multiple EEG signals and multiple facial images to be identified. The judgment module is used to define the dimensions and levels of a three-dimensional emotion model, and uses statistical indicators to calculate the dimensions and levels of each EEG signal and facial image to be identified in the three-dimensional emotion model. The scoring module uses the number of recognitions of each level in each dimension of the three-dimensional emotion model as input, and constructs an analytic hierarchy process (AHP) using the three-dimensional emotion model to score the learning status of trainees.
[0024] A third aspect of the present invention relates to a terminal, including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to perform the steps of the method according to the first aspect of the present invention.
[0025] A fourth aspect of the present invention relates to a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method of the first aspect of the present invention.
[0026] The beneficial effects of this invention are that, compared with the prior art, the multimodal network training system's student learning status evaluation method, system, terminal, and computer-readable storage medium of this invention obtain the trigger count of each dimension and level in a three-dimensional emotion model through facial images of acquired EEG signals, and realize the scoring of the trainees' learning status. Using the method of this invention, not only can the learning status of different trainees be obtained, but different training courses can also be analyzed to evaluate the quality of course content. Furthermore, this method can subsequently support the correlation between EEG signals and facial expressions, optimizing the evaluation method based on inherent correlation features to accurately judge the learning status solely through facial expressions.
[0027] The beneficial effects of the present invention also include:
[0028] 1. The study comprehensively uses EEG signals and facial image information of trainees to judge their learning status. A three-dimensional emotion model is constructed to link EEG signals and image information that cannot be correlated and analyzed. The study uses three different dimensions to determine the trainees' level of reception of the course, ensuring the rationality and accuracy of the evaluation results.
[0029] 2. The method is based on a three-dimensional emotion model, which categorizes the EEG signals and image features associated with each dimension. According to the assessment requirements, the method precisely calculates and analyzes one or more statistical features of the EEG signals in each dimension, enabling the method to fully support the practical meaning and effectiveness of the three-dimensional emotion assessment indicators through technical means.
[0030] 3. The method uses a multi-level analytic hierarchy process (AHP) to personalize the weights of each level of indicators, ensuring the effectiveness of the AHP. The indicators are then rationally ranked according to their importance, ultimately ensuring the accuracy of the evaluation results. Attached Figure Description
[0031] Figure 1 This is a schematic diagram illustrating the steps of a method for evaluating the learning status of trainees in a multimodal network training system according to the present invention.
[0032] Figure 2 This is a schematic diagram showing the placement and names of the EEG electrodes in a multimodal network training system student learning status evaluation method according to the present invention.
[0033] Figure 3 This is a schematic diagram of facial feature points in a learning status evaluation method for a multimodal network training system according to the present invention.
[0034] Figure 4 This is a schematic diagram of the feature vector formed by the Euclidean distance between facial feature points in a multimodal network training system learning status evaluation method of the present invention.
[0035] Figure 5 This is a block diagram of the three-dimensional emotion recognition system based on facial expressions in a multimodal network training system student learning status evaluation method of the present invention;
[0036] Figure 6 This is a schematic diagram of a three-dimensional emotion model in a learning status evaluation method for a multimodal network training system according to the present invention.
[0037] Figure 7 This is a schematic diagram of the student learning status evaluation index system based on a three-dimensional emotion model in the student learning status evaluation method of a multimodal network training system of the present invention;
[0038] Figure 8 This is an information table of the three-dimensional emotion model in the learning status evaluation method of a multimodal network training system of the present invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The embodiments described in this invention are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments not described in this invention obtained by those skilled in the art based on the embodiments described in this invention without creative effort should fall within the protection scope of this invention.
[0040] Figure 1 This is a schematic diagram illustrating the steps of a method for evaluating the learning status of trainees in a multimodal network training system according to the present invention. Figure 1 As shown, the first aspect of the present invention relates to a method for evaluating the learning status of trainees in a multimodal network training system, the method comprising steps 1 to 3.
[0041] Step 1: Collect EEG signals and facial images of trainees within a preset time period, and preprocess the EEG signals and facial images respectively to extract multiple EEG signals and multiple facial images to be identified.
[0042] This invention provides a method for recognizing the learning status of trainees in a power grid dispatching network training system based on a three-dimensional emotion model and utilizing electroencephalography (EEG) and facial expression recognition. To this end, the method first collects EEG signals and facial expression signals from trainees while they are learning online courses.
[0043] Preferably, the process of collecting EEG signals from trainees within a preset time period includes: placing multiple EEG acquisition electrodes at preset positions on the trainees' heads, with the EEG acquisition electrodes being divided into a first electrode, a second electrode, and a third electrode according to the preset head position.
[0044] Research has found a lateralization of emotional satisfaction in the prefrontal cortex; when people feel happy, the left frontal lobe is more active, while when they feel sad, the right frontal lobe is more active. Furthermore, the level of arousal activation is correlated with the activity levels of both the left and right prefrontal cortexes; and the level of interest is closely related to the EEG signals Fz, Cz, and Pz. Therefore, the system selects eight EEG electrode signals (AF3, AF4, F3, F4, F7, F8, FC5, FC6) in the prefrontal cortex to identify satisfaction, eleven EEG electrode signals (AF3, AF4, F3, F4, F7, F8, FC5, FC6, Fz, Cz, Pz) to identify arousal, and three EEG electrode signals (Fz, Cz, Pz) to identify interest. In addition to the eleven measuring electrodes, two ear electrodes are also placed as reference signals.
[0045] Understandably, this invention is capable of collecting EEG signals from multiple leads of trainees and sending the digitized EEG signals to a processor for further analysis.
[0046] Preferably, the acquisition of facial images of trainees within a preset time period includes: acquiring facial depth images of the trainees using a depth camera and generating a three-dimensional point cloud, and acquiring RGB two-dimensional images using an RGB area array camera.
[0047] Meanwhile, the method can also use a depth camera to acquire facial depth image data and 3D point cloud data of trainees, and use an RGB area array camera to acquire RGB 2D images.
[0048] Preferably, the EEG signal is preprocessed, including: using a low-pass filtering method to filter out high-frequency interference above 70Hz and power frequency interference from the EEG lead signals acquired by the first, second, and third electrodes to obtain a noise-filtered signal; using independent component analysis to decompose the noise-filtered signal into independent components, removing the electrooculogram artifacts, and then performing db4 wavelet decomposition and reconstruction to filter out the DC component and low-frequency drift less than 0.1Hz, thereby obtaining the real EEG signal; and truncating the real EEG signal into a fixed-length time dimension to obtain multiple EEG signals to be identified.
[0049] In one embodiment of the present invention, EEG lead signals from the first electrode, the second electrode, and the third electrode can be collected respectively, and a low-pass filtering method is used to filter out high-frequency interference and power frequency interference from the intercepted 3-second EEG signal. Subsequently, independent component analysis is used to separate ocular artifacts in the EEG signal and to identify ocular artifacts.
[0050] Furthermore, the method can utilize the db4 wavelet in wavelet transform to perform 8-level decomposition and reconstruction of the preprocessed EEG signals from the leads, filtering out the DC component and low-frequency drift in the EEG signals, thus obtaining the preprocessed true EEG signals. Depending on the different definitions of the first, second, and third electrodes, the method can perform different analyses and statistics on multiple EEG signals, obtaining the identification of different dimensions of emotion levels in subsequent steps.
[0051] Figure 2 This is a schematic diagram illustrating the placement and names of the EEG electrodes in a learning status evaluation method for a multimodal network training system according to the present invention. Figure 2 As shown, positions AF3, AF4, F3, F4, F7, F8, FC5, FC6, Fz, Cz, and Pz are the first electrodes; Fz, Cz, and Pz are the second electrodes; and AF3, AF4, F3, F4, F7, F8, FC5, and FC6 are the third electrodes. The definition of the first to third electrodes here is merely related to the reference process of different emotional model dimensions and has no other practical significance.
[0052] Preferably, the facial images are preprocessed, including: defining facial feature points for the trainees, with the facial feature points being T1 group feature points, T2 group feature points, and T3 group feature points; extracting depth data of the T1 group feature points from the facial depth image, and extracting the two-dimensional coordinates of the T1 group feature points from the RGB two-dimensional image; converting the depth data and two-dimensional coordinates into three-dimensional coordinates of the T1 group feature points; directly extracting the coordinates of the T2 group feature points from the three-dimensional point cloud, and calculating the three-dimensional coordinates of the T3 group feature points using the three-dimensional coordinates of the T2 group feature points and the three-dimensional coordinates of the T1 group feature points; calculating the Euclidean distance between preset feature points to obtain multiple feature vectors, performing principal component analysis on the multiple feature vectors, and extracting principal component features.
[0053] Similarly, the method also preprocesses the raw image data related to facial expressions to obtain feature points on the image and feature vectors composed of these feature points.
[0054] Figure 3 This is a schematic diagram of facial feature points in a student learning status evaluation method for a multimodal network training system according to the present invention. For example... Figure 3As shown, the method first 1) extracts T1 group of 3D feature points using depth data and RGB 2D images. First, the 2D coordinates of the feature points are extracted using the RGB 2D images. Then, the depth data of the feature points is obtained using the depth image data. Further, the 3D spatial coordinates of the feature points are obtained using the transformation relationship between the depth data and the 3D point cloud. The obtained T1 group of 3D feature points includes points 8-27, defined as: 16--right outer canthus, 17--right inner canthus, 18--left inner canthus, 19--left outer canthus, 20--right... 21 - Center point of upper eyelid margin; 22 - Center point of lower right eyelid margin; 23 - Center point of lower left eyelid margin; 8 - Outer boundary point of right eyebrow; 9 - Inner boundary point of right eyebrow; 10 - Inner boundary point of left eyebrow; 11 - Outer boundary point of left eyebrow; 12 - Center point of upper right eyebrow; 13 - Center point of lower right eyebrow; 14 - Center point of upper left eyebrow; 15 - Center point of lower left eyebrow; 24 - Right corner of mouth; 25 - Left corner of mouth; 26 - Midpoint of upper lip edge; 27 - Midpoint of lower lip edge.
[0055] Subsequently, 2) the three-dimensional feature points of group T2 were extracted using three-dimensional point cloud data. The obtained three-dimensional feature points of group T2 include 1-5. The feature points are defined as: 1--the deepest midline point of the nasal corner, 2--the tip of the nose, 3--the lower edge of the nasal septum, 4--the right outermost point of the nasal ala, and 5--the left outermost point of the nasal ala.
[0056] Finally, 3) using the three-dimensional feature points 1 and 2 obtained from group T2 and the three-dimensional feature points 24 and 25 obtained from group T1, extract the feature points of group T3, including: 6--chin point, 28--right face edge point, 29--left face edge point, 7--hairline midpoint.
[0057] Figure 4 This is a schematic diagram of the feature vector formed by the Euclidean distance between facial feature points in a student learning status evaluation method of a multimodal network training system according to the present invention. Figure 4 As shown, the Euclidean distances between feature points constitute 77 three-dimensional feature vectors.
[0058] Figure 5 This is a block diagram of the facial expression-based three-dimensional emotion recognition system in a multimodal network training system student learning status evaluation method of the present invention. Figure 5 As shown, after acquiring and analyzing images to obtain feature vectors, this invention can perform principal component analysis and feature extraction on the feature vectors, and then realize the analysis of separate emotion models.
[0059] Step 2: Define the dimensions and levels of the three-dimensional emotion model. Statistical indicators are used to calculate the dimensions and levels of each EEG signal and facial image to be identified in the three-dimensional emotion model.
[0060] Preferably, the dimensions and levels of the three-dimensional emotion model are defined, including: the dimensions of the three-dimensional emotion model include arousal, interest and satisfaction, and each dimension is defined with dimensional indicators based on EEG signals and facial images; each dimensional indicator includes the same number of levels, and each level is assigned a level weight.
[0061] Figure 6 This is a schematic diagram of a three-dimensional emotion model in a student learning status evaluation method for a multimodal online training system according to the present invention. Figure 6 As shown in the figure, the three-dimensional coordinate axes are the satisfaction dimension coordinate axis, the arousal dimension coordinate axis, and the interest dimension coordinate axis, respectively. The range of the emotional satisfaction dimension is from low disappointment to high satisfaction, the range of the emotional arousal dimension is from low lethargy to high excitement, and the range of the emotional interest dimension is from low aversion to high interest.
[0062] Furthermore, the three-dimensional emotion model defined in this invention includes the following dimensions and levels:
[0063] Arousal dimension: Defined as the degree to which a learner's emotional activation during online training ranges from minimum emotional lethargy to maximum emotional excitement. Specifically, in the field of EEG recognition, arousal dimension is divided into five different levels: very excited – excited – neutral – lethargic – very lethargic; in the field of facial expression recognition, arousal dimension is divided into five different levels: excited – active – calm – tired – drowsy.
[0064] Interest dimension: Defined as the attitude and inclination of learners towards training activities when participating in online training, specifically an indicator of the degree of liking or concern for the training activities, ranging from the lowest level of aversion to the highest level of strong interest. In the field of EEG recognition, the interest dimension is divided into five different levels: very interested, interested, neutral, no interest, and very no interest; in the field of facial expression recognition, the interest dimension is divided into five different levels: focused, attentive, indifferent, perfunctory, and averse.
[0065] Satisfaction dimension: Defined as the trainees' heartfelt evaluation of the training effect during online training. It represents the trainees' level of pleasure after their needs are met, reflecting the trainees' perceived effectiveness of the training compared to their expectations, ranging from disappointment to satisfaction. In the field of EEG recognition, the satisfaction dimension is divided into five different levels: very satisfied – satisfied – neutral – disappointed – very disappointed. In the field of facial expression recognition, the satisfaction dimension is divided into five different levels: understanding – doubt – listening – resistance – disdain.
[0066] Specifically, the present invention can define the dimensions and levels of the three-dimensional emotion model according to the above method, and then determine the dimension and level to which each EEG signal and facial image to be identified belongs based on its features.
[0067] Preferably, statistical indicators are used to calculate the dimensions and levels of each EEG signal to be identified in the three-dimensional emotion model. This includes: calculating the mean frequency area, the mean frequency corresponding to the peak power, and the mean median frequency of each real EEG signal collected by the first electrode in multiple preset frequency ranges to form a spectral feature vector for each real EEG signal; calculating the autoregressive coefficient of each real EEG signal collected by the first electrode; standardizing the spectral feature vector and the autoregressive coefficient, using them as input data for EEG arousal, and using the nearest neighbor regression algorithm to identify the arousal level of each real EEG signal.
[0068] Specifically, the method calculates 22 basic EEG feature data related to arousal recognition. These include five data points: the average area under the spectrum of the signals in each of the δ, θ, α, β, and γ frequency ranges for 11 leads; the average peak power frequency (MF) for 5 data points; the average median frequency (MF) for 5 data points (the frequency range where the entire spectrum is divided into two regions, each representing 50% of the total power); and the autoregressive coefficients of the 11 EEG leads. δ, θ, α, β, and γ represent commonly used frequency ranges for EEG signals: δ waves are typically between 0.2 and 3 Hz, θ waves between 3 and 8 Hz, α waves between 8 and 12 Hz, β waves between 12 and 27 Hz, and γ waves above 27 Hz.
[0069] The above 26 features are standardized and used as input data for the wakefulness recognition model.
[0070] In one embodiment of the present invention, five professional technicians evaluate the emotional arousal of the subjects and obtain arousal evaluation values for the corresponding information. The arousal types are divided into five categories: very lethargic, lethargic, neutral, excited, and very excited. The classification results are used as label data, and an EEG-based arousal recognition model is established using the nearest neighbor regression algorithm to achieve arousal recognition.
[0071] Preferably, statistical indicators are used to calculate the dimensions and levels of each EEG signal to be identified in the three-dimensional emotion model. This includes: performing a db4 wavelet transform on each real EEG signal collected by the second electrode and extracting the α rhythm signal and β rhythm signal; calculating the average and variance of the power spectral density of each rhythm signal; using the average and variance of the power spectral density of each rhythm signal as features of each real EEG signal; and using the nearest neighbor regression algorithm to identify the level of interest of each real EEG signal.
[0072] In the above calculation process, the method uses the db4 wavelet in wavelet transform to perform 8-level decomposition and reconstruction on the EEG signal after filtering out low frequencies, extracting the alpha and beta rhythm signals from the EEG signal. Subsequently, the fast Fourier transform algorithm is used to solve for the power spectral density (PSD) of each frequency of the alpha and beta rhythm signals in each lead (Fz, Cz, Pz). i The mean F1 and variance F2.
[0073] The formulas for calculating the mean and variance are as follows:
[0074]
[0075]
[0076] Using the above method, a total of 12 features of the EEG signal were obtained.
[0077] Preferably, statistical indicators are used to calculate the dimensions and levels of each EEG signal to be identified in the three-dimensional emotion model, including: the area and peak power of the power spectrum of each real EEG signal collected by the third electrode within a set frequency range; using the area and peak power as input data for EEG satisfaction, the nearest neighbor regression algorithm is used to identify the level of satisfaction of each real EEG signal.
[0078] It is understood that in this invention, after preprocessing the eight leads AF3, AF4, F3, F4, F7, F8, FC5, and FC6, 16 basic EEG feature data related to satisfaction recognition are calculated, including: the area and peak power of the power spectrum of the four leads of the left frontal lobe (AF3, F3, F7, and FC5) and the area and peak power of the power spectrum of the four leads of the right frontal lobe (AF4, F4, F8, and FC6).
[0079] Subsequently, the 16 features were standardized and used as input data for the satisfaction recognition model. Finally, five professional technicians evaluated the participants' emotional satisfaction, obtaining corresponding satisfaction ratings. The satisfaction levels were categorized into five types: very disappointed, disappointed, neutral, satisfied, and very satisfied. The classification results were used as label data, and an EEG-based satisfaction recognition model was established using the nearest neighbor regression algorithm to achieve satisfaction recognition.
[0080] Preferably, statistical indicators are used to calculate the dimension and level of each facial image to be identified in the three-dimensional emotion model, including: using a support vector machine model to classify the principal component features.
[0081] Understandably, this invention may include principal component analysis for arousal recognition, principal component analysis for interest recognition, and principal component analysis for satisfaction recognition. The method utilizes support vector machines (SVMs) to recognize emotions across three dimensions: SVM arousal recognition, SVM interest recognition, and SVM satisfaction recognition, ultimately identifying the dimension and level corresponding to each facial expression.
[0082] Step 3: Using the number of recognitions for each level in each dimension of the three-dimensional emotion model as input, construct the hierarchical analysis method using the three-dimensional emotion model to score the learning status of trainees.
[0083] After implementing step 2, the present invention also implements the evaluation of trainees' learning status based on multi-level fuzzy comprehensive evaluation.
[0084] Specifically, the method is based on a three-dimensional emotion model and EEG and facial expression recognition technology to construct an evaluation system and method for evaluating students' learning status, and sets up three levels of evaluation indicators.
[0085] Figure 7 This is a schematic diagram of the student learning status evaluation index system based on a three-dimensional emotion model in the student learning status evaluation method of a multimodal network training system of the present invention. Figure 8 This is an information table of the three-dimensional emotion model in the learning status evaluation method of a multimodal network training system of the present invention.
[0086] like Figure 7 and Figure 8 As shown, the first-level evaluation indicators include arousal level B1, interest level B2, and satisfaction level B3; the second-level evaluation indicators include: EEG arousal level C, which belongs to the first-level evaluation indicator arousal level B1. 11 and facial expression arousal C 12 EEG interest level C, which belongs to the first-level evaluation index B2, is related to interest level. 21 and facial expression interest C 22 EEG satisfaction C, which belongs to the first-level evaluation index B3, is... 31 and facial expression satisfaction C 32 The third-level evaluation indicators include: EEG arousal level C, which belongs to the second-level evaluation indicator. 11 Number of times of extreme excitement D 111 Number of excitations D 112 Neutral frequency D 113 Number of times of lethargy D 114 and the number of times of extreme lethargy D 115 This belongs to the second-level evaluation index, facial expression arousal C. 12 Number of excitations D 121 Number of active users D 122 Number of calm times D 123 Number of times of fatigue D 124 and the number of times of drowsiness D 125 It belongs to the second-level evaluation index, EEG interest level C. 21 Number of times of great interest D 211 Number of times of interest D 212 Neutral frequency D 213 Number of times I wasn't interested (D) 214 Number of times with very little interest (D) 215 This belongs to the second-level evaluation index, facial expression interest level C. 22 Number of times of focus D 221 Number of views D 222 Number of times of normalcy D 223 Number of times of perfunctory service (D) 224 And the number of times of aversion D 225 This belongs to the second-level evaluation index, EEG satisfaction C. 31 Number of times very satisfied D 311 Number of times satisfied (D) 312 General number D 313 Number of disappointments (D) 314 And the number of times I was very disappointed (D) 315 This belongs to the second-level evaluation index, facial expression satisfaction (C). 32 Number of times of comprehension D 321 Number of doubts (D) 322 Number of times of listening (D) 323 Number of times of resistance D 324 and the number of times of disdain D 325The "number" represents the total number of times a certain level of a three-dimensional emotion appeared during the monitoring period, as identified using EEG or facial expression recognition.
[0087] This invention uses the analytic hierarchy process (AHP) to determine the weight coefficients of each level of indicators.
[0088] Preferably, a three-dimensional emotion model is used to construct an analytic hierarchy process (AHP), which includes a two-layer model. The first layer model is an evaluation matrix for arousal, interest, and satisfaction. The second layer model is an evaluation matrix for EEG and facial expression for each dimension. The construction process of the first layer model is as follows: the importance of arousal, interest, and satisfaction is compared pairwise, and the evaluation value is obtained based on the comparison results. The evaluation value is used as the value of each item in the evaluation matrix.
[0089] The construction process of the second-layer model is as follows: EEG signals and facial images are acquired and scored by experts based on arousal level; the average of these expert scores is calculated to generate an arousal level dataset; the arousal level of EEG signals and facial images is obtained through the EEG and facial expression evaluation matrices under arousal level; the matching degree between the arousal level dataset and the arousal level is used to calculate the EEG signal recognition accuracy 'a' and the facial image signal recognition accuracy 'b', thereby obtaining the weight coefficients of EEG signals and facial expressions in the second-layer model.
[0090] Specifically, the Analytic Hierarchy Process (AHP) was used to determine the weight coefficients of the first-level indicators: arousal level (B1), interest level (B2), and satisfaction level (B3). Experts from the power system and training fields were invited to conduct pairwise comparisons of the importance of these first-level indicators, evaluating their relative importance and establishing an evaluation matrix A.
[0091]
[0092] Then, using the Analytic Hierarchy Process (AHP), the weight coefficients W1 = (W1, W2, W3, W4) of the first-level indicators are calculated. 11 W 12 W 13 Simultaneously, the method can also solve for the largest eigenvalue λ. max A consistency check is performed, requiring CR < 0.1.
[0093] For the second layer, the method sets the EEG arousal level C. 11 and facial expression arousal C 12 The weighting coefficient is W 21 (j)(j=1:2), EEG interest level C 21 and facial expression interest C 22 The weighting coefficient is W 22(j)(j=1:2), EEG satisfaction C 31 and facial expression satisfaction C 32 The weighting coefficient is W 23 (j)(j=1:2), where W 21 (1)+W 21 (2) = 1, W 22 (1)+W 22 (2) = 1, W 23 (1)+W 23 (2) = 1.
[0094] Among them W 21 (1) and W 21 (2) The proportional relationship is determined based on the recognition accuracy of EEG arousal and facial expression arousal, W 22 (1) and W 22 (2) The proportional relationship is determined based on the recognition accuracy of EEG interest and facial expression interest. 23 (1) and W 23 (2) The proportional relationship is determined based on the recognition accuracy of EEG arousal and facial expression arousal. Assuming the actual recognition accuracy of EEG arousal and facial expression arousal are a and b respectively, then the EEG arousal C... 11 and facial expression arousal C 12 The weighting coefficient is
[0095] In addition to the two levels of indicators mentioned above, the method also defines different levels of weights, which can be understood as the level weights of the third-level indicators. The evaluation indicators of each dimension in the third level are assigned corresponding score weight coefficients W3 (1:5), from low to high: 20, 40, 60, 80, and 100.
[0096] Preferably, the number of recognitions for each level in each dimension of the three-dimensional emotion model is input into the first-layer model and the second-layer model respectively, thereby obtaining the learning status score of the trainees.
[0097] The method determines the frequency of occurrence of each indicator, and then calculates the second-level evaluation indicator C using the second-level evaluation indicators. ij The values of (i = 1:3, j - 1:2) are:
[0098]
[0099] Among them, D ijk (i=1∶3,j-1∶2,k=1∶5) represents the evaluation results of the 3-level indicators.
[0100] Subsequently, the values of each indicator in the first-level evaluation index are calculated, such as brainwave arousal level, interest level, and satisfaction level, which are assigned the value B. iThe formula for calculating (i = 1:3) is:
[0101] B i =C i1 ×W 2i (1)+C i2 ×W 2i (2)
[0102] The formula for calculating the evaluation result F of the overall objective in this invention is as follows:
[0103] F = W 11 ×B1+W 12 ×B2+W 13 ×B2
[0104] Finally, the method sets the evaluation set for the learning status of trainees in the training resources as V = (v1, v2, v3, v4, v5), with scores divided into 1-5 levels: Excellent (90 points), Very Good (80 points), Good (70 points), Average (60 points), and Poor (50 points). Based on the scores of the above evaluation levels, the method determines the final evaluation level of the learning status.
[0105] Example 1
[0106] In one embodiment of the present invention, an evaluation is performed every minute within a 20-minute monitoring period, and the number of times each level is obtained under each dimension is as follows:
[0107] EEG: Number of times of extreme excitement D 111 =3, number of excitations D 112 =6, Neutral degree D 113 =8. Number of times of lethargy D 114 =2 and the number of times of extreme lethargy D 115 =1;
[0108] Facial expression: Number of times excited (D) 121 =3, Number of active users D 122 =7. Number of calm times D 123 =9. Number of times fatigued (D) 124 =1 and the number of times of drowsiness D 125 =0;
[0109] EEG: Number of times: Very interested D 211 =2, Number of times of interest D 212 =2, neutral degree D 213 =10, Number of times no interest was found (D) 214 =2 and the number of times D was not very interested 215 =4;
[0110] Facial expressions: Number of times focused (D) 221 =3, Number of times followed (D) 222 =2, Number of times of normalcy D 223=9. Number of times perfunctory responses (D) 224 =2 and the number of times of aversion D 225 =3;
[0111] EEG: Number of times very satisfied D 311 =4, Number of times satisfied D 312 =7, General Degree D 313 =6. Number of disappointments (D) 314 =2 and the number of times of great disappointment D 315 =1
[0112] Facial expression: Number of comprehension attempts (D) 321 =3, Number of doubts D 322 =8. Number of times listening (D) 323 =6. Number of times resistance D 324 =3 and the number of times D was disdainful 325 =0.
[0113] Therefore, the values of the second-level evaluation indicators can be calculated as follows:
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120] Furthermore, the method calculates the weights for EEG and facial expression as follows:
[0121] W 21 (1) = 0.55, W 21 (2) = 0.45
[0122] W 22 (1) = 0.5, W 22 (2) = 0.5
[0123] W 23 (1) = 0.45, W 23 (2) = 0.55
[0124] Therefore, according to the first-level evaluation index, EEG arousal level B i The formula for calculating (i = 1:3): B i =C i1 ×W 2i (1)+C i2 ×W2i (2) We get:
[0125] B1 = C 11 ×W 21 (1)+C 12 ×W 21 (2) = 68 * 0.55 + 72 * 0.45 = 69.8
[0126] B2 = C 21 ×W 22 (1)+C 22 ×W 22 (2) = 56 * 0.5 + 57 * 0.5 = 56.5
[0127] B3 = C 31 ×W 23 (1)+C 32 ×W 23 (2) = 71 * 0.45 + 71 * 0.55 = 71
[0128] Based on this, the method calculates the first-level index weight coefficient as W1 = (0.163, 0.297, 0.54), and at this time, the largest eigenvalue λ... max =3.01, CR=0.0088<0.1.
[0129] Based on the evaluation results of the overall goal, we have F = W 11 ×B1+W 12 ×B2+W 13 ×B2, therefore we get:
[0130] F=0.163×69.8+0.297×56.5+0.54×71=66.5
[0131] Based on the five levels of the learner's learning status evaluation set V, this learner's learning status assessment result for this 20-minute period is "average".
[0132] A second aspect of this invention relates to a learning status evaluation device for a multimodal network training system. The device includes a data acquisition module, a judgment module, and a scoring module. The data acquisition module is used to acquire EEG signals and facial images of trainees within a preset time period, and preprocess the EEG signals and facial images to extract multiple EEG signals and multiple facial images to be identified. The judgment module is used to define the dimensions and levels of a three-dimensional emotion model, and uses statistical indicators to calculate the dimensions and levels of each EEG signal and facial image to be identified in the three-dimensional emotion model. The scoring module uses the number of recognitions of each level in each dimension of the three-dimensional emotion model as input, and constructs an analytic hierarchy process (AHP) using the three-dimensional emotion model to score the learning status of trainees.
[0133] Example 2
[0134] In this embodiment, the acquisition module may specifically include an EEG acquisition module, an EEG signal preprocessing module, a facial image acquisition module, and an image preprocessing module. The EEG acquisition module acquires EEG signals from multiple leads of the trainee and sends the digitized EEG signals to the processor. The EEG signal preprocessing module uses low-pass filtering to remove high-frequency and power frequency interference and utilizes independent component analysis (ICA) to separate ocular artifacts from the EEG signals, thus enabling the identification of ocular artifacts. The facial image acquisition module acquires facial depth image data, 3D point cloud data, and RGB 2D images of the trainee. The image preprocessing module extracts the 3D spatial coordinates of facial expression feature points.
[0135] The judgment module specifically includes modules for EEG arousal dimension feature extraction and expression recognition, EEG interest dimension feature extraction and expression recognition, EEG satisfaction dimension feature extraction and expression recognition, facial expression arousal dimension recognition, facial expression interest dimension recognition, and facial expression satisfaction dimension recognition. Specifically, the EEG arousal dimension feature extraction and expression recognition module first extracts EEG features related to arousal dimension expressions, and then uses the extracted EEG signal features to identify the type of arousal dimension expression. The EEG interest dimension feature extraction and expression recognition module first extracts EEG features related to interest dimension expressions, and then uses the extracted EEG signal features to identify the type of interest dimension expression. The EEG satisfaction dimension feature extraction and expression recognition module first extracts EEG features related to satisfaction dimension expressions, and then uses the extracted EEG signal features to identify the type of satisfaction dimension expression. The facial expression arousal dimension recognition module uses the three-dimensional spatial coordinates of facial expression feature points of different types of arousal dimension to establish an arousal recognition dataset, thus realizing the recognition of facial expression arousal dimension. The facial expression interest dimension recognition module uses the three-dimensional spatial coordinates of facial expression feature points of different types of interest dimension to establish an interest recognition dataset, thus realizing the recognition of facial expression interest dimension. The facial expression satisfaction dimension recognition module uses the three-dimensional spatial coordinates of facial expression feature points of different types of satisfaction dimension to establish a satisfaction recognition dataset, thereby realizing the recognition of facial expression satisfaction dimension.
[0136] The scoring module is specifically a training participant learning status evaluation module based on multi-level fuzzy comprehensive evaluation. Based on a three-dimensional emotion model, a multi-level fuzzy comprehensive evaluation system and method are established, and EEG and facial expression recognition technologies are used to evaluate the learning status of trainees in the power grid dispatching network training system.
[0137] It is understood that the learning status evaluation device of the multimodal network training system includes hardware structures and / or software modules corresponding to the execution of each function in order to achieve the various functions provided in the embodiments of this application. Those skilled in the art should readily recognize that, in conjunction with the algorithm steps of the examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0138] This application embodiment can divide the learning status evaluation of trainees in a multimodal network training system into functional modules based on the above method example. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0139] The multimodal network training system's student learning status evaluation device includes at least one processor, a bus system, and at least one communication interface. The processor comprises a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other hardware. The memory comprises read-only memory (ROM), random access memory (RAM), etc. The memory can be independent and connected to the processor via a bus. Alternatively, the memory can be integrated with the processor. The hard disk can be a mechanical hard disk (HDD) or a solid-state drive (SSD), etc. This embodiment of the invention does not limit the specific implementation. The above embodiments are typically implemented using software and hardware. When implemented using software programs, it can be implemented in the form of a computer program product. This computer program product includes one or more computer instructions.
[0140] A third aspect of the present invention relates to a terminal, including a processor and a storage medium; the storage medium is used to store instructions; the processor is used to operate according to the instructions to perform the steps of the method according to the first aspect of the present invention.
[0141] A fourth aspect of the present invention relates to a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method in the first aspect of the present invention.
[0142] Example 3
[0143] When computer program instructions are loaded and executed on a computer, the corresponding functions are implemented according to the process provided in the embodiments of this invention. The computer program instructions involved may be assembly instructions, machine instructions, or code written in a programming language, etc.
[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A method for evaluating the learning status of trainees in a multimodal online training system, characterized in that, The method includes the following steps: Step 1: Collect EEG signals and facial images of trainees within a preset time period, and preprocess the EEG signals and facial images respectively to extract multiple EEG signals and multiple facial images to be identified. Step 2: Define the dimensions and levels of the three-dimensional emotion model. Statistical indicators are used to calculate the dimensions and levels of each of the EEG signals and facial images to be identified in the three-dimensional emotion model. Step 3: Using the number of recognitions for each level in each dimension of the three-dimensional emotion model as input, construct the hierarchical analysis method using the three-dimensional emotion model to score the learning status of the trainees. The analytic hierarchy process (AHP) consists of a two-layer model. The first layer is an evaluation matrix for arousal, interest, and satisfaction. The second layer is an evaluation matrix for EEG and facial expression in each dimension. The number of recognitions for each level in each dimension of the three-dimensional emotion model is input into the first and second layer models respectively to obtain the learning status score of the trainees.
2. The method for evaluating the learning status of trainees in a multimodal online training system according to claim 1, characterized in that: The collection of EEG signals from trainees within a preset time period includes: Multiple EEG acquisition electrodes are placed at preset positions on the trainee's head. The EEG acquisition electrodes are divided into a first electrode, a second electrode, and a third electrode according to the preset positions on the head.
3. The method for evaluating the learning status of trainees in a multimodal online training system according to claim 2, characterized in that: The preprocessing of the EEG signal includes: Low-pass filtering was used to filter out high-frequency interference above 70Hz and power frequency interference from the EEG lead signals acquired by the first, second and third electrodes to obtain noise-filtered signals. Independent component analysis was used to decompose the noise-filtered signal into independent components. After removing the electrooculogram artifacts, db4 wavelet decomposition and reconstruction were performed to filter out the DC component and low-frequency drift less than 0.1 Hz, thereby obtaining the real EEG signal. The real EEG signal is truncated to a fixed length in the time dimension to obtain the multiple EEG signals to be identified.
4. The method for evaluating the learning status of trainees in a multimodal online training system according to claim 1, characterized in that: Collect facial images of trainees within a preset time period, including: A depth camera was used to acquire facial depth images of the trainees and generate a 3D point cloud. An RGB area array camera was used to acquire RGB 2D images.
5. The method for evaluating the learning status of trainees in a multimodal online training system according to claim 4, characterized in that: The preprocessing of the facial images includes: Define the facial feature points of the trainees, namely, feature points of group T1, feature points of group T2 and feature points of group T3; Depth data of the T1 group of feature points are extracted from the facial depth image, and two-dimensional coordinates of the T1 group of feature points are extracted from the RGB two-dimensional image; The depth data and the two-dimensional coordinates are converted into the three-dimensional coordinates of the T1 set of feature points; The coordinates of the T2 group of feature points are directly extracted from the three-dimensional point cloud, and the three-dimensional coordinates of the T3 group of feature points are calculated using the three-dimensional coordinates of the T2 group of feature points and the three-dimensional coordinates of the T1 group of feature points. Calculate the Euclidean distance between preset feature points to obtain multiple feature vectors, perform principal component analysis on the multiple feature vectors, and extract principal component features.
6. A method for evaluating the learning status of trainees in a multimodal online training system according to claim 3 or 5, characterized in that: The dimensions and levels of the defined three-dimensional emotion model include: The three-dimensional emotion model includes arousal, interest, and satisfaction, and each dimension has its own dimensional indicators based on EEG signals and facial images. Each of the aforementioned dimensional indicators includes the same number of levels, and each level is assigned a corresponding level weight.
7. The method for evaluating the learning status of trainees in a multimodal online training system according to claim 6, characterized in that: The step involves using statistical indicators to calculate the values of multiple EEG signals to be identified, thereby determining the dimension and level of each EEG signal in the three-dimensional emotion model. This includes: The mean frequency area, mean frequency corresponding to peak power, and mean median frequency of each real EEG signal collected by the first electrode are statistically analyzed in multiple preset frequency ranges to form a spectral feature vector for each real EEG signal. Calculate the autoregressive coefficient of each of the real EEG signals acquired by the first electrode; After standardizing the spectral feature vector and autoregressive coefficients, they are used as input data for EEG arousal. The nearest neighbor regression algorithm is then used to identify the arousal level of each real EEG signal.
8. The method for evaluating the learning status of trainees in a multimodal online training system according to claim 6, characterized in that: The step involves using statistical indicators to calculate the values of multiple EEG signals to be identified, thereby determining the dimension and level of each EEG signal in the three-dimensional emotion model. This includes: Each real EEG signal acquired by the second electrode was subjected to a db4 wavelet transform again, and the extracted signals were then analyzed. Rhythm signals and For rhythmic signals, calculate the average and variance of the power spectral density for each rhythmic signal; The average value and variance of the power spectral density of each rhythm signal are used as features of each real EEG signal, and the nearest neighbor regression algorithm is used to identify the level of interest of each real EEG signal.
9. The method for evaluating the learning status of trainees in a multimodal online training system according to claim 6, characterized in that: The step involves using statistical indicators to calculate the values of multiple EEG signals to be identified, thereby determining the dimension and level of each EEG signal in the three-dimensional emotion model. This includes: The area and peak power of the power spectrum of each real EEG signal acquired by the third electrode within a set frequency range; Using the area and peak power as input data for EEG satisfaction, the nearest neighbor regression algorithm is used to identify the satisfaction level of each real EEG signal.
10. The method for evaluating the learning status of trainees in a multimodal online training system according to claim 6, characterized in that: Statistical indicators are used to calculate the dimension and level of each facial image to be identified in the three-dimensional emotion model, including: The principal component features are classified using a support vector machine model.
11. The method for evaluating the learning status of trainees in a multimodal online training system according to claim 10, characterized in that: The construction process of the first layer model is as follows: A pairwise importance comparison is performed on arousal level, interest level, and satisfaction level, and an evaluation value is obtained based on the comparison results. The evaluation value is used as the value of each item in the evaluation matrix.
12. The method for evaluating the learning status of trainees in a multimodal online training system according to claim 11, characterized in that: The construction process of the second-layer model is as follows: The electroencephalogram (EEG) signals and facial images are acquired, and expert scores based on arousal are obtained. The average value of the expert scores is calculated to generate an arousal dataset. The arousal level of the EEG signal and the facial image is obtained through the EEG and facial expression evaluation matrix under the arousal level. The accuracy of EEG signal recognition (a) and facial image recognition (b) are calculated based on the matching degree between the arousal dataset and the arousal level, respectively, thereby obtaining the weight coefficients of EEG signals and facial expressions in the second-layer model. , .
13. A learning status evaluation device for a multimodal network training system utilizing the method described in any one of claims 1-8, characterized in that: The device includes a data acquisition module, a judgment module, and a scoring module; wherein... The acquisition module is used to acquire the EEG signals and facial images of trainees within a preset time period, and to preprocess the EEG signals and facial images respectively to extract multiple EEG signals and multiple facial images to be identified. The determination module is used to define the dimensions and levels of the three-dimensional emotion model, and to use statistical indicators to calculate the dimensions and levels of each of the EEG signals and facial images to be identified in the three-dimensional emotion model. The scoring module is used to score the learning status of the trainees by using the number of recognitions of each level in each dimension of the three-dimensional emotion model as input and constructing a hierarchical analysis method using the three-dimensional emotion model.
14. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the method according to any one of claims 1-12.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-12.
Citation Information
Patent Citations
Method and Apparatus for Recognizing Emotion Based on Image Converted from Brain Signal
KR1020190035368A
A non-invasive multimodal screening and assessment system for human health monitoring and a method thereof
WO2023012818A1