A method for objectively measuring cognitive load based on multi-channel face video
By using multi-channel facial video acquisition and signal processing technology, combined with mental arithmetic tasks and the XGBoost classifier, the problems of subjectivity in measuring cognitive load and inconvenience of contact devices were solved, and accurate cognitive load assessment was achieved in motion.
Patent Information
- Application Number
- CN202310590363.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-23
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-05-23
AI Technical Summary
Existing methods for measuring cognitive load in the brain suffer from high subjectivity and inconvenience due to contact devices. Single-channel facial video measurement requires subjects to remain still, leading to psychological stress and measurement inconvenience.
We employed multi-channel facial video acquisition technology, combined with mental arithmetic tasks to induce different levels of cognitive load, extracted physiological indicators and task accuracy from the multi-channel facial videos, used an improved signal processing method to denoise and synchronize the pulse wave signals, and trained an XGBoost classifier to quantify and assess cognitive load.
It achieves non-contact, comfortable, and accurate cognitive load measurement while the subject is in motion, and provides an objective and real-time cognitive load assessment by combining physiological indicators and task performance.
Smart Images

Figure CN116849604B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology and smart health technology, and particularly relates to a cognitive load objective measurement method based on multi-channel face video. BACKGROUND
[0002] The cognitive load theory considers that people consume certain cognitive resources when performing tasks. There are three commonly used methods for measuring cognitive load: subjective measurement method, task performance measurement method and physiological index measurement method. The subjective measurement method requires the subject to accurately express the feelings of each cognitive load evaluation dimension in the cognitive process. This method is highly subjective. The task performance measurement method judges the cognitive load brought to the subject by the subject's performance in completing the specified task. However, the task performance is highly related to the nature of the task, so it is not universal. The physiological index measurement method indirectly evaluates the cognitive load by measuring the physiological indexes of the subject during the task completion process. Since the physiological response is difficult to control by subjective consciousness, the physiological index measurement method has good objectivity, and the physiological index also changes in real time with the different levels of cognitive load of the human body, which can better reflect the detailed trend of cognitive load.
[0003] Studies have shown that blood oxygen saturation decreases faster under high cognitive load, and pulse variability is related to the activity level of the sympathetic and parasympathetic nervous systems of the human body, and can reflect the mental stress level of the human body, so these two physiological indexes can be used to measure brain cognitive load. In addition to these, commonly used physiological indexes also include brain waves, skin conductance, eye movements, etc. However, the measurement of these indexes often requires the subject to wear heavy equipment, which not only has a high testing cost, but also brings additional psychological pressure to the subject, affecting the measurement of brain cognitive load. In contrast, using a non-contact camera to collect the face video of the subject can extract blood oxygen saturation and pulse variability, and then measure cognitive load, which has the characteristics of low cost and little psychological pressure to the subject. However, the single-channel measurement method requires the subject to remain still, which still brings certain psychological pressure to the subject. The single channel here refers to using a single camera to capture face video. SUMMARY
[0004] The purpose of the present application is to solve the problems of subjectivity of the general brain cognitive load self-assessment scale and inconvenience of contact device measurement in the prior art, and to provide a cognitive load objective measurement method based on multi-channel face video. The scene of performing mental arithmetic tasks is designed to induce different levels of cognitive load, and different levels of cognitive load are induced by the subject performing mental arithmetic tasks of different difficulties. In this scenario, the physiological indexes extracted from the multi-channel face video and the correct rate of the mental arithmetic task recorded are used to objectively measure the cognitive load of the subject.
[0005] The object of the present application can be achieved by adopting the following technical solutions
[0006] A cognitive load objective measurement method based on multi-channel face video, the cognitive load objective measurement method comprises the following steps:
[0007] S1, the subject faces a plurality of independent cameras to carry out cognitive load experiment, collects the multi-channel face video of the subject in the process of performing mental arithmetic task under different cognitive load levels, and records the mental arithmetic task accuracy; wherein, the cognitive load is the total amount of cognitive resources called by the brain of the subject, and different cognitive load levels are induced by the subject performing different difficulty mental arithmetic tasks; each camera obtains a channel of face video, and a plurality of cameras obtain multi-channel face video;
[0008] Collecting multi-channel face video is used for subsequent physiological index extraction, and the cognitive load level is measured from the perspective of human physiological state; the mental arithmetic task accuracy measures the cognitive load level by reflecting the performance of the subject in performing the task;
[0009] S2, the collected multi-channel face video is subjected to face recognition respectively, and the region of interest is located, the gray mean value of the region of interest and the area of the region of interest are obtained, and the original pulse wave signal and the area sequence of the region of interest are extracted; wherein, the region of interest is a skin area for extracting the original pulse wave signal, located at the forehead of the face; the area sequence of the region of interest is a sequence composed of the area size of the region of interest in each frame of picture in the face video;
[0010] S3, the original pulse wave signal is subjected to improved adaptive noise ensemble empirical mode decomposition, and combined with wavelet threshold denoising and Butterworth band-pass filter to remove noise, to obtain the denoised pulse wave signal;
[0011] S4, the pulse wave of the G channel in the denoised pulse wave signal is extracted, the delay between multiple channels is calculated by using cross-correlation, and the time synchronization of the multi-channel pulse wave signal and the area sequence of the region of interest is realized according to the delay; the G channel pulse wave is the green channel pulse wave obtained by color channel separation of the pulse wave signal;
[0012] S5, for the same moment, the pulse wave data of the channel with the largest area of the corresponding region of interest in the multiple pulse waves is selected as the data at this moment, all moments are traversed, and the fused pulse wave is obtained;
[0013] S6, extract the blood oxygen saturation and pulse variability of the fused pulse wave signal, the fused pulse wave signal contains R, G and B three color channels, R represents red, G represents green and B represents blue, the blood oxygen saturation is obtained by fitting the pulse wave of B and R channels, and the pulse variability is obtained by extracting the time interval between the wave peaks of the pulse wave of the G channel;
[0014] Since the pulse variability can reflect the brain cognitive load from the perspective of the nervous excitement of the human body, and the blood oxygen saturation reflects the oxygen consumption of the brain under different cognitive loads, the level of cognitive load can be indirectly reflected, therefore, the two physiological indicators are selected;
[0015] S7, repeat steps S1-S6 for the mental calculation tasks of the multiple subjects under low, medium and high cognitive load levels, obtain the mental calculation task accuracy, blood oxygen saturation and pulse variability indicators of the multiple subjects under different cognitive load levels, form a data set required for cognitive load level classification, and then train the XGBoost classifier based on the data set;
[0016] S8, repeat steps S1-S6 for the mental calculation tasks of the to-be-tested object under low, medium and high cognitive load levels, obtain the mental calculation task accuracy, blood oxygen saturation and pulse variability indicators of the to-be-tested object under low, medium and high cognitive load levels; input the mental calculation task accuracy, blood oxygen saturation and pulse variability indicators of the to-be-tested object under low, medium and high cognitive load levels into the trained XGBoost classifier, and estimate the posterior probability of classification;
[0017] S9, use the posterior probability of classification to quantitatively evaluate the cognitive load level of the to-be-tested object.
[0018] Further, the process of using different difficulty mental calculation tasks to induce different cognitive load levels of the subject and recording the multi-channel facial video and the mental calculation task accuracy under the corresponding cognitive load level in step S1 is as follows:
[0019] S101, the subject sits in front of a display screen controlled by an external notebook computer, which displays the mental calculation task to be calculated, N independent cameras are uniformly placed at an angle of θ degrees from the subject at a distance of 50 centimeters, and the angle range is (N-1)θ degrees on a circular arc, and the multi-channel facial video of the subject is collected, and the head of the subject can freely rotate within the range of (N-1)θ degrees; wherein N is the number of independent cameras, that is, the number of channels; the display screen of this step is used to display the mental calculation task to be calculated by the subject, and the cognitive resources are called through the mental calculation task to generate cognitive load; the multi-channel facial video collected by the camera in this step is used for extracting physiological indicators and measuring cognitive load level subsequently;
[0020] S102, display a mental arithmetic task to be calculated on the display screen and start the countdown, the subject faces the display screen, calculates the mental arithmetic task, and the N cameras record the facial video of the subject at the same time; this step allows the subject to freely rotate within the range of (N-1) degrees, so that the measurement experience of the subject is more comfortable, at the same time, the psychological pressure of the subject is reduced, and the measurement of cognitive load is more accurate;
[0021] S103, after the countdown ends, the mental arithmetic task displayed on the screen disappears, and the subject speaks out the calculation result within the limited time, and the calculation result is recorded as a calculation error if it is not given within the limited time, and after the limited time ends, the display screen automatically cuts to the next mental arithmetic task, a total of five mental arithmetic tasks with the same difficulty are executed, and the mental arithmetic task accuracy is recorded; the mental arithmetic task accuracy recorded in this step reflects the level of cognitive load from the perspective of task performance, and avoids that the cognitive load is not comprehensive enough by only using physiological indexes.
[0022] Further, the process of extracting the original pulse wave and the area sequence of the region of interest according to the photographed multi-channel facial video in the step S2 is as follows:
[0023] S201, frame the collected facial video of channel 1 to obtain a continuous color picture sequence, and the format of the picture is RGB; the facial video of channel 1 refers to the facial video collected by the first camera; this step frames the facial video into a picture sequence, which is convenient for extracting the original pulse wave from the picture sequence;
[0024] S202, for each frame of the color picture sequence, a regression tree method based on gradient boosting learning is used for face key point detection, the face key point is detected, and the forehead part of the face is located as the region of interest; this step selects the forehead part of the face as the region of interest because the signal-to-noise ratio of the pulse wave signal extracted from here is high, and the physiological index calculated by using the pulse wave is accurate;
[0025] S203, color channel separation is performed on the image of the located region of interest of each frame of picture to obtain images of R, G and B color channels, and the area of the region of interest is calculated; for each color channel image, the gray mean value of all the pixel points of the region of interest is used as the amplitude of the pulse wave signal; the image sequence obtained by frame division of a video is processed to obtain the original pulse wave signals of R, G and B channels and a region of interest area sequence in time sequence, wherein the original pulse wave signals of R, G and B channels and the region of interest area sequence are extracted from a channel of the face video, and they are time-synchronized; this step performs color channel separation on the image, and for the pulse variability that can be extracted by using only a single color channel signal, the G channel with the highest signal-to-noise ratio is selected as the channel for subsequent index extraction. For the blood oxygen saturation that can be extracted by using two different color channel signals, based on the theory that the absorption of hemoglobin to red light and blue light is different, the R and B channels are selected for the extraction of blood oxygen saturation;
[0026] S204, steps S201-S203 are performed on the collected face videos of channels 2 to N respectively to obtain the original pulse wave signals of R, G and B channels and a region of interest area sequence for each channel of the face video; finally, 3N original pulse wave signals including N original pulse wave signals of R channels, N original pulse wave signals of G channels and N original pulse wave signals of B channels and N region of interest area sequences are obtained. This step completes the extraction of the original pulse wave signals and the region of interest area sequences of all channels, facilitating subsequent signal synchronization and signal fusion.
[0027] Further, the process of denoising the obtained original pulse wave signal to obtain a pure pulse wave signal in step S3 is as follows:
[0028] S301, the original pulse wave signal restored in step S2 is processed by using improved adaptive noise set empirical mode decomposition to obtain a plurality of components. The non-contact video signal has a low signal-to-noise ratio, and the improved adaptive noise set empirical mode decomposition can solve the mode aliasing problem of traditional empirical mode decomposition;
[0029] S302, the first three components are denoised by using wavelet threshold denoising. The first three components contain more high-frequency noise, and the wavelet threshold denoising can remove the high-frequency noise therein;
[0030] S303, the first three denoised components and the remaining components are superimposed to obtain a reconstructed signal. The reconstructed signal removes the high-frequency noise, and thus has a higher signal-to-noise ratio than the original signal, but is still insufficient for extracting physiological indexes and needs to be further processed;
[0031] S304, band-pass filter the reconstructed signal using a Butterworth band-pass filter. The Butterworth band-pass filter is used to filter out noise outside the frequency band of the pulse wave signal, further improving the signal quality.
[0032] Further, the process of time synchronization between the pulse wave signals of multiple channels and the area sequence of the region of interest in step S4 is as follows:
[0033] S401, select N G channel signals from the 3N pulse wave signals obtained in step S3, and use the cross-correlation method to calculate the delay between the N channels. Since the R, G, and B channel pulse waves extracted from the face video of one channel and the area sequence of the region of interest are time-synchronized, the delay between the N G channels is also the delay between the N channels.
[0034] This step calculates the delay between the channels because the N independent cameras cannot start and end recording at the same time, so the signals collected by different cameras will have a certain delay and need to be synchronized.
[0035] This step uses cross-correlation to calculate the delay between the channels because even if the pulse wave signals measured from the same person's skin area by different devices are similar, according to the cross-correlation principle, the time delay at the maximum value of the cross-correlation function of the two signals to be synchronized is the delay between the signals.
[0036] This step selects G channel pulse wave signals to calculate the delay between the channels because the pulse wave signals extracted from the G channel have large amplitude and small noise.
[0037] S402, time synchronize the 3N filtered pulse wave signals extracted from the face video of N channels and the N area sequences of the region of interest according to the delay between the N channels. Among them, the 3 filtered pulse wave signals and 1 area sequence of the region of interest extracted from the face video of the same channel are time-synchronized.
[0038] The specific method of time synchronization in this step is as follows: if the G channel pulse wave signal of channel 1 is 10 data points ahead of the G channel pulse wave signal of channel 2, that is, the signal of channel 1 is 10 data points ahead of the signal of channel 2, then remove the first 10 data of the G channel, B channel, and R channel pulse wave signals of channel 1 and the first 10 data of the area sequence of the region of interest, and remove the last 10 data of the G channel, B channel, and R channel pulse wave signals of channel 2 and the last 10 data of the area sequence of the region of interest, to align the data length and synchronize channel 1 and channel 2. Similarly, the synchronization of multiple channels can be achieved.
[0039] Further, the step S5 in accordance with the size of the region of interest area to the time synchronization of the multi-channel pulse wave signal fusion process as follows:
[0040] S501, the time synchronization of the N channel pulse wave obtained in step S4, for the same time, select the largest area of the region of interest in the sequence of the region of interest area of the pulse wave data of the channel as the data at this time, all the time, eventually get a fusion of R channel pulse wave;
[0041] The principle of this step signal fusion is: the signal-to-noise ratio of the pulse wave extracted from the region of interest is proportional to the area of the region of interest, for the same region of interest, in the case of head rotation, the size of the region of interest will change with the angle between the subject's face and the camera, at any time, there will be a camera with the best shooting angle in multiple cameras, so that the area of the region of interest is the largest, select the face picture of the channel with the largest area of the region of interest at each time to extract the pulse wave, which can get the pulse wave with the highest signal-to-noise ratio under the motion condition.
[0042] S502, the time synchronization of the N channel pulse wave obtained in step S4, execute step S501, get a fusion of G channel pulse wave;
[0043] S503, the time synchronization of the N channel pulse wave obtained in step S4, execute step S501, get a fusion of B channel pulse wave.
[0044] Further, the step S6 in accordance with the fusion of the pulse wave signal to extract the two types of physiological indicators of pulse variability and blood oxygen saturation process as follows:
[0045] S601, the fusion of the G channel pulse wave obtained in S5, peak detection, extract all the time of the wave peak, for the extraction of pulse variability; this step detects the peak point by judging the relative size of the adjacent data points of the pulse wave signal, the peak point is characterized by the amplitude of the data points on the left and right of it is smaller than it, which meets the condition of the detected peak point; the time of the peak point can be used for the extraction of pulse variability signal;
[0046] S602, difference the time of adjacent wave peak, get the pulse interval sequence, use cubic spline interpolation method to interpolate the interval sequence, get the pulse variability signal; this step gets the original signal by difference, which belongs to non-uniform sampling signal, the frequency domain information of pulse variability requires that the signal is uniformly sampled, so the cubic spline interpolation method is used to interpolate the signal to get the uniformly sampled signal;
[0047] S603, extracting a plurality of pulse variability indexes: performing feature extraction on the pulse variability signal obtained in step S602, preferably six pulse variability indexes, wherein, in the time domain, the mean MEAN and the standard deviation SDPP are extracted, the power spectrum of the pulse variability signal is solved, and the frequency domain features are extracted, obtaining the total power TP, the low-frequency power LF, the high-frequency power HF, and the power ratio LF / HF of the low and high frequencies; the indexes selected in this step can reflect the activity of the sympathetic and parasympathetic nerves of the human body, and reflect the high and low of the cognitive load level from the perspective of the human nervous condition;
[0048] S604, calculating blood oxygen saturation: using the fused R channel pulse wave and B channel pulse wave obtained in S5, calculating the independent variables in the blood oxygen saturation empirical formula, using the measured data of the standard finger clip oximeter as the dependent variable, using the least squares method to linearly fit the blood oxygen saturation empirical formula, and fitting the unknown parameters of the blood oxygen saturation empirical formula; using the fitted blood oxygen saturation empirical formula to calculate the current blood oxygen saturation.
[0049] Since the blood oxygen saturation empirical formula contains two unknown parameters, two different signals must be used for fitting. Therefore, the R channel and B channel pulse wave signals extracted in step S5 are used to linearly fit the blood oxygen saturation empirical formula with the measured data of the standard finger clip oximeter, and the unknown parameters of the blood oxygen saturation empirical formula are solved. The fitted blood oxygen saturation empirical formula and the time domain parameters of the R channel and B channel pulse wave signals, including the mean and standard deviation, are used to jointly calculate the current blood oxygen saturation. Blood oxygen saturation can well reflect the use of oxygen resources by the human body under different cognitive loads, and thus can be used to measure the size of cognitive load.
[0050] It should be noted that: after the unknown parameters of the blood oxygen saturation empirical formula are fitted, they do not need to be fitted again in the subsequent test and use, and the blood oxygen saturation empirical formula can be directly used to calculate the current blood oxygen saturation.
[0051] Further, the process of constructing a cognitive load level classification data set and training a classifier according to the data set in step S7 is as follows:
[0052] S701, collecting data: repeating steps S1-S6 for a plurality of subjects under low, medium and high cognitive load levels in mental calculation tasks, obtaining the mental calculation task accuracy, blood oxygen saturation, and pulse variability indexes of the plurality of subjects under different cognitive load levels. This step collects indexes of multiple subjects under different cognitive load levels to obtain sufficient sample data for subsequent training of the classifier;
[0053] S702, constructing a classification dataset: the collected dataset is labeled, wherein the data under low cognitive load level is labeled as 0, the data under medium cognitive load level is labeled as 1, and the data under high cognitive load level is labeled as 2, to obtain the dataset required for training the classifier. This step considers the subsequent step of classifying the cognitive load level, and labels the cognitive load level, thereby obtaining labeled data representing different cognitive load levels;
[0054] S703, training the classifier: using the dataset obtained in S702, an XGBoost classifier is trained. XGBoost (eXtreme Gradient Boosting) is an algorithm based on gradient boosting trees, from the paper "XGBoost: A Scalable Tree Boosting System", Chen T, 2016, which is widely used in the field of data science. The function of the XGBoost classifier is to judge the class to which a new sample data belongs on the basis of training data with labeled classes, and can output the probability of the new sample data being classified into each class.
[0055] Further, the step S8 process is as follows:
[0056] S801, obtaining the index of the object to be measured: repeating steps S1-S6, the correct rate of mental arithmetic task, blood oxygen saturation, and pulse variability index of the object to be measured under low, medium, and high cognitive load levels are obtained.
[0057] S802, obtaining the posterior probability of classification: inputting the mental arithmetic task accuracy, blood oxygen saturation, and pulse variability obtained in S801 into the trained XGBoost classifier, outputting the posterior probability estimate of classification, obtaining the probability P l that the cognitive load level is classified as low load level, the probability P m that the cognitive load level is classified as medium load level, and the probability P h that the cognitive load level is classified as high load level. Among them, the probability P m P h is used to measure the relative strength of low cognitive load level, the probability P l P h is used to measure the relative strength of medium cognitive load level, and the probability P l P m is used to measure the relative strength of high cognitive load level.
[0058] Further, the quantification and evaluation process of the cognitive load level of the to-be-tested object in step S9 is as follows:
[0059] S901, quantifying three cognitive load levels: using the posterior probability of the classification obtained in S8, respectively, quantifying and scoring the cognitive load when the to-be-tested object is in a low cognitive load level, a medium cognitive load level, and a high cognitive load level, wherein the quantification score of the low cognitive load level is represented by S l , the quantification score of the medium cognitive load level is represented by S m , and the quantification score of the high cognitive load level is represented by S h .
[0060] S902, calculating the total cognitive load score: averaging the scores obtained in step S902 to obtain the brain cognitive load score of the current to-be-tested object.
[0061] The present application has the following advantages and effects over the prior art:
[0062] 1) The present application uses physiological indicators to measure the brain cognitive load level. Since physiological responses are difficult to control by subjective consciousness and will change in real time with changes in cognitive load level, the present application has good objectivity and real-time performance compared to current subjective measurement methods. Using a camera to collect facial video can achieve non-contact measurement, solving the problem of additional psychological stress on the subject caused by the use of contact devices in current objective measurement methods. Using multiple independent cameras to collect facial video can measure cognitive load in the case of subject motion, avoiding the additional psychological stress on the subject caused by the requirement that the subject remain still in single-camera measurement. The method of measuring cognitive load by combining heart rate calculation task accuracy and physiological indicators creates a cognitive load measurement environment that is comfortable to measure and objective and accurate to measure.
[0063] 2) The present application proposes a cognitive load objective measurement method based on classification posterior probability. The XGBoost classifier is trained using the blood oxygen saturation, pulse variability physiological indicators extracted from the multi-channel facial video, and the heart rate calculation task accuracy indicator. The posterior probability output by the classifier is used to quantitatively score the cognitive load level, which can objectively reflect the brain cognitive load level. The blood oxygen saturation, pulse variability series of indicators reflect the changes in cognitive load from the perspective of changes in physiological indicators, and the heart rate calculation task accuracy indicator reflects the changes in cognitive load from the perspective of task performance, avoiding the incompleteness of single physiological indicator measurement.
[0064] 3) The application proposes a method of multi-channel face video extraction, pulse wave signal cross-correlation time synchronization, and splicing and fusion according to the area of the region of interest, to obtain a pulse wave with good signal integrity and high signal-to-noise ratio, realize the extraction of blood oxygen saturation and pulse variability under the motion state, and make the subjects more comfortable and free during the physiological index extraction process. BRIEF DESCRIPTION OF DRAWINGS
[0065] The drawings described herein are used to provide further understanding of the application, form a part of the application, and the illustrative embodiments of the application and the description thereof are used to explain the application, and do not constitute an improper limitation on the application. In the drawings:
[0066] Figure 1 is a flowchart of a cognitive load objective measurement method based on multi-channel face video disclosed in embodiment 1 of the application;
[0067] Figure 2 is a flowchart of extracting original pulse wave and region of interest area sequence from the region of interest (ROI) in embodiment 1 of the application;
[0068] Figure 3 is a flowchart of original pulse wave signal denoising in embodiment 1 of the application;
[0069] Figure 4 is a flowchart of multi-channel time synchronization in embodiment 1 of the application;
[0070] Figure 5 is a flowchart of extracting physiological indexes in embodiment 1 of the application;
[0071] Figure 6 is a principle diagram of quantifying three different cognitive load levels in embodiment 1 of the application;
[0072] Figure 7 is a raw pulse wave waveform and region of interest area sequence diagram in embodiment 2 of the application;
[0073] Figure 8 is a denoised pulse wave waveform diagram in embodiment 2 of the application;
[0074] Figure 9 is a fused pulse wave waveform diagram in embodiment 2 of the application. DETAILED DESCRIPTION
[0075] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0076] Embodiment 1
[0077] The embodiment discloses a cognitive load objective measurement method based on multi-channel face video, as shown in the following steps: Figure 1 The specific steps are as follows:
[0078] S1, the subject faces three independent cameras to perform a cognitive load experiment, collects three-channel face videos of the subject in the process of performing a mental calculation task under different cognitive load levels, and records the mental calculation task accuracy; the face video is used to extract relevant physiological indicators to measure the cognitive load level from the perspective of human physiological state; the mental calculation task accuracy measures the cognitive load level through the performance of the task.
[0079] S101, the subject sits in front of a display screen controlled by an external notebook computer, the display screen displays a mental calculation task to be calculated, three independent cameras are directed to the face of the subject, and are evenly placed at an angle range of 60 degrees on an arc at a distance of 50 centimeters from the subject at an angle of 30 degrees each other, multi-channel face videos of the subject are collected, and the head of the subject can be freely rotated within the range of 60 degrees;
[0080] S102, a mental calculation task to be calculated is displayed on the display screen and the countdown starts, the subject faces the display screen, calculates the mental calculation task, and the three cameras simultaneously record the face video of the subject;
[0081] S103, after the countdown ends, the mental calculation task displayed on the screen disappears, the subject speaks out the calculation result within 5 seconds, and the calculation error is recorded within 5 seconds without giving the calculation result, and after the limited time ends, the display screen automatically cuts to the next mental calculation task, a total of five mental calculation tasks with the same difficulty are performed, and the mental calculation task accuracy is recorded;
[0082] The difficulty of the mental calculation task is divided into three categories, the low difficulty refers to the addition and subtraction operation without carry and borrow of a two-digit number and a one-digit number, the medium difficulty refers to the addition and subtraction operation without carry and borrow of two two-digit numbers, and the high difficulty refers to the addition and subtraction operation with carry and borrow of two two-digit numbers.
[0083] S2, face recognition is performed on the collected face videos of channels 1-3 respectively, the region of interest is located, the gray mean value of the region of interest and the area of the region of interest are obtained, and the original pulse wave signal and the area sequence of the region of interest are extracted; wherein the region of interest is a skin region for extracting the original pulse wave signal, located at the forehead of the face; the area sequence of the region of interest is a sequence composed of the area size of the region of interest in each frame of the face video;
[0084] Figure 2 The flowchart for extracting the original pulse wave signal and the area sequence of the region of interest.
[0085] S201, the collected face video of channel 1 is framed to obtain a continuous color picture sequence, and the format of the picture is RGB, wherein R represents the red channel, G represents the green channel, and B represents the blue channel; the face video of the channel 1 is the face video collected by the first camera;
[0086] S202, for each frame of the color picture sequence obtained in S201, a regression tree method based on gradient improved learning is used for face key point detection, the face key point is detected, and the forehead part is located as the region of interest;
[0087] S203, the image of the region of interest located in each frame of the picture obtained in S202 is color channel separated to obtain the images of R, G and B color channels, and the area of the region of interest is calculated; for each color channel image, the gray mean value of all pixel points in the region of interest is used as the amplitude of the pulse wave signal; the image sequence after framing a video is processed, sorted by time, and the original pulse wave signals of R, G and B channels and a region of interest area sequence are obtained;
[0088] S204, steps S201-S203 are performed on the collected face videos of channels 2 and 3 respectively, and the original pulse wave signals of R, G and B channels and the region of interest area sequence of each channel face video are obtained; finally, three R channel original pulse wave signals, three G channel original pulse wave signals, three B channel original pulse wave signals and three region of interest area sequences are obtained.
[0089] S3, the nine original pulse wave signals of channels 1-3 obtained in S2 are respectively subjected to improved adaptive noise ensemble empirical mode decomposition, combined with wavelet threshold denoising and Butterworth band-pass filter to remove noise, and denoised pulse wave signals are obtained.
[0090] Figure 3The denoising flowchart of the original pulse wave signal.
[0091] S301, using the improved adaptive noise set empirical mode decomposition on the original pulse wave signal obtained in step S2 to obtain a plurality of components;
[0092] S302, using wavelet threshold denoising to denoise the first three components;
[0093] S303, superimposing the denoised first three components and the remaining components to obtain a reconstructed signal;
[0094] S304, using a Butterworth band-pass filter to band-pass filter the reconstructed signal.
[0095] The method of improved adaptive noise set empirical mode decomposition is based on the improvement of the classical empirical mode decomposition algorithm, and the execution steps of the classical empirical mode decomposition algorithm are as follows:
[0096] 1) find all extreme points of the original signal x(t);
[0097] 2) using cubic spline interpolation method to fit all maximum points and minimum points respectively, to obtain the upper envelope line and the lower envelope line of the signal, and to obtain the signal m(t) by averaging the upper and lower envelope lines;
[0098] 3) calculate the difference between the original signal x(t) and m(t), as shown in the following formula:
[0099] h(t) = x(t) - m(t) formula (1)
[0100] Determine whether h(t) meets the convergence condition. The convergence condition generally has three conditions, which are that m(t) approaches 0, the number of local extreme points and zero-crossing points of h(t) is not more than 1, and the set iteration threshold is reached. As long as one of the three conditions is met, it is considered that the convergence condition is met;
[0101] If the convergence condition is not met, repeat the previous operation, otherwise obtain the intrinsic component IMF1 = m(t) and the residual res1 = h(t);
[0102] 4) take the residual res1 as the new original signal x(t), execute the previous steps, until the decomposition cannot continue. After EMD decomposition, the original signal x(t) becomes the sum of N IMF components and the final residual term, as shown in the following formula:
[0103]
[0104] Where i represents the i-th IMF component, N represents the number of IMF components, res N is the final residual term.
[0105] Because the decomposition of EMD to signal will cause the problem of modal aliasing, so here uses the method of improved adaptive noise ensemble empirical mode decomposition to decompose the original pulse wave. Its basic idea is the same as EMD, only in the calculation of each layer IMF component plus adaptive complementary white noise to suppress modal aliasing, the algorithm flow as follows:
[0106] 1) The first order intrinsic mode component ξ1E1(n j (t)) of L group white noise is obtained by EMD decomposition, and the original signal x(t) is added respectively to obtain:
[0107] x j (t) = x(t) + ξ1E1(n j (t)), j = 1, 2, …, L Formula (3)
[0108] Wherein E1(·) represents the first IMF component of EMD decomposition of a signal, ξ1 represents the noise amplitude coefficient added in the first decomposition, the value range is 0.1-0.3, n j (t) represents the jth group of zero mean white noise with variance 1, a total of L groups;
[0109] The local mean of x j (t) is calculated, and then the first order residual component res1 is obtained by ensemble average:
[0110]
[0111] Wherein M(·) represents the local mean of the signal;
[0112] 2) Calculate the first order intrinsic mode component
[0113]
[0114] 3) The second order intrinsic mode component ξ2E2(n j (t)) of L group white noise is obtained by EMD decomposition, and res1 is added respectively, and the second order residual component res2 is obtained by calculating the local mean and ensemble average:
[0115]
[0116] Wherein E2(·) represents the second IMF component of EMD decomposition of a signal, ξ2 represents the noise amplitude coefficient added in the second decomposition, the value range is 0.1-0.3;
[0117] Finally, the second order intrinsic mode component
[0118]
[0119] 4) Calculate the kth intrinsic mode component Using steps (5) to (6) to calculate:
[0120] 5) Calculate the residual component res k :
[0121]
[0122] where E k (k) represents the kth IMF component of a signal EMD decomposition, ξ k represents the noise amplitude coefficient added in the kth decomposition, with a value range of 0.1 to 0.3, K represents the number of IMF components, k = 3, 4, …, K;
[0123] 6) Calculate the intrinsic mode component
[0124]
[0125] 7) Perform the process of the previous steps until the residual term res K cannot be further decomposed, at which time the res K is the residual term. The original signal x(t) becomes the sum of the K IMF components and the residual term, as shown in the following formula:
[0126]
[0127] The purpose of wavelet threshold denoising is to remove high-frequency noise in the signal. In this embodiment, a hard threshold function and db9 wavelet are selected for decomposition, with 3 layers of decomposition.
[0128] S4, on the pulse wave signals of channel 1 to channel 3 obtained after denoising in S3, select 3 G channel pulse waves, use cross-correlation to find the delay between channel 1 to channel 3, and realize time synchronization of the three-channel pulse wave signals and the area sequence of the region of interest according to the delay.
[0129] The specific process of realizing time synchronization of the pulse wave signals and the area sequence of the region of interest among the three channels is as shown in Figure 4 .
[0130] S401, on the 9 pulse wave signals after denoising in step S3, select 3 G channel signals and use the cross-correlation method to find the delay between the channels;
[0131] S402, time synchronization of the nine filtered pulse wave signals and the three area sequences of the regions of interest is achieved according to the three inter-channel time delays obtained in S401;
[0132] S5, for the same time, the pulse wave data of the channel with the largest area of the region of interest in the pulse wave of channel 1 to channel 3 is selected as the data at this time, and all time points are traversed to finally obtain the fused pulse wave.
[0133] S501, for the same time, the pulse wave data of the channel with the largest area of the region of interest in the pulse wave of channel 1 to channel 3 is selected as the data at this time, and all time points are traversed to finally obtain the fused pulse wave.
[0134] S502, for the same time, the pulse wave data of the channel with the largest area of the region of interest in the pulse wave of channel 1 to channel 3 is selected as the data at this time, and all time points are traversed to finally obtain the fused pulse wave.
[0135] S503, for the same time, the pulse wave data of the channel with the largest area of the region of interest in the pulse wave of channel 1 to channel 3 is selected as the data at this time, and all time points are traversed to finally obtain the fused pulse wave.
[0136] S6, the blood oxygen saturation and the pulse variability are extracted from the fused pulse wave signal obtained in S5, wherein the blood oxygen saturation is obtained by fitting the pulse waves of B and R channels, and the pulse variability is obtained by extracting the peak-to-peak time interval of the pulse wave of G channel.
[0137] The flowchart of the above pulse variability and blood oxygen extraction process is shown in Figure 5
[0138] S601, the fused pulse wave signal obtained in S5 is selected to detect the peak value of the signal of G channel, and the continuous wave peaks are extracted, and the amplitude of the peak value is required to be greater than 0, and the interval between adjacent pulse wave peaks is required to be greater than 0.3 seconds.
[0139] S602, the difference between the continuous peak-to-peak values is obtained to obtain the beat-to-beat interval sequence, and the cubic spline interpolation method is used to interpolate the beat-to-beat interval sequence to obtain the pulse variability signal.
[0140] S603, extracting a plurality of pulse variability indexes: performing time domain feature extraction on the pulse variability signal obtained in step S602, preferably six pulse variability indexes, wherein, in the time domain, the mean MEAN and the standard deviation SDPP are obtained, the power spectrum of the pulse variability signal is solved, and the frequency domain features are extracted, and the total power TP, the low frequency power LF, the high frequency power HF, and the power ratio LF / HF of the low frequency and the high frequency are obtained.
[0141] S604, using the pulse wave signals of the R channel and the B channel extracted in step S501, calculating the independent variables in the empirical formula of blood oxygen saturation, using the measured data of the standard finger clip oximeter as the dependent variable, using the least square method to linearly fit the empirical formula of blood oxygen saturation, and fitting the unknown parameters in the empirical formula of blood oxygen saturation. Using the fitted empirical formula of blood oxygen saturation and the pulse wave parameters of the R channel and the B channel to calculate the current blood oxygen saturation;
[0142] wherein the empirical formula of blood oxygen saturation is:
[0143] SPO2=-A1*T+A2 Formula (11)
[0144] SPO2 represents blood oxygen saturation, A1 and A2 are the first and second constants of the empirical formula for calculating blood oxygen saturation, obtained by linear fitting, and T is the independent variable, calculated by combining the R channel pulse wave and the B channel pulse wave;
[0145] The calculation formula of T is:
[0146]
[0147] wherein stdr represents the standard deviation of the R channel, stdb represents the standard deviation of the B channel, meanb represents the mean of the B channel, and meanr represents the mean of the R channel.
[0148] It should be noted that: after the first and second constants A1 and A2 of the empirical formula of blood oxygen saturation are fitted, they do not need to be fitted again in the later test and use, and the current blood oxygen saturation can be directly calculated using the empirical formula (11) and formula (12) of blood oxygen saturation.
[0149] S7, repeating steps S1-S6 for the mental calculation tasks of the 40 subjects under low, medium and high cognitive load levels, obtaining the mental calculation task accuracy, blood oxygen saturation and pulse variability indexes of the 40 subjects under different cognitive load levels, forming a data set required for cognitive load level classification, and then training the XGBoost classifier based on the data set.
[0150] S701, repeat steps S1-S6 to the 40 subjects under low, medium and high cognitive load levels of mental arithmetic tasks, get the 40 subjects' mental arithmetic task accuracy, blood oxygen saturation, pulse variability indicators under different cognitive load levels, a total of 120 groups of data;
[0151] S702, label the collected data set, where the 40 groups of data under low cognitive load level are labeled as 0, the 40 groups of data under medium cognitive load level are labeled as 1, and the 40 groups of data under high cognitive load level are labeled as 2, to obtain the data set required for training the classifier;
[0152] S703, use the data set obtained in S702 to train the XGBoost classifier, and split the data set into a training set and a test set, wherein the training set is used to train the classifier, and the test set is used to verify the performance of the classifier, wherein the proportion of the training set is 0.8 and the proportion of the test set is 0.2.
[0153] S8, execute steps S1-S6 to the test object to obtain the mental arithmetic task accuracy, blood oxygen saturation, and pulse variability indicators; input the mental arithmetic task accuracy, blood oxygen saturation, and pulse variability indicators of the test object into the trained XGBoost classifier, and output the posterior probability estimate of the classification.
[0154] S801, repeat steps S1-S6 to the test object under low, medium and high cognitive load levels of mental arithmetic tasks, to obtain the test object's mental arithmetic task accuracy, blood oxygen saturation, and pulse variability indicators under low, medium and high cognitive load levels;
[0155] S802, input the mental arithmetic task accuracy, blood oxygen saturation, and pulse variability obtained in S801 into the trained XGBoost classifier, and output the posterior probability estimate of the classification, to obtain the probability P l that the cognitive load level is classified as low, the probability P m that the cognitive load level is classified as medium, and the probability P h that the cognitive load level is classified as high.
[0156] S9, use the posterior probability of the classification output in S8 to quantitatively evaluate the cognitive load level of the test object.
[0157] S901, quantify the three cognitive load levels: use the probability P l that the cognitive load level is classified as low, the probability P m that the cognitive load level is classified as medium, and the probability P hThe cognitive load of the test subjects was quantitatively scored when they were at low, medium, and high cognitive load levels. The quantitative score for the low cognitive load level was represented by S. l The quantitative score of cognitive load level is represented by S. m The quantitative score representing a high level of cognitive load is used to represent S. h express.
[0158] The cognitive load in this embodiment is quantified and scored according to the following principles:
[0159] For data categorized as low cognitive load, such as Figure 6 As shown, the case where it will definitely be classified as a low cognitive load level and will not be classified as any other level corresponds to the case where the relative intensity is at the left end of the low cognitive load segment, that is, the relative intensity is 0, and the probability P of being classified as any other case is... m +P h This means that the relative intensity should be increased by a certain percentage, P. m +P h Therefore, the quantitative score for low cognitive load is:
[0160] S l =100×(P) m +P h )=100×(1-P l ) Formula (13)
[0161] For data categorized as moderate cognitive load, such as Figure 6 As shown, it will definitely not be classified into other levels, corresponding to the case where the relative intensity is at the midpoint of the cognitive load segment, that is, the relative intensity is 0.5, P l P represents the proportion by which the relative intensity should decrease at this point. h This represents the proportion by which the relative intensity should be increased. Therefore, the quantitative score for the cognitive load level is:
[0162]
[0163] For data categorized as having a high level of cognitive load, such as Figure 6 As shown, the case where it will definitely not be classified into other levels corresponds to the case where the relative intensity is at the right end of the high cognitive load segment, that is, the relative intensity is 1, and the probability P of being classified into other levels. l +P m This means that the relative intensity should decrease by a certain percentage, which is P. l +P m Therefore, the quantitative score for high cognitive load is:
[0164] S h= 100 x (1 - (P l + P m )) = 100 x P h Equation (15)
[0165] The cognitive load quantification scoring principle can also use other criteria, which will not be described here.
[0166] S902, calculate the total cognitive load score: average the scores obtained in step S902 to obtain the brain cognitive load score of the current subject.
[0167] The embodiment uses multi-channel face video to measure cognitive load, that is, multiple independent cameras are used to simultaneously collect the face video of the subject to obtain multi-channel face video. This method allows the subject to have a certain range of movement during the measurement process, which makes the test experience of the subject more comfortable, further reduces the psychological pressure brought to the subject by the test scene, and makes the measurement of cognitive load more accurate.
[0168] Embodiment 2
[0169] The embodiment continues to disclose a brain cognitive load objective measurement method based on multi-channel face video, comprising the following steps:
[0170] S1, the subject faces three independent cameras to perform a cognitive load experiment, collects three-channel face videos of the subject during the execution of the mental calculation task under different cognitive load levels, and records the mental calculation task accuracy; the face video is used to extract relevant physiological indicators to measure the cognitive load level from the perspective of human physiological state; the mental calculation task accuracy measures the cognitive load level through the performance of the task.
[0171] S2, face recognition is performed on the collected face videos of channel 1 to channel 3, the region of interest is located, the gray mean value of the region of interest and the area of the region of interest are obtained, and the original pulse wave signal and the area sequence of the region of interest are extracted; wherein the region of interest is a skin region used to extract the original pulse wave signal, located at the forehead of the face; the area sequence of the region of interest is a sequence composed of the area size of the region of interest in each frame of the face video; the original pulse wave signal and the area sequence of the region of interest extracted from channel 1 are as shown in Figure 7
[0172] S3, the original pulse wave signals of channel 1 to channel 3 obtained in S2 are respectively subjected to improved adaptive noise ensemble empirical mode decomposition, and combined with wavelet threshold denoising and Butterworth band-pass filter to remove noise, to obtain denoised pulse wave signals, and the waveform of the original pulse wave signal extracted from channel 1 after denoising in S3 is as shown in Figure 8 S3, the original pulse wave signals of channel 1 to channel 3 obtained in S2 are respectively subjected to improved adaptive noise ensemble empirical mode decomposition, and combined with wavelet threshold denoising and Butterworth band-pass filter to remove noise, to obtain denoised pulse wave signals, and the waveform of the original pulse wave signal extracted from channel 1 after denoising in S3 is as shown in Figure 8 S3, the original pulse wave signals of channel 1 to channel 3 obtained in S2 are respectively subjected to improved adaptive noise ensemble empirical mode decomposition, and combined with wavelet threshold denoising and Butterworth band-pass filter to remove noise, to obtain denoised pulse wave signals, and the waveform of the original pulse wave signal extracted from channel 1 after denoising in S3 is as shown in Figure 8
[0173] S4, the pulse wave signals of channels 1-3 obtained after de-noising in S3 are selected, the pulse waves of the three G channels are used to calculate the time delay between channels 1-3 by cross-correlation, and time synchronization of the pulse wave signals of channels 1-3 and the area sequence of the region of interest is realized according to the time delay.
[0174] S5, for the time-synchronized pulse wave signals of channels 1-3 and the area sequence of the region of interest obtained in step S4, for the same time, the pulse wave data of the channel with the largest area in the area sequence of the region of interest in channels 1-3 is selected from the pulse waves of channels 1-3 as the data at this time, and all time points are traversed to finally obtain the fused pulse wave; the fused pulse wave is as shown in Figure 9
[0175] S6, the blood oxygen saturation and pulse variability are extracted from the fused pulse wave signal obtained in step S5, wherein the blood oxygen saturation is obtained by fitting the pulse waves of B and R channels, and the pulse variability is obtained by extracting the peak-peak time interval of the pulse wave of the green channel.
[0176] The equation to be fitted is shown in formula (11) of Embodiment 1. According to the fitting results of the measured data and the blood oxygen saturation data collected by the oximeter, the fitted equation is:
[0177] SPO2 = -3.6359*T + 102.6168 Formula (16)
[0178] S7, for the mental calculation tasks of the subjects under low, medium and high cognitive load levels, steps S1-S6 are repeated to obtain the mental calculation task accuracy, blood oxygen saturation and pulse variability indicators of the subjects under different cognitive load levels, form the data set required for cognitive load level classification, and then train the XGBoost classifier based on the data set.
[0179] S8, execute steps S1-S6 on the test object to obtain the mental calculation task accuracy, blood oxygen saturation and pulse variability indicators; input the mental calculation task accuracy, blood oxygen saturation and pulse variability indicators of the test object into the trained XGBoost classifier, and output the posterior probability estimate of the classification. The classification posterior probability of the test object under different cognitive load levels is shown in Table 1:
[0180] Table 1. Classification posterior probability table of the test object in Embodiment 2
[0181]
[0182] S9, use the posterior probability of the classification output by S8 to quantify and score the cognitive load level of the test object.
[0183] S901, calculating the low cognitive load level quantitative score S of the to-be-tested object by formula (13) l is 16 points; calculating the medium cognitive load level quantitative score S of the to-be-tested object by formula (14) m is 36 points; calculating the high cognitive load level quantitative score S of the to-be-tested object by formula (15) h is 87 points.
[0184] S902, calculating the total cognitive load quantitative score S of the to-be-tested object by averaging S l , S m , S h , and S
[0185] The above embodiment is a preferred embodiment of the present application, but the embodiment of the present application is not limited by the above embodiment, and any change, modification, substitution, combination, simplification made without departing from the spirit and principle of the present application should be an equivalent replacement mode, and all are included in the protection scope of the present application.
Claims
1. A method for objectively measuring cognitive load based on multi-channel facial video, characterized in that, The objective measurement method for cognitive load includes the following steps: S1. Subjects face multiple independent cameras to conduct a cognitive load experiment. Multi-channel facial videos are collected during the subjects' mental arithmetic tasks at different cognitive load levels, and the accuracy rate of the mental arithmetic tasks is recorded. The cognitive load is the total amount of cognitive resources used by the subject's brain. Different cognitive load levels are induced by the subjects performing mental arithmetic tasks of different difficulty. Each camera acquires one channel of facial video, and multiple cameras obtain multi-channel facial videos. S2. Perform face recognition on the acquired multi-channel facial videos, locate the region of interest (ROI), obtain the average grayscale value and area of the ROI, and extract the original pulse wave signal and the ROI area sequence. The ROI is the skin region used to extract the original pulse wave signal, located on the forehead of the face. The ROI area sequence is a sequence composed of the area sizes of the ROI in each frame of the facial video. The process is as follows: S201. The facial video of channel 1 is captured and framed to obtain a continuous color image sequence. The image format is RGB. The facial video of channel 1 refers to the facial video captured by the first camera. S202. For each frame of the color image sequence, a regression tree method based on gradient boosting learning is used to detect facial landmarks. Once facial landmarks are detected, the forehead of the face is located as the region of interest. S203. For each frame of the image, the region of interest (ROI) is separated into color channels to obtain the R, G, and B channel images, and the area of the ROI is calculated. For each color channel image, the average gray value of all pixels in the ROI is used as the amplitude of the pulse wave signal. The image sequence after a video segment is framed is processed and sorted by time to obtain the original pulse wave signals of the R, G, and B channels and a sequence of ROI area. The original pulse wave signals of the R, G, and B channels and the ROI area sequence are extracted from the facial video of one channel, and the four are time-synchronized. S204. Perform steps S201-S203 on the facial videos acquired from channels 2 to N respectively, to obtain the original pulse wave signals of channels R, G, and B and one region of interest area sequence for each channel's facial video; finally, obtain 3N original pulse wave signals including N original pulse wave signals of channel R, N original pulse wave signals of channel G, and N original pulse wave signals of channel B, and N region of interest area sequences; S3. The original pulse wave signal is subjected to improved adaptive noise ensemble empirical mode decomposition, and noise is removed by combining wavelet threshold denoising and Butterworth bandpass filter to obtain the denoised pulse wave signal; the process is as follows: S301. The original pulse wave signal obtained in step S2 is reconstructed and subjected to improved adaptive noise ensemble empirical mode decomposition to obtain multiple components, as follows: S3011, The first-order intrinsic mode component ξ1E1(n) obtained by EMD decomposition of group L white noise is... j Adding the original signal x(t) to each of the given signals, we get: x j (t)=x(t)+ξ1E1(n j (t)),j=1,2,…,L Where E1(·) represents finding the first IMF component of a signal's EMD decomposition, ξ1 represents the noise amplitude coefficient added in the first decomposition, with a value ranging from 0.1 to 0.3, and n j (t) represents the j-th group of zero-mean white noise with a variance of 1, and there are a total of L groups; For x j (t) Calculate the local mean, and then perform ensemble averaging to obtain the first-order residual component res1: Where M(·) represents the local mean of the signal; S3012, Calculate the first-order natural mode components. S3013, The second-order intrinsic mode component ξ2E2(n) of the white noise group L is decomposed by EMD. j Add res1 to (t) respectively, calculate the local mean and ensemble mean to obtain the second-order residual component res2: Where E2(·) represents finding the second IMF component of a signal EMD decomposition, and ξ2 represents the noise amplitude coefficient added in the second decomposition, with a value range of 0.1 to 0.3; Finally, the second-order intrinsic mode component is obtained by subtracting the second-order residual component from the first-order residual component. S3014. Calculate the remaining component res. k : Where E k (·) represents finding the k-th IMF component of an EMD decomposition of a signal, ξ k The noise amplitude coefficient added in the k-th decomposition is represented, with a value ranging from 0.1 to 0.
3. K represents the number of IMF components, k = 3, 4, ..., K. S3015. Calculate the k-th intrinsic mode component. S3016. Repeat steps S3011-S3015 until the remaining item res is reached. K It cannot be decomposed any further; at this point, res K The residual term is the sum of the K IMF components and the residual term of the original signal x(t), as shown in the following equation: S302. Use wavelet threshold denoising to denoise the first three components; S303. The first three noise-reduced components are superimposed with the remaining components to obtain the reconstructed signal; S304. Use a Butterworth bandpass filter to perform bandpass filtering on the reconstructed signal; S4. Extract the G-channel pulse wave from the denoised pulse wave signal, calculate the delay between multiple channels through cross-correlation, and use the delay to achieve time synchronization between the multi-channel pulse wave signal and the area sequence of the region of interest; wherein, the G-channel pulse wave is the green channel pulse wave obtained by color channel separation of the pulse wave signal; the process is as follows: S401. For the 3N filtered pulse wave signals obtained in step S3, select N G channel signals and use the cross-correlation method to calculate the delay between the N G channel pulse wave signals, and use it as the delay between the N channels. S402. Based on the delay between N channels, synchronize the 3N filtered pulse wave signals extracted from the N-channel facial video with the N area sequences of regions of interest in time. S5. For the same moment, select the pulse wave data of the channel with the largest corresponding region of interest from multiple pulse waves as the data for that moment, and iterate through all moments to obtain the fused pulse wave; the process is as follows: S501. For the pulse waves of the N R channels after time synchronization obtained in step S4, for the same moment, select the pulse wave data of the channel with the largest region of interest area in the N region of interest area sequence as the data of this moment, traverse all moments, and finally obtain 1 fused R channel pulse wave. S502. For the pulse waves of the N G channels after time synchronization obtained in step S4, execute step S501 to obtain a fused G channel pulse wave. S503. For the pulse waves of the N B channels after time synchronization obtained in step S4, execute step S501 to obtain a fused B channel pulse wave. S6. Extract the blood oxygen saturation and pulse variability of the fused pulse wave signal. The fused pulse wave signal contains three color channels: R, G, and B. R represents red, G represents green, and B represents blue. The blood oxygen saturation is obtained by fitting the pulse waves of the B and R channels. The pulse variability is obtained by extracting the time interval between the peaks of the pulse wave of the G channel. S7. Perform mental arithmetic tasks at low, medium and high cognitive load levels on multiple subjects. Repeat steps S1-S6 to obtain the accuracy, blood oxygen saturation and pulse variability of the mental arithmetic tasks of multiple subjects at different cognitive load levels. This forms the dataset required for cognitive load level classification. Then, based on the dataset, train the XGBoost classifier. S8. Perform mental arithmetic tasks at low, medium, and high cognitive load levels on the test subject. Repeat steps S1-S6 to obtain the test subject's accuracy, blood oxygen saturation, and pulse variability index at the low, medium, and high cognitive load levels. Input the test subject's accuracy, blood oxygen saturation, and pulse variability index at the low, medium, and high cognitive load levels into the trained XGBoost classifier to estimate the posterior probability of the classification. S9. Use the posterior probability of the classification to quantitatively assess the cognitive load level of the test subject.
2. The method for objectively measuring cognitive load based on multi-channel facial video according to claim 1, characterized in that, The process of step S1 is as follows: S101. The subject sits in front of a display screen controlled by an external laptop, showing a mental arithmetic task to be calculated. N independent cameras are positioned evenly in pairs at an angle of θ degrees on an arc 50 cm away from the subject, capturing multi-channel facial video of the subject. The subject's head can rotate freely within the range of (N-1)θ degrees; where N is the number of independent cameras and also the number of channels. S102. A mental arithmetic task to be calculated is displayed on the screen and a countdown begins. The subject faces the screen and calculates the mental arithmetic task. N cameras simultaneously record the subject's facial video. S103. After the countdown ends, the mental arithmetic task displayed on the screen disappears. The subject must state the calculation result within a limited time. If no calculation result is given within the limited time, it is recorded as a calculation error. After the limited time ends, the screen will automatically switch to the next mental arithmetic task. A total of 5 mental arithmetic tasks of the same difficulty are performed, and the accuracy rate of the mental arithmetic tasks is recorded.
3. The method for objectively measuring cognitive load based on multi-channel facial video according to claim 1, characterized in that, The process of step S6 is as follows: S601. Perform peak detection on the fused G-channel pulse wave obtained in S5, and extract the time of all peaks for pulse variability extraction. S602. Difference the time of adjacent peaks to obtain the pulse beat interval sequence. Use cubic spline interpolation to interpolate this beat interval sequence to obtain the pulse variability signal. S603. The pulse variability signal is subjected to feature extraction to obtain six pulse variability indices. Among them, the mean (MEAN) and standard deviation (SDPP) are obtained by extracting features in the time domain. The power spectrum of the pulse variability signal is solved and the frequency domain features are extracted to obtain the total signal power (TP), low frequency power (LF), high frequency power (HF), and the power ratio of low frequency to high frequency (LF / HF). S604. Using the fused R-channel and B-channel pulse waves obtained in S5, calculate the independent variables in the empirical formula for blood oxygen saturation. Using the measured data from a standard finger-clip pulse oximeter as the dependent variable, use the least squares method to linearly fit the empirical formula for blood oxygen saturation and fit the unknown parameters of the empirical formula for blood oxygen saturation. Use the fitted empirical formula for blood oxygen saturation to calculate the current blood oxygen saturation.
4. The method for objectively measuring cognitive load based on multi-channel facial video according to claim 1, characterized in that, The process of step S7 is as follows: S701. Perform mental arithmetic tasks at low, medium and high cognitive load levels on multiple subjects, repeat steps S1-S6, and obtain the accuracy rate, blood oxygen saturation and pulse variability index of the mental arithmetic tasks of multiple subjects at different cognitive load levels. S702. Label the collected dataset, where the data under low cognitive load level is labeled 0, the data under medium cognitive load level is labeled 1, and the data under high cognitive load level is labeled 2, thus obtaining the dataset required for training the classifier. S703. Use the dataset obtained in S702 to train the XGBoost classifier.
5. The method for objectively measuring cognitive load based on multi-channel facial video according to claim 1, characterized in that, The process of step S8 is as follows: S801. The test subject is given mental arithmetic tasks at three cognitive load levels: low, medium, and high. Steps S1-S6 are repeated to obtain the test subject's accuracy rate, blood oxygen saturation, and pulse variability index at the three cognitive load levels. S802. Input the mental arithmetic task accuracy, blood oxygen saturation, and pulse variability obtained in S801 into the trained XGBoost classifier, and output the posterior probability estimate of the classification to obtain the probability P that the cognitive load level is classified as low load level. l The probability P that the cognitive load level is classified as moderate load level m The probability P that the cognitive load level is classified as high load level h .
6. The method for objectively measuring cognitive load based on multi-channel facial video according to claim 1, characterized in that, The process of step S9 is as follows: S901. Using the posterior probability of the classification obtained in S8, the cognitive load of the test subject at low, medium, and high cognitive load levels is quantified and scored. The quantification score for the low cognitive load level is represented by S... l The quantitative score of cognitive load level is represented by S. m The quantitative score for high cognitive load is represented by S. h express; S902. Average the scores obtained in step S902 to obtain the brain cognitive load score of the current test subject.
Citation Information
Patent Citations
Non-contact continuous blood pressure measurement method and system based on video stream
CN114569096A
Brain cognitive load quantification method based on facial video physiological index extraction
CN114983414A