Method for evaluating working memory load based on multi-channel video in non-stationary state

By using multi-channel facial video technology and signal processing methods, the subjectivity and static limitations of traditional working memory load assessment are overcome, enabling efficient and accurate assessment in non-static states.

CN116616709BActive Publication Date: 2026-04-21SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2023-05-23
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing methods for assessing working memory load suffer from issues such as high subjectivity, long processing time, expensive equipment, and the need for a static state, resulting in insufficient comfort and accuracy.

Method used

Using multi-channel facial video technology, facial videos are captured through multi-channel network cameras, physiological parameters are extracted, and signal processing techniques such as variational mode decomposition, wavelet threshold filtering, and multi-harmonic signal-to-noise ratio weighted fusion are combined to calculate heart rate and heart rate variability indices, and to construct a physiological model of working memory load.

Benefits of technology

It enables efficient, non-destructive, and non-invasive assessment of working memory load in a non-static state, reducing subjectivity and equipment inconvenience, and improving comfort and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116616709B_ABST
    Figure CN116616709B_ABST
Patent Text Reader

Abstract

This invention discloses a method for assessing working memory load based on physiological indicators from multi-channel facial videos under non-static conditions. It employs multi-channel recording of facial videos with varying working memory loads, extracts indicators such as heart rate, and assesses the working memory load level, overcoming the problem of requiring the subject to remain still during measurement. The steps are as follows: Multi-channel acquisition of facial videos under different loads; division of a single-channel video into multiple regions of interest (ROIs) and extraction of pulse wave signals for denoising; then fusing the pulse wave signals from multiple ROIs based on multi-harmonic signal-to-noise ratio (MNR) to obtain a single-channel pulse wave signal; repeating the above steps to obtain multi-channel pulse wave signals and aligning them; then fusing the multi-channel pulse wave signals to obtain the final pulse wave signal, extracting indicators such as heart rate, and assessing the working memory load level. This invention provides a non-invasive and non-destructive method for assessing working memory load.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information processing technology, and specifically to a method for evaluating the working memory load of multi-channel video under non-static conditions. Background Technology

[0002] Common methods for assessing working memory load involve indirect measurements, such as behavioral, psychological, and physiological observations. These indirect methods can be broadly categorized into subjective and objective approaches. Subjective assessments are cumbersome, time-consuming, and suffer from significant subjectivity and time lag. Objective assessments can be divided into contact-based parametric measurements and non-contact physiological parameter measurements. Contact-based measurements primarily include electrocardiograms (ECG), electroencephalograms (EEG), and photoplethysmography (PPG). While accurate physiological parameters can be obtained using wearable medical devices, these devices are expensive, bulky, and require professional operation, leading to poor comfort and discomfort for test subjects, making them unsuitable for daily working memory load assessment and prevention. Traditional non-contact parametric measurement methods typically require the test subject to remain relatively still, further reducing comfort. Summary of the Invention

[0003] The purpose of this invention is to improve upon the subjectivity and inconvenience of traditional non-contact methods for measuring physiological parameters and working memory load identification methods, as well as the limitations of contact-based measurements. It provides a working memory load assessment method based on multi-channel facial video in a non-static state. This method uses a multi-channel webcam to record facial video, extracts physiological parameters, and then assesses working memory load. It uses non-contact devices to measure physiological parameters and does not require the subject to remain still. It has advantages such as being non-invasive, comfortable, efficient, economical, and objective, providing objective auxiliary analysis for working memory load assessment and possessing significant practical value.

[0004] The objective of this invention can be achieved by adopting the following technical solutions:

[0005] A working memory load assessment method based on multi-channel facial video in a non-static state is proposed to improve upon the subjectivity and inconvenience of traditional non-contact methods for measuring physiological parameters and identifying working memory load levels using contact devices. The identification method includes the following steps:

[0006] S1. Construct a test scenario, collect multi-channel facial videos under different working memory loads, and calculate the test accuracy. The different working memory load levels are set to three different levels: high load, medium load, and low load. The test accuracy refers to the accuracy of reproducing the content of the first c rounds after c ≤ 3 rounds of memory. In the test scenario, facial videos are collected by N cameras placed around the radius of the test subject's circle. The facial video collected by each camera is called one channel video.

[0007] S2. For each channel's video, segment the video into frames, locate multiple regions of interest (ROIs), and extract the original pulse wave signal for each ROI, then denoise it. For each channel's facial video under different working memory loads, segment the facial video into frames to obtain a facial video image sequence, then perform face recognition and facial feature point detection to locate multiple ROIs. Separate the color channels of each ROI and extract the average grayscale value of the green channel as the original pulse wave signal for that ROI. Then, preprocess the original pulse wave signal, perform variational mode decomposition, wavelet thresholding, and Butterworth bandpass filtering to denoise it.

[0008] S3. Repeat step S2 to obtain the denoised pulse wave signal for each region of interest, and calculate the multi-harmonic signal-to-noise ratio and weight respectively. Then, the pulse wave signals of multiple regions of interest are fused according to the weight to obtain the pulse wave signal of each channel.

[0009] S4. Repeat steps S2 to S3 to obtain N channels of pulse wave signals. Calculate the multi-harmonic signal-to-noise ratio and weight of each channel of pulse wave signal. Then, perform cross-correlation and translation alignment on the N channels of pulse wave signals. Finally, weighted fuse of the multi-channel pulse wave signals according to the weights to obtain the N-channel fused pulse wave signal.

[0010] S5. Perform time-domain and frequency-domain analysis on the N-channel fused pulse wave signal to calculate heart rate and heart rate variability index;

[0011] S6. Test several subjects with questions of different working memory load levels, repeat steps S1 to S5, obtain the subjects' test accuracy, heart rate and heart rate variability characteristic indicators, form the dataset required for model construction, and then perform feature elimination, feature selection and load level classification to construct a working memory load physiological model.

[0012] S7. In the actual test of the working memory load of the test subjects, the test accuracy and the heart rate and heart rate variability indicators extracted from the multi-channel facial video after eliminating individual differences are input into the working memory load physiological model to assess the current working memory load level of the test subjects.

[0013] Further, step S1 is as follows:

[0014] S101. Under sufficient external lighting conditions, multi-channel facial videos under different working memory loads are acquired using N camera channels placed around the radius of the test subject's circle. The shooting distance is 50 cm, and the shooting durations for low, medium, and high load levels are 10 seconds, 15 seconds, and 20 seconds, respectively. This step is based on the continuous variation of working memory load, with the acquired multi-channel facial videos corresponding to low, medium, and high working memory load scenarios.

[0015] S102. Display the memorized content on the screen, gradually increasing the working memory load level by requiring the test subject to recall content from previous rounds; starting from the fourth round, the test subject memorizes the current round and writes down the content from the previous three rounds. This step ensures the continuous and variable working memory load by applying a continuous and parameter-variable load to the working memory, thereby continuously stimulating the brain and ensuring the scientific rigor of the test.

[0016] S103. For low workload, recall the content of the previous round; for medium workload, recall the content of the previous two rounds; for high workload, recall the content of the previous three rounds. This step is to induce different levels of working memory load by increasing the difficulty of recalling previous rounds. The higher the difficulty, the more tense the brain nerves become, thereby inducing different levels of working memory load.

[0017] S104. Starting from the 4th round, record the content pressed by the test subject in the previous 3 rounds, and calculate the test accuracy rate by comparing it with the questions. This step takes into account the limited capacity of working memory. Usually, the memory of one thing will cover another, and memory loss is easy to occur. Therefore, the test accuracy rate is regarded as an important indicator to identify the working memory load level.

[0018] Furthermore, step S2 is as follows:

[0019] S201. First, the acquired N-channel facial video is framed to obtain a facial video image sequence. The purpose of this step is to obtain a facial video image sequence by framing the video, which facilitates subsequent face recognition and facial feature point detection;

[0020] S202. Perform face recognition and facial feature point detection on the facial video image sequence, and locate multiple regions of interest based on the facial feature points. This step takes into account that the signal-to-noise ratio of the pulse wave is different in different regions of the face. Based on face recognition and facial feature point detection, regions with high signal-to-noise ratio and quality are selected as regions of interest to reduce errors when calculating heart rate and heart rate variability indicators.

[0021] S203. Separate the color channels of each region of interest and extract the average grayscale value of the green channel as the original pulse wave signal for that region of interest. This step is to extract the original pulse wave signal, and considering that the green channel contains the most heart rate signal, is most sensitive to changes in heart rate, and best reflects the information of heartbeat;

[0022] S204. Perform preprocessing on the extracted raw pulse wave, including detrending and Z-score normalization. This step aims to eliminate interference from light variations over time that could affect the raw pulse wave signal waveform in the green channel, and to eliminate DC components and scale factors in the signal, while preserving the morphological characteristics of the signal so that the signals can be compared or analyzed within the same range.

[0023] S205. Perform variational mode decomposition on the preprocessed pulse wave signal to obtain multiple intrinsic mode function components. The intrinsic mode function is abbreviated as IMF. Perform fast Fourier transform on each IMF component to calculate the frequency corresponding to the maximum peak in the spectrum. Select the first λ high-frequency IMF components with higher frequencies to facilitate subsequent noise reduction.

[0024] S206. Perform wavelet denoising on the selected λ high-frequency IMF components. This step takes into account that the high-frequency component signals obtained by variational mode decomposition contain high-frequency noise such as motion artifacts, while wavelet denoising has a good removal effect on environmental noise, power frequency noise, and other noises.

[0025] S207. The high-frequency IMF component that has undergone wavelet denoising is superimposed with the remaining IMF component to reconstruct the pulse wave signal.

[0026] S208. Perform Butterworth bandpass filtering on the reconstructed pulse wave signal. The purpose of this step is to remove signals outside the normal human heart rate range. Under normal circumstances, the human heart rate ranges from 42 to 180 beats per minute, corresponding to a frequency of 0.7-3 Hz. Therefore, the bandpass frequency is set to 0.7-3 Hz, and filtering is performed to obtain the final pulse wave signal for a specific region of interest, called the region of interest pulse wave signal. Assuming that the number of regions of interest obtained in step S202 is M, then steps S202 to S208 are performed on each region of interest to obtain M region of interest pulse wave signals Y. m m = 1, 2, 3, ..., M.

[0027] Furthermore, step S3 is as follows:

[0028] S301. Calculate the pulse wave signal Y of the m-th region of interest using Fast Fourier Transform. m Multiharmonic signal-to-noise ratio (SNR) ROJ (m), the calculation formula is as follows:

[0029]

[0030] Among them, F (m) (.) indicates that the pulse wave signal Y m The spectrum diagram, where i represents the frequency index of the spectrum diagram, Fs represents the sampling rate of the pulse wave signal, and a max The fundamental frequency is represented by A[k] = [a], where A[k] represents the frequency corresponding to the peak value in the spectrum. max ,2a max ,3a max ,..ka max ] indicates less than The multi-harmonic frequency array, where b represents the frequency corresponding to a point of adjacent fundamental frequency or adjacent harmonic.

[0031] The multiharmonic signal-to-noise ratio (SNR) of the pulse wave signals in the M regions of interest is calculated. ROJ (m);

[0032] S302. Based on the signal-to-noise ratio of the M regions of interest, calculate the weights of the pulse wave signals in the M regions of interest using the following formula:

[0033]

[0034] Among them, SNR ROI (m) represents the signal-to-noise ratio of the pulse wave signal in the m-th region of interest, and M represents the number of regions of interest. The sum of the signal-to-noise ratios of the M regions of interest is represented, and r is an empirical constant, representing the nonlinear stretching of the weights by the stretching parameter.

[0035] S303. The pulse wave signals of the M regions of interest are weighted and fused according to the above weights to obtain the pulse wave signal of a single channel. The calculation formula is as follows:

[0036]

[0037] Among them, IPPG cam This represents the pulse wave signal after fusion of individual channels.

[0038] Furthermore, step S4 is as follows:

[0039] S401. Process the N-channel facial video using steps S2 to S3 to obtain N-channel pulse wave signals IPPG. cam (n), n=1, 2, 3,..., N;

[0040] S402. Calculate the multiharmonic signal-to-noise ratio (SNR) of the nth channel pulse wave signal using Fast Fourier Transform.cam (n);

[0041] The formula for calculating the multiharmonic signal-to-noise ratio of the pulse wave signal in this channel is as follows:

[0042]

[0043] Calculate the multi-harmonic signal-to-noise ratio (SNR) for N channels to obtain the SNR of the N channel pulse wave signals. cam (n);

[0044] S403. Based on step S402, obtain the signal-to-noise ratio of the N channels, and then calculate the weights of the N channel pulse wave signals. The calculation formula is as follows:

[0045]

[0046] Among them, SNR cam (n) represents the signal-to-noise ratio of the pulse wave signal in the nth channel, and N represents the number of channels. This represents the sum of the signal-to-noise ratios of the N channels;

[0047] S404. Cross-correlate and align the pulse wave signals of each channel;

[0048] S405. The aligned pulse wave signals are weighted and fused according to the weights calculated in step S403 to obtain an N-channel fused pulse wave signal. The fusion calculation formula is as follows:

[0049]

[0050] Among them, IPPG cam (n) represents the pulse wave signal of the nth channel, and Signal represents the pulse wave signal after the fusion of N channels.

[0051] Furthermore, step S5 is as follows:

[0052] S501. Convert the pulse wave signal Signal to the frequency domain, and regard the frequency corresponding to the maximum peak of the spectrum as the frequency of heartbeat; multiply the obtained frequency by 60 to get the heart rate.

[0053] S502. Perform narrowband pass filtering on the pulse wave signal Signal. The purpose of this step is to further improve the quality of the fused pulse wave signal and reduce interference from heart rate variability feature extraction.

[0054] S503. Perform cubic spline interpolation on the narrowband-pass filtered pulse wave signal. The purpose of this step is to increase the number of sampling points and improve the accuracy of subsequent peak detection.

[0055] S504. Peak point detection is performed on the pulse wave signal after cubic spline interpolation. The purpose of this step is to extract the time interval between all adjacent points for the extraction of heart rate variability.

[0056] The S505 pulse rate variability (PRV) signal and heart rate variability (HRV) signal exhibit a high degree of consistency, thus allowing the calculation of heart rate variability indices from the pulse wave signal. Based on the time interval between adjacent peak points, the HRV indices are calculated, including the standard deviation of heartbeat intervals (SDNN), the standard deviation of the difference between adjacent heartbeat intervals (SDSD), the root mean square (RMSSD) of the difference between adjacent heartbeat intervals, the mean heart rate (Mean_HR), the maximum heart rate (Max_HR), the minimum heart rate (Min_HR), the standard deviation of heart rate (STD_HR), the total signal power (TP), the low-frequency power (LF), the high-frequency power (HF), the ratio of low-frequency to high-frequency power (LF / HF), and the very low-frequency power (VLF).

[0057] Furthermore, step S6 is as follows:

[0058] S601. Collect multi-channel facial videos of the test subjects under different working memory loads and calculate the test accuracy.

[0059] S602. Repeat steps S2 to S5 to extract heart rate and heart rate variability characteristic indicators;

[0060] S603. Eliminate individual differences in heart rate and heart rate variability characteristics. The purpose of this step is to eliminate the influence of the individual's baseline physiological parameters on the extracted physiological parameter characteristics, thereby improving the accuracy of assessment of working memory load levels;

[0061] S604. Combine the heart rate and heart rate variability feature indicators after individual differences have been eliminated with the test accuracy to form the feature dataset required to build the model, and label the feature dataset, where the low load label is 0, the medium load label is 1, and the high load label is 2.

[0062] S605. Treat each indicator as a feature, and through feature selection from the feature dataset, eliminate irrelevant and redundant features to form the optimal subset of feature data. The purpose of this step is to reduce the time and learning difficulty of model building, and improve the efficiency and classification accuracy of the model;

[0063] S606. Divide the feature data subset obtained in step S605 into a training set and a test set; wherein, the training set is used to train the random forest model, the random forest model is an existing technology, from "Breiman L. Random forests[J].Machine learning,2001,45:5-32.", input the training set and labels to obtain the working memory load physiological classification model through training, and the test set is used to verify the performance indicators of the model;

[0064] S607. Construct a physiological classification model for working memory load. First, train the model based on the training set data. Then, use a classifier for classification and obtain the hyperparameters of the classifier through K-fold cross-validation. Finally, obtain the physiological classification model for working memory load. The purpose of this step is to evaluate the performance of the classification model and reduce the risk of model overfitting.

[0065] S608. Verify the recognition accuracy of the physiological classification model of working memory load using a test set. The purpose of this step is to establish a mapping relationship between various features and different levels of working memory load through accurate classification by the model, thereby achieving the assessment of working memory load levels.

[0066] Furthermore, step S7 is as follows:

[0067] S701. Collect multi-channel facial videos of the test subjects and calculate the test accuracy rate.

[0068] S702. Heart rate and heart rate variability characteristic indicators are obtained through steps S2 to S5.

[0069] S703. Input the test accuracy, heart rate, and heart rate variability into the working memory load physiological classification model to assess the subject's current working memory load level.

[0070] The present invention has the following advantages and effects compared with the prior art:

[0071] (1) The present invention is based on extracting physiological parameters from pulse wave signals from multi-channel facial video to assess working memory load level. It is an objective, non-contact method that improves the shortcomings of subjective assessment methods, such as subjective bias and time lag, as well as the inconvenience caused by contact measurement in objective evaluation methods. The present invention improves the traditional non-contact method of measuring physiological parameters by means of multi-channel facial video, which is usually measured when the subject is stationary or the subject's range of motion is greatly limited, and the measurement results are easily affected by motion interference, thus improving the comfort of the subject.

[0072] (2) This invention selects multiple facial regions of interest that are closely related to the pulse wave signal to improve the quality and signal-to-noise ratio of the original pulse wave signal and reduce the error when calculating heart rate and heart rate variability index; This invention combines a joint denoising algorithm of signal preprocessing, signal decomposition and wavelet thresholding to denoise the original pulse wave signal, eliminate the noise of the pulse wave signal, and obtain a pulse wave signal with high quality and signal-to-noise ratio; This invention defines the multi-harmonic signal-to-noise ratio from the perspective of energy, takes the energy near the fundamental frequency and each harmonic as the energy of the pulse wave, and the energy in other places as the noise signal-to-noise ratio calculation formula, and weighted fuses the pulse wave signals of multiple regions of interest to obtain a high-quality pulse wave signal.

[0073] (3) This invention proposes a multi-channel signal cross-correlation method for multi-channel camera pulse wave signals, which reduces the data misalignment caused by inconsistent opening times of multiple cameras due to various reasons such as camera device hardware, software drivers, or computer performance and complexity, resulting in errors in subsequent heart rate and heart rate variability calculations; This invention also proposes a multi-channel signal fusion method based on multi-harmonic signal-to-noise ratio for multi-channel camera pulse wave signal fusion, which achieves signal alignment between different channels and then performs weighted fusion based on multi-harmonic signal-to-noise ratio to obtain a high-quality pulse wave signal, thereby improving the accuracy of extracting heart rate and heart rate variability.

[0074] (4) The present invention uses the extracted physiological parameters as features, and then uses a recursive feature elimination algorithm to select features, reducing the influence of irrelevant and redundant features on the classification model, improving the efficiency and accuracy of the model, constructing a physiological classification model of working memory load, establishing the correspondence between different levels of working memory load and multiple features, and realizing the effective classification of different levels of working memory load based on multi-channel facial video. Attached Figure Description

[0075] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0076] Figure 1 This is a flowchart of a working memory load assessment method based on multi-channel facial video in a non-static state disclosed in this invention;

[0077] Figure 2 This is a flowchart of the original pulse wave denoising process in Embodiment 1 of the present invention;

[0078] Figure 3 This is a flowchart illustrating the process of weighted fusing of pulse wave signals from each camera to obtain a multi-channel fused pulse wave signal in Embodiment 1 of the present invention.

[0079] Figure 4 This is a flowchart of the process for extracting heart rate and pulse variability in Embodiment 1 of the present invention;

[0080] Figure 5 This is a schematic diagram of the green channel signals for the three regions of interest in Embodiment 2 of the present invention;

[0081] Figure 6 This is a schematic diagram of the pulse wave signal after denoising the three regions of interest in Embodiment 2 of the present invention;

[0082] Figure 7 This is a schematic diagram of the pulse wave signal obtained by fusing three regions of interest in Embodiment 2 of the present invention;

[0083] Figure 8 This is a schematic diagram of the pulse wave signals of the three cameras in Embodiment 2 of the present invention;

[0084] Figure 9 This is a schematic diagram of the pulse wave signal obtained by fusing three cameras in Embodiment 2 of the present invention;

[0085] Figure 10 This is a schematic diagram of the pulse wave signal after passing through a narrow bandpass filter in Embodiment 2 of the present invention. Detailed Implementation

[0086] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0087] Example 1

[0088] Figure 1 This is a flowchart of a working memory load assessment method based on multi-channel facial video physiological indicators under non-static conditions, provided by an embodiment of the present invention. This embodiment is implemented in a piano playing scenario, enabling non-contact assessment of working memory load levels. The specific steps are as follows:

[0089] S1. Construct a test scenario, collect 3-channel facial videos with different working memory loads and calculate the test accuracy; In the test scenario, facial videos are collected by 3 cameras placed around the radius of the test subject's circle, and the facial video collected by each camera is called 1 channel video;

[0090] S101. Under sufficient external lighting conditions, multi-channel facial videos under different working memory loads are acquired through three camera channels placed around the radius of the test subject's circle. The shooting distance is 50 cm, and the shooting durations for low load, medium load, and high load levels are 10 seconds, 15 seconds, and 20 seconds, respectively.

[0091] S102. Display 7 randomly generated Arabic numerals on the screen. You have 15 seconds to memorize them. After 15 seconds, they will be replaced with 7 different numerals. This process will continue for 5 rounds. Starting from the 4th round, you will memorize the current round and write down the contents of the previous 3 rounds.

[0092] S103. For low load, recall the content of the previous wheel; for medium load, recall the content of the previous two wheels; for high load, recall the content of the previous three wheels.

[0093] S104. In this embodiment, the 88-key electronic keyboard has been modified. A circuit board with multiple sensors is installed under the keys. When a key is pressed, the sensors record the key played by the test subject via a serial port. After finishing playing the current piece of music, the test subject presses key 88 at the end of the keyboard to indicate the end of playing the current piece. After playing all pieces of music, the data recorded by the serial port is output as a txt file.

[0094] S105. Starting from the 4th round, record the content pressed by the test subject in the previous 3 rounds, and calculate the test subject's accuracy rate in playing the music by comparing it with the questions.

[0095] S2. The video is frame-by-framed to locate three regions of interest (ROIs), and the raw pulse wave signals of the three ROIs are extracted and then denoised. The specific steps are as follows:

[0096] like Figure 2 The flowchart for extracting the raw pulse wave and denoising.

[0097] S201. First, the 3-channel facial video is divided into frames to obtain a facial video image sequence;

[0098] S202. Perform face recognition and face feature point detection on the facial video image sequence. Based on the face feature points, locate three parts: the forehead, the left cheek, and the right cheek, and use these as regions of interest.

[0099] S203. Separate the color channels of each region of interest and extract the average gray value of the green channel as the original pulse wave signal of that region of interest;

[0100] S204. Preprocessing of the extracted raw pulse wave, including detrending and Z-score normalization; wherein, the Z-score normalization formula is as follows:

[0101]

[0102] Where x represents the initial data, μ represents the mean, σ represents the standard deviation, and z represents the normalization result.

[0103] S205. Variational mode decomposition is performed on the preprocessed pulse wave signal to obtain multiple intrinsic mode function components. The intrinsic mode function is abbreviated as IMF. Fast Fourier transform is performed on each IMF component to calculate the frequency corresponding to the maximum peak value. It is found that there are 8 IMF components with relatively high frequencies.

[0104] S206. Wavelet denoising is performed on the selected 8 high-frequency IMF components. The wavelet basis function in the wavelet threshold denoising is the "db8" wavelet, the wavelet decomposition level is 3, and the threshold is a global threshold and a hard threshold function. The formula for the global threshold is as follows:

[0105]

[0106] Where σ represents the standard deviation of the noise, and N represents the length of the signal.

[0107] S207. The 8 IMF components that have undergone wavelet denoising are superimposed with the remaining 2 IMF components to reconstruct the pulse wave signal.

[0108] S208. The reconstructed pulse wave signal is processed using a Butterworth 5th order bandpass filter to obtain a higher quality pulse wave signal.

[0109] S3. Repeat step S2 to obtain the denoised pulse wave signal for each region of interest, and calculate the multi-harmonic signal-to-noise ratio and weights respectively. Then, the pulse wave signals of multiple regions of interest are fused according to the weights to obtain the pulse wave signal of each channel. The specific steps are as follows:

[0110] S301. Calculate the multiharmonic signal-to-noise ratio of the pulse wave signals of the three regions of interest after step S208 using fast Fourier transform;

[0111] S302. Calculate the weights of the pulse wave signals in the three regions of interest based on their signal-to-noise ratios.

[0112] S303. The pulse wave signals of the three regions of interest are weighted and fused according to the above weights to obtain a single-channel fused pulse wave signal.

[0113] S4. Repeat steps S2 to S3 to obtain pulse wave signals from three channels. Calculate the multiharmonic signal-to-noise ratio and weight of each channel's pulse wave signal. Then, perform cross-correlation and shifting alignment on the three channels' pulse wave signals. Finally, weightedly fuse the three channels' pulse wave signals according to their weights to obtain a fused three-channel pulse wave signal. The specific steps are as follows:

[0114] like Figure 3 This is a flowchart for multi-channel fusion.

[0115] S401. Process the facial video of the three channels using steps S2 to S3 to obtain the pulse wave signals of the three channels.

[0116] S402. Calculate the multiharmonic signal-to-noise ratio of the pulse wave signals of the three channels respectively using fast Fourier transform;

[0117] S403. Obtain the signal-to-noise ratio of the three channels according to step S402, and then calculate the weight of the pulse wave signal of the three channels.

[0118] S404. Cross-correlate and align the pulse wave signals from the three channels;

[0119] S405. The aligned pulse wave signals are weighted and fused according to the weight ratio calculated in step S403 to obtain three pulse wave signals with the same combination.

[0120] S5. Perform time-domain and frequency-domain analysis on the fused 3-channel pulse wave signal to calculate the heart rate and heart rate variability index. The specific steps are as follows:

[0121] like Figure 4 This is a flowchart of the process for extracting heart rate and pulse variability.

[0122] S501. Convert the pulse wave signal to the frequency domain and find the frequency corresponding to the highest peak of the amplitude spectrum, which is the number of heartbeats per second of the subject; multiply the obtained number of heartbeats per second by 60 to get the heart rate;

[0123] S502. Perform narrowband pass filtering on the fused pulse wave signal;

[0124] S503. Perform cubic spline interpolation on the narrowband pass filtered pulse wave signal;

[0125] S504. Detect the peak point of the pulse wave signal after cubic spline interpolation;

[0126] S505. Pulse rate variability (PRV) and heart rate variability (HRV) signals exhibit high consistency, therefore, HRV indicators can be calculated using pulse wave signals. Based on the time interval between adjacent peak points obtained in step S504, HRV indicators are calculated. These indicators include 25 metrics such as: standard deviation of heartbeat interval (SDNN), standard deviation of the difference between adjacent heartbeat intervals (SDSD), root mean square (RMSSD) of the difference between adjacent heartbeat intervals, mean heart rate (Mean_HR), maximum heart rate (Max_HR), minimum heart rate (Min_HR), standard deviation of heart rate (STD_HR), total signal power (TP), low-frequency power (LF), high-frequency power (HF), low-frequency to high-frequency power ratio (LF / HF), and very low-frequency power (VLF).

[0127] S6. Test several subjects with questions of different working memory load levels, repeating steps S1 to S5 to obtain the subjects' test accuracy, heart rate, and heart rate variability characteristics, forming the dataset required for model construction. Then, eliminate individual differences in features, select features, and classify load levels to construct a physiological model of working memory load. The specific steps are as follows:

[0128] S601. Collect multi-channel facial videos of the test subjects under different working memory loads and calculate the test accuracy.

[0129] S602. Repeat steps S2 to S5 to extract heart rate and heart rate variability characteristic indicators;

[0130] S603. Eliminate individual differences in heart rate and heart rate variability indicators;

[0131] S604. Combine the heart rate and heart rate variability feature indicators after individual differences have been eliminated with the test accuracy to form the feature dataset required to build the model, and label the feature dataset, where the low load label is 0, the medium load label is 1, and the high load label is 2.

[0132] S605. Treat each indicator as a feature, and eliminate irrelevant and redundant features from the feature dataset through feature selection to form the optimal feature data subset. Feature selection is performed using a recursive feature elimination algorithm.

[0133] S606. Divide the feature data subset obtained in step S605 into a training set and a test set. The training set is used to train the random forest model. The random forest model is from "Breiman L. Random forests[J].Machine learning,2001,45:5-32." The working memory load physiological classification model is obtained by inputting the training set and labels and training. The test set is used to verify the performance indicators of the model.

[0134] S607. Construct a physiological classification model for working memory load. First, train the model based on the training set data. Then, use a classifier for classification and obtain the hyperparameters of the classifier through 5-fold cross-validation. Finally, obtain the physiological classification model for working memory load. Specifically, the physiological classification model for working memory load is constructed using a random forest algorithm, and the classification results are validated using 5-fold cross-validation. Experimental results show that the mean value of the 5-fold cross-validation is 0.65.

[0135] S608. Validate the recognition accuracy of the working memory load physiological classification model using a test set. The formula for recognition accuracy is as follows:

[0136]

[0137] In the formula, TP represents the number of minority class samples correctly predicted as minority class, FN represents the number of minority class samples incorrectly predicted as majority class, FP represents the number of majority class samples incorrectly predicted as minority class, and TN represents the number of majority class samples correctly predicted as majority class. The accuracy of the model for classifying each type of load is shown in Table 1.

[0138] Table 1. Accuracy of load level identification in Example 1 S7. In the actual test of the working memory load of the test subjects, the test accuracy and the heart rate and heart rate variability indicators extracted from the multi-channel facial video after eliminating individual differences are input into the working memory load physiological model to assess the current working memory load level of the test subjects.

[0139] S701. Collect multi-channel facial videos of the test subjects and calculate the test accuracy rate.

[0140] S702. Heart rate and heart rate variability characteristic indicators are obtained through steps S2 to S5.

[0141] S703. Input the test accuracy, heart rate, and heart rate variability into the working memory load physiological classification model to assess the subject's current working memory load level.

[0142] Example 2

[0143] Figure 1 This is a flowchart of a method for assessing working memory load based on physiological indicators of multi-channel facial videos under non-static conditions, provided by an embodiment of the present invention. The method uses heart rate and heart rate variability extracted from facial videos with high working memory load as reference indicators to assess working memory load. The specific steps are as follows:

[0144] S1. Collect 20 seconds of three-channel facial video of the subject under high working memory load and calculate the accuracy of playing and reproducing notes.

[0145] S101. Under conditions of sufficient lighting and no facial damage or concealment, collect 3-channel facial videos of the subject under high working memory load, at a shooting distance of 50 cm and a video duration of 20 seconds.

[0146] S102. Refer to the corresponding steps in Example 1, which will not be repeated here;

[0147] S103. Refer to the corresponding steps in Example 1, which will not be repeated here;

[0148] S104. Refer to the corresponding steps in Example 1, which will not be repeated here;

[0149] S105. Refer to the corresponding steps in Example 1, which will not be repeated here;

[0150] S2. Refer to the corresponding steps in Example 1, which will not be repeated here. The green channel signals for the three regions of interest are as follows: Figure 5 As shown, the denoised pulse wave signals of the three regions of interest are as follows: Figure 6 As shown;

[0151] S3. Refer to the corresponding steps in Example 1, which will not be repeated here. The pulse wave signal obtained by fusing the three regions of interest is as follows: Figure 7 As shown in Table 2, the frequencies corresponding to the maximum peak values ​​of each IMF component in the spectrum are shown in Table 3, and the signal-to-noise ratios and weights of the three regions of interest in camera channel 1 are shown in Table 3.

[0152] Table 2. Frequency table corresponding to the maximum peak value of the IMF component in Example 2.

[0153]

[0154] Table 3. Signal-to-noise ratio and weights of the three regions of interest of camera 1 in Example 2

[0155]

[0156] S4. Refer to the corresponding steps in Example 1, which will not be repeated here. The pulse wave signals of the three channels are as follows: Figure 8 As shown, the pulse wave signal obtained by fusing the three channels is as follows: Figure 9 As shown in the figure; the signal-to-noise ratio and weights of the three channels are shown in Table 4.

[0157] Table 4. Signal-to-noise ratio and weights of the three channels in Example 2

[0158]

[0159] S5. Refer to the corresponding steps in Example 1, which will not be repeated here. The pulse wave signal after narrowband pass filtering is as follows: Figure 10 As shown;

[0160] S6. Refer to the corresponding steps in Example 1, which will not be repeated here;

[0161] S7. Refer to the corresponding steps in Example 1, which will not be repeated here; the recognition results are shown in Table 5.

[0162] Table 5. Identification Results in Example 2

[0163]

[0164] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A method for assessing working memory load based on multi-channel facial video in a non-static state, characterized in that, The evaluation method includes the following steps: S1. Construct a test scenario, collect multi-channel facial videos under different working memory loads, and calculate the test accuracy. The different working memory load levels are set to three different levels: high load, medium load, and low load. The test accuracy refers to the accuracy of reproducing the content of the first c rounds after c ≤ 3 rounds of memory. In the test scenario, facial videos are collected by N cameras placed around the radius of the test subject's circle. The facial video collected by each camera is called one channel video. S2. For each channel's video, segment the video into frames, locate multiple regions of interest (ROIs), extract the original pulse wave signal for each ROI, and then denoise it. For each channel's facial video under different working memory loads, segment the facial video into frames to obtain a facial video image sequence, and then perform face recognition and facial feature point detection to locate multiple ROIs. Separate the color channels of each ROI and extract the average grayscale value of the green channel as the original pulse wave signal for that ROI. Then, preprocess the original pulse wave signal, perform variational mode decomposition, wavelet thresholding, and Butterworth bandpass filtering to denoise it. S3. Repeat step S2 to obtain the denoised pulse wave signal for each region of interest, and calculate the multi-harmonic signal-to-noise ratio and weights respectively. Then, the pulse wave signals of multiple regions of interest are fused according to the weights to obtain the pulse wave signal of each channel. The process of step S3 is as follows: S301. Calculate the pulse wave signal Y of the m-th region of interest using Fast Fourier Transform. m Multiharmonic signal-to-noise ratio (SNR) ROI (m), the calculation formula is as follows: Among them, F (m) (·) indicates that the pulse wave signal Y m The spectrum diagram, where i represents the frequency index of the spectrum diagram, Fs represents the sampling rate of the pulse wave signal, and a max The fundamental frequency is represented by A[k] = [a], where A[k] represents the frequency corresponding to the peak value in the spectrum. max ,2a max ,3a max ,..ka max ] indicates less than The multi-harmonic frequency array, where b represents the frequency corresponding to a point of adjacent fundamental frequency or adjacent harmonic; The multiharmonic signal-to-noise ratio (SNR) of the pulse wave signals in the M regions of interest is calculated. ROI (m); S302. Based on the signal-to-noise ratio of the M regions of interest, calculate the weights of the pulse wave signals in the M regions of interest using the following formula: Among them, SNR ROI (m) represents the signal-to-noise ratio of the pulse wave signal in the m-th region of interest, and M represents the number of regions of interest. The sum of the signal-to-noise ratios of the M regions of interest is represented, and r is an empirical constant, representing the nonlinear stretching of the weights by the stretching parameter. S303. The pulse wave signals of the M regions of interest are weighted and fused according to the above weights to obtain the pulse wave signal of a single channel. The calculation formula is as follows: Among them, IPPG cam This represents the pulse wave signal after fusion of individual channels; S4. Repeat steps S2 to S3 to obtain pulse wave signals of N channels. Calculate the multi-harmonic signal-to-noise ratio and weight of each channel pulse wave signal. Then, perform cross-correlation and translation alignment on the N channels pulse wave signals. Finally, weighted fuse the multi-channel pulse wave signals according to the weights to obtain the N-channel fused pulse wave signal. S5. Perform time-domain and frequency-domain analysis on the N-channel fused pulse wave signal to calculate heart rate and heart rate variability index; S6. Test several subjects with questions of different working memory load levels, repeat steps S1 to S5, obtain the subjects' test accuracy, heart rate and heart rate variability characteristic indicators, form the dataset required for model construction, and then perform feature elimination, feature selection and load level classification to construct a working memory load physiological model. S7. In the actual test of the working memory load of the test subjects, the test accuracy and the heart rate and heart rate variability indicators extracted from the multi-channel facial video after eliminating individual differences are input into the working memory load physiological model to assess the current working memory load level of the test subjects.

2. The method for assessing working memory load based on multi-channel facial video in a non-static state according to claim 1, characterized in that, The process of step S1 is as follows: S101. Under sufficient external lighting conditions, multi-channel facial videos under different working memory loads are collected by N cameras placed around the radius of the test subject's circle. The shooting distance is 50 cm, and the shooting durations for low load, medium load, and high load levels are 10 seconds, 15 seconds, and 20 seconds, respectively. S102. Display the memorized content on the screen, and increase the working memory load level by continuously increasing the amount of content the test subject recalls from previous rounds; starting from the 4th round, memorize the current round and write down the content from the previous 3 rounds; S103. For low load, recall the content of the previous wheel; for medium load, recall the content of the two wheels before the current wheel; for high load, recall the content of the three wheels before the current wheel. S104. Starting from the 4th round, record the content pressed by the test subject in the previous 3 rounds, and calculate the test accuracy rate by comparing it with the questions.

3. The method for assessing working memory load based on multi-channel facial video in a non-static state according to claim 1, characterized in that, The process of step S2 is as follows: S201. First, the acquired N-channel facial video is divided into frames to obtain a facial video image sequence; S202. Perform face recognition and facial feature point detection on the facial video image sequence, and locate multiple regions of interest based on the facial feature points; S203. Separate the color channels of each region of interest and extract the average gray value of the green channel as the original pulse wave signal of that region of interest; S204. Perform preprocessing on the extracted raw pulse wave, including detrending and Z-score normalization. S205. Perform variational mode decomposition on the preprocessed pulse wave signal to obtain multiple intrinsic mode function components. The intrinsic mode function is abbreviated as IMF. Perform fast Fourier transform on each IMF component to calculate the frequency corresponding to the maximum peak in the spectrum. Select the first λ high-frequency IMF components with higher frequencies. S206. Perform wavelet denoising on the selected λ high-frequency IMF components. S207. The high-frequency IMF component that has undergone wavelet denoising is superimposed with the remaining IMF component to reconstruct the pulse wave signal. S208. Perform Butterworth bandpass filtering on the reconstructed pulse wave signal to obtain M pulse wave signals Y representing regions of interest. m m = 1, 2, 3, ..., M.

4. The method for assessing working memory load based on multi-channel facial video in a non-static state according to claim 1, characterized in that, The process of step S4 is as follows: S401. Process the N-channel facial video using steps S2 to S3 to obtain N-channel pulse wave signals IPPG. cam (n), n=1, 2, 3,...,N; S402. Calculate the multiharmonic signal-to-noise ratio (SNR) of the nth channel pulse wave signal using Fast Fourier Transform. cam (n); The formula for calculating the multiharmonic signal-to-noise ratio of the pulse wave signal in this channel is as follows: Calculate the multi-harmonic signal-to-noise ratio (SNR) for N channels to obtain the SNR of the N channel pulse wave signals. cam (n); S403. Based on step S402, obtain the signal-to-noise ratio of the N channels, and then calculate the weights of the N channel pulse wave signals. The calculation formula is as follows: Among them, SNR cam (n) represents the signal-to-noise ratio of the pulse wave signal in the nth channel, and N represents the number of channels. This represents the sum of the signal-to-noise ratios of the N channels; S404. Cross-correlate and align the pulse wave signals of each channel; S405. The aligned pulse wave signals are weighted and fused according to the weights calculated in step S403 to obtain an N-channel fused pulse wave signal. The fusion calculation formula is as follows: Among them, IPPG cam (n) represents the pulse wave signal of the nth channel, and Signal represents the pulse wave signal after the fusion of N channels.

5. The method for assessing working memory load based on multi-channel facial video in a non-static state according to claim 1, characterized in that, The process of step S5 is as follows: S501. Convert the pulse wave signal Signal to the frequency domain, find the frequency corresponding to the highest peak of the amplitude spectrum, which is the number of heartbeats per second of the subject; multiply the obtained number of heartbeats per second by 60 to get the heart rate; S502. Perform narrowband pass filtering on the pulse wave signal Signal; S503. Perform cubic spline interpolation on the narrowband pass filtered pulse wave signal; S504. Detect the peak point of the pulse wave signal after cubic spline interpolation; S505. Based on the obtained time interval between adjacent peak points, calculate the heart rate variability index, which includes the standard deviation of heartbeat interval SDNN, the standard deviation of the difference between adjacent heartbeat intervals SDSD, the root mean square of the difference between adjacent heartbeat intervals RMSSD, the mean heart rate Mean_HR, the maximum heart rate Max_HR, the minimum heart rate Min_HR, the standard deviation of heart rate STD_HR, the total signal power TP, the low-frequency power LF, the high-frequency power HF, the ratio of low-frequency power to high-frequency power LF / HF, and the very low-frequency power VLF.

6. The method for assessing working memory load based on multi-channel facial video in a non-static state according to claim 1, characterized in that, The process of step S6 is as follows: S601. Collect multi-channel facial videos of several subjects under different working memory loads and calculate the test accuracy. S602. Repeat steps S2 to S5 to extract heart rate and heart rate variability characteristic indicators; S603. Eliminate individual differences in heart rate and heart rate variability indicators; S604. Combine the heart rate and heart rate variability feature indicators after individual differences have been eliminated with the test accuracy to form the feature dataset required to build the model, and label the feature dataset, where the low load label is 0, the medium load label is 1, and the high load label is 2. S605. Treat each indicator as a feature, and through feature selection, eliminate irrelevant and redundant features from the feature dataset to form the optimal feature data subset. S606. Divide the feature data subset obtained in step S605 into a training set and a test set; wherein, the training set is used to train the random forest model, and the working memory load physiological classification model is obtained by training the input training set and labels; the test set is used to verify the performance indicators of the model. S607. Construct a physiological classification model for working memory load. First, train the model based on the training set data. Then, use a classifier to classify the data and obtain the hyperparameters of the classifier through K-fold cross-validation. Finally, obtain the physiological classification model for working memory load. S608. Verify the recognition accuracy of the working memory load physiological classification model using the test set.

7. The method for assessing working memory load based on multi-channel facial video in a non-static state according to claim 1, characterized in that, The process of step S7 is as follows: S701. Collect multi-channel facial videos of the test subjects and calculate the test accuracy rate; S702. Heart rate and heart rate variability characteristic indicators are obtained through steps S2 to S5. S703. Input the test accuracy, heart rate, and heart rate variability into the working memory load physiological classification model to assess the working memory load level of the test subjects.

Citation Information

Patent Citations

  • Working memory load level identification method based on face video

    CN114983415A

  • Video acquisition system and method for monitoring a subject for a desired physiological function

    US20140378842A1