A task-driven fatigue data collection method in an intelligent interaction scenario
By designing progressive fatigue induction steps and a multimodal verification mechanism in intelligent interaction scenarios, the problems of lag and low label confidence in existing visual fatigue detection datasets are solved, and a high-confidence visual fatigue detection dataset suitable for intelligent interaction scenarios is constructed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
- Filing Date
- 2026-06-09
- Publication Date
- 2026-07-10
AI Technical Summary
Existing visual fatigue detection datasets struggle to capture early cognitive fatigue in intelligent interaction scenarios, and the labels lack objective validation, resulting in low confidence levels.
The design incorporates a progressive active fatigue induction process that includes visual load, working memory load, and eye-tracking search load. It combines subjective scale scores, objective heart rate features, and eye-tracking features with multimodal cross-validation to construct a pure visual modality fatigue detection dataset.
It accurately captures facial microscopic eye movement features in the early and middle stages of fatigue, and constructs a high-confidence visual fatigue detection dataset with rigorous physiological benchmarks, which is suitable for lightweight deployment.
Smart Images

Figure CN122365100A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of human-computer interaction and computer vision technology, specifically to a task-driven fatigue data acquisition method in intelligent interaction scenarios, which is particularly suitable for constructing a high-confidence pure visual fatigue detection dataset. Background Technology
[0002] In the field of human-computer interaction, the user's state directly affects the quality and efficiency of the interaction. With the widespread use of smart devices, users often experience fatigue when interacting with these devices for extended periods, leading to decreased attention, slower reaction times, and reduced decision-making abilities, significantly impacting the overall user experience. Especially in tasks requiring high concentration, the accumulation of latent fatigue can easily lead to misoperations and may also pose health and safety risks. Therefore, how to monitor user fatigue in real-time and seamlessly has become a key research direction for improving the quality of smart device interactions.
[0003] Traditional fatigue detection methods mainly include two categories: subjective evaluation and objective detection. Subjective evaluation usually relies on drowsiness scales or fatigue scales, such as the subjective drowsiness evaluation method proposed by Åkerstedt and Gillberg (Åkerstedt T, Gillberg M. Subjective and objective sleepiness in the active individual[J]. International Journal of Neuroscience, 1990, 52(1-2): 29-37), and the fatigue assessment scale proposed by Lee et al. (LeeK A, Hicks G, Nino-Murcia G. Validity and reliability of a scale to assess fatigue[J]. Psychiatry Research, 1991, 36(3): 291-298). These methods are simple to operate, but rely on active user feedback and have poor real-time performance.Objective detection methods mainly utilize physiological signals such as EEG and EOG. For example, Zheng and Lu (Zheng WL, Lu BL. A multimodal approach to estimating vigilance using EEG and forehead EOG[J]. Journal of Neural Engineering, 2017, 14(2): 026017) estimated alertness levels based on EEG and forehead EOG. Cao et al (Cao Z, Chuang CH, King JK, et al. Multi-channel EEG recordings during asustained-attention driving task[J]. Scientific Data, 2019, 6(1): 19) published multi-channel EEG data for continuous driving tasks. Massoz et al (Massoz Q, Langohr T, François C, et al. The ULg multimodality drowsiness database (called DROZY) and examples of use[C] / / 2016 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, (2016: 1-7) A DROZY multimodal fatigue dataset containing video and physiological signals was constructed. This type of method can reflect fatigue state relatively accurately, but it usually requires wearing sensors, which are complex and somewhat invasive, making it difficult to widely apply to everyday lightweight human-computer interaction scenarios.
[0004] To lower the application threshold, visual fatigue detection based on camera video has gradually attracted attention. For example, Ghoddoosian et al. (Ghoddoosian R, Galib M, Athitsos V. A realistic dataset and baseline temporal model for early drowsiness detection[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition Workshops.2019: 0-0) proposed the RLDD dataset for fatigue detection in real-world scenarios, while Abtahi et al. (Abtahi S, Omidyeganeh M, Shirmohammadi S, et al. YawDD: A yawning detection dataset[C] / / Proceedings of the 5th ACM multimedia systems conference. 2014: 24-28) proposed the YawDD dataset, which is mainly used for yawning behavior detection. However, existing visual fatigue detection datasets and acquisition methods have the following significant limitations in practical applications: Behavioral representation suffers from severe lag, making it difficult to capture early cognitive fatigue. Existing visual fatigue detection datasets often rely on significant, overt behaviors during collection and annotation. For example, NTHU-DDD (Weng CH, Lai YH, Lai SH. Driver drowsiness detection via a hierarchical temporal deep belief network[C] / / Asian Conference on Computer Vision. Cham: Springer, 2016: 117-133.) is a video dataset for driver drowsiness detection, including typical drowsy behaviors such as yawning, closing eyes, and nodding; YawDD is mainly used for detecting driver yawning, further demonstrating the dependence of existing datasets on strong overt behaviors. However, in real-world intelligent interaction scenarios, when users exhibit such dramatic movements, they are often already in a state of extreme fatigue. This collection paradigm, which relies solely on exaggerated behaviors for judgment, ignores the gradual changes in eye movement features and cannot represent "latent cognitive fatigue" in the early to mid-stages.
[0005] The labels lack objective validation and have extremely low confidence. Some existing datasets, such as RLDD, mainly rely on subjects recording videos themselves in specific states and subjectively reporting their fatigue levels as data labels. This method, based on purely subjective feedback, is divorced from medical cross-validation using objective physiological indicators and is highly susceptible to individual perceptual differences, emotions, and psychological suggestions, thus introducing label noise into the dataset and limiting the fitting accuracy and robustness of deep learning models. Summary of the Invention
[0006] Purpose of the Invention: To address the problems existing in the above-mentioned background technology, this invention provides a task-driven fatigue data acquisition method in intelligent interactive scenarios, aiming to solve the following technical problems:
[0007] First, in response to the problem that existing methods have lagging behavioral representation and difficulty in capturing early cognitive fatigue, this invention designs a progressive active fatigue induction step that includes visual load, working memory load and eye-tracking search load, gradually increasing the cognitive consumption of the subject to obtain the facial visual features of the subject in the early and middle stages of fatigue, thus overcoming the limitation of traditional methods that can only detect severe fatigue.
[0008] Second, addressing the issues of existing datasets lacking objective validation and having low confidence levels, this invention proposes a triple-subjective and objective fatigue induction validity validation mechanism. During the data acquisition phase, the subject's subjective scale scores, objective heart rate characteristics, and eye movement characteristics are acquired simultaneously. Fatigue labels are assigned to visual feature sequences only when all three characteristics pass the consistency test. This invention uses multimodal cross-validation for data filtering, ultimately constructing a pure visual modality fatigue detection dataset supported by objective physiological benchmarks.
[0009] Technical solution: To achieve the above objectives, the technical solution adopted by this invention is as follows:
[0010] A task-driven fatigue data acquisition method in an intelligent interactive scenario includes the following steps:
[0011] Step S1: Obtain the subject's subjective fatigue scale score at baseline, and collect facial video sequences and continuous heart rate signals during the subject's performance of the visual interaction task;
[0012] Step S2: Perform progressive active fatigue induction on the subject by sequentially applying visual load, cognitive and flicker load, and eye movement and search load to induce fatigue in the subject.
[0013] Step S3: After fatigue induction, the subject's subjective fatigue scale score is obtained again, and facial video sequences and continuous heart rate signals are collected during the subject's performance of the same visual interaction task as in step S1.
[0014] Step S4: Extract subjective fatigue scale scores, heart rate features, and facial visual features including eye movement features from the data collected in Step S1 and Step S3, respectively.
[0015] Step S5: Compare and analyze the subjective scores, heart rate characteristics, and eye movement characteristics obtained in Step S1 and Step S3 to perform a triple verification of the effectiveness of subjective and objective fatigue induction.
[0016] Step S6: If the verification in step S5 is valid only, mark the facial video sequence acquired in step S1 as a conscious state and the facial video sequence acquired in step S3 as a fatigued state, and output a fatigue detection dataset in pure visual modality.
[0017] Preferably, the method further includes setting a multi-period data collection plan to obtain the subject's baseline sleep quality indicators. The multi-period data collection plan includes: a morning node: assessing the subject's sleep quality index the previous night using a sleep quality questionnaire; and morning and afternoon nodes: conducting two complete fatigue induction and data collection processes each day, in the morning and afternoon respectively, to capture the evolution of the subject's fatigue performance at different time periods. This baseline sleep quality indicator is used to eliminate extreme abnormal physiological baselines caused by severe sleep deprivation in subsequent analyses.
[0018] Preferably, the detailed steps of the data acquisition process in steps S1 and S3 are as follows:
[0019] Subjective questionnaire collection: Subjects first completed a subjective fatigue questionnaire, which specifically included the Visual Analog Scale (VAS-F), the Visual Fatigue Scale, and the Visual Fatigue Subjective Evaluation Scale (VFSES) for assessing overall fatigue and energy levels.
[0020] Performing a visual interaction task (“Z-Shape” gaze task): After completing the questionnaire, under controlled conditions with sufficient ambient light and no direct glare, and with the eyes about 60 cm away from the interactive screen, the subjects visually followed the “Z” shaped movement trajectory of the visual target on the screen within a set time.
[0021] Synchronous multimodal data acquisition: During the process of the subject performing the visual interaction task, the system synchronously records the subject's facial image video to extract eye movement features, and uses wearable physiological sensors to synchronously and continuously record the subject's continuous heart rate signal and adjacent heartbeat interval data.
[0022] Preferably, the detailed steps of the progressive active fatigue induction in step S2 are as follows:
[0023] Step S2.1, Visual Load Induction Stage: Subjects watch a 3D video sequence, which increases binocular convergence and accommodation load, reduces blinking frequency, and induces local visual fatigue and dry eye symptoms.
[0024] Step S2.2, Mild Cognitive and Flicker Load Induced Phase: Subjects perform a 1-Back working memory task. This task is divided into a flicker-free adaptation period and a flicker recording period on the timeline to gradually increase the load; in a spatial grid, target characters are presented alternately at intervals with additional random jitter; subjects need to perform immediate temporal memory retrieval, judge whether the current character is consistent with the previous character and make key input feedback;
[0025] Step S2.3, Eye Movement and Search Load Induction Phase: Subjects perform a visual search task. Within a set spatial grid, the system presents a visual array containing a single target item and multiple distracting items; subjects must quickly complete the visual search and click to locate the target item before it disappears.
[0026] Preferably, the specific steps of feature extraction in step S4 are as follows:
[0027] Step S4.1, Subjective scale scoring: Subjective scores were obtained and calculated using VAS-F, visual fatigue scale and VFSES.
[0028] Step S4.2, Heart Rate Characteristic Calculation:
[0029] Extracting artifacts after artifact removal and filtering The sequence of adjacent heartbeat intervals (RR intervals) is denoted as the nth interval. The duration of each interval is Calculate the average heartbeat interval :
[0030]
[0031] in: This represents the total number of valid RR intervals after artifact removal; Represents the first in the sequence The duration of each RR interval, in milliseconds.
[0032] The standard deviation of the normal RR interval (SDNN) reflects the overall variability of autonomic nervous system regulation.
[0033]
[0034] in: This represents the total number of valid RR intervals after artifact removal; Represents the first in the sequence The duration of each RR interval, in milliseconds. Representing all The average of the effective RR intervals.
[0035] The root mean square (RMSSD) of the difference between adjacent RR intervals reflects the level of parasympathetic activity.
[0036]
[0037] in: This represents the total number of valid RR intervals after artifact removal; Represents the first in the sequence The duration of each RR interval, in milliseconds. Representing the adjacent number The 1st RR interval after the 1st RR interval The duration of each RR interval.
[0038] Percentage of adjacent RR intervals exceeding 50 milliseconds (pNN50):
[0039]
[0040] in: This represents the total number of valid RR intervals after artifact removal; This represents the number of times the absolute value of the difference between two adjacent RR intervals exceeds 50 milliseconds.
[0041] Step S4.3, Eye movement feature calculation:
[0042] The system extracts global facial feature points based on facial video streams using a facial keypoint detection algorithm. Building upon this, it further extracts the coordinates of six core key points of the eye contour, which are denoted as follows: to And calculate the eye aspect ratio (EAR) based on the coordinates of the key points:
[0043]
[0044] in, This represents the Euclidean distance between the two-dimensional coordinates of two key points. and The key point is the horizontal corner of the eye. , This is a key point on the upper eyelid. , This is a key point on the lower eyelid.
[0045] Set continuous time window The total number of valid blinks can be detected by determining whether the average EAR value of consecutive frames falls below a preset threshold (which can be set to 50% of the average EAR value; when the EAR is less than this number, it is determined as closed eyes, and when it is greater than this number, it is determined as open eyes). Passing the exam Duration of one blink Further calculations were performed on the following indicators:
[0046] blink rate per unit time :
[0047]
[0048] in, This represents the set length of the continuous observation time window; Represents the time window The total number of valid blinks detected by the internal system;
[0049] Average blink time :
[0050]
[0051] in, Represents the time window The total number of valid blinks detected by the internal system; Representing the The duration of a blink;
[0052] Standard deviation of blink duration :
[0053]
[0054] in, Represents the time window The total number of valid blinks detected by the internal system; Representing the The duration of a blink; The average blink duration within a time window;
[0055] Perclosing ratio (PERCLOS):
[0056]
[0057] in, This represents the total duration of a set test segment; Representative at Within, the total cumulative time during which the eye aspect ratio (EAR value) is below a set threshold.
[0058] Preferably, the specific implementation method for verifying the effectiveness of the subjective and objective triple fatigue induction in step S5 is as follows:
[0059] Step S5.1: Assess the fatigue trend of subjective scale scores: Compare and analyze the subjective fatigue scale scores in the baseline stage before the task and the control stage after the task, and calculate the change in scores.
[0060] Step S5.2: Assess objective heart rate characteristics and fatigue trends: Compare and analyze the objective heart rate characteristics obtained in the two stages to assess the regulatory changes in sympathetic nerve activation and parasympathetic nerve inhibition.
[0061] Step S5.3: Assess the fatigue trend of objective eye movement features: Compare and analyze the eye movement features extracted from the facial video in the two stages, and assess changes in blink frequency, blink duration and eyelid closure ratio;
[0062] Step S5.4: Perform consistency test and judgment: The progressive active fatigue induction of the sample is deemed effective only when the subjective scale score, objective heart rate characteristics, and objective eye movement characteristics all indicate a significant increase in fatigue level. Specifically, fatigue induction is deemed effective only when the subjective fatigue scale score increases by more than a preset threshold after the task compared to before the task, the average heart rate increases and the percentage of adjacent RR interval differences exceeding 50 milliseconds decreases, the blink frequency increases and the average blink duration increases.
[0063] Furthermore, the allocation of state labels and the construction of the dataset in step S6 are as follows:
[0064] Step S6.1: Label the facial video sequence acquired in step S1 as a conscious state tag;
[0065] Step S6.2: Label the facial video sequence acquired in step S3 as a fatigue state tag;
[0066] Step S6.3: Output the feature sequences with control labels as a pure visual modality fatigue detection dataset. This dataset only contains facial video sequences and corresponding awake or fatigued state labels, and does not contain continuous heart rate signal data.
[0067] Beneficial effects:
[0068] (1) This invention increases the continuous cognitive consumption of subjects through a progressive active fatigue induction step that includes visual load, working memory, and eye movement search, while accurately capturing the facial micro-eye movement features in the early and middle stages of fatigue. Previous methods relied heavily on subjects' overt behaviors such as yawning and extremely slow blinking, which may cause serious representational lag in actual intelligent interaction scenarios, resulting in missed opportunities for fatigue warning and failure of prevention. This invention reduces the lag in feature detection through a composite induction strategy that continuously increases cognitive load.
[0069] (2) By introducing a subjective and objective triple fatigue induction effectiveness verification mechanism, this invention enables the final constructed pure visual feature sequence to have rigorous objective physiological benchmark support. Existing visual data acquisition mechanisms mostly use subjective self-report as the label basis, which has a certain effect in simple scenarios with obvious physiological fatigue, but has great limitations in complex intelligent interaction tasks and is easily affected by individual perceptual differences, resulting in a large number of false labels. The multimodal cross-validation of this invention makes the pure visual fatigue dataset suitable for wide-ranging lightweight deployment and also has extremely high label confidence. Attached Figure Description
[0070] Figure 1 This is an overall flowchart of a task-driven fatigue data acquisition method in an intelligent interactive scenario according to the present invention.
[0071] Figure 2 This is a schematic diagram of the overall facial feature points extracted based on video stream in an embodiment of the present invention;
[0072] Figure 3 This is a schematic diagram of the coordinates of six key points of the eye contour used to calculate the aspect ratio of the eye in an embodiment of the present invention. Detailed Implementation
[0073] The present invention will be further described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0074] like Figure 1 As shown, the present invention provides a task-driven fatigue data acquisition method in an intelligent interactive scenario. The overall process includes: setting the test environment and multi-time period acquisition plan, baseline data acquisition before the task, execution of three-stage progressive active fatigue induction, post-task control data acquisition, multimodal feature extraction and calculation, subjective and objective triple validity verification, label allocation and output of pure visual modality fatigue detection dataset.
[0075] The preliminary steps include setting up the testing environment and a multi-time period data collection plan. Specifically, this includes: setting up a controlled interactive environment, ensuring sufficient lighting without direct glare, and maintaining the screen approximately 60 cm from the subject's eyes. The multi-time period data collection plan involves: in the morning session, a short sleep quality questionnaire is used to assess the subject's sleep quality the previous night. This step aims to obtain baseline sleep quality indicators to eliminate extreme abnormal physiological baselines caused by severe sleep deprivation in subsequent analyses. In the morning and afternoon sessions, two complete fatigue induction and data collection processes are conducted each day, one in the morning and one in the afternoon, to capture the evolution of the subject's fatigue performance at different time periods.
[0076] Next, perform the following steps S1 to S6 in sequence.
[0077] Step S1: Perform baseline data acquisition before fatigue induction: Obtain the subject's subjective fatigue scale score while awake, and collect facial video sequences and continuous heart rate signals during the subject's performance of the visual interaction task. This includes the following steps:
[0078] Step S1.1, Subjective Questionnaire Collection: Subjects first complete a subjective fatigue questionnaire, which includes the Visual Analogue Scale (VAS-F), the Visual Fatigue Scale, and the Visual Fatigue Subjective Evaluation Scale (VFSES), used to assess overall fatigue and energy levels.
[0079] Step S1.2: Perform the visual interaction task (“Z-Shape” gaze task): After completing the questionnaire, under controlled conditions, the subject performs visual tracking scans according to the “Z” shaped motion trajectory of the visual target on the screen.
[0080] Step S1.3, Synchronous Multimodal Data Acquisition: During the process of the subject performing the visual interaction task, the system synchronously records the subject's facial image video at a frame rate of 30fps to extract eye movement features, and uses a wearable physiological sensor (Polar heart rate chest belt) to synchronously and continuously record the subject's heart rate signal and adjacent heartbeat interval data.
[0081] Step S2: Perform progressive active fatigue induction. This involves designing multiple interactive tasks with different load characteristics, and then performing a three-stage progressive active fatigue induction process, from simple to complex. Specifically, this includes:
[0082] Step S2.1, Visual Load Induction Stage: Subjects watch a 3D video sequence, which increases binocular convergence and accommodation load, reduces blinking frequency, and induces local visual fatigue and dry eye symptoms.
[0083] Step S2.2, Mild Cognitive and Flicker Load Induced Phase: Subjects perform a 1-Back working memory task. This task is divided into a flicker-free adaptation period and a flicker recording period on the timeline; in a spatial grid, target characters are presented alternately at time intervals with additional random jitter; subjects need to perform immediate temporal memory retrieval, determine whether the current character is consistent with the previous character, and make a key press response.
[0084] Step S2.3, Eye Movement and Search Load Induction Phase: Subjects perform a visual search task. Within a set spatial grid, the system presents a visual array containing a single target item and multiple distracting items; subjects must quickly complete the visual search and click to locate the target item before it disappears.
[0085] Step S3: Perform post-fatigue induced control data collection: Obtain the subject's subjective score in a fatigued state, and simultaneously record the subject's facial video sequence and continuous heart rate signal during the process of performing the same visual interaction task as in step S1.
[0086] Step S4: Extract subjective scale scores, heart rate features, and eye movement features obtained at the pre-task baseline and post-task comparison stages, respectively:
[0087] Step S4.1, Subjective scale scoring: Subjective scores were obtained and calculated using VAS-F, visual fatigue scale and VFSES.
[0088] Step S4.2, Heart Rate Characteristic Calculation:
[0089] Extracting artifacts after artifact removal and filtering The sequence of adjacent heartbeat intervals (RR intervals) is denoted as the nth interval. The duration of each interval is Calculate the average heartbeat interval :
[0090]
[0091] in: This represents the total number of valid RR intervals after artifact removal; Represents the first in the sequence The duration of each RR interval, in milliseconds.
[0092] The standard deviation of the normal RR interval (SDNN) reflects the overall variability of autonomic nervous system regulation.
[0093]
[0094] in: This represents the total number of valid RR intervals after artifact removal; Represents the first in the sequence The duration of each RR interval, in milliseconds. Representing all The average of the effective RR intervals.
[0095] The root mean square (RMSSD) of the difference between adjacent RR intervals reflects the level of parasympathetic activity.
[0096]
[0097] in: This represents the total number of valid RR intervals after artifact removal; Represents the first in the sequence The duration of each RR interval, in milliseconds. Representing the adjacent number The 1st RR interval after the 1st RR interval The duration of each RR interval.
[0098] Percentage of adjacent RR intervals exceeding 50 milliseconds (pNN50):
[0099]
[0100] in: This represents the total number of valid RR intervals after artifact removal; This represents the number of times the absolute value of the difference between two adjacent RR intervals exceeds 50 milliseconds.
[0101] Step S4.3, Eye movement feature calculation:
[0102] The system extracts global facial feature points (such as...) based on facial video sequences using facial key point detection algorithms. Figure 2 As shown); based on this, the coordinates of 6 core key points of the eye contour are further extracted (e.g., Figure 3 As shown), these are denoted as P1 to P6 respectively. The eye aspect ratio (EAR) is then calculated based on the coordinates of these key points.
[0103]
[0104] in, P1 and P6 represent the Euclidean distance between the two-dimensional coordinates of two key points, where P1 and P6 are horizontal eye corner key points. , This is a key point on the upper eyelid. , This is a key point on the lower eyelid.
[0105] Set continuous time window The total number of valid blinks can be detected by determining whether the average EAR value of consecutive frames falls below a preset threshold (which can be set to 50% of the average EAR value; when the EAR is less than this number, it is determined as closed eyes, and when it is greater than this number, it is determined as open eyes). Passing the exam Duration of one blink Further calculations were performed on the following indicators:
[0106] blink rate per unit time :
[0107]
[0108] in, This represents the set length of the continuous observation time window; Represents the time window The total number of valid blinks detected by the internal system;
[0109] Average blink time :
[0110]
[0111] in, Represents the time window The total number of valid blinks detected by the internal system; Representing the The duration of a blink;
[0112] Standard deviation of blink duration :
[0113]
[0114] in, Represents the time window The total number of valid blinks detected by the internal system; Representing the The duration of a blink; The average blink duration within a time window;
[0115] Perclosing ratio (PERCLOS):
[0116]
[0117] in, This represents the total duration of a set test segment; Representative at Within, the total cumulative time during which the eye aspect ratio (EAR value) is below a set threshold.
[0118] Step S5: Compare and analyze the extracted three-dimensional features to verify their effectiveness in inducing fatigue:
[0119] Step S5.1: Assess the fatigue trend of subjective scale scores: Calculate the change in subjective scores after the task compared to before the task. For example, in a valid induced sample, the Visual Analogue Scale (VAS-F) score significantly increased from approximately 30.5 points before the task to approximately 44.8 points (a higher score indicates greater fatigue), and both the Visual Analogue Scale and VFSES scores showed statistically significant increases, indicating a significant increase in the subject's perceived fatigue level.
[0120] Step S5.2: Assess Objective Heart Rate Characteristics and Fatigue Trends: Compare objective heart rate and heart rate variability characteristics before and after the task to assess the sympathetic activation and parasympathetic inhibition phenomena induced by increased cognitive load. For example, under effective induction, the subjects' average heart rate significantly increased from approximately 76 bpm at baseline to approximately 80 bpm; simultaneously, heart rate variability (HRV) indicators reflecting cardiac rhythm stability showed a consistent decrease, such as a decrease of approximately 5 milliseconds in the root mean square difference of adjacent RR intervals (RMSSD), and significant reductions in both SDNN and pNN50. These systemic physiological changes of "increased heart rate and decreased HRV" objectively confirm that the subjects had entered a state of cognitive fatigue.
[0121] Step S5.3: Assess objective eye movement characteristics of fatigue trends: Extract and compare eye movement characteristic parameters from facial videos of the two phases. For example, as task load and fatigue intensify, the subjects' blinking frequency per unit time significantly increases (e.g., from approximately 8.8 times per minute to 11.4 times per minute), the average blink duration significantly prolongs (e.g., an increase of approximately 14 milliseconds), and both the standard deviation of blink duration and the per-eyelid closure ratio (PERCLOS) show an increasing trend. These changes in eye behavior patterns provide direct visual behavioral evidence for fatigue induction.
[0122] Step S5.4: Perform consistency test and judgment: The progressive active fatigue induction of the sample is deemed effective only when the subjective scale score, objective heart rate characteristics, and objective eye movement characteristics all indicate a significant increase in fatigue level. Specifically, fatigue induction is deemed effective only when the subjective fatigue scale score increases by more than a preset threshold after the task compared to before the task, the average heart rate increases and the percentage of adjacent RR interval differences exceeding 50 milliseconds decreases, the blink frequency increases and the average blink duration increases.
[0123] Step S6: After successful verification, remove physiological sensor data such as heart rate used for cross-validation and construct a pure visual modal fatigue detection dataset.
[0124] Step S6.1: Label the facial video sequence acquired in step S1 as "awake state";
[0125] Step S6.2: Label the facial video sequence acquired in step S3 as "fatigue state";
[0126] Step S6.3: Output the feature sequences with control labels as a pure visual modality fatigue detection dataset. This dataset only contains facial video sequences and corresponding awake or fatigued state labels, and does not contain continuous heart rate signal data. This dataset can be directly used for training and testing deep learning visual network models.
Claims
1. A task-driven fatigue data acquisition method in an intelligent interactive scenario, characterized in that, include: Step S1: Obtain the subject's subjective fatigue scale score at baseline, and collect facial video sequences and continuous heart rate signals during the subject's performance of the visual interaction task; Step S2: Perform progressive active fatigue induction on the subject by sequentially applying visual load, cognitive and flicker load, and eye movement and search load to induce fatigue in the subject; Step S3: After fatigue induction, the subject's subjective fatigue scale score is obtained again, and facial video sequences and continuous heart rate signals are collected during the subject's performance of the same visual interaction task as in Step S1. Step S4: Extract subjective fatigue scale scores, heart rate features, and facial visual features including eye movement features from the data collected in Step S1 and Step S3, respectively. Step S5: Compare and analyze the subjective scores, heart rate characteristics, and eye movement characteristics obtained in Step S1 and Step S3 to verify the effectiveness of the subjective and objective triple fatigue induction. Step S6: Only if the verification in step S5 is valid, mark the facial video sequence acquired in step S1 as a conscious state, mark the facial video sequence acquired in step S3 as a fatigued state, and output a fatigue detection dataset in pure visual modality.
2. The method according to claim 1, characterized in that, The method also includes setting a multi-segment data collection plan to obtain baseline sleep quality indicators of the subjects, wherein the multi-segment data collection plan includes: Morning node: Assess the subject's sleep quality index of the previous night using a sleep quality questionnaire; Morning and afternoon nodes: Two complete fatigue induction and data collection processes were conducted each day in the morning and afternoon to capture the evolution of subjects' fatigue performance at different time periods.
3. The method according to claim 1, characterized in that, The visual interaction task described in steps S1 and S3 is as follows: In a controlled environment, the subject performs visual following scans according to the "Z"-shaped motion trajectory of the visual target on the screen.
4. The method according to claim 3, characterized in that, The controlled environment includes: sufficient ambient light without direct glare, and the subject's eyes being 60 centimeters away from the interactive screen.
5. The method according to claim 1, characterized in that, The progressive active fatigue induction described in step S2 specifically includes: visual load induction: watching a 3D video sequence; cognitive and flicker load induction: performing a 1-Back working memory task; and eye-tracking and search load induction: performing a visual search task.
6. The method according to claim 1, characterized in that, The subjective fatigue scale scoring in step S4 includes: a visual analog scale for assessing overall fatigue and energy levels, a visual fatigue scale, and a subjective evaluation scale for visual fatigue.
7. The method according to claim 1, characterized in that, The heart rate characteristics described in step S4 include: average heart rate, standard deviation of normal RR interval, root mean square of the difference between adjacent RR intervals, and percentage of adjacent RR interval differences exceeding 50 milliseconds.
8. The method according to claim 1, characterized in that, The eye movement features mentioned in step S4 include: eye aspect ratio, blink frequency per unit time, average blink duration, standard deviation of blink duration, and eyelid closure ratio.
9. The method according to claim 1, characterized in that, The effectiveness verification of the subjective and objective triple fatigue induction in step S5 is as follows: fatigue induction is deemed effective only when the subjective fatigue scale score in step S3 increases by more than a preset threshold compared to step S1, the average heart rate increases and the percentage of the root mean square difference of the difference between adjacent RR intervals or the difference between adjacent RR intervals exceeding 50 milliseconds decreases, the blinking frequency increases and the average blinking duration increases.
10. The method according to claim 1, characterized in that, The pure visual modal fatigue detection dataset output in step S6 only contains facial video sequences and corresponding awake or fatigued state labels, and does not contain continuous heart rate signal data.