Cognitive load evaluation method based on multi-modal data and edge computing device

By using multimodal data fusion and edge computing devices, personalized cognitive load assessment and real-time interface adjustment in complex HCI scenarios were achieved, solving the problems of inaccurate assessment and insufficient closed-loop intervention in existing technologies, and improving operational efficiency and security.

CN121997117APending Publication Date: 2026-05-08KINGFAR INTERNATIONAL INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KINGFAR INTERNATIONAL INC
Filing Date
2025-12-25
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies for assessing cognitive load in complex HCI scenarios suffer from insufficient individual variability and dynamism, leading to inaccurate assessment results, inability to achieve effective closed-loop intervention, and impacting operational efficiency and safety.

Method used

A multimodal data fusion method is adopted, including human physiological signals, behavioral data and interface feature data, to conduct personalized cognitive load assessment through edge computing devices, construct a personalized cognitive baseline and adjust the human-computer interaction interface in real time, so as to achieve accurate quantification and dynamic response of cognitive load.

Benefits of technology

It improves the accuracy and efficiency of cognitive load assessment, reduces operational risks, and enhances the adaptability and security of the user interface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997117A_ABST
    Figure CN121997117A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cognitive load evaluation method based on multi-modal data and edge computing equipment. The method comprises the following steps: determining individual cognitive baseline data of a user; determining a personalized cognitive load grading threshold value of the user according to the individual cognitive baseline data of the user; obtaining multi-modal data of the user in the working environment, wherein the multi-modal data comprises human body physiological signal data and human body behavior data of the user and interface feature data of a human-computer interaction interface; and determining a cognitive load evaluation result of the user according to the multi-modal data and the personalized cognitive load grading threshold value of the user. According to the technical scheme provided by the embodiment of the invention, cognitive load evaluation can be effectively carried out, and the operation efficiency and safety are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction (HCI) technology, and more particularly to a cognitive load assessment method and edge computing device based on multimodal data. Background Technology

[0002] With the rapid development of smart terminals and complex HCI interfaces, scenarios such as smart cockpits, virtual reality (VR) / augmented reality (AR) immersive systems, medical surgical robots, and industrial central control rooms are becoming increasingly common. These HCI interfaces typically feature multimodal (e.g., visual, auditory, and tactile fusion) interaction, high information density (e.g., parallel display of multi-source data), and dynamic task switching (e.g., simultaneously handling navigation and hazard avoidance while driving). Users need to continuously allocate their attention to process information and make decisions during operation, and the resulting cognitive load directly affects operational efficiency and safety.

[0003] Excessive cognitive load can lead to delayed user response and decision-making errors (such as pilots misinterpreting instrument data), while insufficient cognitive load may cause distraction (such as monitoring personnel missing abnormal signals). How to effectively assess cognitive load to improve operational efficiency and safety is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a cognitive load assessment method and an edge computing device based on multimodal data, which can effectively assess cognitive load and improve operational efficiency and security.

[0005] On one hand, embodiments of the present invention provide a cognitive load assessment method based on multimodal data, including: Determine the user's individual cognitive baseline data; The user's personalized cognitive load grading threshold is determined based on the user's individual cognitive baseline data; Acquire multimodal data of the user in the work environment, including the user's human physiological signal data, human behavior data, and interface feature data of the human-computer interaction interface; The cognitive load assessment result of the user is determined based on the multimodal data and the user's personalized cognitive load classification threshold.

[0006] On the other hand, embodiments of the present invention provide a computer-readable storage medium including a stored program, wherein the program controls the device where the computer-readable storage medium is located to execute the above-described method when it is executed.

[0007] On the other hand, embodiments of the present invention provide an edge computing device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions, wherein the program instructions are loaded and executed by the processor to implement the steps of the above method.

[0008] The technical solution provided in this invention involves: determining the user's individual cognitive baseline data; determining the user's personalized cognitive load grading threshold based on the user's individual cognitive baseline data; acquiring multimodal data of the user in their work environment, including the user's human physiological signal data, human behavioral data, and interface feature data of the human-computer interaction interface; and determining the user's cognitive load assessment result based on the multimodal data and the user's personalized cognitive load grading threshold. The technical solution provided in this invention can effectively assess cognitive load, improving operational efficiency and safety. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram of a cognitive load assessment system based on multimodal data provided in an embodiment of the present invention; Figure 2 A flowchart illustrating a cognitive load assessment method based on multimodal data, provided in an embodiment of the present invention; Figure 3 A flowchart for determining a user's individual cognitive baseline data provided in an embodiment of the present invention; Figure 4 A flowchart illustrating how to update a general cognitive baseline to obtain a user's individual cognitive baseline data, as provided in an embodiment of the present invention; Figure 5 A flowchart for determining a user's cognitive load assessment result based on multimodal data and the user's personalized cognitive load classification threshold, provided as an embodiment of the present invention; Figure 6 This is a schematic diagram of an edge computing device provided in an embodiment of the present invention. Detailed Implementation

[0011] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0012] It should be understood that the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0013] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0014] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0015] In the process of practicing the embodiments of this disclosure, the inventors of this disclosure have found that cognitive load assessment technology has gradually evolved from traditional standardized cognitive load tests (such as the National Aeronautics and Space Administration Task Load Index, or NASA TLX) to objective quantitative methods based on human physiological signal data (such as electroencephalogram (EEG), heart rate variability (HRV)) and human behavioral data (eye movement trajectory, operation response lag). However, there are still technical bottlenecks in adapting to the dynamism, individual differences, and closed-loop intervention capabilities of complex HCI scenarios.

[0016] For example, in scenarios where users operate intelligent driving cockpits, if only isolated collection of human physiological signal data (such as relying solely on HRV) or human behavioral data (such as operation error rate) is made, it is easy for the evaluation results to become disconnected from the actual interface operation scenario. For instance, when multiple alarm windows suddenly pop up in the intelligent driving cockpit, judging cognitive load solely based on heart rate changes is prone to misjudgment and has low accuracy.

[0017] In the process of cognitive load assessment, a general threshold (such as eye fixation duration >2 seconds being considered high cognitive load) can be used for assessment. However, this method does not take into account the differences in cognitive load tolerance among different individuals. For example, for a drone swarm control interface, expert users can handle the status information of 10 drones simultaneously, while novice users are already at a high cognitive load level under the same interface. The general threshold will lead to the distortion of the assessment for these two types of users, resulting in low accuracy of cognitive load assessment.

[0018] Furthermore, if only the cognitive load value is output during the cognitive load assessment process without forming a closed-loop linkage with the interface system, its contribution to improving operational safety is limited. For example, in scenarios with high cognitive load levels (such as when a surgeon experiences fatigue during the operation of a surgical robot), the inability to adjust the human-computer interaction interface in real time (such as simplifying the display of unnecessary parameters) to alleviate cognitive load results in higher operational risks.

[0019] To address one or more of the aforementioned technical problems, embodiments of the present invention provide a cognitive load assessment system based on multimodal data. Its specific objectives include: achieving multi-source dynamic coupling of "interface feature data - human physiological signal data - human behavioral data" to improve the fit between cognitive load assessment and actual HCI scenarios; establishing a personalized cognitive load assessment benchmark based on the user's cognitive profile to eliminate cognitive load assessment errors caused by individual differences; constructing a real-time closed loop of "assessment-intervention-feedback" to proactively alleviate high cognitive load through adaptive adjustments of the human-computer interaction interface; and adapting to 3D / immersive human-computer interaction interfaces and highly dynamic task scenarios to achieve accurate quantification of cognitive load and rapid response to sudden changes.

[0020] Figure 1 A schematic diagram of a cognitive load assessment system based on multimodal data provided in an embodiment of the present invention is shown below. Figure 1 As shown, the system includes: a personalized cognitive baseline construction module 1, a multimodal data acquisition module 2, and a dynamic coupling evaluation module 3. The multimodal data acquisition module 2 is connected to the personalized cognitive baseline construction module 1, and the multimodal data acquisition module 2 is connected to the dynamic coupling evaluation module 3.

[0021] The Personalized Cognitive Baseline Construction Module 1 is used to determine the user's individual cognitive baseline data; and to determine the user's personalized cognitive load grading threshold based on the user's individual cognitive baseline data.

[0022] The multimodal data acquisition module 2 is used to acquire multimodal data of the user in the working environment. The multimodal data includes the user's human physiological signal data, human behavior data, and interface feature data of the human-computer interaction interface.

[0023] The dynamic coupling assessment module 3 is used to determine the cognitive load assessment result of the user based on multimodal data and the user's personalized cognitive load classification threshold.

[0024] In this embodiment of the invention, the multimodal data acquisition module 2 integrates hardware interfaces and software protocols, and can simultaneously acquire human physiological signal data, human behavior data, and interface feature data of the human-computer interaction interface.

[0025] In this embodiment of the invention, the system may further include a data preprocessing module 4. The data preprocessing module 4 is connected to the multimodal data acquisition module 2 and the dynamic coupling evaluation module 3, respectively.

[0026] The data preprocessing module 4 is used to perform spatiotemporal alignment and noise filtering on the multimodal data of the user in the working environment acquired by the multimodal data acquisition module 2, and send the processed multimodal data to the dynamic coupling evaluation module 3. The dynamic coupling evaluation module 3 can determine the user's cognitive load evaluation result based on the processed multimodal data and the user's personalized cognitive load classification threshold.

[0027] In this embodiment of the invention, the system further includes an adaptive intervention module 5 and a mutation detection and early warning module 6. The dynamic coupling evaluation module 3 is connected to the adaptive intervention module 5, and the mutation detection and early warning module 6 is connected to the multimodal data acquisition module 2.

[0028] The adaptive intervention module 5 is used to adjust the content displayed on the human-computer interaction interface and / or the interaction method when the user's cognitive load assessment results indicate that the user's cognitive load level exceeds the set cognitive load level threshold.

[0029] The mutation detection and early warning module 6 is used to calculate the multimodal data based on the isolated forest algorithm when the multimodal data is within the set abnormal multimodal data range, and generate abnormal data; when the abnormal data is greater than the set abnormal threshold and the abnormal data corresponds to abnormal interface feature data, it is determined to be a cognitive load mutation state, and the edge computing device is controlled and alarmed and the human-machine interface is adjusted according to the set cognitive load mutation reminder rules.

[0030] Based on the aforementioned cognitive load assessment system based on multimodal data, this invention provides a cognitive load assessment method based on multimodal data. Figure 2 A flowchart of a cognitive load assessment method based on multimodal data provided in an embodiment of the present invention is shown below. Figure 2 As shown, the method includes: Step 102: Determine the user's individual cognitive baseline data.

[0031] In this embodiment of the invention, each step is executed by an edge computing device, which includes: a vehicle (such as a car, airplane, or ship), industrial manufacturing equipment (such as a CNC machine tool or a central control room), an immersive interactive device, or a medical robot. The edge computing device includes a processor module configured within it, the processor module being selected from at least one of the following: a vehicle control system (including a driving control unit for a car, airplane, or ship), an industrial manufacturing control system (including an operating terminal for a CNC machine tool or a central control room), an immersive interactive device (including a processing unit for a VR / AR system), or a medical robot control platform.

[0032] Figure 3 A flowchart for determining a user's individual cognitive baseline data provided in an embodiment of the present invention, such as... Figure 3 As shown, step 102 includes: Step 1022: Determine the user's cognitive profile based on the obtained user attribute information and the results of the cognitive load standardization test.

[0033] In this embodiment of the invention, a user's cognitive profile can be determined by the user's attribute information (such as age, professional field, and operational experience) and the results of cognitive load standardization tests (such as NASA TLX multitasking capability assessment and multi-target tracking test).

[0034] Step 1024: Obtain the general cognitive baseline corresponding to the user's cognitive profile from the pre-acquired general cognitive model.

[0035] In this embodiment of the invention, the general cognitive model is a cognitive load assessment framework trained on a massive sample (covering users of different ages, occupations, and cognitive abilities).

[0036] Step 1026: Update the general cognitive baseline to obtain the user's individual cognitive baseline data.

[0037] Figure 4 A flowchart for updating a general cognitive baseline to obtain a user's individual cognitive baseline data, as provided in an embodiment of the present invention, is shown below. Figure 4 As shown, step 1026 includes: Step S1: Based on the user's cognitive profile, construct at least two types of tasks with different levels of human-computer interaction complexity.

[0038] In this embodiment of the invention, at least two types of tasks include a first task, a second task, and a third task whose human-computer interaction complexity increases sequentially.

[0039] In this embodiment of the invention, task pairs (at least 5 sets) of the first task and the third task can be constructed for different scenarios to avoid users developing a proficiency bias in a single task.

[0040] For example, the first task is to operate on a human-computer interaction interface with low information density (e.g., ≤5 interface elements and no redundant information), few operation steps (≤3 steps), and no distractions (e.g., no pop-ups / alarms). For example, the first task could be to start a program and set a countdown (e.g., 10 minutes) on a human-computer interaction interface with only three buttons: "Start," "Pause," and "End."

[0041] For example, the third task involves operating on a high-information-density human-computer interaction interface (e.g., ≥15 interface elements with redundant information), multiple operation steps (≥8 steps), and interference (e.g., randomly popping up prompts, flashing background data). For instance, in a human-computer interaction interface that simultaneously displays real-time data curves, multiple parameter setting boxes, and system status prompts, the task involves completing a series of operations such as starting the program, filtering specific data, setting a 10-minute countdown, and confirming 3 related parameters.

[0042] Step S2: Obtain personnel capability data of users during the execution of at least two types of tasks.

[0043] For example, by setting up a comparative experiment between the first task and the third task (such as having users complete the same operation on a low-information-density interface and a high-information-density interface), the personnel ability data corresponding to the user when performing the first task and the personnel ability data corresponding to the third task can be obtained.

[0044] In this embodiment of the invention, personnel capability data may include information processing rate and / or error rate generated based on human physiological signal data, human behavioral data, and subjective feedback data. The information processing rate, in bits per second (bit / s), can be calculated using the formula: Total information contained in the task / Completion time = Information processing rate. The error rate, representing operational accuracy, can be calculated using the formula: Number of erroneous operations / Total number of operations = Error rate.

[0045] Human behavioral data can serve as a core indicator for determining personalized cognitive load grading thresholds. For example, human behavioral data can include information processing speed and operational accuracy.

[0046] Human physiological signal data can serve as auxiliary validation for determining personalized cognitive load grading thresholds. For example, human physiological signal data includes autonomic nervous system indicators (e.g., heart rate variability, skin conductance response) and electroencephalogram (EEG) indicators (e.g., prefrontal theta wave power). Specifically, the higher the user's cognitive load, the stronger the prefrontal theta wave activity.

[0047] In this embodiment of the invention, subjective feedback data can serve as a calibration reference for determining the threshold for personalized cognitive load grading. For example, immediately after the user completes the first and third tasks, the user's ratings (1-10 points) of "mental demand" and "frustration" can be collected using the NASA TLX scale.

[0048] Step S3: Update the general cognitive baseline based on the personnel ability data corresponding to at least two types of tasks to obtain the individual cognitive baseline data of the user.

[0049] In this embodiment of the invention, the differences in personnel ability data corresponding to the user when performing the first task and the third task can be determined, and the obtained general cognitive baseline can be updated according to the differences; the validity of the adjusted general cognitive baseline can be verified according to the personnel ability data corresponding to the user when performing the second task, and when the verification passes, the adjusted general cognitive baseline is determined to be the user's individual cognitive baseline data.

[0050] For example, the difference in personnel capability data when performing the first and third tasks could include: calculating the decrease in information processing rate (Δ rate = information processing rate of complex tasks - information processing rate of simple tasks). It could also include: calculating the increase in error rate (Δ error rate = error rate of complex tasks - error rate of simple tasks).

[0051] In this embodiment of the invention, the general cognitive baseline may include the information processing rate of users with different cognitive profiles. For example, the information processing rate of expert users is ≥80 bits / s, and the information processing rate of elderly users is ≥40 bits / s.

[0052] In this embodiment of the invention, if a user's Δ rate is less than or equal to a preset minimum rate threshold (e.g., the preset minimum rate threshold is -20%, meaning the rate decrease does not exceed 20%) and Δ error rate is less than or equal to a preset minimum error rate threshold (e.g., the preset minimum error rate threshold is 15%) in the third task, it indicates that their current general cognitive baseline can withstand the complexity; otherwise, the general cognitive baseline needs to be lowered to generate individual cognitive baseline data. Here, the lowered general cognitive baseline can be determined based on the median difference between the user's general cognitive baseline and the user's information processing rate in the third task.

[0053] For example, the general cognitive baseline for expert users is "information processing rate ≥ 80 bit / s". In the comparative experiment, the information processing rate of the first task is 100 bit / s, and the information processing rate of the third task is stable at 85 bit / s (Δ rate = -15%). If the information processing rate of the third task drops to 70 bit / s (Δ rate = -30%), the general cognitive baseline is lowered to 75 bit / s.

[0054] For example, the general cognitive baseline for elderly users is "information processing rate ≥ 40 bits / s". If the information processing rate of the third task decreases from 50 bits / s to 35 bits / s (Δ rate = -30%), then the general cognitive baseline will be lowered to 38 bits / s.

[0055] In this embodiment of the invention, the acquired general cognitive baseline can be updated based on differences. The update triggering conditions are: 1. Periodic triggering (e.g., setting an update to be triggered periodically after 20 user interactions); 2. Anomaly triggering (e.g., triggering an update when abnormal user behavior is detected during daily operations, such as a sudden 30% increase in the error rate). Update frequency: In the initial stage (first 100 interactions), the initial individual cognitive baseline data can be updated once every 5 experiments; in the stable period (after 100 interactions), it can be updated once every 20 experiments to avoid frequent fluctuations. A sliding window algorithm can be used (e.g., 70% weighting for the most recent 3 experiments and 30% for historical data) to balance recent performance and long-term trends in updating the general cognitive baseline.

[0056] In this embodiment of the invention, the personnel ability data corresponding to the user when performing the second task is obtained, and the effectiveness of the adjusted general cognitive baseline is verified.

[0057] In this embodiment of the invention, if the personnel ability data corresponding to the user when performing the second task is within the range of the adjusted general cognitive baseline, the adjusted general cognitive baseline can be determined as the individual cognitive baseline data.

[0058] In this embodiment of the invention, if the user's performance in the second task (such as information processing speed and error rate) falls within the range of the adjusted general cognitive baseline, and the subjective score is ≥5 points, then the adjusted general cognitive baseline is confirmed to be valid, and the adjusted general cognitive baseline is determined as the individual cognitive baseline data. If the user's performance in the second task does not fall within the range of the adjusted general cognitive baseline (such as information processing speed or error rate exceeding the range of the adjusted general cognitive baseline, or a low subjective score), then the reasons for the data anomaly are analyzed retrospectively (such as task design deviations in at least two types of tasks), the reasons for the anomaly are corrected, and the process returns to step 1026.

[0059] By following the steps above, individual cognitive baseline data can always be aligned with the user's current cognitive ability characteristics, providing an accurate individual benchmark for the subsequent generation of personalized cognitive load grading thresholds.

[0060] Step 104: Determine the user's personalized cognitive load grading threshold based on the user's individual cognitive baseline data.

[0061] In this embodiment of the invention, the user's personalized cognitive load classification threshold includes a low load threshold, a medium load threshold, a high load threshold, and a dangerous load threshold.

[0062] In this embodiment of the invention, the acquired target evaluation model can be adapted to an individual based on a transfer learning algorithm, outputting individual cognitive baseline data to generate a personalized cognitive load grading threshold for the user. The target evaluation model is a cognitive load assessment framework trained on a massive dataset (covering users of different ages, occupations, and cognitive abilities). Its core parameters (target evaluation model parameters) are general regular parameters describing the mapping relationship between "multimodal data and cognitive load values," mainly including three categories: feature extraction layer parameters, weight adaptation layer parameters, and output layer mapping parameters.

[0063] In this embodiment of the invention, the parameters of the feature extraction layer correspond to the feature fusion layer in the dynamic coupling evaluation model, including: the convolution kernel parameters of the multi-scale convolutional neural network (CNN) (weights used to extract features of different dimensions, such as the importance weights of features like interface element density and physiological signal fluctuation amplitude); the temporal parameters of the long short-term memory network (LSTM) (such as the slope weights of temporal data such as heart rate and reaction time, reflecting the general law of "cognitive load accumulation effect in the time dimension"); and the spatial encoding parameters of the Transformer model (such as the attention weights of spatial features like object depth difference and operation path length in a 3D interface, describing the group commonality of "correlation between spatial complexity and cognitive load").

[0064] In this embodiment of the invention, the weight adaptive layer parameters correspond to the general allocation ratio of the attention mechanism in the task stage. For example: in the stage of sudden increase in interface information: the default weight of interface feature data, human behavior data, and human physiological signal data (determined based on the average performance of the group, such as a weight ratio of 60%:30%:10%); in the decision-making stage: human physiological signal data (such as the default weight of EEG features), human stress data, and other features (reflecting the common dependence of the group on EEG signals when making decisions, such as a weight ratio of 40%:30%:30%).

[0065] It should be noted that the weighting ratios shown here are just examples, and users can adjust them according to their actual needs.

[0066] In this embodiment of the invention, the output layer mapping parameters correspond to the mapping coefficients of “feature combination value → cognitive load score (0-100)” in the fully connected network, as well as the thresholds corresponding to the cognitive load score and the four-level personalized cognitive load grading thresholds (such as the general standard: 0-25 corresponds to “low load threshold”, 26-50 corresponds to “medium load threshold”, etc.).

[0067] In this embodiment of the invention, the core of generating personalized cognitive load grading thresholds is to "adapt" the parameters of the target evaluation model to the individual through transfer learning, so that the personalized cognitive load grading thresholds fit the user's cognitive characteristics (such as age, professional background, operating habits, etc.). The transfer learning operation includes: freezing parameters in the general target evaluation model that have "strong group commonality" (such as the CNN convolutional kernels of the feature extraction layer parameters), and only fine-tuning parameters that are "related to individual differences" (such as the stage attention ratio or output layer mapping parameters corresponding to the weight adaptation layer parameters), so that the model prioritizes learning the patterns in individual data. For example, expert users perform more stably in high-information-density interfaces; after fine-tuning, the proportion of their behavioral data weights during the "sudden increase in interface information stage" can be increased accordingly (e.g., from the general 30% to 45%), indicating that expert behavioral data is more reliable.

[0068] In this embodiment of the invention, cognitive load values ​​less than or equal to the product of individual cognitive baseline data and a preset first weighting coefficient can be set as a low load threshold; cognitive load values ​​greater than the product of individual cognitive baseline data and the first weighting coefficient but less than or equal to the product of individual cognitive baseline data and a preset second weighting coefficient can be set as a medium load threshold; cognitive load values ​​greater than the product of individual cognitive baseline data and the second weighting coefficient but less than or equal to the product of individual cognitive baseline data and a preset third weighting coefficient can be set as a high load threshold; and cognitive load values ​​greater than the product of individual cognitive baseline data and the third weighting coefficient can be set as a dangerous load threshold. Wherein, the first weighting coefficient is less than the second weighting coefficient, and the second weighting coefficient is less than the third weighting coefficient.

[0069] Specifically, based on threshold classification using individual cognitive baseline data, four levels of personalized cognitive load grading thresholds can be defined: Low load threshold: Cognitive load value ≤ individual cognitive baseline data × first weighting coefficient (extremely low cognitive resource consumption, smooth and stress-free operation); Medium load threshold: Individual cognitive baseline data × first weighting coefficient < cognitive load value ≤ individual cognitive baseline data × second weighting coefficient (moderate cognitive resource consumption, task can be completed stably); High load threshold: Individual cognitive baseline data × second weighting coefficient < cognitive load value ≤ individual cognitive baseline data × third weighting coefficient (cognitive resources are close to saturation, error rate begins to rise); Dangerous load threshold: Cognitive load value > individual cognitive baseline data × third weighting coefficient (cognitive resources are overloaded, operational stability is significantly reduced, emergency intervention is required). For example, the first weighting coefficient is 50%, the second weighting coefficient is 80%, and the third weighting coefficient is 120%.

[0070] For example, the information processing rate of an expert user's individual cognitive baseline data is 80 bits / s. Their personalized cognitive load grading thresholds might be: 0-40 for low load, 41-64 for medium load, 65-96 for high load, and >96 for dangerous load. As another example, the information processing rate of an elderly user's individual cognitive baseline data is 40 bits / s. Their personalized cognitive load grading thresholds might be: 0-20 for low load, 21-32 for medium load, 33-48 for high load, and >48 for dangerous load.

[0071] In this embodiment of the invention, when a user's individual cognitive baseline data changes, the personalized cognitive load grading threshold is adjusted according to a set ratio. Specifically, when the individual cognitive baseline data changes due to factors such as training / aging (e.g., by comparing simple and complex tasks and finding that the information processing rate of the individual cognitive baseline data decreases from 80 bit / s to 75 bit / s), the personalized cognitive load grading threshold is automatically adjusted proportionally.

[0072] In this embodiment of the invention, if the deviation between the personalized cognitive load grading threshold and the cognitive load standardized test result (such as the NASA TLX score) is greater than a set deviation threshold (such as 15%), the output layer mapping parameters of the target assessment model are corrected to update the user's individual cognitive baseline data, and step 104 is continued.

[0073] Step 106: Obtain multimodal data of the user in the work environment. The multimodal data includes the user's human physiological signal data, human behavior data, and interface feature data of the human-computer interaction interface.

[0074] In this embodiment of the invention, interface feature data of the human-computer interaction interface can be obtained through the application programming interface (API) of the interface log. For example, interface feature data includes: the number of information modules (such as the number of navigation / entertainment / vehicle status windows in a smart cockpit), the complexity of interaction steps (such as the number of clicks required to complete a task), the dynamic frequency of visual elements (such as the flashing frequency of alarm pop-ups), and three-dimensional spatial parameters (such as the spatial coordinates and density of virtual objects in a VR scene).

[0075] In this embodiment of the invention, human physiological signal data can be collected by wearable devices (such as EEG caps, wristband sensors, etc.). For example, human physiological signal data includes EEG alpha / β waves, HRV, and galvanic skin response (GSR).

[0076] In this embodiment of the invention, human behavior data can be acquired using an eye tracker, including, for example, gaze trajectory and scan rate. Human behavior data can also be acquired using an operating device (such as a mouse, VR controller, steering wheel, or joystick), including, for example, response lag and click error.

[0077] In this embodiment of the invention, the multimodal data may also include human pressure data, which can be acquired through a camera, microphone and / or pressure sensor. For example, human pressure data includes: micro-expression features (frowning frequency) collected by a camera, speech prosody (speech rate, pause interval) collected by a microphone and / or changes in operational force collected by a pressure sensor.

[0078] In this embodiment of the invention, spatiotemporal alignment and noise filtering can be performed on multimodal data. Spatiotemporal alignment can be performed by matching interface feature data, human physiological signal data, human behavior data, and human stress data in time sequence based on high-precision timestamps (error ≤ 10ms) to ensure the correlation of data under the same interactive event (e.g., aligning "the appearance of the alarm pop-up of interface feature data" with "the change in EEG amplitude of human physiological signal data"). Noise filtering can be performed by using wavelet transform to remove electromyographic artifacts in EEG signals, optimizing eye movement trajectory drift through Kalman filtering, and interpolating and repairing outliers (such as signal interruption caused by sensor detachment).

[0079] Step 108: Based on multimodal data and the user's personalized cognitive load grading threshold, determine the user's cognitive load assessment result, which includes the cognitive load level.

[0080] In this embodiment of the invention, the cognitive load assessment result of a user can be determined based on a dynamic coupling assessment model according to multimodal data and the user's personalized cognitive load classification threshold. The dynamic coupling assessment model includes a feature fusion layer and a cognitive load assessment layer. The feature fusion layer is used to extract features of different dimensions using a multi-scale CNN, including temporal dimension, spatial dimension, and weight adaptive layer.

[0081] In this embodiment of the invention, when the user's cognitive load assessment result indicates that the user's cognitive load level exceeds a set cognitive load level threshold, the content displayed on the human-computer interaction interface and / or the interaction method are adjusted.

[0082] Figure 5 A flowchart illustrating how a user's cognitive load assessment result is determined based on multimodal data and the user's personalized cognitive load grading threshold, as provided in one embodiment of the present invention, is shown below. Figure 5 As shown, step 108 includes: Step 1082: Determine the temporal variation characteristics of the human physiological signal data and the human behavior data based on the human physiological signal data and human behavior data in the multimodal data.

[0083] In this embodiment of the invention, the temporal variation characteristics of human physiological signal data and human behavioral data in multimodal data (such as the trend of heart rate as interface complexity increases) can be learned through LSTM networks.

[0084] In this embodiment of the invention, the cognitive load value is not a static value, but rather changes dynamically over time (e.g., the cognitive load is low at the beginning of the task and accumulates to increase with the increase in operational complexity / duration). The core purpose of the LSTM network is to capture the evolution of human physiological signal data (such as heart rate and EEG) and human behavioral data (such as operation interval and error rate) over time, and to establish a mapping relationship between temporal change characteristics and the trend of cognitive load change.

[0085] For example, if a user operates continuously in a highly complex interface for 10 minutes, their heart rate may gradually increase from 60 beats / minute at rest to 85 beats / minute, while the error rate increases from 5% to 20%. The LSTM network needs to learn this "correlation between the slope of the heart rate increase and the rate of increase of the error rate" to determine whether the cognitive load is slowly accumulating or suddenly surging.

[0086] In this embodiment of the invention, the learning process using an LSTM network includes: Data preprocessing: Human physiological signal data (sampled by timestamp, such as recording heart rate and brainwave theta power every 100ms) and human behavior data (such as the time point, operation type, and error mark of each operation) are aligned to a unified time axis to form a time series (such as a time step sequence of length 100, each time step containing features such as heart rate, error rate, and operation interval).

[0087] LSTM Network Design: Input Layer: Feature vectors for each time step (e.g., heart rate, EEG theta wave power, operation interval, number of errors); Hidden Layer: Preserves key temporal information (e.g., sudden increases in heart rate over three consecutive time steps) and forgets irrelevant noise (e.g., accidental operational errors) through the LSTM network's gating mechanism (input gate, forget gate, output gate); Output Layer: Intermediate cognitive load trend values ​​for each time step (e.g., cognitive load growth rate), serving as input features for subsequent fully connected networks. Training and Optimization: The model is trained using historical data labeled with time and cognitive load values ​​(e.g., knowing that users are at a high cognitive load level during a certain period). The weight parameters of the LSTM network are optimized by minimizing the loss function (e.g., mean-square error (MSE)) between the predicted trend and the actual cognitive load change.

[0088] In this embodiment of the invention, compared with static features (such as heart rate at a single time point), the temporal learning of LSTM networks can identify dynamic patterns such as the cumulative effect of cognitive load and sudden fluctuations (such as a continuous increase in heart rate + an extended operation interval indicating that cognitive load is about to be overloaded), thereby improving the real-time performance of cognitive assessment.

[0089] In this embodiment of the invention, the gating mechanism of the LSTM network can filter out accidental interference (such as brief fluctuations in heart rate caused by sudden limb movements of the user), focus on meaningful temporal trends, and improve the stability of cognitive assessment.

[0090] In this embodiment of the invention, by learning the evolution of cognitive load over time, future changes in cognitive load can be predicted (e.g., based on the current trend, it will enter a dangerous cognitive load level in 5 minutes), providing a basis for early intervention (e.g., simplifying the interface).

[0091] Step 1084: Determine the spatial distribution characteristics of the interface feature data of the human-computer interaction interface based on the interface feature data of the human-computer interaction interface.

[0092] In this embodiment of the invention, the spatial distribution characteristics of the interface feature data of the virtual human-computer interaction interface can be encoded using the Transformer model.

[0093] In this embodiment of the invention, for a three-dimensional human-computer interaction interface, the spatial distribution characteristics of the interface feature data of the human-computer interaction interface (such as the impact of the jump frequency between depth layers on cognitive load) can be virtually encoded using the Transformer model.

[0094] In this embodiment of the invention, in a three-dimensional human-computer interaction interface (such as an industrial control 3D interface), the spatial layout of virtual objects (such as position, depth, and hierarchical relationships) directly affects the user's cognitive load (e.g., frequent jumps between depth layers increase spatial cognitive cost). The core purpose of the Transformer model is to transform the spatial features of objects in three-dimensional space, such as relative positions, hierarchical structures, and interaction paths, into quantifiable vector representations, capturing the correlation between spatial complexity and cognitive load. For example, in a virtual cockpit, the dashboard is distributed across three depth layers, and the key indicators are too far from the operation buttons. This spatial distribution increases the user's visual search cost, and the Transformer model needs to encode this relationship between "distance, hierarchical difference, and operation time."

[0095] In this embodiment of the invention, the encoding process of the Transformer model includes: spatial feature extraction: defining the core spatial attributes of the virtual object, including: absolute coordinates (x, y, z); relative relationships (such as the distance and depth difference between object A and object B); hierarchical structure (such as the nesting relationship of "main interface → sub-menu → function button"); and interaction path (such as the average movement trajectory length of the user from object A to object B).

[0096] In this embodiment of the invention, the encoding mechanism of the Transformer model consists of the following layers: Input layer: converts the spatial attributes of each virtual object into a spatial feature vector and adds position encoding (marking the absolute position of the object in three-dimensional space); Self-attention layer: calculates the attention weights between all objects (e.g., "high attention weights between operation buttons and indicator dashboards" indicates a strong correlation between the two), and captures the dependencies between objects (e.g., "users usually check the dashboard before operating the buttons"); Output layer: outputs the spatial context vector of each object through the encoder, and the aggregates them as a feature reflecting the overall spatial complexity (e.g., "high attention weights concentrated between distant objects → high spatial complexity"). Feature mapping: inputs the encoded spatial feature vectors into subsequent networks, fuses them with other dimensional features (e.g., human physiological signal data, human behavioral data), and finally maps them to the cognitive load value.

[0097] In this embodiment of the invention, compared to measuring complexity solely by the number of objects, the self-attention mechanism of the Transformer model can capture implicit relationships between objects (such as the positive correlation between the frequency of deep layer jumps and the error rate), and can more accurately reflect the impact of spatial layout on cognitive load.

[0098] In this embodiment of the invention, the spatial layout of the three-dimensional interface is flexible and varied (e.g., users can customize the position of objects), and the Transformer model does not require preset rules. It can automatically learn the cognitive load of different spatial patterns through data and adapt to diverse scenarios.

[0099] In this embodiment of the invention, the comprehensiveness of the assessment can be improved. By combining the temporal characteristics of the time dimension and the spatial distribution characteristics of the spatial dimension, the model can simultaneously consider the impact of time and space on cognitive load, forming a more three-dimensional assessment dimension (such as "continuous operation in a complex spatial layout → faster accumulation of cognitive load").

[0100] Step 1086: Based on the current task stage of the user, determine the weight allocation strategy according to the temporal change characteristics and spatial distribution characteristics. The weight allocation strategy is the weight allocation strategy for human physiological signal data, human behavior data and interface feature data in multimodal data. The weight allocation strategy includes the feature vector of each multimodal data and the weight coefficient corresponding to each feature vector.

[0101] In this embodiment of the invention, the design of the weight adaptive layer is based on the attention mechanism of the task stage. For example, the task stage includes the stage of sudden increase in interface information, the stage of proficient operation, or the decision-making stage.

[0102] In this embodiment of the invention, the current task stage of the user can be determined based on the CNN+GRU model and multimodal data, whether it is the stage of sudden increase in interface information, the stage of proficient operation, or the stage of decision-making.

[0103] Among them, the stage of sudden increase in interface information (such as sudden alarm) can increase the weight of interface feature data (e.g., accounting for 60%); the stage of proficient operation can increase the weight of human behavior data (e.g., accounting for 50%). Here, proficient operation refers to users repeatedly performing standardized operations, with stable operation intervals, extremely low error rates, and no significant changes in the interface; the decision-making stage can increase the weight of human physiological signal data and human stress data (e.g., accounting for 70% in total). Here, the decision-making stage refers to situations where users' operation pause time is extended, repeated browsing behavior occurs, and multiple options are faced.

[0104] In this embodiment of the invention, the weight ratio of different types of multimodal data (interface feature data, human behavior data, human physiological signal data, etc.) can be dynamically adjusted according to the user's current task stage (such as the stage of sudden increase in interface information, the stage of proficient operation, or the decision-making stage). This allows the dynamically coupled evaluation model to focus more on the features that have the most significant impact on cognitive load at different stages, thereby improving the accuracy of cognitive evaluation. The specific design can be divided into the following key steps: I. Precise division of task phases and triggering conditions First, it is necessary to clarify the definition and boundaries of the task stages. By monitoring the characteristics of user behavior and interface status in real time, three core stages can be automatically identified (which can be expanded according to the scenario): 1) Stage of sudden increase in interface information Feature indicators: The number of interface elements surges in a short period of time (e.g., ≥10 new elements are added within 1 second), a sudden alarm occurs (e.g., flashing pop-ups, warning sounds), or the interface switches from a low-complexity interface to a high-complexity interface (e.g., jumping from a list page to a data-intensive details page).

[0105] Triggering conditions: This stage is determined when the rate of change of interface elements (Δ number of elements / Δ time) is greater than the preset threshold (e.g., 5 elements / second), or when a system-level alarm signal is detected.

[0106] 2) Proficient Operation Stage Characteristic identifier: Users repeatedly perform standardized operations (such as clicking the same button repeatedly or completing steps according to a fixed process), the operation interval is stable (standard deviation < 0.5 seconds), the error rate is extremely low (< 3%), and the interface does not change significantly.

[0107] Triggering conditions: When more than 5 consecutive operations meet the characteristics of "repetitive operation type + small interval fluctuation + no errors", it is determined to be in this stage.

[0108] 3) Decision-making stage Characteristic markers: prolonged pause time in user operation (>3 seconds), repeated browsing behavior (such as moving the mouse back and forth between multiple options), increased power of prefrontal β waves in EEG signals (reflecting decision-making thinking), and facing multiple options (such as choosing one from more than 3 options).

[0109] Triggering conditions: When the combined features of "operation pause time > threshold + multi-option interface + enhanced EEG beta waves" are detected, this stage is determined.

[0110] II. Feature Weight Allocation Strategies for Each Stage To address the core influencing factors of cognitive load at different stages, a differentiated weighting rule is pre-set (taking a total weight of 100% as an example): 1) Phase of Sudden Increase in Interface Information: Prioritize interface feature data, with the following weighting: Interface feature data > Human behavior data > Human physiological signal data (e.g., 60%, 30%, and 10% respectively). Rationale: Cognitive load at this stage primarily stems from information overload (e.g., rapidly processing new elements and recognizing alarm meanings). The information density, element layout, and alarm salience of the interface directly determine the level of cognitive load. Human behavior data and human physiological signal data may not fully reflect changes in cognitive load due to reaction delays. Specific examples: Interface element density, alarm priority, clarity of information hierarchy, etc.

[0111] 2) Proficient Operation Stage: Prioritize human behavior data, with the following weighting: Human behavior data > Interface feature data > Human physiological signal data (e.g., 50%, 30%, and 20% respectively). The rationale is that once users are familiar with the process, the impact of interface complexity diminishes. At this stage, cognitive load is primarily reflected in operational efficiency (e.g., whether fatigue leads to a decrease in speed) and stability (e.g., whether habitual errors occur). Human behavior data (operation speed, trajectory stability) can more realistically reflect cognitive load. Specific feature examples include: operation interval time, number of consecutive correct operations, mouse movement speed, etc.

[0112] 3) Decision-making stage: Prioritize human physiological signal data and human stress data, with the following weighting: Human physiological signal data and human stress data > Interface feature data > Human behavior data (e.g., 40%, 30%, 20%, and 10% respectively). The rationale is that during decision-making, the user's cognitive load is more reflected in internal thinking (rather than external operations). EEG signals from human physiological signal data (such as prefrontal cortex activity) can reflect the intensity of mental exertion, and human stress data (such as micro-expressions and pupil dilation rate) can reflect the degree of hesitation. Explicit operations (such as clicking) are difficult to reflect the true cognitive load due to their low frequency. Specific feature examples: EEG theta wave power (higher cognitive load is stronger), pupil diameter change rate, and differences in the duration of browsing options, etc.

[0113] III. The technology for dynamic weight adjustment is achieved through a closed-loop process of "stage identification - weight generation - feature weighting," and relies on the following core technologies: Stage recognition network: A lightweight CNN+GRU model is used to analyze input features (such as interface change rate, operation sequence, EEG segments) in real time and output the probability distribution of the current stage (such as the probability of information surge stage = 0.85) to ensure the accuracy of stage judgment.

[0114] Adaptive Weight Generator: A Multilayer Perceptron (MLP) is designed as the weight generation module. It takes "stage recognition results and current feature statistics" (such as the number of interface elements and the operation error rate) as input and outputs the weight coefficients of each feature. For example, if a stage of sudden increase in interface information is identified, the MLP can strengthen the weight coefficients of interface features (e.g., adjust the weight parameter of interface features from the general value of 0.3 to 0.6); if a high error rate is detected simultaneously, the weight of human behavior data is temporarily increased (e.g., from 30% to 35%) to compensate for abnormal situations.

[0115] In this embodiment of the invention, during critical stages such as the sudden increase in interface information, by increasing the weight of core multimodal data, changes in cognitive load can be quickly captured (e.g., identifying overload risk within 0.5 seconds), thus buying time for immediate intervention (e.g., automatically simplifying the interface), resulting in high real-time performance.

[0116] In this embodiment of the invention, robustness can be enhanced, noise features (such as occasional physiological fluctuations) can be weakened during the skilled operation stage, stable behavioral data can be focused on, and evaluation errors caused by irrelevant interference can be reduced.

[0117] Step 1088: Generate a feature matrix based on each feature vector and its corresponding weight coefficient.

[0118] In this embodiment of the invention, each feature vector is multiplied by the weight coefficient corresponding to each feature vector to generate a feature matrix (e.g., vector of interface feature data × 0.6 + vector of human behavior data × 0.3 + vector of human physiological signal data × 0.1).

[0119] Step 1090: Input the feature matrix into the cognitive load value of the fully connected network.

[0120] In this embodiment of the invention, the dynamic coupling evaluation model further includes a cognitive load evaluation layer, which is used to input the feature matrix into the fully connected network to calculate the cognitive load value.

[0121] Step 1092: Determine the user's cognitive load level based on the cognitive load value and the user's personalized cognitive load grading threshold.

[0122] In this embodiment of the invention, if the cognitive load value is within the range of the low load threshold, the user's cognitive load level is determined to be a low cognitive load level; if the cognitive load value is within the range of the medium load threshold, the user's cognitive load level is determined to be a medium cognitive load level; if the cognitive load value is within the range of the high load threshold, the user's cognitive load level is determined to be a high cognitive load level; and if the cognitive load value is within the range of the dangerous load threshold, the user's cognitive load level is determined to be a dangerous cognitive load level.

[0123] For example, if the cognitive load value is between 0 and 25, the cognitive load level is low; if the cognitive load value is between 26 and 50, the cognitive load level is medium; if the cognitive load value is between 51 and 75, the cognitive load level is high; and if the cognitive load value is between 76 and 100, the cognitive load level is dangerous.

[0124] Step 110: Adjust the human-computer interaction interface according to the interface adjustment rules corresponding to the cognitive load level.

[0125] In this embodiment of the invention, interface adjustment rules corresponding to each cognitive load level are pre-set to adjust the human-computer interaction interface.

[0126] For example, the interface adjustment rules for low cognitive load levels are: increase information density (such as displaying extended function buttons) to avoid distraction; the interface adjustment rules for medium cognitive load levels are: optimize information layout (such as aggregating and displaying related parameters) to simplify operation steps; the interface adjustment rules for high cognitive load levels are: hide unnecessary elements (such as closing entertainment windows) and activate auxiliary decision-making (such as recommending operation paths); the interface adjustment rules for dangerous cognitive load levels are: pause non-urgent tasks, trigger strong reminders (such as seat vibration + voice warning), and automatically switch to the "simplified mode" of the human-computer interaction interface (retaining only core parameters).

[0127] Step 112: Continuously monitor the cognitive load level. If the cognitive load level does not decrease within a preset time range, obtain the interface adjustment rules corresponding to the next higher cognitive load level.

[0128] In this embodiment of the invention, after adjusting the human-computer interaction interface, the change in cognitive load level is continuously monitored. If it does not drop to a safe range (such as low or medium cognitive load level) within 10 seconds, the interface adjustment rules corresponding to the next higher cognitive load level are obtained (such as switching from "optimized layout" to "simplified mode" in the human-computer interaction interface).

[0129] For example, if the cognitive load level is high cognitive load, then the next higher cognitive load level is the dangerous cognitive load level.

[0130] Step 114: Adjust the human-computer interaction interface according to the interface adjustment rules corresponding to the higher level of cognitive load.

[0131] In this embodiment of the invention, when multimodal data is within a set range of abnormal multimodal data, the multimodal data is calculated based on the isolated forest algorithm to generate abnormal data. When the abnormal data is greater than the set abnormal threshold and the abnormal data corresponds to abnormal interface feature data, it is determined to be a cognitive load mutation state. The edge computing device is controlled and alarmed according to the set cognitive load mutation reminder rules, and the human-computer interaction interface is adjusted.

[0132] In this embodiment of the invention, the range of abnormal multimodal data can be set according to actual conditions. For example, the range of abnormal multimodal data includes: instantaneous acceleration of human physiological signal data rate >15 times / minute, and operation pauses in human behavior data >300ms.

[0133] In this embodiment of the invention, if an interface event occurs (such as the appearance of an obstacle or a device malfunction alarm), abnormal interface feature data is generated. At this time, the detected human physiological signal data (such as a sudden increase in heart rate > 15 beats / minute) is considered abnormal human physiological signal data, and the detected sudden change in human behavior data (such as an operation pause > 300ms or a change in eye movement fixation point) is considered abnormal human behavior data.

[0134] In this embodiment of the invention, the Isolation Forest algorithm is a fast anomaly detection method based on tree (iTree) ensemble. Its core idea for anomaly detection is that "anomalies are outliers that are easily isolated".

[0135] In this embodiment of the invention, the isolated forest algorithm is deployed using edge computing. An anomaly threshold is set based on the user's historical multimodal data. When the abnormal data exceeds the set anomaly threshold and the abnormal data is strongly correlated with the abnormal interface feature data, it is determined to be a cognitive load mutation state within 500ms. An early warning is triggered on the edge computing device according to the set cognitive load mutation reminder rules (such as partial interface flickering + haptic feedback).

[0136] The above-mentioned cognitive load assessment method based on multimodal data is described below with a specific example.

[0137] Scenario Background: For intelligent driving cockpits (equipped with AR navigation, multi-screen interaction, and voice control functions), drivers need to simultaneously handle road condition recognition, navigation commands, vehicle status monitoring (such as vehicle speed, battery level, and tire pressure), and emergency scenarios (such as obstacle avoidance and system fault alarms). The complexity of the human-machine interface changes dynamically with the driving stage. It is necessary to assess the driver's cognitive load in real time and adaptively adjust the human-machine interface to avoid driving risks caused by excessive cognitive load.

[0138] Hardware configuration: Data acquisition equipment and interface data interface. The data acquisition equipment includes an EEG cap (for acquiring prefrontal alpha / beta waves, sampling rate 250Hz), a wrist heart rate sensor (for monitoring heart rate variability (HRV)), a steering wheel pressure sensor (for acquiring operational force), a cockpit camera (for capturing micro-expressions), an eye tracker (for capturing eye movement trajectories), and an AR-based head-up display (HUD) (for interface rendering and intervention). Interface data interface: Real-time acquisition of cockpit interface status via the vehicle system API (e.g., currently displayed modules: navigation map / entertainment interface / vehicle parameter panel; number of pop-ups: e.g., simultaneously displaying two alerts, "Traffic Congestion Ahead" and "Abnormal Tire Pressure"; interaction steps: e.g., switching navigation destinations requires three steps).

[0139] User cognitive profile: The driver is a 35-year-old male with 5 years of driving experience (intermediate experience). Under normal driving conditions, he can stably handle 2-3 parallel information modules. His average reaction time to sudden alarms is 0.8 seconds. Based on this, the system generates an individual cognitive baseline (cognitive load value ≤ 40, information processing rate ≥ 60 bits / s).

[0140] Acquire multimodal data of users in their work environment. This multimodal data includes human physiological signal data, human behavioral data, human stress data, and interface feature data of the human-computer interaction interface.

[0141] For example, interface feature data: AR-HUD displays a navigation map (1 main module) + vehicle speed / battery panel (1 auxiliary module), no pop-ups (low information density), and low complexity of interaction steps (e.g., adjusting the air conditioning by voice requires only 1 step). Human physiological signal data: EEG alpha wave power ratio 65% (indicating focused attention), heart rate 72 beats / minute, HRV standard deviation 50ms (normal range). Human behavior data: eye movement fixation point switching frequency between navigation and road conditions 1 time / second, stable steering wheel operation force (pressure value 30-40N), no operation delay. Human pressure data: no facial frowning, clear voice commands (speech rate 120 words / minute).

[0142] For example, by aligning the timestamps (error ≤ 5ms), "no pop-up window on the interface", "alpha wave ratio 65%", and "gaze switching frequency 1 time / second" are bound to the same time series data. Wavelet transform removes engine noise from the EEG signal, and Kalman filtering corrects eye movement trajectory drift.

[0143] After the driver has been driving for 10 minutes, it was detected that the driver can stably process the "navigation + vehicle parameters" modules when there are no emergencies. The cognitive load value fluctuates between 20-30 (20-30 is lower than the individual cognitive baseline data of 40). Therefore, the individual cognitive baseline data can be dynamically adjusted to 45 (that is, the cognitive load value ≤ 45 is judged as normal) to match the driver's actual cognitive ability.

[0144] Scenario Trigger: As the driver approaches an intersection, three events suddenly occur: The AR-HUD displays three pop-up windows in the interface feature data: "Pedestrian crossing," "Red light countdown 3 seconds," and "Vehicle overtaking from behind" (information density increases dramatically), and the complexity of the interaction steps rises to high complexity (requiring simultaneous braking and steering). At this time, the proportion of beta waves in the EEG data increases from 30% to 60% (high level of attentional tension), the heart rate surges to 95 beats per minute, and the HRV standard deviation drops to 20 ms. The eye movement fixation point jump frequency in the human behavior data increases to 5 times per second (disorganized saccades), and the brake pedal response delay is 0.5 seconds (exceeding the normal 0.3 seconds). Frowning appears in the human stress data (frequency 2 times per second), and voice commands become hesitant (speech rate drops to 80 words per minute).

[0145] Due to the "sudden increase in interface information + sudden operation", the dynamic coupling assessment model adjusted the weight of interface feature data to 60%, the weight of human physiological signal data to 25%, and the weight of human behavior data to 15%. The cognitive load value calculated by integrating each of the above multimodal data and its corresponding weight was 78 (the cognitive load level was high), triggering the intervention mechanism.

[0146] Initial intervention: Based on the high cognitive load level, hide unnecessary modules (turn off the entertainment interface and battery panel); aggregate pop-up information (merge 3 pop-ups into "Emergency: Pedestrian + Red Light + Overtaking, it is recommended to stop and give way"); highlight the brake prompt in AR-HUD (visual enhancement).

[0147] Feedback monitoring: Within 5 seconds after adjusting the human-computer interaction interface, the data showed that: the proportion of β waves in human physiological signal data decreased to 45%, and the heart rate decreased to 82 beats / minute; the eye movement fixation point in human behavioral data focused on the braking prompt, and the reaction delay was shortened to 0.2 seconds; the cognitive load value decreased to 42 (returning to the normal range), indicating that the intervention was effective.

[0148] If a vehicle suddenly experiences a tire blowout (a red "serious malfunction" alert pops up on the human-machine interface), and transient changes in human physiological signals are detected (e.g., a 200% increase in the P300 amplitude of the event-related potential (ERP) within 300ms, and a 20-beat / minute increase in heart rate within 1 second); or a sudden change in human behavioral data (e.g., a sudden increase in steering wheel effort to 80N, and a 0.4-second freeze in eye fixation), then a sudden change in cognitive load (dangerous cognitive load level) can be identified within 500ms. This triggers: a strong alert (seat vibration + voice prompt "Immediately slow down, turn on hazard lights"); automatic takeover of some functions (e.g., maintaining vehicle stability, reducing steering sensitivity); and the human-machine interface switching to "minimalist mode" (displaying only vehicle speed, fault location, and emergency lane guidance).

[0149] In the technical solution provided by the embodiments of the present invention, the peak cognitive load of the driver in complex scenarios is reduced by 40% (from 78 to 42), the reaction delay in sudden situations is shortened by 30% (from 0.5 seconds to 0.35 seconds), and the interface operation error rate is reduced by 55%, which verifies the effectiveness of the embodiments of the present invention.

[0150] The technical solution provided in this invention involves: determining the user's individual cognitive baseline data; determining the user's personalized cognitive load grading threshold based on the user's individual cognitive baseline data; acquiring multimodal data of the user in their work environment, including the user's human physiological signal data, human behavioral data, and interface feature data of the human-computer interaction interface; and determining the user's cognitive load assessment result based on the multimodal data and the user's personalized cognitive load grading threshold. The technical solution provided in this invention can effectively assess cognitive load, improving operational efficiency and safety.

[0151] In the technical solution provided by the embodiments of the present invention, interface feature data is used as the core input variable to achieve deep coupling between the human-computer interaction interface and the user's cognitive interaction.

[0152] The technical solution provided in this invention addresses the individual adaptation problem of the target evaluation model by calibrating based on individual cognitive baseline data.

[0153] The technical solution provided in this invention upgrades from passive detection to proactive cognitive load management by constructing a closed-loop system of "assessment-intervention-feedback".

[0154] The technical solution provided in this invention integrates three-dimensional spatial features and transient signal detection, adapting to complex human-computer interaction interfaces and highly dynamic task scenarios.

[0155] This invention provides a computer-readable storage medium including a stored program, wherein, when the program runs, it controls the device where the computer-readable storage medium is located to execute the steps of the above-described embodiment of the cognitive load assessment method based on multimodal data. For a detailed description, please refer to the above-described embodiment of the cognitive load assessment method based on multimodal data.

[0156] This invention provides an edge computing device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, they implement the steps of the above-described embodiment of the cognitive load assessment method based on multimodal data. For a detailed description, please refer to the above-described embodiment of the cognitive load assessment method based on multimodal data.

[0157] Figure 6 This is a schematic diagram of an edge computing device provided in an embodiment of the present invention. Figure 6 As shown, the edge computing device 20 of this embodiment includes a processor 21, a memory 22, and a computer program 23 stored in the memory 22 and executable on the processor 21. When the computer program 23 is executed by the processor 21, it implements the cognitive load assessment method based on multimodal data in this embodiment. To avoid repetition, it will not be described in detail here.

[0158] Edge computing device 20 includes, but is not limited to, processor 21 and memory 22. Those skilled in the art will understand that... Figure 6 This is merely an example of edge computing device 20 and does not constitute a limitation on edge computing device 20. It may include more or fewer components than shown, or combine certain components, or different components. For example, edge computing device may also include input / output devices, network access devices, buses, etc.

[0159] The processor 21 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0160] The memory 22 can be an internal storage unit of the edge computing device 20, such as a hard drive or RAM of the edge computing device 20. The memory 22 can also be an external storage device of the edge computing device 20, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the edge computing device 20. Furthermore, the memory 22 can include both internal and external storage units of the edge computing device 20. The memory 22 is used to store computer programs and other programs and data required by the edge computing device. The memory 22 can also be used to temporarily store data that has been output or will be output.

[0161] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0162] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.

[0163] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0164] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0165] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0166] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A cognitive load assessment method based on multimodal data, characterized in that, include: Determine the user's individual cognitive baseline data; The personalized cognitive load grading threshold for the user is determined based on the user's individual cognitive baseline data; Acquire multimodal data of the user in the work environment, including the user's human physiological signal data, human behavior data, and interface feature data of the human-computer interaction interface; The cognitive load assessment result of the user is determined based on the multimodal data and the user's personalized cognitive load classification threshold.

2. The method according to claim 1, characterized in that, Also includes: When the cognitive load assessment result of the user indicates that the user's cognitive load level exceeds the set cognitive load level threshold, the content displayed on the human-computer interaction interface and / or the interaction method are adjusted.

3. The method according to claim 1, characterized in that, The determination of the user's individual cognitive baseline data includes: Based on the obtained user attribute information and the results of the cognitive load standardized test, the user's cognitive profile is determined; Obtain a general cognitive baseline corresponding to the user's cognitive profile from a pre-acquired general cognitive model; Update the general cognitive baseline to obtain the user's individual cognitive baseline data.

4. The method according to claim 3, characterized in that, The process of updating the general cognitive baseline to obtain the user's individual cognitive baseline data includes: Based on the user's cognitive profile, construct at least two types of tasks with different levels of human-computer interaction complexity; Obtain the user's personnel capability data during the execution of the at least two types of tasks; The general cognitive baseline is updated based on the personnel ability data corresponding to the at least two types of tasks to obtain the individual cognitive baseline data of the user.

5. The method according to claim 4, characterized in that, The at least two types of tasks include a first task, a second task, and a third task whose human-computer interaction complexity increases sequentially. The step of updating the general cognitive baseline based on the personnel ability data corresponding to the at least two types of tasks to obtain the user's individual cognitive baseline data includes: Determine the differences in personnel ability data corresponding to the user when performing the first task and the third task, and update the obtained general cognitive baseline based on the differences; Based on the personnel ability data corresponding to the user when performing the second task, the validity of the adjusted general cognitive baseline is verified, and when the verification passes, the adjusted general cognitive baseline is determined to be the user's individual cognitive baseline data.

6. The method according to claim 1, characterized in that, Also includes: When the user's individual cognitive baseline data changes, the personalized cognitive load grading threshold will be adjusted according to a set ratio. If the deviation between the personalized cognitive load grading threshold and the standardized cognitive load test result is greater than the set deviation threshold, then the user's individual cognitive baseline data is updated, and the step of determining the user's personalized cognitive load grading threshold based on the user's individual cognitive baseline data continues.

7. The method according to claim 1, characterized in that, The user's cognitive load assessment result includes the user's cognitive load level. Determining the user's cognitive load assessment result based on the multimodal data and the user's personalized cognitive load grading threshold includes: Based on the human physiological signal data and human behavioral data in the multimodal data, determine the temporal variation characteristics of the human physiological signal data and human behavioral data; Based on the interface feature data of the human-computer interaction interface, determine the spatial distribution characteristics of the interface feature data of the human-computer interaction interface. Based on the current task stage of the user, a weight allocation strategy is determined according to the temporal change characteristics and the spatial distribution characteristics. The weight allocation strategy is a weight allocation strategy for the human physiological signal data, the human behavior data and the interface feature data in the multimodal data. The weight allocation strategy includes the feature vector of each multimodal data and the weight coefficient corresponding to each feature vector. A feature matrix is ​​generated based on each feature vector and its corresponding weight coefficient. The feature matrix is ​​input into the obtained fully connected network to output the cognitive load value; The cognitive load level of the user is determined based on the cognitive load value and the user's personalized cognitive load grading threshold.

8. The method according to claim 1, characterized in that, The user's cognitive load assessment result includes a cognitive load level. After determining the user's cognitive load assessment result based on the multimodal data and the user's personalized cognitive load grading threshold, the process includes: Adjust the human-computer interaction interface according to the interface adjustment rules corresponding to the cognitive load level; The cognitive load level is continuously monitored. If the cognitive load level does not decrease within a preset time range, the interface adjustment rules corresponding to the next higher cognitive load level are obtained. The human-computer interaction interface is adjusted according to the interface adjustment rules corresponding to the higher level of cognitive load.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 8.

10. An edge computing device, comprising a memory and a processor, wherein the memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions, characterized in that, When the program instructions are loaded and executed by the processor, they implement the steps of the method described in any one of claims 1 to 8.