Internet of Things data man-machine interaction visualization system based on artificial intelligence
By using multimodal data acquisition and feedback control loops optimized by deep reinforcement learning agents, the human-machine interface is dynamically adjusted, solving the problem of insufficient adaptability of operator cognitive state in existing technologies. This achieves real-time quantification and stable maintenance of operator cognitive state, improving human-machine collaboration efficiency and safety.
Patent Information
- Application Number
- CN202511034081.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-11
AI Technical Summary
Existing human-machine interfaces cannot adapt to the dynamic changes in operators' cognitive states under different task pressures in complex industrial IoT environments, resulting in insufficient or overloaded information and affecting decision-making quality.
A feedback control loop is constructed by employing a multimodal data acquisition unit, a real-time cognitive load state inference module, an adaptive human-machine interface generation module, and a closed-loop adjustment and control module. This allows for real-time perception of the operator's status, dynamic adjustment of the interface presentation, and optimization of the interface adjustment through a deep reinforcement learning agent.
It enables real-time, accurate quantification of operator cognitive status and stable maintenance within the efficient decision-making range, improving human-machine collaboration efficiency and safety, overcoming noise interference and information bias of single-modal data, and ensuring the robustness and human-friendly adjustment of the system.
Smart Images

Figure CN120928948A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of human-computer interaction technology, industrial Internet of Things (IoT), and artificial intelligence, specifically to an IoT data human-computer interaction visualization system based on artificial intelligence. Background Technology
[0002] In complex industrial IoT environments, existing human-machine interface designs are mostly static and cannot adapt to the dynamic changes in the operator's cognitive state under different task pressures. This often leads to decreased alertness due to insufficient information during routine monitoring, or reduced decision-making quality due to information overload when handling emergencies. The core technical challenge is that it is difficult to quantify the operator's cognitive load in real time, accurately and non-intrusively, and the dynamic adjustment of the interface itself may also constitute a new source of interference.
[0003] To address the aforementioned issues, this invention provides an AI-based IoT data human-computer interaction visualization system. This system aims to infer the operator's cognitive load in real time and, based on this inference, dynamically adjust the information presentation method of the human-computer interaction interface through a closed-loop adaptive mechanism, thereby maintaining the operator's cognitive state within a stable range conducive to efficient decision-making.
[0004] The information disclosed in the background section above is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide an artificial intelligence-based Internet of Things (IoT) data human-computer interaction visualization system to solve the problems mentioned in the background art.
[0006] The technical solution of the present invention includes: a multimodal data acquisition unit, a real-time cognitive load state inference module, a human-computer interaction interface adaptive generation module, and a closed-loop regulation and control module;
[0007] A multimodal data acquisition unit is used to acquire multimodal data related to the operator's status in real time;
[0008] The real-time cognitive load state inference module is used to infer and generate a probability vector representing the operator's current cognitive load state based on the acquired multimodal data.
[0009] The human-computer interaction interface adaptive generation module is used to take the cognitive load state probability vector as the current state in order to generate the optimal interface adjustment action.
[0010] The closed-loop adjustment and control module is used to execute the optimal interface adjustment action to update the visualization interface. The operator status change caused by the updated visualization interface is acquired again by the multimodal data acquisition unit to form a closed-loop control.
[0011] Preferably, the multimodal data includes interaction behavior data, eye-tracking data, and physiological data;
[0012] Interactive behavior data includes the smoothness of mouse trajectory and the rate of keyboard input; eye-tracking data includes changes in pupil diameter and the distribution entropy of fixation point; physiological data consists of heart rate variability indicators collected through wearable devices.
[0013] Preferably, the real-time cognitive load state inference module performs the following processing:
[0014] Feature extraction is performed on multimodal data to generate a unified feature vector, and a pre-trained long short-term memory network model is used to process the unified feature vector to calculate the cognitive load state probability vector.
[0015] Preferably, the cognitive load state probability vector is a three-dimensional vector, and the components of the cognitive load state probability vector represent the probabilities of the operator being in three discrete states: insufficient cognitive load, optimal load, and cognitive overload.
[0016] Preferably, the human-computer interaction interface adaptive generation module has a built-in deep reinforcement learning intelligent agent;
[0017] The deep reinforcement learning agent responds to the cognitive load state probability vector of the input and outputs the optimal interface adjustment action according to the policy network.
[0018] Interface adjustment actions are selected from a predefined set of actions, which includes adjustments to interface information density, abstraction level, and presentation modality.
[0019] Preferably, the deep reinforcement learning agent optimizes the policy network to maximize the long-term accumulated reward value.
[0020] Preferably, the state deviation value is determined;
[0021] The state deviation value is used to quantify the deviation between the newly generated cognitive load state probability vector and the preset target state vector after the interface adjustment action is executed;
[0022] The preset target state vector is a baseline probability vector representing the optimal load state. The baseline probability vector is derived by statistical modeling the cognitive state data of experienced operators when efficiently completing tasks.
[0023] Preferably, the motion disturbance value is determined;
[0024] The motion perturbation value is a scalar used to quantify the degree of perturbation of the interface adjustment motion itself. It is calculated based on the changes in presentation modality, relative changes in information density, and changes in layout caused by the motion.
[0025] Preferably, the reward value is generated by combining the state deviation value and the action disturbance value.
[0026] Preferably, the Long Short-Term Memory (LSTM) network model is trained using a federated learning framework, which allows local data from multiple operators to contribute gradient updates to the LTM network model without leaving the local device.
[0027] This invention provides an improved, artificial intelligence-based IoT data human-computer interaction visualization system, which has the following improvements and advantages compared with the prior art:
[0028] 1. This solution constructs a complete feedback control loop by setting up a multimodal data acquisition unit, a real-time cognitive load state inference module, a human-machine interface adaptive generation module, and a closed-loop adjustment and control module. This loop begins with the continuous perception of the operator's state, proceeds through state inference and decision generation, and ultimately affects the interface update. The effect of the interface update is then perceived again, forming a dynamic process of continuous optimization. This architecture transforms the human-machine interface from a passive display into an intelligent system that can actively participate in and adjust the operator's cognitive process, thereby improving the efficiency and safety of human-machine collaboration.
[0029] 2. The real-time cognitive load state inference module of this solution overcomes the shortcomings of single-modal data being susceptible to noise interference and having one-sided information by integrating three different dimensions of information: interactive behavior data, eye-tracking data, and physiological data. This ensures that the model has high generalization ability and robustness while protecting user privacy.
[0030] 3. The adaptive generation module of the human-computer interaction interface in this solution incorporates a deep reinforcement learning agent. This agent learns the optimal strategy by maximizing a long-term accumulated reward value, ensuring the objectivity of cost assessment. The agent in this solution may learn that a slight action, such as adjusting the data display modality from a complex chart to a simplified list, although improving the state more slowly, has a higher long-term accumulated reward value due to its extremely low perturbation value. This subtle strategy acquired through learning is unparalleled by existing technologies. It ensures that the system's adjustment behavior is efficient and human-centered, maintaining the operator's cognitive state stably within an ideal range conducive to efficient decision-making. Attached Figure Description
[0031] The present invention will be further explained below with reference to the accompanying drawings and embodiments:
[0032] Figure 1 This is a flowchart of the system of the present invention. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0034] Example 1:
[0035] Please see Figure 1 The present invention provides an artificial intelligence-based Internet of Things data human-computer interaction visualization system, comprising: a multimodal data acquisition unit, a real-time cognitive load state inference module, a human-computer interaction interface adaptive generation module, and a closed-loop regulation and control module;
[0036] A multimodal data acquisition unit is used to acquire multimodal data related to the operator's status in real time;
[0037] The real-time cognitive load state inference module is used to infer and generate a probability vector representing the operator's current cognitive load state based on the acquired multimodal data.
[0038] The human-computer interaction interface adaptive generation module is used to take the cognitive load state probability vector as the current state in order to generate the optimal interface adjustment action.
[0039] The closed-loop adjustment and control module is used to execute the optimal interface adjustment action to update the visual interface. The operator status change caused by the updated visual interface is acquired again by the multimodal data acquisition unit to form a closed-loop control.
[0040] This embodiment provides an AI-based IoT data human-computer interaction visualization system. The inherent goal of the system architecture is to overcome the limitations of existing static interfaces that cannot respond to the dynamically changing cognitive states of operators. The system integrates a multimodal data acquisition unit, a real-time cognitive load inference module, an adaptive human-computer interface generation module, and a closed-loop adjustment control module to construct a recursive feedback control loop. The system's operation begins with the multimodal data acquisition unit continuously sensing the operator's state. The acquired data is deeply analyzed by the real-time cognitive load inference module to output a quantitative cognitive state assessment result. This assessment result serves as the basis for decision-making and is transmitted to the adaptive human-computer interface generation module to generate the optimal interface adjustment strategy. The closed-loop adjustment control module is responsible for precisely executing this strategy to update the interface presentation. The impact of the interface update on the operator's state is then captured again by the data acquisition unit, thus forming a continuous adjustment cycle. The technical essence of this design is to achieve real-time, intelligent, and dynamic reconstruction of the human-computer interface, maintaining the operator's cognitive resources within the most efficient range conducive to complex decision-making, thereby improving task performance and safety in complex scenarios such as the Industrial Internet of Things.
[0041] This invention achieves a fundamental shift from static information presentation to dynamic closed-loop regulation. Existing interactive interfaces are essentially one-way, static information output tools. This solution constructs a complete feedback control loop by setting up a multimodal data acquisition unit, a real-time cognitive load state inference module, a human-computer interaction interface adaptive generation module, and a closed-loop regulation control module. This loop begins with the continuous perception of the operator's state, proceeds through state inference and decision generation, and ultimately affects the interface update. The effect of the updated interface is then perceived again, forming a continuously optimizing dynamic process. This architecture transforms the human-computer interaction interface from a passive display into an intelligent system that can actively participate in and regulate the operator's cognitive process, thereby improving the efficiency and safety of human-computer collaboration.
[0042] Example 2
[0043] Multimodal data includes interactive behavior data, eye-tracking data, and physiological data;
[0044] Interactive behavior data includes the smoothness of mouse trajectories and the rate of keyboard input; eye-tracking data includes changes in pupil diameter and the distribution entropy of the fixation point; physiological data consists of heart rate variability indicators collected through wearable devices.
[0045] The real-time cognitive load state inference module performs the following processing:
[0046] Feature extraction is performed on multimodal data to generate a unified feature vector, and a pre-trained long short-term memory network model is used to process the unified feature vector to calculate the cognitive load state probability vector.
[0047] In a preferred embodiment, the pre-trained long short-term memory network model consists of two stacked LSTM layers, each containing 128 hidden units. A fully connected layer containing 3 neurons is connected after the output of the second LSTM layer, and a Softmax activation function is used to output probabilities representing three cognitive load states. To prevent overfitting, a dropout rate of 0.3 is set between the LSTM layers.
[0048] To make the feature extraction process more explicit, those skilled in the art can use the following methods for quantification:
[0049] The smoothness of the mouse trajectory can be quantified by calculating the reciprocal of the arc length of the trajectory's spectrum. This method evaluates the smoothness of the motion through Fourier transform. The rate of keyboard input can be calculated as the number of effective keystrokes in the past minute.
[0050] The gaze distribution entropy can be calculated by dividing the screen into an N×M grid and counting the frequency p of gaze points falling into each grid i within a time window T. i Then, using the Shannon entropy formula:
[0051]
[0052] Where H: the distribution entropy of the gaze point; i: the index number of the grid; N×M: the total number of grids in the screen; p i Within a time window T, the frequency at which the fixation point falls into the i-th grid is a heart rate variability index. Specifically, a time-domain index sensitive to short-term psychological stress can be selected—the root mean square of the difference between adjacent heartbeat intervals—which is calculated from the interpeak period of the PPG signal collected by the wearable device.
[0053] In this embodiment, the multimodal data acquisition unit ensures the comprehensiveness and robustness of state assessment by fusing information sources from three different dimensions. Interactive behavior data, such as mouse trajectory smoothness and keyboard input rate, reflects the operator's explicit operation patterns in performing tasks. Eye-tracking data, such as pupil diameter changes and fixation point distribution entropy, reveals the allocation of visual attention and the depth of cognitive processing. Physiological data, such as heart rate variability, collected through wearable devices, provides objective physiological indicators for measuring the activity of the autonomic nervous system. The real-time cognitive load state inference module receives these unified feature vectors formed after preprocessing and feature extraction, and processes them using a pre-trained long short-term memory network model. The technical reason for choosing the long short-term memory network model is that cognitive states are time-dependent, and this network architecture can effectively capture this temporal dynamic. To solve the inherent uncertainty and noise sensitivity of single data modalities, the module adopts multi-source information fusion to improve the accuracy and anti-interference ability of inference, mapping the multimodal input to the quantitative output of cognitive load state through the following formula:
[0054]
[0055] Among them, Ψ t Let f be the probability vector of cognitive load state output at time t. θ D represents a pre-trained multilayer long short-term memory network model defined by a parameter set θ. b (t),D e (t),D p (t) represents the normalized feature vectors of interactive behavior, eye tracking, and physiological data after feature engineering at time t, and λ represents the feature vectors of these data. b ,λ e ,λ p These are three dimensionless modal weighting coefficients, whose values were determined based on prior experimental data. They are used to adjust the contribution of different data sources and their sum is 1. This is a vector concatenation operation, where t is an instant in time, and θ is the set of parameters defining the weights and biases within the model; t is an instant in time.
[0056] In one embodiment, experimentally calibrated, the aforementioned dimensionless modal weight coefficient can be set as: interaction behavior data weight λ b =0.2, eye-tracking data weight λ e =0.5, physiological data weight λ p =0.3; These values reflect that eye-tracking data generally has high informational value in cognitive load inference;
[0057] The application of this formula lies in the module's input of weighted, fused, and concatenated collected and processed data into the model to calculate the probability vector Ψ. t This lays the data foundation for subsequent adaptive adjustments;
[0058] This invention provides a real-time, objective, and quantitatively accurate assessment of operator cognitive load. Existing technologies lack reliable, non-invasive, real-time quantification methods. The real-time cognitive load state inference module of this solution overcomes the shortcomings of single-modal data, such as susceptibility to noise interference and incomplete information, by integrating interactive behavior data, eye-tracking data, and physiological data from three different dimensions. Its core cognitive state inference model, namely the Long Short-Term Memory (LSTM) network model, is not a simple mathematical derivation but a functional mapper selected based on the temporal characteristics of the problem. This model maps high-dimensional, heterogeneous input data into a normalized cognitive load state probability vector using the following formula:
[0059]
[0060] The practical significance of this formula lies in the fact that it represents the noisy and modally diverse data streams D collected from the physical world. b (t),D e (t),D p (t), through a pre-trained nonlinear time series model f θ The transformation ultimately outputs a dimensionless three-dimensional vector Ψ with a definite probabilistic meaning. t This vector directly quantifies the probability that the operator is under-cognitive, optimally cognitively loaded, or cognitively overloaded. To initially adjust the contribution of each modality feature before inputting it into the model, each feature vector is first multiplied by its preset weight coefficient λ for scaling, and then the scaled vectors are concatenated. Here, the weight coefficient λ reflects the importance of each modality in cognitive load inference in prior knowledge. More importantly, this model is trained based on a federated learning framework that allows multiple operators' local data to contribute to gradient updates without leaving their local devices, ensuring that the model has high generalization ability and robustness while protecting user privacy.
[0061] Example 3
[0062] The cognitive load state probability vector is a three-dimensional vector. The components of the cognitive load state probability vector represent the probabilities of the operator being in three discrete states: insufficient cognitive load, optimal load, and cognitive overload.
[0063] The Long Short-Term Memory (LSTM) network model is trained using a federated learning framework, which allows local data from multiple operators to contribute gradient updates to the LTM network model without leaving the local device.
[0064] To obtain the realistic cognitive load labels required for training, this invention employs a dual verification method: During the experimental phase, after completing a series of calibration tasks with increasing difficulty, operators must immediately fill out the NASA-TLX (National Aeronautics and Space Administration Mission Load Index) subjective assessment scale; simultaneously, their task completion time and error rate are recorded; combining objective task performance data (e.g., error rate exceeding a threshold is defined as overload) and subjective scale scores (e.g., scores below a threshold are defined as underload), the collected multimodal data segments are labeled with discrete labels of underload, optimal load, or cognitive overload, serving as the baseline truth for model training; in a calibration task, when the task error rate exceeds 20%, the corresponding data segment can be labeled as cognitive overload; when the operator's submitted NASA-TLX scale composite score is below 30 out of 100, it is labeled as underload. These thresholds can be adjusted based on the task's baseline test;
[0065] In this embodiment, the representation of cognitive load state is defined in a refined manner; the cognitive load state probability vector Ψ t Designed as a three-dimensional vector <C u C o C x The components represent the probabilities of an operator being in three states: cognitive underload, optimal load, and cognitive overload, with the sum of all three being 1. This probabilistic representation, compared to deterministic classification, better reflects the inherent uncertainty of psychological state assessment. To further clarify the construction method of the pre-trained model, the Long Short-Term Memory (LSTM) network model is trained using a federated learning framework. This framework allows for the aggregation of gradient updates from various clients to train a globally shared model without directly accessing or transmitting the raw local data of multiple operators. This design not only protects user data privacy, especially in industrial environments, but also significantly improves model performance by utilizing diverse data from different individuals. θ Its generalization ability and adaptability to individual differences.
[0066] Example 4
[0067] The human-computer interaction interface adaptive generation module has a built-in deep reinforcement learning intelligent agent.
[0068] The deep reinforcement learning agent responds to the cognitive load state probability vector of the input and outputs the optimal interface adjustment action according to the policy network.
[0069] The interface adjustment actions are selected from a predefined set of actions, which includes adjustments to the interface information density, abstraction level, and presentation modality.
[0070] The deep reinforcement learning agent can use the proximal policy optimization algorithm; the internal policy network and value network both adopt a multilayer perceptron structure, which can be composed of two hidden layers containing 64 neurons each and an output layer. The hidden layers use the ReLU activation function.
[0071] To clarify the agent's decision space, a predefined set of actions includes the following seven discrete actions: Remain unchanged: Do not perform any interface adjustments; Increase information density: Display an additional secondary related information in a preset area; Decrease information density: Hide the lowest priority information in the current interface; Increase abstraction level: Switch from a detailed data list view to an aggregated statistical chart view; Decrease abstraction level: Expand the aggregated statistical chart view into a detailed data list; Switch to graphical modality: Display the selected data as a highlighted portion on a 3D device model; Switch to text modality: Restore the graphical data display to text or numbers.
[0072] In this embodiment, the human-computer interaction interface adaptive generation module has a built-in deep reinforcement learning agent; this agent generates the cognitive load state probability vector Ψ output by the previous module. t As a response to the current environmental state S t The agent perceives this state input; based on this state input, the agent utilizes its internal policy network π. φ (A t |S t Make a decision and output an optimal interface adjustment action A aimed at maximizing the long-term goal. t Available interface adjustment actions A t Predefined within a set of actions, this set covers key dimensions of interface design adjustments, such as increasing or decreasing the number of information items displayed to adjust information density, switching between overview views and detailed data to change the level of abstraction, or switching data from a text list to a 3D model highlight to change the presentation modality. This mechanism enables the system to flexibly and purposefully reconstruct the interactive interface based on the operator's real-time status, thereby proactively managing the operator's cognitive needs.
[0073] Example 5
[0074] Deep reinforcement learning agents optimize policy networks to maximize long-term cumulative rewards;
[0075] Determine the state deviation value;
[0076] The state deviation value is used to quantify the deviation between the newly generated cognitive load state probability vector and the preset target state vector after the interface adjustment action is executed;
[0077] The preset target state vector is a baseline probability vector representing the optimal load state. The baseline probability vector is derived by statistical modeling based on the cognitive state data of experienced operators when efficiently completing tasks.
[0078] Determine the motion disturbance value;
[0079] The motion perturbation value is a scalar used to quantify the degree of perturbation of the interface adjustment motion itself. It is calculated based on the changes in presentation modality, relative changes in information density, and changes in layout caused by the motion.
[0080] A reward value is generated by combining the state deviation value and the action perturbation value.
[0081] In this embodiment, after statistical modeling of data from multiple experienced operators, the predetermined target state vector is determined to be a specific probability distribution: Ψ ref =<0.1,0.8,0.1>; This vector indicates that an ideal optimal load state allows for a slight probability of bias towards the other two states, rather than an absolute <0,1,0>.
[0082] In this embodiment, the policy network optimization process of the deep reinforcement learning agent is guided by a reward function aimed at balancing the control objective and control cost; the generation of this reward value integrates the state deviation value and the action perturbation value; the agent executes action A at time t. t Then, the reward R obtained at the next time t+1 t+1 Calculated using the following formula:
[0083] R t+1 =α·exp(-β||Ψ t+1 -Ψ ref || 2 )-γ·Δ a
[0084] Among them, R t+1 Let exp be the dimensionless scalar reward value returned to the agent at time t+1, and let exp be the exponential function. t+1 To perform action A t The new cognitive load state vector, Ψ, is then calculated by the inference module. ref This is a preset target state vector, a baseline value derived from statistical modeling of the cognitive state data of experienced operators during efficient task execution. It represents the ideal optimal load state. 2To quantify the squared Euclidean distance between the current state and the ideal state, α, β, and γ are dimensionless constant hyperparameters, serving as the positive reward scaling factor, the state deviation penalty sensitivity coefficient, and the interface adjustment cost coefficient, respectively. Their values were determined through extensive experiments in a digital twin simulation environment to maximize the system's macroscopic performance indicators. Δ a Quantify action A t The dimensionless scalar of the degree of self-perturbation is calculated as follows:
[0085] Δ a =ω m δ m +ω d δ d +ω l δ l
[0086] Where, δ m δ is a modal change indicator, a binary variable. If an interface adjustment action triggers a switch in presentation modality, such as switching from text to graphics, then δ... m =1, otherwise δ m =0; δ d The absolute value of the relative change in information density, δ d Defined as the ratio of the absolute value of the change in the number of information items to the number before the change, i.e., |count new -count old | / count old The result is a dimensionless relative rate of change; δ l The layout change δ is a quantitative value representing the degree of change in the layout of key interface elements. l Defined as the weighted average of the movement distances of key elements affected in the interface, normalized by dividing by the screen diagonal length to make it a dimensionless value between 0 and 1; ω m ,ω d ,ω l The preset weights corresponding to the above changes are derived from experimental data based on human factors engineering.
[0087] In one embodiment, the weights used to calculate the motion perturbation value can be set as: modal change weights ω m =0.5, information density change weight ω d =0.3, layout change weight ω l =0.2;
[0088] To balance the goals and costs of regulation, the hyperparameters were optimized to the following reference values in a set of simulation experiments: positive reward scaling factor α = 1.0, state deviation penalty sensitivity coefficient β = 5.0, and interface regulation cost coefficient γ = 0.15;
[0089] The underlying logic of this reward function lies in the fact that by continuously optimizing its policy network to maximize the long-term expected value of the reward, the agent's policy network, after optimization, can form a complex control strategy. This strategy can guide the operator's state Ψ to approach the ideal state Ψ. ref It can also weigh the visual or cognitive disturbance Δ caused by this action. a This generates a series of smooth, efficient, and minimally disruptive interface adjustment actions, keeping the operator stably in the optimal working state.
[0090] This invention achieves adaptive optimization and precise control of interface adjustment strategies. Existing interface adjustments often rely on preset, rigid rules, failing to find optimal solutions in complex and ever-changing situations. The adaptive generation module for the human-computer interaction interface in this solution incorporates a deep reinforcement learning agent, which learns the optimal strategy by maximizing a long-term accumulated reward value. The design of its reward function is one of the key advancements of this solution.
[0091] R t+1 =α·exp(-β||Ψ t+1 -Ψ ref || 2 )-γ·Δ a
[0092] This reward function is not arbitrary; its structure follows the benefit-cost principle in optimal control theory. The first term on the right-hand side of the formula is a Gaussian benefit term used to reward the system state Ψ. t+1 Compared with the preset baseline probability vector Ψ representing the optimal load state ref The closer the approximation, the smaller the state deviation value, and the higher the reward; the second term γ·Δ a This is a linear cost term used to penalize interface adjustment action A. t The disturbance itself; the motion disturbance value Δ here. a It is a quantitative scalar calculated by weighting the changes in presentation modality caused by the action, the relative change in information density, and the degree of layout change, thus ensuring the objectivity of cost assessment;
[0093] The practical significance of this formula lies in providing the agent with a precise, dual-objective optimization direction: not only to effectively guide the operator's cognitive state to the optimal range, but also to complete this process in the smoothest and most interference-free way possible. For example, when faced with operator cognitive overload, a system based on simple rules might crudely remove a large amount of information, but this itself might be a drastic layout change that causes secondary interference. The agent in this scheme, however, might learn to discover that a slight action—adjusting the data display modality from a complex chart to a simplified list—while improving the state more slowly, has a much higher long-term cumulative reward value due to its extremely low perturbation value. This sophisticated strategy acquired through learning is unparalleled by existing technologies. It ensures that the system's regulatory behavior is efficient and human-like, stably maintaining the operator's cognitive state within an ideal range conducive to efficient decision-making.
[0094] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An artificial intelligence-based IoT data human-computer interaction visualization system, characterized in that, include: Multimodal data acquisition unit, real-time cognitive load state inference module, adaptive human-computer interaction interface generation module, and closed-loop regulation and control module; A multimodal data acquisition unit is used to acquire multimodal data related to the operator's status in real time; The real-time cognitive load state inference module is used to infer and generate a probability vector representing the operator's current cognitive load state based on the acquired multimodal data. The human-computer interaction interface adaptive generation module is used to take the cognitive load state probability vector as the current state in order to generate the optimal interface adjustment action. The closed-loop adjustment and control module is used to execute the optimal interface adjustment action to update the visual interface. The operator status change caused by the updated visual interface is acquired again by the multimodal data acquisition unit to form a closed-loop control.
2. The artificial intelligence-based IoT data human-computer interaction visualization system according to claim 1, characterized in that, The multimodal data includes interactive behavior data, eye-tracking data, and physiological data; Interactive behavior data includes the smoothness of mouse trajectory and the rate of keyboard input; eye-tracking data includes changes in pupil diameter and the distribution entropy of fixation point; physiological data consists of heart rate variability indicators collected through wearable devices.
3. The artificial intelligence-based IoT data human-computer interaction visualization system according to claim 1, characterized in that, The real-time cognitive load state inference module performs the following processing: Feature extraction is performed on multimodal data to generate a unified feature vector, and a pre-trained long short-term memory network model is used to process the unified feature vector to calculate the cognitive load state probability vector.
4. The artificial intelligence-based IoT data human-computer interaction visualization system according to claim 3, characterized in that, The cognitive load state probability vector is a three-dimensional vector, and the components of the cognitive load state probability vector represent the probability that the operator is in one of three discrete states: insufficient cognitive load, optimal load, and cognitive overload.
5. The artificial intelligence-based IoT data human-computer interaction visualization system according to claim 1, characterized in that, The human-computer interaction interface adaptive generation module has a built-in deep reinforcement learning intelligent agent. The deep reinforcement learning agent responds to the cognitive load state probability vector of the input and outputs the optimal interface adjustment action according to the policy network. Interface adjustment actions are selected from a predefined set of actions, which includes adjustments to interface information density, abstraction level, and presentation modality.
6. The artificial intelligence-based IoT data human-computer interaction visualization system according to claim 5, characterized in that, The deep reinforcement learning agent optimizes the policy network to maximize the long-term accumulated reward value.
7. The artificial intelligence-based IoT data human-computer interaction visualization system according to claim 6, characterized in that, Determine the state deviation value; The state deviation value is used to quantify the deviation between the newly generated cognitive load state probability vector and the preset target state vector after the interface adjustment action is executed; The preset target state vector is a baseline probability vector representing the optimal load state. The baseline probability vector is derived by statistical modeling the cognitive state data of experienced operators when efficiently completing tasks.
8. The artificial intelligence-based IoT data human-computer interaction visualization system according to claim 6, characterized in that, Determine the motion disturbance value; The motion perturbation value is a scalar used to quantify the degree of perturbation of the interface adjustment motion itself. It is calculated based on the changes in presentation modality, relative changes in information density, and changes in layout caused by the motion.
9. The artificial intelligence-based IoT data human-computer interaction visualization system according to claim 8, characterized in that, A reward value is generated by combining the state deviation value and the action disturbance value.
10. The artificial intelligence-based IoT data human-computer interaction visualization system according to claim 3, characterized in that, The Long Short-Term Memory (LSTM) network model is trained using a federated learning framework, which allows local data from multiple operators to contribute gradient updates to the LTM network model without leaving the local device.
Citation Information
Patent Citations
Self-adaptive human-computer interface configuration method
CN109558005A
Man-machine interaction system and man-machine interaction method
CN118151763A
Cognitive level evaluation system and method applied to caregiver of cerebral infarction patient
CN119700019A
System and method for real-time biodata analysis and personalized response adaptation
US20250225415A1
Cited By
Multi-modal data fusion edge computing gateway and AI processing method
CN121333963A