Open office micro-environment collaborative control system and method based on multi-agent game and cognitive load adaptation

CN122043982BActive Publication Date: 2026-09-15SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610119218.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-28
Publication Date
2026-09-15
Estimated Expiration
2046-01-28

AI Technical Summary

Technical Problem

[0006]本发明的目的在于克服现有技术中存在的缺点与不足,提供一种基于多智能体博弈与认知负荷自适应的开放式办公微环境协同控制系统,该系统有效解决传统PCS系统存在的邻域冲突问题,在保障个体认知状态最优的同时实现多用户满意度的帕累托均衡

Benefits of technology

[0054] (1) This invention quantifies neighbor interference by constructing an environmental coupling matrix and establishes a multi-agent non-cooperative game model that dynamically adjusts the conflict penalty term with the cognitive state vector. It can realize multi-workstation collaborative control in open office spaces, effectively solve the neighborhood conflict problem in traditional PCS systems, and achieve Pareto equilibrium of multi-user satisfaction while ensuring the optimal cognitive state of individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122043982B_ABST
    Figure CN122043982B_ABST
Patent Text Reader

Abstract

The application discloses an open office micro-environment collaborative control system and method based on multi-agent game and cognitive load adaptation, and the system comprises a data sensing module, which is used for collecting user physiological data and workstation environment data and establishing an environment coupling matrix; a state evaluation module, which is used for generating a user real-time cognitive state vector and an environment regulation target function based on the data; a game decision module, which is used for taking each workstation environment regulation target function as an optimization benchmark, taking the environment coupling matrix as a constraint, constructing a multi-agent non-cooperative game model, and obtaining a collaborative control strategy by solving the Nash equilibrium of the model; a compensation execution module, which is used for executing the control strategy and starting cross-sensory modal compensation when physical adjustment fails to reach the target; and an adaptive reinforcement learning feedback module, which is used for adjusting the revenue function parameters according to user feedback. The application realizes multi-target collaborative optimization of the office environment, effectively relieves neighbor conflicts under the premise of protecting privacy, and actively improves user cognitive efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent control technology for smart buildings and human settlements. Specifically, it relates to an open office micro-environment collaborative control system and method based on multi-agent game theory and cognitive load adaptation. Background Technology

[0002] With the development of IoT technology and smart building concepts, modern office environment control systems are gradually shifting from traditional centralized HVAC control to personalized micro-environment regulation. Traditional environmental control mainly relies on the Fanger thermal comfort equation (PMV-PPD model), which makes one-size-fits-all adjustments based on statistical averages. This makes it difficult to adapt to significant differences in individual age, gender, metabolic rate, and clothing habits, leading to frequent instances of catering to diverse needs.

[0003] To improve individual comfort, existing technologies are beginning to explore the introduction of workstation-level microenvironment control systems (PCS), allowing users to improve their local environment using desktop devices such as fans, heating pads, and adjustable lighting. However, the following shortcomings still exist in practical applications:

[0004] First, in open-plan office spaces, there is a significant coupling effect in the physical environment of adjacent workstations. Existing PCS systems typically treat each workstation as an independent entity, adjusting them only according to the instructions of a single user, which can easily lead to conflicts between adjacent workstations. Second, existing systems have a single control objective, mostly based on thermal comfort indicators, without considering the direct impact of environmental parameters on users' cognitive state and work efficiency. Third, environmental adjustment methods are limited in scope and lack cross-modal compensation capabilities. Fourth, there is a conflict between physiological perception and privacy protection.

[0005] Existing research has attempted to use multi-agent game theory to coordinate conflicts among multiple parties and to assess user cognitive load through physiological sensing technology. However, existing game models are mostly based on fixed preferences and perfect rationality assumptions, failing to consider the user's real-time changing cognitive state as a core variable in decision-making, and lacking cross-modal compensation mechanisms when physical adjustment is limited. Therefore, there is an urgent need to research a collaborative workstation control system that can resolve conflicts in the vicinity of multiple workstations, proactively adapt to the user's cognitive state, possess cross-modal compensation capabilities, and protect privacy. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings and deficiencies of the existing technology and provide an open office micro-environment collaborative control system based on multi-agent game theory and cognitive load adaptation. This system effectively solves the neighborhood conflict problem in traditional PCS systems and achieves Pareto equilibrium of multi-user satisfaction while ensuring the optimal cognitive state of individuals.

[0007] The second objective of this invention is to provide a collaborative control method for an open office micro-environment based on multi-agent game theory and cognitive load adaptation.

[0008] The objective of this invention is achieved through the following technical solution: an open office micro-environment collaborative control system based on multi-agent game theory and cognitive load adaptation, comprising:

[0009] The data sensing module is used to collect users' physiological data and workstation environmental data through multimodal sensors deployed at each workstation, and to establish an environmental coupling matrix that characterizes the intensity of environmental interference between workstations through active detection.

[0010] The state assessment module is used to map and generate the user's real-time cognitive state vector and the corresponding environmental regulation objective function based on the physiological data and environmental data, according to the Yerkes-Dodson law.

[0011] The game decision-making module is used to construct a multi-agent non-cooperative game model with the objective function of each workstation's environmental regulation as the optimization benchmark and the environmental coupling matrix as the neighborhood conflict constraint. The cooperative control strategy is obtained by solving the Nash equilibrium of the model. The payoff function of each workstation agent in the model includes: a cognitive performance gain term based on the environmental regulation objective of the workstation, and a neighborhood conflict penalty term calculated based on the environmental coupling matrix, the penalty intensity of which is dynamically adjusted by the real-time cognitive state vector of the adjacent workstation agents.

[0012] The compensation execution module is used to execute the collaborative control strategy and initiate cross-sensory modality compensation adjustment when the execution result does not meet the environmental regulation target based on the current cognitive state vector.

[0013] An adaptive reinforcement learning feedback module is used to adaptively adjust the parameters of the reward function based on the user's physiological and behavioral feedback.

[0014] Preferably, the data sensing module includes:

[0015] The sensing and monitoring unit is used to acquire user heart rate variability, respiratory rate and skin conductance data through millimeter-wave radar, acquire user skin temperature through infrared sensors, and acquire basic environmental parameters of the work area through environmental sensors.

[0016] The edge computing unit is used to perform real-time calculations on the heart rate variability, respiratory rate, and skin conductance data, and to fuse user skin temperature with basic environmental parameters.

[0017] The environmental coupling unit is used to calculate the spillover coefficient of environmental parameters in the adjacent area by the control actions of the equipment at this workstation using the feedback from the sensors at adjacent workstations, and to form an environmental coupling matrix.

[0018] Preferably, the state assessment module includes:

[0019] The cognitive state assessment unit is used to construct a multi-dimensional feature input vector based on the physiological data and environmental data, and to map the input vector to a cognitive state category through a machine learning classification model based on the Yerkes-Dodson law, and output a cognitive state vector representing the corresponding cognitive state category. The cognitive state categories include low arousal / boredom zone, moderate stress / flow zone, and high load / anxiety zone.

[0020] The objective function production unit is used to dynamically generate corresponding environmental regulation objective functions based on the different cognitive state categories represented by the cognitive state vector.

[0021] Preferably, in the game decision-making module, the payoff function of the i-th workstation agent is expressed as:

[0022] ,

[0023] Among them, a i a represents the motion vector of workstation i. -i Let θ represent the set of strategies for all adjacent workstations other than workstation i. i and θ j Let R represent the user cognitive state vectors at workstations i and j, respectively. perf C is the cognitive performance gain function. energy Let P be the energy consumption cost function. conflict Let N be the neighborhood conflict penalty function, where α, β, and γ are adjustable weight coefficients. i Let i represent the set of adjacent workstations of workstation i.

[0024] Preferably, the neighborhood conflict penalty function is expressed as:

[0025] ,

[0026] Among them, Ω ij E represents the environmental coupling coefficient of workstation i to workstation j. i→j (a i ) represents the motion vector a of workstation i. i The change in environmental parameters at workstation j, S sens (θ j θ represents the cognitive state vector of agent j in the adjacent workstation. j Related environmental sensitivity factors.

[0027] Preferably, in the game decision-making module, obtaining the collaborative control strategy by solving the Nash equilibrium of the model specifically includes:

[0028] Iterative solutions are obtained using optimal response dynamics;

[0029] If the solution does not converge to Nash equilibrium within the preset number of iterations, the Pareto suboptimal solution will be used as the final control command.

[0030] For any workstation i, the optimal motion vector corresponding to reaching Nash equilibrium is denoted as a. * -i And it satisfies the Nash equilibrium condition:

[0031] ,

[0032] Among them, a i Let A represent the motion space of workstation i. i For any action vector, a * -i Let represent the set of optimal action vectors for all workstations except workstation i at Nash equilibrium.

[0033] Preferably, the compensation execution module includes:

[0034] The perceived loss calculation unit is used to calculate the perceived loss vector D. i (t):

[0035] ,

[0036] Among them, E opt (θ i ) is based on the user's current cognitive state vector θ i The calculated optimal environment target vector, E actual (a * i (a) is based on the optimal action vector a * i The physical environment vector after actual execution, wherein the perception loss vector includes a thermal sensation loss component δ thermal and cognitive arousal deficit δ arousal ;

[0037] A multimodal compensation unit is used to generate and execute at least one cross-sensory modality compensation instruction based on the perceptual loss vector when the perceptual loss vector is not zero.

[0038] Preferably, the compensation instructions generated by the multimodal compensation unit include:

[0039] The color temperature compensation command is used to correct the user's subjective thermal perception by adjusting the correlated color temperature of the workstation lighting source based on the thermal perception loss component and the hue-thermal hypothesis model.

[0040] The soundscape compensation instruction is used to synthesize an acoustic signal containing specific masking noise or brainwave entrainment frequencies based on the cognitive arousal deficit component, in order to perform auditory frequency domain compensation.

[0041] The proprioceptive compensation command is used to control the electric height-adjustable table to perform micro-float or forced posture changes based on the cumulative fatigue degree obtained by time integration of the cognitive arousal deficit component.

[0042] Preferably, the adaptive reinforcement learning feedback module includes:

[0043] The composite reward calculation unit is used to calculate the immediate reward value R(t) based on the user's explicit interaction behavior and implicit physiological feedback, according to a preset reward function. The reward function quantifies the user's satisfaction with the current environment, and its expression is as follows:

[0044] ,

[0045] in, For the indicator function of artificial intervention, ΔP perf (t) represents the cognitive performance increment assessed based on the change in the cognitive state vector, ΔS stress (t) represents the physiological stress increment calculated based on the derivative of the user's heart rate variability index, where μ1, μ2, and μ3 are weighting coefficients;

[0046] The strategy optimization unit is used to incrementally update the state-action value function according to the reinforcement learning algorithm using the instant reward value, and to map the learning result into a dynamic adjustment of the weight coefficients and environmental sensitivity factors in the reward function.

[0047] A collaborative control method for open office micro-environment based on multi-agent game theory and cognitive load adaptation includes the following steps:

[0048] S1. Collect users' physiological data and workstation environmental data by deploying multimodal sensors at each workstation; establish an environmental coupling matrix characterizing the intensity of interference between workstations by actively detecting environmental interference between workstations.

[0049] S2. Based on the physiological and environmental data, the user's real-time cognitive state vector is mapped and generated according to the Yerkes-Dodson law, and a corresponding environmental regulation objective function is generated based on the cognitive state vector.

[0050] S3. Using the objective function of environmental regulation at each workstation as the optimization benchmark and the environmental coupling matrix as the neighborhood conflict constraint, a multi-agent non-cooperative game model is constructed. The collaborative control strategy is obtained by solving the Nash equilibrium of the model. The payoff function of each workstation agent in the model includes: a cognitive performance gain term based on the environmental regulation objective of the workstation, and a neighborhood conflict penalty term calculated based on the environmental coupling matrix, the penalty intensity of which is dynamically adjusted by the real-time cognitive state vector of the adjacent workstation agents.

[0051] S4. Execute the collaborative control strategy, and when the execution result does not meet the environmental regulation target based on the current cognitive state vector, initiate cross-sensory modality compensation regulation;

[0052] S5. Based on the user's physiological and behavioral feedback after performing step S4, adaptively adjust the parameters of the benefit function described in step S3.

[0053] The present invention has the following advantages and effects compared with the prior art:

[0054] (1) This invention quantifies neighbor interference by constructing an environmental coupling matrix and establishes a multi-agent non-cooperative game model that dynamically adjusts the conflict penalty term with the cognitive state vector. It can realize multi-workstation collaborative control in open office spaces, effectively solve the neighborhood conflict problem in traditional PCS systems, and achieve Pareto equilibrium of multi-user satisfaction while ensuring the optimal cognitive state of individuals.

[0055] (2) Based on the Yerkes-Dodson Law, this invention assesses the user’s cognitive state in real time (low arousal / boredom zone, moderate stress / flow zone, high load / anxiety zone) and dynamically generates personalized environmental control targets, realizing the transformation from passive response to active adaptation, and improving the user’s work efficiency and comfort experience.

[0056] (3) This invention introduces a cross-sensory modal compensation mechanism. When the physical environment is limited, it constructs an equivalent sense of comfort and alertness at the perception level through color temperature adjustment, sound scene synthesis and proprioceptive driving, thus breaking through the limitations of single physical dimension adjustment.

[0057] (4) The present invention adopts an adaptive reinforcement learning feedback mechanism, which integrates explicit user interaction and implicit physiological feedback, dynamically optimizes game strategy parameters, realizes the continuous evolution of the system from general strategy to personalized strategy, and improves long-term applicability and user satisfaction. Attached Figure Description

[0058] Figure 1 This is a block diagram illustrating the logical implementation of the open office micro-environment collaborative control system based on multi-agent game theory and cognitive load adaptation, as described in this invention.

[0059] Figure 2 This is a schematic diagram illustrating the construction of the environmental coupling matrix for the data sensing module of this invention.

[0060] Figure 3 This is a schematic diagram illustrating the cognitive state classification based on the Yerkes-Dodson law by the state assessment module of this invention.

[0061] Figure 4 This is a schematic diagram illustrating the principle of the game decision-making module of the present invention, which makes game decisions through a multi-agent non-cooperative game model.

[0062] Figure 5 This is a schematic diagram illustrating the principle of cross-modal compensation performed by the compensation execution module of this invention.

[0063] Figure 6 This is a schematic diagram of the workflow of the adaptive reinforcement learning feedback module of the present invention.

[0064] Figure 7 This is a schematic diagram of the overall process of the control method of the present invention. Detailed Implementation

[0065] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0066] Example 1

[0067] like Figure 1 As shown, this invention provides a collaborative control system for an open office micro-environment based on multi-agent game theory and cognitive load adaptation. Within an open office area, the system deploys independent workstation agents at each workstation, forming a decentralized collaborative control architecture via a network. The core modules of this system include data perception, state assessment, game theory decision-making, compensation execution, and adaptive reinforcement learning feedback modules.

[0068] I. Data Sensing Module:

[0069] like Figure 2 As shown, this module is responsible for collecting data and quantifying the interference between chemical levels.

[0070] (1) Sensing and monitoring unit:

[0071] Each workstation's intelligent agent is equipped with an integrated sensing terminal. Input consists of raw signals from the physical environment and human physiology, which are collected in real-time by various sensors in the sensing and monitoring unit. Physically deployed on the top of the workstation display or the lower edge of the desktop, it integrates the following core modules:

[0072] Non-contact micro-motion detection module: Employs a 60GHz / 77GHz millimeter-wave radar sensor to receive radar intermediate frequency echo signals (S... radar,i(t) is a time-series data containing Doppler frequency shift information. This module utilizes frequency-modulated continuous wave technology to monitor the user's chest cavity micro-movements and limb micro-movements (such as typing frequency and head rotation) in real time through the Doppler effect without acquiring optical images, effectively protecting user privacy.

[0073] Thermal environment optical flow sensing module: adopts infrared thermopile array (M IR,i (e.g., a 32x24 pixel thermal imaging sensor, represented as M) IR,i =[T pixel (x,y)] 32×24 T pixel (x,y) represents the original radiation temperature value at coordinates (x,y) in the matrix. This module measures the infrared radiation distribution map of the workstation area and uses an image segmentation algorithm to identify and extract the skin temperature of the user's face and neck.

[0074] Environmental basic parameter sensing module: Environmental sensors (including a sound pickup sensor, a high-precision illuminance / color temperature sensor, and an air quality sensor) are responsible for collecting illuminance (L). raw,i Color temperature (K) raw,i ), sound pressure level (A) raw,i CO2 concentration (C) raw,i Basic parameters such as )

[0075] (2) Edge computing unit: Each workstation intelligent agent has a built-in edge computing chip (such as ARM Cortex-M7 or equivalent NPU) to process the raw data input from the sensing and monitoring unit locally and output feature vectors, including:

[0076] Physiological data processing: Bandpass filtering and fast Fourier transform are performed on the echo signal from millimeter-wave radar to calculate heart rate variability (HRV) in real time. i ), respiratory rate (RR) i ), Electrodermal Activity Trends (EDA) i Core physiological data, such as those for cognitive load analysis, are used as the basis for subsequent cognitive load analysis.

[0077] Data fusion and desensitization: The Kalman filter algorithm is used to fuse infrared thermal image data and air temperature data to generate a thermal state vector for the workstation. All image data involving the user's face or body posture is discarded immediately after feature extraction and is not stored or uploaded.

[0078] (3) Environmental Coupling Unit: In order to solve the problem of adjacent interference, the system actively detects (the activation intensity Δu of device k at workstation i) during the initialization phase. i,k For example, briefly turning on the fan, heating pad, and other control equipment at this workstation in sequence, and utilizing the sensor feedback from the intelligent agents at adjacent workstations (the change in environmental parameters Δe detected by the sensors at adjacent workstation j).j ), calculate the spillover coefficient ω of environmental parameters in the adjacent area caused by the operation of the equipment at this workstation. ij ω ij= (Δe) j / (Δu i,k The coupling relationships between all workstation pairs form an environmental coupling matrix Ω. cp The environmental coupling matrix is ​​output to the game decision module as a core constraint.

[0079] II. Status Assessment Module:

[0080] like Figure 3 As shown, this module receives physiological and environmental data from the data perception module, assesses the user's cognitive state, and generates personalized environmental control targets.

[0081] (1) Cognitive State Assessment Unit:

[0082] The input vector X is constructed based on physiological and environmental data. t Input vector X t =[P t B t E t The metrics include three dimensions:

[0083] Physiological dimension P t ={HRV SDNN HRV LF / HF ,T skin ,Resp}, where HRV SDNN The standard deviation of heart rate variability, HRV LF / HF The ratio of low to high frequency is used to characterize the balance of the sympathetic / parasympathetic nervous system and reflect the level of mental stress. T skin Resp is skin temperature; Resp is respiratory rate.

[0084] Behavioral Dimension B t ={Micro m Posture c}, where Micro m The frequency of micro-movements per unit time is used to determine whether a state of stillness and focus or a state of restlessness and hyperactivity. c This represents the attitude change frequency per unit time.

[0085] Environmental Dimension E t ={Temp,Noise dB Lux level}, where Temp amb Noise refers to the dry-bulb temperature within the workstation's microenvironment, expressed in °C. dB Background noise intensity at the workstation, expressed in decibels (lux).level The luminous flux density projected onto the office desk is expressed in lux.

[0086] The module includes a built-in machine learning classification model (such as a random forest or support vector machine, SVM) pre-trained based on the Yerkes-Dodson theorem. This machine learning classification follows an inverted U-shaped performance curve, taking the input vector X... t Real-time classification, and output of the corresponding user's current cognitive state vector θ i =[p low ,p flow ,p high [The numbers represent confidence levels corresponding to three states: low arousal, flow state, and high anxiety, respectively.]

[0087] In this embodiment, the cognitive state includes at least: state I (p low Low arousal / boredom zone), characterized by extremely high HRV, slow and deep breathing, few micro-movements but accompanied by large-scale postural changes, and elevated skin temperature; the system determines that the user's attention is scattered and requires increased environmental stimulation; State II (p flow Moderate pressure / flow zone), characterized by HRV at the baseline median, steady and shallow breathing, minimal micromovements, and stable skin temperature; the system determines the user is in optimal working condition, and the system strategy is protection and maintenance, avoiding disturbance; State III (p high High workload / anxiety zone), characterized by HRV SDNN Significantly reduced HRV LF / HF Increased temperature, rapid breathing, accompanied by frequent micro-movements (such as frequent head touching); the system determines that the user is cognitively overloaded or under too much stress and needs a relaxing environment.

[0088] (2) Objective function generation unit: used to dynamically generate the corresponding environmental regulation objective function based on the cognitive state vector. The core objective of this function is quantified into an optimal environmental objective vector E. opt (θ). For example: when θ represents state I, the E opt The (θ) vector points to a combination of parameters such as low temperature, high color temperature, and high airflow disturbance; when θ represents state II, the E opt The (θ) vector points to a combination of parameters such as thermal neutral temperature, warm color temperature, and low noise; when θ represents state III, the E opt The weight of the noise reduction parameter in the (θ) vector is set to the highest.

[0089] III. Game Theory Decision-Making Module:

[0090] like Figure 4As shown, this module is the core of distributed collaborative decision-making, used to resolve conflicts between multi-workstation agents. This module accepts the environmental control objectives of each workstation agent, uses these as the optimization benchmark, and constructs a multi-agent non-cooperative game model with the environmental coupling matrix as the neighborhood conflict constraint. The collaborative control strategy is obtained by solving the Nash equilibrium of the model. The payoff function of each workstation agent in the multi-agent non-cooperative game model contains three core parts: a cognitive performance gain term based on the environmental control objective of its workstation, an energy consumption cost term based on equipment power consumption, and a neighborhood conflict penalty term based on the environmental coupling matrix and dynamically adjusted by the real-time cognitive state of adjacent workstations.

[0091] Specifically, due to the physical coupling in open-plan office environments (such as airflow diffusion and sound propagation), the optimal solution for a single workstation often harms the interests of adjacent workstations (for example, workstation A turning on strong air conditioning causes workstation B to feel cold). This embodiment constructs a distributed decision-making mechanism that combines non-cooperative game theory and cooperative game theory, with the following specific steps:

[0092] 1. Conflict detection and game domain establishment triggering:

[0093] When workstation agent i generates intention control strategy a i (For example, fan speed = 3, desktop height = 110cm) First, a pre-simulation is performed based on the aforementioned environmental coupling matrix. The intended control strategy a is then calculated. i Changes in environmental parameters ΔE at adjacent workstation j j If ΔE j If the preset perception threshold is exceeded (such as temperature change > 0.5℃ or wind speed > 0.1m / s), it is determined to be a potential conflict. The system automatically locks the intelligent agents at the workstations involved and constructs a temporary local game domain.

[0094] 2. Construct a dynamic payoff function:

[0095] The goal of workstation agent i is to select control strategy a. i ∈A i In order to maximize its profit function U i The profit function constructed in this invention contains three coupling terms, and its mathematical expression is as follows:

[0096] ,

[0097] Among them, a i a represents the motion vector of workstation i. -i Let θ represent the set of strategies for all adjacent workstations other than workstation i. i and θ j Let R represent the user cognitive state vectors at workstations i and j, respectively. perf C is the cognitive performance gain function.energy Let P be the energy consumption cost function. conflict Let N be the neighborhood conflict penalty function, where α, β, and γ are adjustable weight coefficients. i Let i represent the set of adjacent workstations of workstation i.

[0098] The specific definitions of each sub-item in the payoff function are as follows:

[0099] Cognitive performance gain function R perf This function quantifies the effect of environmental parameters on the user's current cognitive state.

[0100] ,

[0101] Among them, E(a) i ) represents the action vector a i The generated environmental parameters; E opt (θ i θ is the user's current cognitive state vector calculated based on the Yerkes-Dodson theorem. i The optimal environmental target value (e.g., when the user is drowsy, the optimal environmental target E) is as follows. opt (θ i (where ) represents low temperature and cold light; σ is the distribution width parameter of the Gaussian function. The closer the current environment is to the optimal environmental target, the higher the benefit.

[0102] Neighborhood conflict penalty function P conflict This function calculates the interference cost to neighbors caused by the current strategy. Unlike traditional algorithms, this invention introduces a neighbor sensitivity term; the expression for the neighborhood conflict penalty function is:

[0103] ,

[0104] Among them, Ω ij The environmental coupling coefficient of workstation i to workstation j is represented by ΔE, which is derived from the environmental coupling matrix. The physical influence attenuation rate of the equipment at workstation i on workstation j is also represented by ΔE. i→j (a i ) represents the motion vector a of workstation i. i The change in environmental parameters at workstation j, S sens (θ j θ represents the cognitive state vector of agent j in the adjacent workstation. j Related environmental sensitivity factors, S sens It is dynamic and depends on the cognitive state of adjacent workstations. For example, when adjacent workstation j is in a focused state, S sens Take the maximum value (e.g., 10.0); when the neighbor is in a resting / away state, S sensTake the minimum value (e.g., 0.1). Using this formula, the system automatically implements intelligent etiquette, maintaining quiet when neighbors are busy and adjusting freely when neighbors are not present.

[0105] Implementation logic: For example, when neighbor j is in state III (high-pressure anxiety), its tolerance for interference is extremely low, and the system will assign it a very high penalty weight S. sens This forces the workstation agent i to abandon highly disruptive strategies (such as abandoning the fan and switching to quiet localized cooling).

[0106] Energy cost function C energy The mathematical expression for this function is:

[0107] C energy (a i )=∑ k p k Power k (a i ),

[0108] Power k p is the power of the k-th device. k Weighted by unit electricity price.

[0109] 3. Iterative solution of Nash equilibrium:

[0110] The system employs algorithms such as optimal response dynamics to iteratively solve for the Nash equilibrium of the game. If convergence fails, a Pareto suboptimal solution is used. Finally, a set of collaborative control strategies is output that allows all workstations to approach their respective goals as closely as possible under conflict constraints. Specifically, this includes:

[0111] The system uses a distributed iterative algorithm to find a combination of strategies (a1). * , ..., a N * ), such that for any workstation i, the following Nash equilibrium condition is satisfied:

[0112] ,

[0113] Among them, a i Let A represent the motion space of workstation i. i For any action vector, a * -i Let represent the set of optimal action vectors for all workstations except workstation i at Nash equilibrium.

[0114] In this embodiment, the solution step employs optimal response dynamics:

[0115] Initialization time t=0, each workstation randomly selects strategy a i (0) ;

[0116] At time t, workstation i follows the strategy a from the previous time step of its neighbor. -i (t-1) Calculate and update its own policy:

[0117] a i t =argmax ai U i (a i ,a -i (t-1) ,θ i ),

[0118] Wherein, the Euclidean distance Δ=||a is used to calculate the policy update. (t) -a (t-1) If Δ < ε (convergence threshold), output the current policy as the control instruction; otherwise, let t = t + 1 and return to the previous step.

[0119] Implementation logic: First, each workstation agent within the game domain describes its ideal environment control strategy. Second, after receiving the ideal environment control strategies of its neighbors, each workstation agent calculates its current payoff U. i If it is found that changing the environmental control strategy can improve its own benefits, the environmental control strategy is adjusted and a new round of environmental control strategy is explained. When the environmental control strategies of all agents remain unchanged in two consecutive iterations, it is considered that a Nash equilibrium point has been reached. At this point, no single agent can gain more benefits by unilaterally changing its strategy. If the iteration count exceeds a preset threshold (e.g., 10 times) and convergence is still not achieved, the system will forcibly switch to a Pareto suboptimal solution, that is, select a compromise solution that minimizes the total loss of satisfaction for all relevant users.

[0120] Once the game process in the game decision-making module reaches a Nash equilibrium, each workstation agent synchronously sends control commands to the underlying execution mechanisms (fans, lights, height-adjustable desks). If the game result results in a user's core need not being met (e.g., the user is extremely hot, but cannot turn on the fan because the neighbor is extremely afraid of drafts), the system generates a "compensation request flag" and automatically triggers the compensation execution module (e.g., using visual cooling instead of physical cooling).

[0121] IV. Compensation Execution Module:

[0122] like Figure 5 As shown, the compensation execution module is the system's fault tolerance and enhancement mechanism. This module receives the coordination control strategy from the game decision module.

[0123] When the game decision-making module reaches Nash equilibrium, and the coordinated control strategy output by the system cannot meet the user's optimal needs (i.e., there is a "perceived loss"), the compensation execution module is activated. It does not change physical environmental parameters (such as air temperature), but instead constructs an equivalent sense of comfort and alertness in the cerebral cortex by adjusting the lighting spectrum, sound field frequency band, and the micro-dynamics of the supporting structure. Specifically, this includes:

[0124] (1) Perceived loss calculation unit, used to calculate the perceived loss vector D i (t):

[0125] To precisely control the compensation intensity, the system first needs to calculate the gap between the ideal state and the post-game reality. The system defines a perceived loss vector D. i (t):

[0126] ,

[0127] Among them, E opt (θ i ) is based on the user's current cognitive state vector θ i The calculated optimal environment target vector, E actual (a * i (a) is based on the optimal action vector a * i The actual physical environment vector after execution (e.g., actual room temperature 26℃, while the user expects 24℃, δ) thermal For thermal sensation loss component; δ arousal This represents a loss in cognitive arousal.

[0128] (2) Multimodal compensation unit: When the perceptual loss vector is not zero, it generates and executes at least one cross-sensory modality compensation instruction based on the perceptual loss vector. The compensation instructions generated by the multimodal compensation unit include: color temperature compensation instruction, which is used to correct the user's subjective thermal sensation by adjusting the correlated color temperature of the workstation lighting source based on the thermal sensation loss component and the hue-heat hypothesis model; sound scene compensation instruction, which is used to synthesize an acoustic signal containing specific masking noise or brainwave entrainment frequency based on the cognitive arousal loss component for auditory frequency domain compensation; and proprioceptive compensation instruction, which is used to control the electric lifting table to perform micro-float or forced posture transformation based on the cumulative fatigue obtained by the time integration of the cognitive arousal loss component.

[0129] Specifically, regarding the thermal perception loss vector δ thermal :

[0130] The system utilizes the Hue-Heat Hypothesis to correct the user's subjective thermal perception by adjusting the color temperature (CCT) of the lighting environment. The system constructs a color temperature compensation model that incorporates circadian rhythm constraints.

[0131] CCT target (t)=CCT base +λ vh ·tanh(δ thermal / σ T ) ·Φ circadian (t),

[0132] The formula introduces the hyperbolic tangent function tanh() as a saturation function, when the thermally perceived deficit vector δ thermal When the value is large (e.g., when the user feels extremely hot), the function value tends to be 1 or -1 to prevent the output color temperature value from exceeding hardware limits or causing visual discomfort; λ vh The visual-thermal coupling coefficient (empirical value approximately 1500K), σ T δ is the normalization factor; if δ thermal > 0 (user feels hot), the correction item is positive, the system will increase the light color temperature to cool white light (e.g., 6000K), transmitting a cooling signal to the brain through retinal ganglion cells; otherwise, it will adjust to warm light; Φ circadian (t) is a time-dependent decay function (0~1). At night (e.g., after 19:00), this factor approaches 0, forcibly turning off cold light compensation to prevent high color temperature from inhibiting melatonin secretion and ensuring that the compensation strategy conforms to the health standards of the human body's biological clock.

[0133] Targeting cognitive arousal deficit δ arousal (For example, a noisy environment makes it difficult to concentrate, or an overly quiet environment causes drowsiness):

[0134] The system synthesizes specific soundscapes, performs frequency domain compensation, and outputs sound signal S. out The synthesis formula for (t) is:

[0135] ,

[0136] Among them, S env (t) represents the basic environmental background sound signal, β mask γ is the acoustic masking intensity coefficient. beat Let f be the beat frequency signal strength coefficient. When interfering noise containing linguistic information (such as conversation between neighbors) is detected, the system calculates the masking threshold curve M(f) and generates filtered pink noise N. pink (f). This represents the inverse Fourier transform. This noise can fill the spectral gaps in ambient sound, reducing the intelligibility of speech and thus lessening its attentional pull.

[0137] When the user is in a "low arousal" state, the system plays a sine wave f with a slightly different frequency through stereo speakers or bone conduction headphones. c ±Δf / 2. The cerebral cortex automatically integrates these two frequencies to generate a beat frequency of Δf (such as a 18Hz Beta wave), which physically enhances concentration through brainwave entrainment.

[0138] When soft compensation from light and sound is insufficient to counteract prolonged fatigue, the system activates the proprioception driver module of the electric height-adjustable desk, executing forced or induced posture changes. The control law for posture height H(T) employs an integral triggering mechanism.

[0139] ,

[0140] Among them, the integral term in the formula Represents cumulative fatigue level; only when the cognitive arousal level is deficient (δ) does it increase. arousal Continues to accumulate and exceeds the recovery threshold ξ rec And the total score exceeds the trigger threshold Ω th The system only determines that the soft compensation has failed when the threshold is reached; if the threshold is not reached, the system controls the desktop to operate at an extremely low frequency (ω). micro ≈0.05Hz), subtle (A≈2mm) implicit sinusoidal fluctuations. This imperceptible dynamic forces the core back muscles to make continuous micro-adjustments, preventing static muscle stiffness. Once the threshold Ω is triggered... th The system determines that the user is about to enter a state of deep fatigue and forces or prompts the user to smoothly raise the desktop to a standing height H. stand It physically awakens the central nervous system by altering the hydrodynamic characteristics of blood.

[0141] V. Adaptive Reinforcement Learning Feedback Module:

[0142] like Figure 6 As shown, the adaptive reinforcement learning feedback module is used to adaptively adjust the parameters of the reward function based on the user's physiological and behavioral feedback.

[0143] Specifically, based on the aforementioned data perception module, state assessment module, game decision-making module, and compensation execution module, an adaptive reinforcement learning feedback module is further constructed. This module quantifies the user's explicit interaction behavior and implicit physiological characteristics, constructs a multi-dimensional reward function, updates the state-action value function in real time, and adaptively adjusts key parameters (such as environmental sensitivity factors) in the game decision-making module accordingly. Simultaneously, it generates a physiological health assessment report based on long-term time-series data. This module specifically includes:

[0144] (1) A composite reward calculation unit, used to calculate the immediate reward value based on the user's explicit interaction behavior and implicit physiological feedback, according to a preset reward function. This reward function is composed of two weighted parts: explicit interaction penalty and implicit physiological reward, used to quantify the user's satisfaction with the current environment. The expression of the reward function is:

[0145] ,

[0146] in, For manual intervention indicator functions, if the user is within a set time window T after the system automatically adjusts... w If environmental parameters are manually modified in reverse (e.g., the system turns on the fan, and the user turns it off) within 5 minutes, then... It is 1 if it is positive, otherwise it is 0. This is the strongest negative feedback signal;

[0147] ΔP perf (t) represents the cognitive performance increment assessed based on the change in the cognitive state vector. If the user successfully transitions from "low arousal" to "moderate stress / flow zone" after adjustment, this item is positive and a reward is given; ΔS stress (t) represents the physiological stress increment calculated based on the derivative of the user's heart rate variability index. If the user's sympathetic nerve activity increases significantly after adjustment (stress increases), this term is positive, resulting in a decrease in the total reward R(t). μ1, μ2, and μ3 are the weighting coefficients of the corresponding terms, and μ1 is set to a maximum value.

[0148] (2) Strategy optimization unit, which uses the instant reward value to incrementally update the state-action value function according to the Bellman equation, and maps the learning result to the dynamic adjustment of the conflict penalty weight coefficient and environmental sensitivity factor of the payoff function in the game decision module.

[0149] Specifically, the system maintains a state-action value function Q(s) t ,a t ), used to store the long-term expected return of taking a specific action a under a specific cognitive state s, the incremental update expression is:

[0150] ,

[0151] Where η is the learning rate, which determines how quickly the model adopts new feedback; and ρ is the discount factor, which balances immediate comfort with long-term health.

[0152] The system maps the updated Q-value back to the game decision-making module, dynamically adjusting the conflict penalty weight ρ and the neighbor's environmental sensitivity factor S. sens (θ jFor example, if a workstation takes highly disruptive actions while a neighbor is active, it will receive a continuous negative reward (extremely low R(T)). The system will then automatically increase the sensitivity factor S corresponding to that neighbor. sens This leads to the conflict penalty term P in the subsequent game equations. conflict The increased weight of the policy forces the system to automatically favor a "more polite" policy in solving the Nash equilibrium, thus achieving true personalized policy evolution.

[0153] In summary, this invention aims to address the technical bottlenecks in existing open office environment control systems, such as neighboring environment coupling interference (i.e., adjacent seat conflict), the inability of a single comfort index to characterize work efficiency, and the difficulty of centralized control in taking individual differences into account. This invention proposes a decentralized multi-agent collaborative architecture that treats each workstation as an independent agent. It quantifies the user's work status through a cognitive load assessment model and uses game theory algorithms to resolve environmental parameter conflicts between adjacent workstations, achieving Pareto optimality by maximizing individual cognitive performance and minimizing global energy consumption and conflict.

[0154] (3) Longitudinal physiological assessment unit:

[0155] Furthermore, the adaptive reinforcement learning feedback module also integrates a longitudinal physiological assessment unit, used for quantitative monitoring and proactive health management of the user's long-term cognitive load and recovery balance based on immediate policy optimization. This unit establishes a cumulative fatigue model to quantify the user's cumulative fatigue over daily, weekly, and monthly dimensions, as shown in the following formula:

[0156] ,

[0157] Among them, L cognitive (t) represents the instantaneous cognitive load intensity output by the state assessment module; Rec rest (t) is the recovery coefficient. This coefficient is positive when the user leaves their workstation or closes their eyes to rest, used to counteract fatigue; e -(T-t) / τ As a forgetting factor, simulating the body's natural recovery process, fatigue from a distant past has a smaller impact on the current state; λ load With λ rec For personalized weighting coefficients.

[0158] The system will calculate F cumulative (T) and the risk threshold F preset according to medical advice th Compare the data and implement two levels of intervention accordingly:

[0159] Level 1 assessment and immediate mandatory intervention: When the system determines F cumulative (T)>0.8F thAt this point, the system detects that the user is about to enter a deep fatigue risk zone. Instead of waiting for user commands, the system forcibly activates a "health intervention mode": raising the electric height-adjustable desk to standing height and simultaneously switching the ambient light source to a blue light-rich mode primarily using the 480nm wavelength, maintaining this for 15 minutes. This combined intervention aims to physically interrupt the continuous accumulation of fatigue from both hydrodynamic and neuroendocrine perspectives by changing posture and stimulating the suprachiasmatic nucleus.

[0160] Secondary Assessment and Periodic Health Reports: The system automatically generates a "Environment-Physiological Coupling Analysis Report" weekly. Based on long-term series data, statistical correlation analysis reveals the potential association between specific environmental parameters (such as the peak CO2 concentration around 2 PM each day) and abnormal physiological indicators of users (such as decreased heart rate variability, reported headaches, etc.), providing users with objective data support to guide them in improving their work habits or environmental settings.

[0161] The system of this invention not only realizes micro-adaptive adjustment of the user's real-time cognitive state, but also constructs a macro-monitoring and management system for the user's long-term workload and health status.

[0162] Example 2

[0163] like Figure 7 The diagram shows a flowchart of the collaborative control method for open office micro-environment based on multi-agent game theory and cognitive load adaptation of the present invention, including the following steps:

[0164] S1. Collect users' physiological data and workstation environmental data by deploying multimodal sensors at each workstation; establish an environmental coupling matrix characterizing the intensity of interference between workstations by actively detecting environmental interference between workstations.

[0165] S2. Based on the physiological and environmental data, the user's real-time cognitive state vector is mapped and generated according to the Yerkes-Dodson law, and a corresponding environmental regulation objective function is generated based on the cognitive state vector.

[0166] S3. Using the objective function of environmental regulation at each workstation as the optimization benchmark and the environmental coupling matrix as the neighborhood conflict constraint, a multi-agent non-cooperative game model is constructed. The collaborative control strategy is obtained by solving the Nash equilibrium of the model. The payoff function of each workstation agent in the model includes: a cognitive performance gain term based on the environmental regulation objective of the workstation, and a neighborhood conflict penalty term calculated based on the environmental coupling matrix, the penalty intensity of which is dynamically adjusted by the real-time cognitive state vector of the adjacent workstation agents.

[0167] S4. Execute the collaborative control strategy, and when the execution result does not meet the environmental regulation target based on the current cognitive state vector, initiate cross-sensory modality compensation regulation;

[0168] S5. Based on the user's physiological and behavioral feedback after performing step S4, adaptively adjust the parameters of the benefit function described in step S3.

[0169] Specifically, this invention can sense and assess the user's cognitive load in real time, coordinate the actions of equipment in each workstation environment through a multi-agent non-cooperative game model, initiate cross-sensory compensation when physical adjustment limits are reached, and utilize reinforcement learning to achieve personalized evolution of strategies. Thus, it dynamically maintains the optimal cognitive working state for each user in a complex and disturbing open environment, achieving multi-objective collaborative optimization of the office environment. It breaks through the averaging bottleneck of traditional centralized control, achieves Pareto optimality of multi-person satisfaction in open spaces through game theory, constructs a cross-modal sensory compensation mechanism, and solves the problem of maintaining comfort under physically limited conditions. It effectively alleviates neighborhood conflicts while protecting privacy, and proactively improves the user's cognitive efficiency and work experience.

[0170] The above embodiments are preferred embodiments of the present invention and are not intended to limit the present invention. Any changes or other equivalent substitutions made without departing from the technical solution of the present invention are included within the protection scope of the present invention.

Claims

1. An open-ended collaborative control system for office micro-environment based on multi-agent game theory and cognitive load adaptation, characterized in that: include: The data sensing module is used to collect users' physiological data and workstation environmental data through multimodal sensors deployed at each workstation, and to establish an environmental coupling matrix that characterizes the intensity of environmental interference between workstations through active detection. The state assessment module is used to map and generate the user's real-time cognitive state vector and the corresponding environmental regulation objective function based on the physiological data and environmental data, according to the Yerkes-Dodson law. The game decision-making module is used to construct a multi-agent non-cooperative game model with the objective function of each workstation's environmental regulation as the optimization benchmark and the environmental coupling matrix as the neighborhood conflict constraint. The cooperative control strategy is obtained by solving the Nash equilibrium of the model. The payoff function of each workstation agent in the model includes: a cognitive performance gain term based on the environmental regulation objective of the workstation, and a neighborhood conflict penalty term calculated based on the environmental coupling matrix, the penalty intensity of which is dynamically adjusted by the real-time cognitive state vector of the adjacent workstation agents. The compensation execution module is used to execute the collaborative control strategy and initiate cross-sensory modality compensation adjustment when the execution result does not meet the environmental regulation target based on the current cognitive state vector. An adaptive reinforcement learning feedback module is used to adaptively adjust the parameters of the reward function based on the user's physiological and behavioral feedback. In the game-theoretic decision-making module, the payoff function of the i-th workstation agent is expressed as: , where a i denotes the action vector of station i, a -i denotes the strategy set of other adjacent stations except station i, θ i and θ j denote the user cognitive state vectors of station i and station j, respectively, R perf is the cognitive performance gain function, C energy is the energy cost function, P conflict is the neighborhood conflict penalty function, and α, β and γ are adjustable weight coefficients, N i denotes the adjacent station set of station i; The neighborhood conflict penalty function is expressed as follows: , Among them, Ω ij E represents the environmental coupling coefficient of workstation i to workstation j. i→j (a i ) represents the motion vector a of workstation i. i The change in environmental parameters at workstation j, S sens (θ j θ represents the cognitive state vector of agent j in the adjacent workstation. j Related environmental sensitivity factors.

2. The open office micro-environment collaborative control system based on multi-agent game theory and cognitive load adaptation as described in claim 1, characterized in that, The data sensing module includes: The sensing and monitoring unit is used to acquire user heart rate variability, respiratory rate and skin conductance data through millimeter-wave radar, acquire user skin temperature through infrared sensors, and acquire basic environmental parameters of the work area through environmental sensors. The edge computing unit is used to perform real-time calculations on the heart rate variability, respiratory rate, and skin conductance data, and to fuse user skin temperature with basic environmental parameters. The environmental coupling unit is used to calculate the spillover coefficient of environmental parameters in the adjacent area by the control actions of the equipment at this workstation using the feedback from the sensors at adjacent workstations, and to form an environmental coupling matrix.

3. The open office micro-environment collaborative control system based on multi-agent game theory and cognitive load adaptation as described in claim 1, characterized in that, The status assessment module includes: The cognitive state assessment unit is used to construct a multi-dimensional feature input vector based on the physiological data and environmental data, and to map the input vector to a cognitive state category through a machine learning classification model based on the Yerkes-Dodson law, and output a cognitive state vector representing the corresponding cognitive state category. The cognitive state categories include low arousal / boredom zone, moderate stress / flow zone, and high load / anxiety zone. The objective function production unit is used to dynamically generate corresponding environmental regulation objective functions based on the different cognitive state categories represented by the cognitive state vector.

4. The open office micro-environment collaborative control system based on multi-agent game theory and cognitive load adaptation as described in claim 1, characterized in that, In the game-theoretic decision-making module, the collaborative control strategy obtained by solving the Nash equilibrium of the model specifically includes: Iterative solutions are obtained using optimal response dynamics; If the solution does not converge to Nash equilibrium within the preset number of iterations, the Pareto suboptimal solution will be used as the final control command. For any workstation i, the optimal motion vector corresponding to reaching Nash equilibrium is denoted as a. * -i And it satisfies the Nash equilibrium condition: , Among them, a i Let A represent the motion space of workstation i. i For any action vector, a * -i Let represent the set of optimal action vectors for all workstations except workstation i at Nash equilibrium.

5. The open office micro-environment collaborative control system based on multi-agent game theory and cognitive load adaptation as described in claim 1, characterized in that, The compensation execution module includes: The perceived loss calculation unit is used to calculate the perceived loss vector D. i (t): , Among them, E opt (θ i ) is based on the user's current cognitive state vector θ i The calculated optimal environment target vector, E actual (a * i (a) is based on the optimal action vector a * i The physical environment vector after actual execution, wherein the perception loss vector includes a thermal sensation loss component δ thermal and cognitive arousal deficit δ arousal ; A multimodal compensation unit is used to generate and execute at least one cross-sensory modality compensation instruction based on the perceptual loss vector when the perceptual loss vector is not zero.

6. The open office micro-environment collaborative control system based on multi-agent game theory and cognitive load adaptation as described in claim 5, characterized in that, The compensation instructions generated by the multimodal compensation unit include: The color temperature compensation command is used to correct the user's subjective thermal perception by adjusting the correlated color temperature of the workstation lighting source based on the thermal perception loss component and the hue-thermal hypothesis model. The soundscape compensation instruction is used to synthesize an acoustic signal containing specific masking noise or brainwave entrainment frequencies based on the cognitive arousal deficit component, in order to perform auditory frequency domain compensation. The proprioceptive compensation command is used to control the electric height-adjustable table to perform micro-float or forced posture changes based on the cumulative fatigue degree obtained by time integration of the cognitive arousal deficit component.

7. The open office micro-environment collaborative control system based on multi-agent game theory and cognitive load adaptation as described in claim 1, characterized in that, The adaptive reinforcement learning feedback module includes: The composite reward calculation unit is used to calculate the immediate reward value R(t) based on the user's explicit interaction behavior and implicit physiological feedback, according to a preset reward function. The reward function quantifies the user's satisfaction with the current environment, and its expression is as follows: , in, For the indicator function of artificial intervention, ΔP perf (t) represents the cognitive performance increment assessed based on the change in the cognitive state vector, ΔS stress (t) represents the physiological stress increment calculated based on the derivative of the user's heart rate variability index, where μ1, μ2, and μ3 are weighting coefficients; The strategy optimization unit is used to incrementally update the state-action value function according to the reinforcement learning algorithm using the instant reward value, and to map the learning result into a dynamic adjustment of the weight coefficients and environmental sensitivity factors in the reward function.

8. A collaborative control method for an open office micro-environment based on multi-agent game theory and cognitive load adaptation, applied to the system described in claim 1, characterized in that, Including the following steps: S1. Collect users' physiological data and workstation environmental data by deploying multimodal sensors at each workstation; establish an environmental coupling matrix characterizing the intensity of interference between workstations by actively detecting environmental interference between workstations. S2. Based on the physiological and environmental data, the user's real-time cognitive state vector is mapped and generated according to the Yerkes-Dodson law, and a corresponding environmental regulation objective function is generated based on the cognitive state vector. S3. Using the objective function of environmental regulation at each workstation as the optimization benchmark and the environmental coupling matrix as the neighborhood conflict constraint, a multi-agent non-cooperative game model is constructed. The collaborative control strategy is obtained by solving the Nash equilibrium of the model. The payoff function of each workstation agent in the model includes: a cognitive performance gain term based on the environmental regulation objective of the workstation, and a neighborhood conflict penalty term calculated based on the environmental coupling matrix, the penalty intensity of which is dynamically adjusted by the real-time cognitive state vector of the adjacent workstation agents. S4. Execute the collaborative control strategy, and when the execution result does not meet the environmental regulation target based on the current cognitive state vector, initiate cross-sensory modality compensation regulation; S5. Based on the user's physiological and behavioral feedback after performing step S4, adaptively adjust the parameters of the benefit function described in step S3.