Intelligent system man-machine collaborative decision task allocation method, system and device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING JIAOTONG UNIV
- Filing Date
- 2026-05-09
- Publication Date
- 2026-08-07
AI Technical Summary
1)难以刻画操作员感知-认知-决策全过程的动态演化特性:现有操作员绩效评估方法多基于静态指标或最终行为结果进行分析,未能对操作员在任务执行过程中信息感知、认知加工及决策行为的连续变化过程进行统一建模,导致绩效评估结果难以真实反映操作员状态随时间的变化特征
[0010]采用上述技术方案所产生的有益效果在于:1)所述方法通过将态势感知、人为失误概率与绩效统一建模为操作员多维度时变状态,突破了现有技术多基于单一指标进行任务分配的局限,使人机协作决策任务分配策略能够真实反映操作员状态随时间与任务演化的变化特征。
Smart Images

Figure CN122529296A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-machine collaboration technology in intelligent systems, and in particular to a method, system, and device for human-machine collaboration decision-making and task allocation in intelligent systems. Background Technology
[0002] In complex human-machine collaboration and intelligent scheduling applications, operators typically need to complete tasks such as information perception, cognitive judgment, and decision execution in a dynamically changing information environment. Their operational performance directly affects the system's operational efficiency and security. Existing research usually measures operator performance by evaluating task execution results. Common methods include statistically analyzing behavioral indicators such as task completion time, accuracy, and reaction time, or combining subjective scales and physiological signals to analyze operator status. Based on this, some human-machine collaboration systems allocate tasks to operators based on their performance or preset capability parameters, determining whether a task should be performed by a human or a machine through rule constraints, weight calculations, or optimization algorithms. In addition, some studies have attempted to introduce cognitive models or data-driven methods to model operator decision-making behavior and use the model outputs to assist in task allocation or collaborative decision-making processes.
[0003] The existing technical solutions have the following drawbacks: 1) Difficulty in depicting the dynamic evolution of the operator's perception-cognition-decision process: Existing operator performance evaluation methods are mostly based on static indicators or final behavioral results for analysis, and fail to uniformly model the continuous changes in the operator's information perception, cognitive processing and decision-making behavior during task execution. As a result, the performance evaluation results are difficult to truly reflect the changes in the operator's state over time.
[0004] 2) Lack of effective linkage mechanism between operator's multidimensional cognitive state and task allocation strategy: Existing human-machine collaborative decision-making task allocation methods usually rely on fixed rules, experience weights or single performance indicators, and fail to dynamically map the operator's situational awareness level, performance changes and other multidimensional cognitive states to the task allocation strategy, making it difficult to achieve adaptive human-machine task adjustment with state changes.
[0005] 3) Insufficient stability of task allocation optimization methods in complex human-computer collaboration scenarios: In collaborative scenarios with high state space dimensionality and complex task coupling relationships, existing methods based on optimization algorithms or reinforcement learning often suffer from slow convergence speed, sensitivity to parameters, or large policy fluctuations, which affect the stability and repeatability of task allocation results. Summary of the Invention
[0006] The technical problem to be solved by this invention is how to provide a human-machine collaborative decision-making task allocation method for intelligent systems that can improve the convergence efficiency and stability of the task allocation optimization process in complex scenarios, thereby achieving more reasonable and reliable human-machine collaborative decision-making.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a method for allocating human-machine collaborative decision-making tasks in an intelligent system, comprising the following steps: S1, Multi-source data acquisition: Receives operator physiological data, behavioral data, and task scenario data; S2, Situational Awareness Modeling: Situational awareness is defined as the operator's ability to perceive key information during collaboration. An initial calculation model for situational awareness is constructed. The basic stability level of situational awareness (SA) is corrected by parameters. The interference of cognitive state fluctuations is reduced by time-series smoothing. The normalized real-time situational awareness value is output. S3, Performance Modeling: Select core quantitative indicators, determine the three-parameter Weibull distribution function to fit time performance according to relevant criteria; define situational awareness, introduce situational awareness into the performance distribution function to construct the performance variability distribution function, dynamically characterize the impact of situational awareness on time performance, and output the normalized real-time performance value. S4, Human Error Probability Modeling: Determine the situational and personal factors and specific indicators of performance influencing factors; obtain the multiplier values of each performance influencing factor through expert evaluation and form a multiplier matrix; fit the task completion time to a Gaussian distribution; use the convolution method to quantify the error probability caused by task uncertainty; combine the performance influencing factor multiplier values and the error probability of task uncertainty; correct the time decay characteristics of individual differences through the variability function; model the influence of external member status on cognitive behavior through the S-shaped function; and output the final human error probability value. S5, Determine the human-machine collaborative decision-making task allocation model: Define the core elements of the Markov decision process model, construct the state transition equation based on the extended decision field theory; construct the reward function by combining the number of correct human and machine decisions to quantify the effectiveness of task allocation; S6, FFOA-SAC model optimization solution: The hyperparameters and policy parameters of the soft actor-critic algorithm are globally optimized by the fennec fox optimization algorithm; S7, Dynamic Task Allocation and Execution: Real-time collection and preprocessing of operator situational awareness, probability of human error, and performance data, input into the trained FFOA-SAC model to obtain the optimal task allocation ratio suggestion value, setting an adjustment threshold, and triggering adjustment when the difference between the optimal ratio and the actual ratio exceeds the threshold, dynamically dividing the scope of human-machine task assignment.
[0008] This invention also discloses an intelligent system human-machine collaborative decision-making task allocation system, comprising: Data acquisition module: used to collect operator physiological data, cognitive behavioral data, performance data, and task environment data, providing raw data support; Data preprocessing module: Used to perform data cleaning, standardization, and feature extraction operations to ensure data quality and usability; Operator Time-Varying Status Modeling Module: This module integrates the situational awareness calculation model, HEP calculation model, and performance calculation model sub-modules to output real-time operator status assessment results. MDP Model Building Module: Used to define the state, actions, transition equations and reward function of the MDP model, transforming the human-computer collaborative task allocation problem; FFOA-SAC Algorithm Module: Used to solve the MDP model using the FFOA-SAC fusion algorithm, including parameter optimization and policy learning sub-modules; Task allocation module: Based on model solution results and real-time status data, dynamically adjust the human-machine task allocation ratio; Feedback and update module: Used to monitor the task execution effect, dynamically update the state model and algorithm parameters, and form a closed-loop optimization.
[0009] The present invention also discloses an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the aforementioned intelligent system human-machine collaborative decision-making task allocation method.
[0010] The beneficial effects of adopting the above technical solution are as follows: 1) The method breaks through the limitation of existing technologies that mostly allocate tasks based on a single indicator by unifying situational awareness, probability of human error and performance into the multi-dimensional time-varying state of the operator, so that the human-machine collaborative decision-making task allocation strategy can truly reflect the changes in the operator's state over time and task evolution.
[0011] 2) The method improves the reliability and robustness of state assessment in human-machine collaborative task allocation by using an improved accelerated failure time model and introducing a disturbance smoothing mechanism to perform continuous and stable temporal assessment of operator situational awareness. This transforms the noise-sensitive cognitive state into a stable input that can be used for state modeling.
[0012] 3) By introducing a human error probability assessment mechanism that couples time performance distribution modeling with performance shaping factors and task uncertainty, the present invention can simultaneously characterize efficiency changes and error risk evolution, avoiding the shortcomings of existing methods that only focus on task completion efficiency while ignoring safety and reliability.
[0013] 4) The method introduces the operator's time-varying state into the Markov decision process model and integrates the fennec fox optimization algorithm and the soft actor-critic algorithm to solve the task allocation strategy. This improves the adaptability of the task allocation strategy and enhances the convergence stability of the algorithm, realizing a smooth and dynamic adjustment of the human-machine task allocation ratio. It is suitable for complex and uncertain collaborative scenarios.
[0014] 5) The method establishes a closed-loop mechanism of "data acquisition - state modeling - model construction - algorithm optimization - dynamic allocation - feedback update", and dynamically adjusts the human-machine task allocation ratio based on the operator's real-time situational awareness, the probability of human error, and performance data to achieve adaptive task allocation.
[0015] 6) The method uses the task allocation ratio as the continuous action space output, and dynamically adjusts the human-machine task assignment boundary according to the operator's real-time situational awareness, probability of human error and performance status, so as to realize the adaptive adjustment of human-machine collaborative task allocation, rather than a static switching mechanism based on rules or thresholds. Attached Figure Description
[0016] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0017] Figure 1 This is a flowchart illustrating the method described in Embodiment 1 of the present invention; Figure 2 This is a flowchart of the data acquisition and data processing process in the method described in Embodiment 1 of the present invention; Figure 3 This is a diagram showing the three core dimensions and their coupling relationships in the method described in Embodiment 1 of the present invention: situational awareness, probability of human error, and performance. Figure 4 This is a schematic diagram of the core elements of the human-machine collaborative decision-making task allocation model in the method described in Embodiment 1 of the present invention; Figure 5 This is a schematic diagram of the architecture of the FFOA algorithm in the method described in Embodiment 1 of the present invention; Figure 6 This is a schematic diagram of the architecture of the SAC algorithm in the method described in Embodiment 1 of the present invention; Figure 7 This is a schematic diagram of the system described in Embodiment 2 of the present invention; Figure 8 This is a schematic diagram of the device described in Embodiment 3 of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0019] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0020] Example 1 This embodiment discloses a human-machine collaborative decision-making task allocation method for intelligent systems based on operator time-varying states and FFOA-SAC. The method acquires operator physiological signals, behavioral data, and task execution data through a multi-source data acquisition module. These data are preprocessed to remove noise, fill in missing values, and standardize, providing high-quality data support for subsequent modeling. A multi-dimensional time-varying state model of the operator is then constructed. An accelerated failure time model is adopted, and a perturbation smoothing layer is introduced to accurately assess situational awareness. The probability of human error is calculated by combining performance influencing factors and task uncertainty errors. A three-parameter Weibull distribution function is used to describe time performance, and the suppression effect of situational awareness is introduced to obtain the performance variation distribution. Based on this, a Markov decision process model for human-machine collaborative decision-making task allocation is established. Situational awareness, the probability of human error, and performance are considered as system states, and the task allocation ratio is considered as actions. A state transition equation is constructed based on extended decision field theory, and a reward function is constructed by combining the number of correct human-machine decisions. A FFOA-SAC algorithm, formed by fusing the Fennec Fox Optimization (FFOA) algorithm and the Soft Actor-Critic (SAC) algorithm, is used to solve the Markov decision process model. FFOA is used to globally optimize the hyperparameters and policy parameters of SAC, while KL divergence smoothing constraints are introduced into the SAC Actor network to improve policy stability. Finally, based on the operator's real-time updated situational awareness, probability of human error, and performance data, the task allocation ratio between the intelligent system and the operator is dynamically adjusted to achieve precise task matching.
[0021] The method mainly includes the following steps: S1, Multi-Source Data Acquisition: Receives three core data categories: operator physiological data, behavioral data, and task scenario data, covering indicators such as real-time heart rate, eye movement tracking, attention span, task complexity, and environmental noise. Acquisition devices are selected based on the application scenario, including heart rate monitors, eye trackers, and industrial sensors, ensuring that the data is multi-dimensional, real-time, and scenario-adaptable.
[0022] S2, Situational Awareness Modeling: Situational awareness is defined as the operator's ability to perceive key information during collaboration. Based on collected eye-tracking data and task collaboration time data, the definition of "probability of deviating from key information" is clarified. An initial calculation model of situational awareness following the Weibull distribution is constructed, and covariate coefficients are introduced to reflect the impact of fatigue resistance on situational awareness. The basic stability level of situational awareness is corrected by parameters to prevent it from approaching zero infinitely over time. Time-series smoothing is applied to reduce the interference of cognitive state fluctuations, and the normalized real-time situational awareness value is output.
[0023] S3, Performance Modeling: Select core quantitative indicators such as task response time, decision accuracy, and task completion efficiency; determine the use of a three-parameter Weibull distribution function to fit time performance based on relevant information criteria; clarify the inhibitory effect of situational awareness on performance; introduce situational awareness into the performance distribution function to construct a performance variability distribution function, dynamically characterize the influence of situational awareness on time performance, and output normalized real-time performance values.
[0024] S4, Human Error Probability Modeling: This involves identifying the contextual and individual factors and their specific indicators that influence performance; obtaining the values of each performance factor through expert evaluation and forming a multiplier matrix; fitting the task completion time to a Gaussian distribution; and using a convolution method to quantify the error probability caused by task uncertainty; combining performance factors with the error probability due to task uncertainty; and correcting for the time decay characteristics of individual differences using a variability function. S The model uses a function to model the influence of external personnel's state on cognitive behavior and outputs the final probability value of human error.
[0025] S5. Establish a human-machine collaborative decision-making task allocation model: Define the core elements of the Markov decision process model: with situational awareness, human error probability, and performance as the system state, and the task allocation ratio (0-1 interval) as the action space. Construct a state transition equation based on extended decision field theory; and construct a reward function by combining the number of correct human-machine decisions to quantify the effectiveness of task allocation.
[0026] S6, FFOA-SAC Algorithm Optimization Solution: Global optimization of the hyperparameters and policy parameters of SAC is performed using FFOA.
[0027] FFOA Iterative Optimization: Simulate the behavior of "hunting for prey and avoiding predators" to select the optimal combination of parameters (policy network learning rate, temperature coefficient, etc.).
[0028] SAC policy learning: Introducing KL divergence smoothing constraints, constructing an Actor-Critic architecture, and improving policy stability through experience recycling and batch updates.
[0029] S7, Dynamic Task Allocation and Execution: Real-time collection and preprocessing of operator situational awareness, human error probability, and performance data, inputting them into the trained FFOA-SAC model to obtain the optimal task allocation ratio suggestion. An adjustment threshold is set; when the difference between the optimal ratio and the actual ratio exceeds the threshold, adjustment is triggered to dynamically divide the human-machine task assignment range.
[0030] The above steps will be explained in detail below with reference to specific content: Basic data collection, such as Figure 2 As shown: Collect diverse data from operators in human-machine collaboration scenarios, covering three core data categories: Cognitive and physiological data: This includes situational awareness-related data and physiological data collected by electrocardiogram and eye tracker, including data such as operator heart rate variability, duration of fixation on key information, duration of single fixation, and number of times deviating from key information in different task scenarios.
[0031] Performance data includes indicators that directly reflect operational capabilities, such as task response time, decision accuracy, and task completion efficiency. Task and environment data: such as task complexity, information quality, and work environment parameters. Specific examples of data collection items are shown in Table 1 below, illustrating the design of contextual and personal factors.
[0032] Table 1 Key Performance Influencing Factors and Explanations
[0033] Data preprocessing: The invention employs the 3σ criterion to remove outliers in response time, fills in missing physiological signal data through linear interpolation, and quantizes and encodes non-numerical data (such as task complexity levels) to ensure that the data meets the requirements for subsequent modeling. This invention does not specifically limit the data acquisition equipment; industrial sensors, eye trackers, and other devices can be selected according to the application scenario.
[0034] It should be noted that the specific data collection methods, equipment, indicator dimensions, and quantification forms are not limited to the examples mentioned above. Any multi-data that can characterize the operator's situational awareness level, performance, and probability of human error, and can support subsequent status modeling and analysis, falls within the scope of this step.
[0035] Situational awareness modeling: Based on the accelerated failure time model, situational awareness is defined as "cooperation time". T Exceeding time t "The probability that the operator will continuously monitor key information." This is first defined based on the collected data. T probability density function g(t) and cumulative distribution function G(t) The formula is as follows: ; Then, using survival analysis, SA is expressed as the survival probability, as shown in the following formula: ; In the formula λ These are positional parameters, expressed as functions of covariates. p It is a proportional parameter, and p>0 means that the probability of survival will decrease as time increases.
[0036] ; g(t) Following the Weibull distribution, covariate coefficients are introduced. v i (e.g., fatigue resistance) Calculate location parameters λ and through parameters z (Basic situational awareness level) f (Lower bound of situational awareness) Adjust the survival probability to avoid situational awareness approaching zero infinitely. The final situational awareness calculation formula is as follows: ; In addition, to suppress cognitive state fluctuations, a sliding window filter is introduced for temporal smoothing, with a window length of [missing information]. k Dynamically adjusts with disturbance variance: ; Performance modeling: Based on task response time and decision accuracy, a three-parameter Weibull distribution is selected as the performance fitting distribution, and its probability density function is shown in the formula: ; In the formula, t For collaboration time, a This is a scale parameter that reflects the performance benchmark level (e.g., for operators with high fatigue resistance, the performance benchmark is higher). β For shape parameters (reflecting the trend of performance degradation). γ This is a location parameter (task start delay time).
[0037] At the same time, situational awareness operators are introduced. ε i , ( ε i (Calculated from situational awareness values, and adjustments made to baseline performance). ε i A score >0.5 indicates strong decision-making ability, while a score <0.5 indicates strong decision-making ability. ε i<1 indicates weak decision-making ability. A performance variation distribution function is constructed to reflect the dynamic impact of situational awareness on performance. The formula is as follows: ; Modeling the probability of human error: The probability of human error is calculated based on the cumulative multiplier of performance impact factors. Referring to Table 1, eight categories of key performance impact factors (situational factors: WPU, IQ, TC, HCI; personal factors: ET, AC, SS, CL) are selected. The values of each performance impact factor (range 0-1, higher values contribute more to error) are determined through questionnaires and expert evaluation. The probability of human error is then quantitatively calculated using the relevant information of the performance impact factors, using the following formula: ; In the formula, P This indicates the contribution of the objective environment to human error. w i and w q This represents the weighting coefficient of PSF.
[0038] Treating task completion time as a Gaussian distribution, we differentiate and define the contribution of task uncertainty to the probability of human error. Based on the assumption that the task completion time of the collected data follows a Gaussian distribution, the uncertainty error is calculated by integrating the probability density function of the Gaussian distribution. A convolution method is used to calculate the probability of error caused by task uncertainty. P t The calculation formula is: ; in F t The probability density function representing the end of the current task. f T This represents the probability density function indicating the start of the next task. The distribution characteristics of the task start and end times are determined by the collected data.
[0039] The probability of human error exhibits a variable distribution, influenced by the operator's situational awareness capabilities. The calculation formula is as follows: ; The operator makes a mistaken cognitive behavior after being influenced by the state of external members. S Type function modeling: ; in, C j It's about communication performance. f i It is important to convey information. HEP IThis represents the probability that an operator will perform a failed action at a specific stage of cognitive function without external interaction. The response strength.
[0040] The interaction among the three core dimensions: The three dimensions are coupled and linked through a "positive promotion-negative inhibition" mechanism. The specific relationships are as follows: The coupling of situational awareness and performance: Situational awareness has a positive effect on performance; the higher the situational awareness value, the lower the performance variation distribution. ε i The larger the value, the slower the performance decline; conversely, a decrease in situational awareness value will accelerate performance decline. The coupling between situational awareness and the probability of human error: Situational awareness has an inverse inhibitory effect on the probability of human error; the higher the situational awareness value, the more fully the operator perceives key information, and the lower the probability of error. The relationship between the two is quantified by a formula: ; Model Output Layer: Operator Time-Varying State Assessment Results The model output layer outputs a comprehensive assessment result of the operator's real-time state, including: Single-dimensional indicators: Real-time performance value (normalized range 0-1), situational awareness value SA(t) (Normalized range 0.1-0.9) and human error probability value (range 0-1) provide an intuitive basis for subsequent task allocation. The three core dimensions—situational awareness, human error probability, and performance—and their coupling relationships are as follows: Figure 3 As shown.
[0041] Human-machine collaborative decision-making task allocation model: Step 1: State space definition. System state S t It covers key indicators of operator time-varying status, directly mapping the operator's multi-dimensional time-varying status, including key indicators such as real-time performance value, situational awareness value, and probability of human error. S t The dynamic changes with task execution time and task type are the core basis for formulating task allocation strategies. The model updates the current state using the state data of the previous task cycle.
[0042] Step 2: Define the action space. Action a t The task allocation ratio between the intelligent system and the operator is defined, with the action space encompassing all possible ratios between 0 and 1 (0 represents no task assigned to the operator, and 1 represents the intelligent system undertaking all tasks). At the start of each task cycle, the model determines the task allocation ratio based on the current state. S t Select the optimal action and determine the respective task scope of the human and machine.
[0043] Step 3: State transition equation. The state transition equation describes the current state. S t In performing the action a t Then transition to the next state S t+1 The probabilistic relationship is determined. Based on extended decision field theory, operator states and intelligent system states are incorporated into the mathematical framework of extended decision field theory to adapt to the description of operator cognitive behavior over long time scales, overcoming the limitation of traditional decision field theory models in failing to characterize continuous cognitive processes. Operator states and system states are introduced, and states are updated according to a performance dynamic model. The calculation formula is as follows: in, η It is the growth-decay rate parameter. It is a random variable. This is the system state, calculated using the following formula: in, θ It is a correction factor. It is the percentage of operators who make correct decisions in tasks suggested by the intelligent system. A t It is the proportion of times that intelligent systems make correct decisions. It is the proportion of tasks that are automatically suggested to humans for execution. It represents the proportion of tasks that the intelligent system performs on its own.
[0044] Step 4: Define the reward function. Reward function R t The formula for maximizing the number of correct decisions and minimizing the probability of human error, based on the overall performance of human-machine collaboration, is as follows: In the formula, , and These represent the proportion of correct decisions made by humans and the proportion of correct decisions made by intelligent systems under different circumstances, respectively.
[0045] Figure 4 This is a schematic diagram of the core elements of the human-machine collaborative decision-making task allocation model provided by the present invention, which clarifies the state ( S ),action( A ), state transition ( P ), reward function ( R The specific connotations and relationships of the four core elements.
[0046] FFOA-SAC Algorithm Optimization and Task Allocation Strategy Generation: The optimal strategy for the MDP model is solved by fusing the FFOA and SAC algorithms, which includes two stages: FFOA parameter optimization phase: The FFOA parameter optimization module, as the upper-level optimization unit, simulates the fennec fox's "initialization-finding prey-avoiding predators" behavior to optimize SAC hyperparameters (policy network learning rate, ... Q Global search optimization is performed using factors such as network learning rate, temperature coefficient, and discount factor. The specific steps are as follows: Step 1: Optimize parameter definitions The set of parameters to be optimized is determined, including SAC hyperparameters and Actor network policy parameters, as shown in Table 2: Table 2 FFOA Optimization Parameters
[0047] like Figure 5 As shown, the FFOA optimization process is divided into three stages, each corresponding to a natural behavior of the fennec fox, as detailed below: Phase 1: Population Initialization and Generation N Fennec fox (candidate parameter combinations), each fennec fox corresponds to one D Dimensional parameter vector, corresponding to Table 3 D (Number of parameters), the position initialization formula is shown in the formula: ; in x i,j For the first i Only the fennec fox in the first j The position of the dimension lb j For the first j Lower bound of dimension parameter ub j For the first j Upper limit of dimension parameter, r A random number in the interval [0,1]; Phase 2: Hunting Prey. Centered on the current optimal fennec fox location (optimal parameter combination), within a radius... R A local search is performed within the neighborhood to update the position and find better parameters, as shown in the following formula: ; ; ; In the formula, This indicates that according to the first phase update, the i indivual FF Location, It indicates that it is in the first j The position in the middle, It corresponds to , Yes, the neighborhood radius. t It is an iterative counter. T It is the total number of iterations. r It is a random number in the interval [0, 1]. a Set to a constant value.
[0048] Phase 3: Predator Avoidance. This phase simulates the random directional movements of a fennec fox when avoiding predators, enabling a global search to avoid getting trapped in local optima. The position update formula is as follows: ; ; ; In the formula, Indicates the first i The fennec fox's escape destination It indicates that it is in the first j The position in the middle, The fitness value represents the location. This is the second phase update. i The location of the fennec fox ear. It is in the first j The bit value in the dimension, It is the fitness value of the location. I and r It is a random number in the interval [0, 1].
[0049] Fitness function and optimization termination condition: The fitness value is the average cumulative reward (ACR) of the SAC algorithm on the validation set. F i =ACR(xi) The higher the ACR, the better the fitness. Termination condition: When the number of iterations reaches the maximum number of iterations, or the difference between the optimal fitness values of two adjacent iterations is less than or equal to the threshold, the optimization stops and the optimal parameter combination is output.
[0050] SAC Policy Learning Phase: The SAC algorithm module, as the core decision-making unit, learns the policy based on optimal parameters. It constructs an "Actor-Critic-Experience Recovery" architecture based on the optimal parameters output by FFOA to learn the task allocation policy, as detailed below: Network structure: Contains 1 Actor network and 2 Critic networks, with Dropout regularization introduced to avoid overfitting; Objective function: The policy objective function maximizes reward and entropy, as shown in the following formula: ; in The entropy represents the policy entropy. The larger the policy entropy, the higher the randomness of the policy and the greater the exploration of the "state-action" space. α The temperature coefficient is used as the entropy weight in the dynamic control function, and a larger one... α Increase the weight of the entropy term to encourage further exploration. Temperature coefficient. α The algorithm is dynamically updated throughout the training process by minimizing the entropy loss function. ; in, π It is a parameter η The parameterized strategy function, It is the target entropy, used to control the randomness of the strategy.
[0051] Q The network provides the target through the Bellman equation. Q Value estimation: ; in, γ It is a discount factor.
[0052] SAC usage restrictions (double) Q Network techniques to minimize Q Value, to solve Q The problem of overestimating its value. Q The network's loss function is given by the mean square Bellman error function: ; In the formula, This indicates a small batch of sampling in the experience pool. d A marker indicating the end of a sequence. θ j It is parameterized Q Function, target The formula provides the result.
[0053] Policy smoothing constraint: introduced in Actor network optimization KL Divergence constraints limit the magnitude of policy updates and improve stability. ; in, The weights of the KL constraint terms control the smoothing strength. π ref This is a strategy for reference. Calculate the difference between the current policy and the reference policy in the state. s t The smaller the difference in motion distribution, the smoother the motion output. Figure 6 This is a schematic diagram of the SAC algorithm architecture.
[0054] Training process initialization: Load the FFOA-optimized parameters, initialize the network weights of Actor, Critic, and target Critic; Interactive sampling: The agent interacts with the environment to collect data. S t , A t , R t , S t+1 , d t Store data in the experience recycling buffer pool; Batch update: When the buffer pool data volume is greater than or equal to the batch size, randomly sample batch data and update the Critic, Actor networks and temperature coefficients; Target network update: Update the target Critic network parameters every 5-10 training steps; Termination condition: When the training rounds reach 300-500, or the validation set ACR no longer improves, stop training and output the optimal policy.
[0055] Data flow relationship FFOA→SAC: The optimal parameters (learning rate, temperature coefficient, etc.) output by FFOA are input into the SAC algorithm module to initialize the network parameters; SAC→FFOA: The cumulative reward value of the SAC algorithm is used as the fitness value of FFOA to guide the optimization direction of FFOA parameters; KL Divergence constraint layer → Actor network: KL The divergence constraint term is fed back to the Actor network loss function in real time, dynamically adjusting the policy update magnitude.
[0056] The method described herein unifies SA, HEP, and time performance into a multi-dimensional time-varying state of the operator. This invention breaks through the limitations of existing technologies that mostly allocate tasks based on a single indicator, enabling human-machine collaborative decision-making on task allocation to truly reflect the changing characteristics of the operator's state over time and task evolution.
[0057] The proposed method uses an improved accelerated failure time model and introduces a disturbance smoothing mechanism to perform continuous and stable temporal evaluation of operator situational awareness. This transforms the originally noise-sensitive cognitive state into a stable input that can be used for decision-making, thereby improving the reliability and robustness of state evaluation in human-machine collaborative task allocation.
[0058] By introducing a human error probability assessment mechanism that couples time performance distribution modeling with performance shaping factors and task uncertainty, the present invention can simultaneously characterize efficiency changes and error risk evolution, avoiding the shortcomings of existing methods that only focus on task completion efficiency while ignoring safety and reliability.
[0059] The proposed method introduces the operator's time-varying state into a Markov decision process and integrates the fennec fox optimization algorithm and the soft actor-critic algorithm for strategy solving. This improves the adaptability of the task allocation strategy while enhancing the algorithm's convergence stability, achieving smooth and dynamic adjustment of the human-machine task allocation ratio. It is suitable for complex and uncertain collaborative scenarios.
[0060] Example 2 Corresponding to the method described in Example 1, such as Figure 7 As shown, Embodiment 2 of the present invention discloses an intelligent system human-machine collaborative decision-making task allocation system, the system comprising: Data acquisition module 121: Collects operator physiological data, cognitive behavior data, performance data and task environment data to provide raw data support.
[0061] Data preprocessing module 122: Performs data cleaning, standardization, feature extraction and other operations to ensure data quality and usability.
[0062] Operator Time-Varying Status Modeling Module 123: Integrates situational awareness calculation, human error probability calculation, and performance modeling sub-modules to output real-time operator status assessment results.
[0063] MDP Model Building Module 124: Defines the state, action, transition equations and reward function of the MDP model, transforming the human-computer collaborative task allocation problem.
[0064] FFOA-SAC Algorithm Module 125: This module uses the FFOA-SAC fusion algorithm to solve the MDP model, and includes sub-modules for parameter optimization and policy learning.
[0065] Task allocation module 126: Based on the model solution results and real-time status data, dynamically adjust the human-machine task allocation ratio.
[0066] Feedback Update Module 127: Monitors task execution performance, dynamically updates state model and algorithm parameters, and forms closed-loop optimization.
[0067] The system receives operator physiological, behavioral, and task scenario data collected by multi-source sensors through the data acquisition module 121. After data cleaning, standardization, and smoothing by the data preprocessing module 122, it provides high-quality data support for subsequent modeling. In the state modeling stage, the SA modeling submodule, performance modeling submodule, and human error probability modeling submodule under the operator time-varying state modeling module 123 respectively quantify the three core dimensions and output the operator's real-time state vector. The MDP model construction module 124 defines the system state and action space based on the state vector, and constructs the state transition equation and reward function by combining extended decision field theory, forming a complete Markov decision process model. The FFOA-SAC algorithm module 125, as the core computing module, first performs global optimization of the hyperparameters and policy parameters of the SAC algorithm through FFOA, selecting the optimal parameter combination; then, based on the optimal parameters, it constructs an Actor-Critic architecture, completes policy training, and outputs a stable optimal task allocation policy. The task allocation unit 126 outputs the task allocation ratio according to the real-time state and the optimal policy, adjusting the human-machine task range. The feedback update module 127 monitors the task execution effect in real time and sends back data for iterative optimization of model and algorithm parameters.
[0068] It should be noted that the specific implementation methods of each module in the system described in this embodiment two can refer to the methods described in embodiment one, and will not be repeated here.
[0069] Example 3 In one exemplary embodiment, the present invention also provides a computer device, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 8 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements the intelligent system human-machine collaborative decision-making task allocation method described in Embodiment 1.
[0070] Those skilled in the art will understand that Figure 8The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0071] In one exemplary embodiment, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0072] In one exemplary embodiment, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0073] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0074] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0075] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units, etc., and are not limited to these.
[0076] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0077] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for assigning tasks in a human-machine collaborative decision-making system for intelligent systems, characterized in that, Includes the following steps: S1, Multi-source data acquisition: Receives operator physiological data, behavioral data, and task scenario data; S2, Situational Awareness Modeling: Define situational awareness as the operator’s ability to perceive key information during collaboration. Construct an initial situational awareness calculation model, correct the basic level of situational awareness through parameters, reduce the interference of cognitive state fluctuations through time-series smoothing, and output the normalized real-time situational awareness value. S3, Performance Modeling: Select core quantitative indicators, determine the three-parameter Weibull distribution function to fit time performance according to relevant criteria; introduce situational awareness into the distribution function to construct a performance variability distribution function, dynamically characterize the impact of situational awareness on time performance, and output the normalized real-time performance value. S4, Human Error Probability Modeling: This involves identifying the contextual and individual factors and their specific indicators that influence performance; obtaining the values of each performance factor through expert evaluation; fitting the task completion time to a Gaussian distribution; and using a convolution method to quantify the error probability caused by task uncertainty; combining the performance factor values with the error probability due to task uncertainty; and correcting for individual differences in time decay characteristics using a variability function. S The model uses a type function to model the influence of external member states on cognitive behavior and outputs the final probability value of human error. S5, Determine the human-machine collaborative decision-making task allocation model: Define the core elements of the Markov decision process model, construct the state transition equation based on the extended decision field theory, construct the reward function by combining the number of correct human and machine decisions, and quantify the effectiveness of task allocation. S6, FFOA-SAC model optimization solution: The hyperparameters and policy parameters of the soft actor-critic algorithm are globally optimized using the fennec fox optimization algorithm; S7, Dynamic Task Allocation and Execution: Real-time collection and preprocessing of operator situational awareness, probability of human error, and performance data, input into the trained FFOA-SAC model to obtain the optimal task allocation ratio, setting an adjustment threshold, and triggering adjustment when the difference between the optimal ratio and the actual ratio exceeds the threshold, dynamically dividing the scope of human-machine task assignment.
2. The intelligent system human-machine collaborative decision-making task allocation method as described in claim 1, characterized in that, Step S1 includes: S1-1, Multi-source data acquisition: Collect diverse data from operators in human-machine collaboration scenarios, including three core data categories: Cognitive and physiological data: including situational awareness-related data, and physiological data collected by electrocardiogram and eye tracker, including the operator's heart rate variability, key information fixation duration, single fixation duration and number of deviations from key information in different task scenarios; Performance data includes task response time, decision accuracy, and task completion efficiency. Task and environment data: including task complexity, information quality, and working environment parameters; S1-2, Data preprocessing: The 3σ criterion is used to remove outliers in response time, missing physiological signal data is filled in by linear interpolation, and non-numerical data is quantized and encoded.
3. The intelligent system human-machine collaborative decision-making task allocation method as described in claim 1, characterized in that, Constructing the initial computational model for situational awareness includes the following steps: Based on the accelerated failure time model, situational awareness is defined as the collaboration time. T Exceed t The probability that the operator will continuously monitor key information; Define the probability density function of T based on the collected data. g(t) and cumulative distribution function G(t) The formula is as follows: ; Situational awareness is represented as survival probability using survival analysis methods, as shown in the following formula: ; In the formula, λ is the position parameter, expressed as a function of the covariates. p It is a proportional parameter, and p>0 This indicates that the probability of survival decreases as time goes on; ; g(t) Following the Weibull distribution, covariate coefficients are introduced. v i Calculate position parameters λ And through basic situational awareness level z Situational awareness lower limit f After adjusting the survival probability, the final situational awareness calculation formula is as follows: ; A sliding window filter is introduced for time series smoothing, with a window length of [missing information]. k Dynamically adjusts with disturbance variance: 。 4. The intelligent system human-machine collaborative decision-making task allocation method as described in claim 1, characterized in that, The performance modeling step S3 includes the following steps: Based on task response time and decision accuracy, a three-parameter Weibull distribution function is selected to construct the performance distribution function, and its probability density function is shown in the formula: ; In the formula, t For collaboration time, a This is a scale parameter that reflects the performance benchmark level; β The shape parameter reflects the performance degradation trend. γ This is a location parameter that reflects the task start delay time; At the same time, situational awareness is introduced. ε i , ε i A score >0.5 indicates strong decision-making ability, while a score <0.5 indicates strong decision-making ability. ε i A value less than 1 indicates weak decision-making ability. A performance variability distribution function is constructed to reflect the dynamic impact of situational awareness (SA) on performance. The formula is as follows: 。 5. The intelligent system human-machine collaborative decision-making task allocation method as described in claim 1, characterized in that, Modeling the probability of human error includes the following steps: The probability of human error is calculated based on the cumulative performance impact factors. The values of each performance impact factor are determined through questionnaires and expert evaluations to form a performance impact factor multiplier matrix. The probability of human error is then quantitatively calculated using relevant information from these performance impact factors. The calculation formula is as follows: ; In the formula, P This indicates the contribution of the objective environment to human error. w i and w q Represents the linear weighting coefficient of performance influencing factors; Treating task completion time as a Gaussian distribution, we define its differential to quantify the contribution of task uncertainty to the probability of human error. Based on the assumption that the task completion time of the collected data follows a Gaussian distribution, the uncertainty error is calculated by integrating the probability density of the Gaussian distribution. A convolution method is then used to calculate the probability of error caused by task uncertainty. P t The calculation formula is: ; in F t The probability density function representing the end of the current task. f T The probability density function represents the start of the next task, and the distribution characteristics of the task start and end times are obtained from data collection and statistics. The probability of human error exhibits a variable distribution, influenced by the operator's situational awareness capabilities. The calculation formula is as follows: ; The operator makes a mistaken cognitive behavior after being influenced by the state of external members. S Type function modeling: ; in, C j It's about communication performance. f i It is important to convey information. HEP I This represents the probability that an operator will perform a failed action at a specific stage of cognitive function without external interaction. The response strength.
6. The intelligent system human-machine collaborative decision-making task allocation method as described in claim 1, characterized in that, The method for constructing a human-machine collaborative decision-making task allocation model includes the following steps: State space definition: System state S t This includes real-time performance values, situational awareness values, probability of human error, fatigue resistance, and system status. S t The current state in the model changes dynamically with the task execution time and task type, and is updated using the state data from the previous task cycle. Action space definition: Action a t The task allocation ratio between the intelligent system and the operator is defined, and the action space covers all possible ratios between 0 and 1; at the beginning of each task cycle, the model determines the task allocation ratio based on the current state. S t Select the optimal action and determine the respective task scope of the human and machine; State transition equation: The state transition equation is used to describe the current state. S t In performing the action a Then transition to the next state S t+1 The probability relationship; Based on extended decision field theory, operator states and intelligent system states are incorporated into the mathematical framework of extended decision field theory to adapt to the description of operator cognitive behavior over long time scales. Operator states and system states are introduced into the model, and the states are updated according to a dynamic performance model. The calculation formula is as follows: ; in, η It is the growth-decay rate parameter. It is a random variable. This is the system state, calculated using the following formula: ; in, θ It is a correction factor. It is the percentage of operators who make correct decisions in tasks suggested by the intelligent system. A t It is the proportion of times that intelligent systems make correct decisions. It is the proportion of tasks that are automatically suggested to humans for execution. It represents the proportion of tasks that the intelligent system performs on its own; Reward function definition: Reward function R t The formula for maximizing the number of correct decisions and minimizing the probability of human error, based on the overall performance of human-machine collaboration, is as follows: ; In the formula, , and These represent the proportion of correct decisions made by humans and the proportion of correct decisions made by intelligent systems under different circumstances, respectively.
7. The intelligent system human-machine collaborative decision-making task allocation method as described in claim 1, characterized in that, The fennec fox optimization algorithm for optimizing the parameters of the soft actor-critic SAC includes the following steps: Population initialization phase: generation N One fennec fox, each fennec fox corresponds to one D Dimensional parameter vector, corresponding to Table 3 D The position initialization formula for the parameters is shown in the formula below: ; in x i,j For the first i Only the fennec fox in the first j The position of the dimension lb j For the first j Lower bound of dimension parameter ub j For the first j Upper limit of dimension parameter, r A random number in the interval [0, 1]; Hunting phase: Centered on the current optimal fennec fox location, within a radius... R A local search is performed within the neighborhood to update the position and find better parameters, as shown in the following formula: ; ; ; In the formula, This indicates that according to the first phase update, the i The location of the fennec fox ear. It indicates that it is in the first j The position in the middle, It corresponds to , Yes, the neighborhood radius. t It is an iterative counter. T It is the total number of iterations. r It is a random number in the interval [0, 1]. a Set to a constant value; Predator avoidance phase: Simulates the random directional movement of a fennec fox when avoiding predators to achieve a global search. The position update formula is as follows: ; ; ; In the formula, Indicates the first i The fennec fox's escape destination It indicates that it is in the first j The position in the middle, The fitness value represents the location. This is the second phase update. i The location of the fennec fox ear. It is in the first j The bit value in the dimension, It is the fitness value of the location. I and r It is a random number in the interval [0, 1]. Fitness function and optimization termination condition: The fitness value is the average cumulative reward of the SAC algorithm on the validation set. F i = ACR(xi) The higher the ACR, the better the fitness. Termination condition: When the number of iterations reaches the maximum number of iterations, or the difference between the optimal fitness values of two adjacent iterations is less than or equal to the threshold, the optimization stops and the optimal parameter combination is output.
8. The intelligent system human-machine collaborative decision-making task allocation method as described in claim 1, characterized in that, SAC strategy learning includes the following steps: Network structure: Includes 1 Actor network and 2 Critic networks, with Dropout regularization introduced to avoid overfitting; Objective function: The policy objective function maximizes reward and entropy, as shown in the following formula: ; in, This represents policy entropy. The larger the policy entropy, the higher the randomness of the policy and the greater the exploration of the "state-action" space. α The temperature coefficient is used as the entropy weight in the dynamic control function. α The algorithm is dynamically updated throughout the training process by minimizing the entropy loss function. ; in, π It is a parameter η The parameterized strategy function, It is the target entropy, used to control the randomness of the strategy; Q-networks provide estimates of the target Q-value through the Bellman equation: ; in, γ It is a discount factor; SAC usage restrictions (double) Q Network techniques to minimize Q value, Q The network's loss function is given by the mean square Bellman error function: ; In the formula, This indicates a small batch of sampling in the experience pool. d A marker indicating the end of a sequence. θ j It is parameterized Q Function, target The formula is given. Policy smoothing constraint: introduced in Actor network optimization KL Divergence constraint limits the policy update magnitude: ; in, for KL The weights of the constraint terms control the smoothing strength; π ref For reference strategy; Calculate the difference between the current policy and the reference policy in the state. s t The smaller the difference in motion distribution, the smoother the motion output.
9. A human-machine collaborative decision-making task allocation system for an intelligent system, characterized in that, include: Data acquisition module: used to collect operator physiological data, cognitive behavioral data, performance data, and task environment data to provide data support; Data preprocessing module: Used to perform data cleaning, standardization, and feature extraction operations to ensure data quality and usability; Operator Time-Varying Status Modeling Module: This module integrates the situational awareness calculation model, the human error probability calculation model, and the performance calculation model sub-modules to output real-time operator status assessment results. Situational Awareness Model Construction Module: Used to define the state, actions, transition equations and reward function of the situational awareness model, transforming the human-machine collaborative task allocation problem; FFOA-SAC Algorithm Module: Used to solve the MDP model using the FFOA-SAC fusion algorithm, including parameter optimization and policy learning sub-modules; Task allocation module: Based on model solution results and real-time status data, dynamically adjust the human-machine task allocation ratio; Feedback and update module: Used to monitor the task execution effect, dynamically update the state model and algorithm parameters, and form a closed-loop optimization.
10. An electronic device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program and the processor runs the computer program to enable the electronic device to perform a human-machine collaborative decision-making task allocation method for an intelligent system as described in any one of claims 1 to 8.