Human-computer interaction safety control method based on finite rational game and reachability analysis

Through limited rational game and accessibility analysis, the opponent's action space is cut, and the conservative problem of security control strategies in human-computer interaction is solved, and safe and efficient interaction control is achieved.

CN120276257APending Publication Date: 2025-07-08UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510429275.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, in human-computer interaction, the Nash equilibrium game model cannot accurately predict non-optimal behavior, resulting in the security control strategy being too conservative and reducing efficiency; traditional accessibility analysis frequently triggers conservative control in close-range interaction scenarios, affecting efficiency.

Method used

The opponent's behavior is modeled through finite rational game, and different rational level strategies of the agent are trained through recursive reasoning and reinforcement learning, combined with Hamilton-Jacobian reachability analysis, the opponent's action space is cut, and the low-conservative backward reachable tube is calculated to achieve safe and efficient control.

Benefits of technology

Accurate and differentiated modeling of interactive opponents is achieved, the frequency of security control strategies is reduced, and the efficiency and security of human-computer interaction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276257A_ABST
    Figure CN120276257A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of man-machine interaction safety control, and discloses a man-machine interaction safety control method based on a finite rationality game and accessibility analysis, which is used for modeling finite rationality behaviors of opponents in a man-machine cooperation system. And a control framework in which an efficient reinforcement learning control strategy and a low-conservative safety control strategy based on backward reachable management are combined is formed. The man-machine interaction safety control method comprises the following steps: determining a model of an intelligent robot participating in man-machine interaction; modeling interaction between the bounded rationality agents by using recursive reasoning; designing a reward function for guiding each agent to perform reinforcement learning training, and training control strategies of different rationality levels of each agent step by step; designing an interactive opponent rationality grade inference method; based on Hamilton-Jacobian accessibility analysis, a backward accessibility management and safety control strategy of the low-conservative man-machine cooperation system is obtained; and human-computer interaction safety control is realized by using the determined efficient reinforcement learning control strategy and the determined low-conservative safety control strategy, so that the robot can perform optimal action selection on the basis of identifying human rational levels and predicting human non-optimal behaviors, and safe and efficient interaction and collaboration of a human-computer collaboration system are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human - machine interaction safety control, and particularly to a human - machine interaction safety control method based on bounded - rationality game and reachability analysis. Background Art

[0002] Human - machine interaction is a technology in which robots and humans cooperate to complete the same work task in the same shared workspace. This technology abandons the traditional physical - isolation - form safety guarantee. At the same time, the addition of humans increases the complexity of the robot's working environment, which poses challenges to the efficiency and safety of the robot's command execution.

[0003] Considering intelligent robots with perception, autonomous decision - making, and execution capabilities, that is, intelligent agents with physical forms, modeling the behavior of humans and other robots, and evaluating all possible response actions of themselves according to the predicted opponent's behavior to make the optimal decision at present, helps to achieve safer and more efficient cooperation. Nash equilibrium is often used to model the interaction between "fully rational" intelligent agents, and the strategy of each player is the best response to the current strategy of other players. However, in real life, due to the incompleteness of information, the limitations of participants' cognitive abilities, the limitations of computing power, etc., the assumption of "fully rational" is difficult to meet. Therefore, the concept of "bounded rationality" is proposed to describe the various limitations faced by players in making decisions in reality. The bounded - rationality game model can make up for the limitations of Nash equilibrium. It takes into account the bounded rationality of players, the dynamic changes of the environment, and the mutual influence between players' behaviors.

[0004] Reachability analysis is a method for verifying the safety of intelligent agents. By calculating the backward reachable tube, it judges whether the human - machine collaborative system will enter the unsafe state set at a future moment and decides whether to enable the safety - first control strategy. However, this method conducts safety detection under the assumption of the worst - case scenario, which will cause conservative safety control strategies to be frequently triggered in the operation scenarios where humans and machines need to interact closely, reducing efficiency. Therefore, the present disclosure designs a method for predicting the behavior of opponents with different rationality levels based on bounded - rationality game. After adaptively judging the rationality level of the opponent, it calculates the low - conservative backward reachable tube that matches it, ignoring the extremely low - probability dangerous situations that do not conform to the rationality level, so as to achieve the balance between efficiency and safety in the human - machine interaction process. Summary of the Invention

[0005] To solve the above - mentioned technical problems, the present invention provides an efficient human - machine interaction safety control method for modeling the behavior of bounded - rational opponents based on bounded - rationality game and achieving collision avoidance through reachability analysis, so as to solve the problems that the equilibrium - game solutions represented by Nash equilibrium cannot predict non - optimal behaviors and the traditional reachability analysis is too conservative and lacks pertinence.

[0006] To solve the above technical problems, the present invention adopts the following technical solutions:

[0007] A human-machine interaction safety control method based on bounded rationality game and reachability analysis, which is used to model the behavior of opponents in robot modeling interaction and intervene in a timely manner before reaching an unsafe state, and switch the control strategy obtained through reinforcement learning training to a safety control strategy; the human-machine interaction safety control method includes:

[0008] Based on the dynamic equation of the robot, construct a human-machine collaborative system model composed of multiple robots with autonomous decision-making capabilities;

[0009] Use recursive reasoning to model the interaction between bounded rational agents;

[0010] Design a reward function to guide the reinforcement learning training of each agent and train the control strategies of different rational levels of each agent step by step;

[0011] By initializing the belief probability of the opponent agent and dynamically updating the belief probability based on the difference between the observed action and the multi-level rational prediction action, and combining normalization processing to infer the rational level of the opponent agent in real time;

[0012] Obtain the backward reachable tube and safety control strategy of the low-conservative human-machine collaborative system through Hamilton-Jacobi reachability analysis;

[0013] The intelligent robot executes the control strategies of different rational levels, infers the rational level of the opponent robot, calculates the backward reachable tube, and executes the safety control strategy when the state of the human-machine collaborative system reaches the boundary of the backward reachable tube.

[0014] In one embodiment, it is necessary to determine the dynamic equations of the agents participating in the human-machine interaction: the human-machine collaborative system is composed of n intelligent robots, and the state of each robot consists of information such as the position, speed, acceleration, and yaw angle of the robot in space. The dynamic equation of each robot i can be expressed as:

[0015]

[0016] Where, represents the first derivative of x i a i is the control input of the agent, and f is a non-linear function.

[0017] In one embodiment, the use of recursive reasoning to model the interaction between bounded rational agents specifically includes:

[0018] Set the recursive rules of the recursive reasoning game model;

[0019] Set the strategy of the agent with a rational level of 0;

[0020] Train the policy of the agent with rationality level k through reinforcement learning.

[0021] In one of the embodiments, setting the recursive rules of the recursive inference game model specifically includes:

[0022] Recursive inference means that in the decision-making process, each agent will consider the reasoning processes of other agents and make the best response based on the reasoning; the number of steps of the agent's reasoning is used as the rationality level of the agent. The agent with rationality level k assumes that the rationality level of the opponent agent is k ∈ {0, 1,..., k - 1} and follows a Poisson distribution. Then the agent with rationality level k believes that a certain opponent agent has a rationality level of k or the proportion of agents with level k among all opponent agents is:

[0023]

[0024] The optimal policy when the rationality level of the i-th agent is k is:

[0025]

[0026] where -i is the set of opponent agents, is the optimal policy when the rationality level of the opponent agent is κ, and κ follows the distribution x is the state of the human-machine collaborative system, J i is the expected cumulative reward function when the opponent agent executes a certain policy π -i and the agent i executes the policy π i :

[0027]

[0028] where p is the state transition function, r i (x t , π i (x), π -i (x)) is the reward function of the agent i, γ is the discount factor, and x t is the state of the human-machine collaborative system at time t;

[0029] Training the policy of the agent with a set rationality level k through reinforcement learning specifically includes: The policies of other agents in the reinforcement learning environment are obtained by sampling .

[0030] In one of the embodiments, setting the policy of the agent with rationality level 0 specifically includes:

[0031] The reasoning step of an agent with rationality level 0 is 0, and it does not reason about the behavior of the opponent agent during the interaction. Considering the human-machine interaction scenario, when the robot cannot model the behavior of the human it is interacting with, to ensure safety, the most conservative Max-Min decision model is adopted. It is assumed that all other agents will take actions that minimize the current agent's immediate reward function. The best response of an agent with rationality level 0 is:

[0032]

[0033] where x is the state of the human-machine cooperation system, and a i represents the action of agent i, A i represents the action space of agent i, and a -i represents the joint action of agent i's opponent, and k -i represents the joint action space of agent i's opponent.

[0034] In one embodiment, the strategy of setting an agent with rationality level k through reinforcement learning training specifically includes: The k-level strategy of the agent is obtained through reinforcement learning training. The strategies of other agents in the reinforcement learning environment are sampled with probability to obtain.

[0035] In one embodiment, the design of the reward function that guides each agent's reinforcement learning training and the step-by-step training of the control strategies of each agent at different rationality levels specifically includes:

[0036] Design reward function terms according to each agent's task objective, safety limit, and behavior smoothness respectively, and assign weight values to each reward function term to obtain the reward function of each agent;

[0037] Use the SAC algorithm to train the control strategies corresponding to each rationality level of each agent step by step: The 0-level strategy of each agent is known. When training the k-level strategy of agent i, any other agent j∈-i in the environment samples its own rationality level according to probability and executes the strategy Initialize the k-level strategy of agent i to randomly select actions, and use the SAC algorithm for training. The obtained strategy is the optimal strategy for the opponent agent whose rationality level follows After training the k-level strategies of all agents, the rationality level can be increased and then traverse the agents for training until the maximum rationality level K; is the proportion that an agent with rationality level k believes that a certain opponent agent has rationality level κ or the proportion of agents with level κ among all opponent agents, κ∈{0,1,…,k - 1}, and -i is the set of opponent agents. ​

[0038] In one embodiment, designing reward function terms according to the task objectives, safety constraints, and behavior smoothness of each agent respectively, and assigning weight values to each reward function term to obtain the reward function of each agent, specifically including:

[0039]

[0040] Among them, are the weight values of the three parts of the reward function, and r i is the reward function of the i-th agent.

[0041] In one embodiment, initializing the belief probability of the opponent agent and dynamically updating the belief probability based on the difference between the observed action and the multi-level rational prediction action, and combining the normalization process to infer the rationality level of the opponent agent in real time, specifically including:

[0042] Using T (K=k) [j](t) represents the belief probability of modeling the opponent agent j as level k, which is initialized as a probability function with a uniform distribution, and the belief of the opponent agent is updated at each moment:

[0043] The update rule is to observe the action taken by the opponent agent j at the current moment t Obtain the current state observation value from the perspective of the opponent agent j Predict the actions that the opponent agent will take when executing different rationality level strategies Calculate the difference between the predicted action of each rationality level and the real action taken by the opponent agent at the current moment to obtain the most likely rationality level κ of the opponent agent at the current moment * :

[0044]

[0045] Increase the belief of κ * by ΔP and normalize it:

[0046]

[0047]

[0048] ΔP represents the belief increment of the most likely rationality level of the opponent agent, represents the updated belief of the opponent agent with a rationality level of κ * and then update the belief at the current moment through normalization.

[0049] In one embodiment, obtaining the backward reachable tube and safety control strategy of the low-conservative human-machine cooperation system through Hamilton-Jacobi reachability analysis, specifically including:

[0050] Establish a relative system model between the controlled agent and a certain interactive agent in sequence;

[0051] Divide the sampling values of the relative system state vector according to the set grid size, and discretize the continuous state space of the relative system;

[0052] Use the zero - level subset of a bounded and Lipschitz - continuous signed - distance function with respect to the joint state of the system to represent the unsafe state set of the system;

[0053] Set the rules for trimming the action space of the opponent agent;

[0054] Solve the backward reachable tube and the safety control strategy through the level - set method;

[0055] Calculate the backward reachable tubes of all relative systems in sequence and take the union as the backward reachable tube of the human - machine collaborative system.

[0056] In one embodiment, the establishing a relative system model between the controlled agent and a certain interactive agent in sequence specifically includes:

[0057] Consider the interaction between the controlled agent \(i\) and any other agent \(j\in - i\), where \(-i\) is the set of opponent agents. Then the dynamic equation of the relative system composed of these two agents is:

[0058]

[0059] where \(x\) rel is the relative state vector, \(a\) i , \(a\) j are the control inputs of agent \(i\) and agent \(j\) respectively. Assume that for fixed \(a\) i , \(a\) j is uniformly continuous, bounded and Lipschitz - continuous.

[0060] In one embodiment, the using the zero - level subset of a bounded and Lipschitz - continuous signed - distance function with respect to the joint state of the system to represent the unsafe state set of the system specifically includes:

[0061] The backward reachable set represents a state set at the current moment. Starting from this state set, the relative system will definitely enter the unsafe state set within the next \(T\) time. Use the zero - level subset of a bounded and Lipschitz - continuous signed - distance function \(l(x\) rel ) to represent the unsafe state set of the system:

[0062]

[0063] \(x\) relis the relative state vector of the agents in the relative system.

[0064] In one of the embodiments, the rule for trimming the action space of the opponent agent specifically includes:

[0065] When solving the backward reachable set, it is assumed that the goal of the controlled agent i is to keep the system away from the unsafe state set, while the goal of agent j is to make the system enter the unsafe state set. Under the action of the relative system dynamics and the above-determined unsafe state set, the backward reachable tube can be defined as

[0066]

[0067] where ξ f is the system trajectory calculated from the system dynamics equation for the given control inputs of the two agents, and γ represents the non-predictive strategy that agent j takes in response to α i .

[0068] In the calculation of the backward reachable tube, the worst-case scenario is assumed: agent j will take actions that make the system enter the unsafe state set the fastest. However, in real human-machine interactions, robots or humans will not maliciously damage the system. Therefore, the calculation process of the backward reachable tube is too conservative, which results in the system initial states included in the backward reachable tube being too large. To reduce the range of the backward reachable tube and lower the triggering frequency of the safety control strategy, after modeling the behavior of the opponent using the game model and the policy model, the action space of the opponent is trimmed.

[0069] The trimming rule is to ignore actions with probabilities lower than the threshold ∈, and obtain the trimmed action space A j of the opponent agent j at the current state:

[0070] A j (x rel , a i ):={a j : P(a j |x rel , a i ) > ∈};

[0071] a i and a j are the control inputs of agent i and agent j respectively, and P(a j |x rel , a i ) represents the probability that the opponent agent j selects action a rel at the current state x j and the action of agent i; the trimmed action space of the opponent agent is represented by the instantaneous action boundary , where is the upper bound of the action, τ ∈ {0, …, T}, where T represents the prediction duration of the backward reachable tube. A j (x rel , a i ) = min a j (x rel τ , a i τ ) is the lower bound of the action.

[0072] In one of the embodiments, solving the backward reachable tube and the safety control strategy by the level set method specifically includes:

[0073] Using the level set method, determine the range of the backward reachable tube by solving the viscosity solution of the time-varying Hamilton-Jacobi-Isaacs partial differential equation:

[0074] The set of states where the value function v(x rel , t) is zero for the grid points of the discrete state space of the system is the zero level set, and the zero level set corresponds to the safety state boundary of the system. The negative level set is the backward reachable tube of the system. The Hamilton-Jacobi-Isaacs partial differential equation is expressed as:

[0075] min[D t v(x rel , t) + H(x rel , t, D xrel v(x rel , t)), l(x rel ) - v(x rel , t)] = 0;

[0076] v(x rel , 0) = l(x rel ), t ∈ [-T, 0];

[0077] Where D xrel v(x rel , t) is the spatial derivative of the value function, is the Hamilton operator; x rel is the relative state vector, and -T represents the duration of backward prediction of the backward reachable tube;

[0078] The backward reachable tube V(t) is represented by the zero lower subset of the value function:

[0079]

[0080] At the same time, obtain the optimal safety control action of agent i:

[0081]

[0082] In one embodiment, the intelligent robot executes the control strategies of different rationality levels, infers the rationality level of the opponent robot, calculates the backward reachable tube, and executes the safety control strategy, specifically including:

[0083] During the operation of the human-robot collaborative system, each robot executes control strategies of different rationality levels. Before executing an action at each moment, it infers the rationality level of the opponent robot, that is, obtains the control strategy of the opponent robot, then cuts the action space of the opponent, calculates the less conservative backward reachable tube, and when the state of the human-robot collaborative system reaches the boundary of the backward reachable tube, the robot starts to execute the safety control strategy.

[0084] The opponent robot in the present invention refers to other robots that interact with the current robot.

[0085] Compared with the prior art, the beneficial technical effects of the present invention are:

[0086] (1) Considering the bounded rational behavior of humans and robots due to the inability to obtain global information, cognitive ability limitations, computing power limitations, etc., using bounded rationality games to model the interaction between agents and obtaining the strategies of different rationality levels of each agent through a recursive reasoning model, which is more in line with the real situation of the real world.

[0087] (2) Inferring the rationality level of the opponent agent through the historical data of the interaction, predicting the non-optimal behavior of the opponent agent in the current state, and realizing accurate and differentiated modeling of the interacting opponent agent.

[0088] (3) Based on the result of modeling the opponent agent, cutting the action space of the interacting opponent agent, ignoring the actions with high safety hazards but extremely low occurrence probabilities, reducing the range of the backward reachable tube, reducing the frequency of triggering the safety control strategy, and reducing conservatism. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 It is a schematic flowchart of the human-robot interaction safety control method based on bounded rationality games and reachability analysis in the embodiment of the present invention.

[0090] Figure 2 It is a schematic diagram of the kinematic model of the vehicle in the embodiment of the present invention and the ranges of the uncertainty area, the unsafe area, and the collision area. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0091] The present disclosure provides a human-machine collaborative safety control method based on bounded rationality game and reachability analysis. The safety control method takes into account the bounded rationality of the robots participating in the interaction and the differences in rationality levels, solves the problem that the Nash equilibrium game model inaccurately models human behavior in the real world, and designs a conservativeness that can improve the backward reachable tube to ensure interaction safety based on opponent modeling, thereby improving efficiency on the basis of ensuring safety.

[0092] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the following further elaborates on the present disclosure in detail with reference to specific embodiments and the accompanying drawings.

[0093] In an embodiment of the present disclosure, a method for interactive control of an autonomous vehicle based on bounded rationality game and reachability analysis is provided, which is applied to a highway scenario of multiple vehicles. In this scenario, the task of the vehicle is to move forward quickly and safely. To meet the requirement of high speed, the vehicle will perform complex behaviors such as lane changing and overtaking according to the obstacles and congestion of the vehicles ahead; to ensure driving safety, it is necessary to observe the behaviors of the interactive vehicles, infer the driving styles of the interactive vehicles, predict the future driving trajectories of the interactive vehicles, and make decisions based on this. This method models the human-driven vehicles within the field of view, enabling the autonomous vehicle to adjust its decisions according to the predicted behaviors of the interactive vehicles, and calculates the backward reachable tube through reachability analysis. When the state of the interactive vehicle reaches the boundary of the backward reachable tube, the autonomous vehicle can timely switch to a safety control strategy. The specific implementation process is as Figure 1 shown, and the human-machine interaction safety control method includes:

[0094] Step A: Determine the models of the intelligent agents participating in the human-machine interaction;

[0095] Step B: Use recursive reasoning to model the interaction between bounded rational intelligent agents;

[0096] Step C: Design a reward function to guide the reinforcement learning training of each intelligent agent and train the control strategies of different rationality levels of each intelligent agent step by step;

[0097] Step D: Design a method for inferring the rationality level of the interaction opponent;

[0098] Step E: Obtain the backward reachable tube and safety control strategy of the low-conservative human-machine collaborative system through Hamilton-Jacobi reachability analysis;

[0099] Step F: Achieve safe, efficient interaction and collaboration of the human-machine collaborative system through the control strategies of different rationality levels, the inference method, and the safety control strategy.

[0100] In one embodiment, determining the models of the intelligent agents participating in the human-machine interaction in Step A specifically includes:

[0101] Determine the dynamic models of autonomous vehicles and interacting human-driven vehicles. In a high-speed scenario where vehicles interact, a total of n vehicles are running together. The state of each vehicle consists of its position on the road, yaw angle, and speed. The dynamic equation of each vehicle i can be expressed as:

[0102]

[0103] where the state z = (x, y, v, ψ), x and y are the abscissa and ordinate of the vehicle's position on the road respectively, v is the vehicle speed, ψ is the vehicle yaw angle, and the control input u = (a, δ), a is the acceleration, and δ is the front wheel angle of the vehicle. respectively represent the first-order derivatives of the abscissa x and ordinate y of the vehicle's position on the road, and f is a non-linear function. l f ,l r are the distances from the vehicle's mass center to the front axle and rear axle respectively; z i represents the state of vehicle i, and u i represents the control input of vehicle i. is the first-order derivative of the vehicle yaw angle ψ, represents the first-order derivative of the state of vehicle i.

[0104] In one embodiment, the use of recursive reasoning to model the interaction between bounded-rational agents in step B specifically includes:

[0105] Step B1: Set the recursive rules of the recursive reasoning game model;

[0106] Step B2: Set the policy model of the level-0 agent;

[0107] Step B3: Set the policy model of the level-k agent.

[0108] In one embodiment, the setting of the recursive rules of the recursive reasoning game model in step B1 specifically includes:

[0109] In a multi-vehicle highway scenario, different drivers have different driving strategies. The method of the present disclosure only distinguishes different driving strategies by the level of rationality, that is, drivers with the same level of rationality have the same driving strategy. Assume that the level of rationality of the vehicle to be controlled is k, and the levels of rationality κ of other vehicles on the road are different but all lower than that of the vehicle to be controlled, and the levels follow a Poisson distribution Then the vehicle to be controlled assumes that the levels of rationality of other agents satisfy the probability based on the Poisson distribution:

[0110]

[0111] The optimal strategy of the controlled vehicle with a rationality level of k when facing a cluster of opponent vehicles with a rationality level distribution of is as follows: That is:

[0112]

[0113] where is the expected cumulative reward function of the controlled vehicle when the opponent vehicle samples and uses the κ-level strategy.

[0114] In one embodiment, setting the policy model of the level 0 agent in step B2 specifically includes:

[0115] The level 0 agent adopts a Max-Min decision model. Assuming that other vehicles will take actions that cause the greatest damage to the reward function of the controlled vehicle, the best response of the controlled vehicle under this assumption is:

[0116]

[0117] In one embodiment, setting the policy model of the level k agent in step B3 specifically includes:

[0118] The level k driving strategy is obtained through reinforcement learning training. The training policy rationality level gradually increases. When training the k, level policy, the rationality levels of other vehicles in the environment are sampled according to the probability Since κ < k and the driving strategies of all vehicles are only distinguished by the rationality level, the strategies of other vehicles in the environment are all pre-trained and can be directly loaded.

[0119] In one embodiment, designing the reward function for guiding each agent to perform reinforcement learning training in step C and training the control strategies of each agent at different rationality levels level by level specifically includes:

[0120] Step C1: Design the reward function of each agent;

[0121] Step C2: Set the reinforcement learning algorithm;

[0122] Step C3: Train the control strategies corresponding to each rationality level of each agent level by level.

[0123] In one embodiment, designing the reward function of each agent in step C1 specifically includes:

[0124] In the multi-vehicle highway scenario, the task objectives of the vehicles are the same, and high-speed operation is achieved on the premise of safe driving.

[0125] To ensure safety and avoid collisions between vehicles, three rectangular regions are set according to the vehicle body shape, such asFigure 2 As shown, the rectangular areas from large to small are the uncertainty area, the unsafe area, and the collision area. The three-layer progressive setting can make the vehicle have a sense of crisis before a collision and take measures in advance. When the corresponding areas of the vehicle overlap respectively, an uncertainty penalty w1 = -1, an unsafe penalty w2 = -1, and a collision penalty w3 = -1 are given, and no penalty is given in other cases.

[0126] To achieve high-speed operation, a linear reward w4 = v is given to the vehicle speed within the vehicle's speed range.

[0127] To improve the smoothness of the vehicle's actions and avoid dangerous behaviors such as sudden acceleration, sudden deceleration, and sharp turns, a linear penalty w5 = -|a| - |δ| is given to the vehicle's control input.

[0128] Finally, based on the above three aspects, the reward function of the vehicle is:

[0129] r = w1*r1 + w2*r2 + w3*r3 + w4*r4 + w5*r5.

[0130] In one embodiment, the setting of the reinforcement learning algorithm in step C2 specifically includes:

[0131] Considering the importance of the smoothness of the vehicle's actions, two control inputs of the vehicle, the acceleration and the front wheel angle, are set to be continuously variable. Therefore, the SAC algorithm for complex control problems in the continuous action space is selected to train the k-level driving strategy.

[0132] In one embodiment, the step of gradually training the control strategy corresponding to each rational level of each agent in step C3 specifically includes:

[0133] Start training from the 1-level strategy, gradually increase the size of the rational level, and when training the k-level strategy, sample the strategies of other vehicles in the environment according to the probability distribution P k (k), and use the SAC algorithm determined in step C2 to learn the k-level strategy following the reward function designed in step C1.

[0134] In one embodiment, the method of designing to infer the rational level of the interaction opponent in step D specifically includes:

[0135] Observe the real actions of the opponent vehicle within the interaction field of view at each moment Based on the observed values of the opponent vehicle Calculate the best actions corresponding to the k ∈ {0, 1,..., k - 1} level strategies Calculate the difference between the best action at each level and the real action taken by the opponent at the current moment to obtain the most likely rational level k of the opponent at the current moment * , which can be expressed as

[0136]

[0137] For k * increase the belief in k by ΔP and normalize it to obtain the belief T about the opponent vehicle (K=k) [j](t):

[0138]

[0139] In one embodiment, obtaining the backward reachable tube and the safety control strategy of the low-conservative human-machine collaborative system through Hamilton-Jacobi reachability analysis in step E specifically includes:

[0140] Step E1: successively establish the relative system model between the controlled agent and a certain interactive agent;

[0141] Step E2: discretize the state space of the relative system into grids according to a set size;

[0142] Step E3: set the range of the unsafe set;

[0143] Step E4: set the rules for pruning the opponent's action space;

[0144] Step E5: solve the backward reachable tube and the safety control strategy through the level set method.

[0145] Step E6: successively calculate the backward reachable tubes of all relative systems and take the union as the backward reachable tube of the human-machine collaborative system.

[0146] In one embodiment, successively establishing the relative system model between the controlled agent and a certain interactive agent in step E1 specifically includes:

[0147] Establish the relative system between the controlled vehicle and any opponent vehicle. To obtain an affine relative system, simplify the vehicle model in the backward reachable tube calculation part as:

[0148]

[0149] where the state vector of each vehicle is still z = (x, y, v, ψ), x and y are the coordinates of the vehicle on the road, v is the vehicle speed, ψ is the vehicle yaw angle, and the control input u = (a, w), a is the acceleration, and w is the angular velocity.

[0150] Define the origin of the relative system coordinate system at the center of the controlled vehicle and keep it consistent with the direction of the controlled vehicle coordinate system, then the relative position x rel , y rel is defined as the position of the opponent vehicle on the horizontal and vertical axes of the relative system coordinate system:

[0151] xrel = cosψ1(x2 - x1) + sinψ1(y2 - y1);

[0152] y rel = -sinψ1(x2 - x1) + cosψ1(y2 - y1);

[0153] The relative yaw angle can be expressed as: ψ rel = ψ2 - ψ1.

[0154] Then the relative system between the two vehicles can be modeled as:

[0155]

[0156] In one embodiment, discretizing the state space of the relative system into a grid according to a set size in step E2 specifically includes:

[0157] Discretize each dimension of the five-dimensional relative system state vector z rel = (x rel , y rel , ψ rel , v1, v2) according to the set grid size respectively.

[0158] In one embodiment, setting the range of the unsafe set in step E3 specifically includes:

[0159] For the highway scenario of multiple vehicles, the unsafe state of the relative system between two vehicles is the state set when the vehicles collide, and the vehicle collision can be expressed as the overlap of the positions occupied by the two vehicles, that is, the absolute values of x rel , y rel are respectively less than the length and width of the vehicle. Therefore, define the function l(z rel ) representing the unsafe state set of the system as:

[0160]

[0161] where len is the vehicle length and wd is the vehicle width.

[0162] The unsafe state set of the relative system is:

[0163]

[0164] In one embodiment, setting the rule for trimming the opponent's action space in step E4 specifically includes:

[0165] Since the strategy of the interacting vehicle is obtained by reinforcement learning, it can be based on the current observation value obs of the opponent vehicle 2Obtain the probability density function of the opponent's continuous action space at the current moment, so as to obtain the probability P(u 2 ) of the opponent taking each action u 2 |obs 2 ). Set the clipping threshold ∈ = 0.001, then the clipped opponent action space is:

[0166] U 2 (obs 2 ):={u 2 :P(u 2 |obs 2 )>0.001};

[0167] The opponent's instantaneous action boundary is:

[0168]

[0169] where is the action upper bound, U 2 (obs 2 ) = min U 2 (obs 2 τ ), τ ∈ {0, …, T} is the action lower bound.

[0170] In one embodiment, in step E5, by using the level set method to solve the backward reachable tube and the safety control strategy, it specifically includes:

[0171] Solve the Hamilton-Jacobi-Isaacs partial differential equation:

[0172]

[0173] v(z rel , 0) = l(z rel ), t ∈ [-T, 0];

[0174] where, is the spatial derivative of the value function, is the Hamilton operator.

[0175] Then the backward reachable tube can be represented by the zero-level subset of the value function:

[0176]

[0177] At the same time, the optimal safety control action of the controlled vehicle can be obtained:

[0178]

[0179] In one embodiment, the step of successively calculating the backward reachable tubes of all relative systems in step E6 and taking the union as the backward reachable tube of the human-machine cooperation system specifically includes:

[0180] Use V 1j (t) to represent the backward reachable tubes of the relative systems of the controlled vehicle and an opponent vehicle j∈-i within any interaction field of view in turn. Finally, from the perspective of the controlled vehicle at the current moment, the backward reachable tube can be represented as the union of all V 1j (t), that is:

[0181] V 1 (t) = V 11 (t) ∪ V 12 (t) ∪ … ∪ V 1n (t);

[0182] Where n represents the number of opponent vehicles within the interaction field of view.

[0183] In one embodiment, the safe, efficient interaction and cooperation of the human-machine cooperation system are realized through the control strategies of different rationality levels, the inference method, and the safety control strategy, specifically including:

[0184] Use the efficient reinforcement learning control strategy determined in step C and the low-conservative safety control strategy determined in step E to realize the safe and efficient interaction and cooperation of the human-machine cooperation system. Specifically, during the operation of the human-machine cooperation system, each robot executes the reinforcement learning control strategy in step C. Before performing an action at each moment, infer the rationality level of the interaction opponent through step D, that is, obtain the control strategy of the opponent, and then trim the action space of the opponent through step E to calculate the low-conservative backward reachable tube. When the state of the human-machine cooperation system reaches the boundary of the backward reachable tube, the robot starts to execute the safety control strategy in step E.

[0185] It should be understood that although the steps in the flowchart of the accompanying drawings of the specification are shown in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowchart of the accompanying drawings of the specification may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0186] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.

[0187] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A human-computer interaction security control method based on bounded rationality game and reachability analysis, characterized in that Including: Based on the dynamic equations of the robot, construct a human-robot collaborative system model composed of multiple robots with autonomous decision-making capabilities, namely agents; Use recursive reasoning to model the interactions between bounded-rational agents; Design a reward function to guide the reinforcement learning training of each agent and gradually train the control strategies of each agent at different rationality levels; By initializing the belief probability of the opponent agent and dynamically updating the belief probability based on the difference between the observed action and the multi-level rationality prediction action, combined with normalization processing, infer the rationality level of the opponent agent in real-time; Obtain the backward reachable tube and safety control strategy of the low-conservative human-robot collaborative system through Hamilton-Jacobi reachability analysis; The robot executes the control strategies at different rationality levels, infers the rationality level of the opponent robot, calculates the backward reachable tube, and when the state of the human-robot collaborative system reaches the boundary of the backward reachable tube, executes the safety control strategy.

2. The human-computer interaction security control method based on bounded rationality game and reachability analysis according to claim 1, characterized in that The use of recursive reasoning to model the interactions between bounded-rational agents specifically includes: Set the recursive rules of the recursive reasoning game model; Set the strategy of the agent with a rationality level of 0; Train the strategy of the agent with a rationality level of k through reinforcement learning.

3. The human-computer interaction security control method based on bounded rationality game and reachability analysis according to claim 2, wherein The setting of the recursive rules of the recursive reasoning game model specifically includes: Take the number of steps of the agent's reasoning as the rationality level of the agent. An agent with a rationality level of k assumes that the rationality levels of its opponent agents are κ ∈ {0, 1, …, k - 1}, and follows a Poisson distribution. Then an agent with a rationality level of k believes that a certain opponent agent has a rationality level of κ or the proportion of agents with level k among all opponent agents is: The optimal strategy when the rationality level of the i-th agent is k is as follows: where -i is the set of opponent agents, is the optimal strategy when the rationality level of the opponent agent is k, and k follows a distribution x is the state of the human-machine collaborative system, and J i is the expected cumulative reward function when the opponent agent executes a certain strategy π -i and agent i executes strategy π i : Among them, p is the state transition function, r i (x t , π i (x), π -i (x)) is the reward function of agent i, γ is the discount factor, x t is the state of the human-machine collaborative system at time t; The policy of the agent with the rationality level k set by reinforcement learning training specifically includes: The policies of other agents in the reinforcement learning environment are obtained by sampling and performing sampling.

4. The human-computer interaction security control method based on bounded rationality game and reachability analysis according to claim 2, characterized in that The setting of the strategy of the agent with a rationality level of 0 specifically includes: An agent with a rationality level of 0 has 0 reasoning steps and does not reason about the behavior of opponent agents during interaction; assuming that all other agents will take actions that minimize the current agent's immediate reward function, the best response of an agent with a rationality level of 0 is: Among them, x is the state of the human-machine collaboration system, a i represents the action of agent i, A i represents the action space of agent i, a -i represents the joint action of agent i's opponent, A -i represents the joint action space of agent i's opponent.

5. The human-computer interaction security control method based on bounded rationality game and reachability analysis according to claim 1, characterized in that The design of the reward function to guide the reinforcement learning training of each agent and the gradual training of the control strategies of each agent at different rationality levels specifically includes: Design reward function terms according to the task objectives, safety restrictions, and behavior smoothness of each agent and assign the weight values of each reward function term to obtain the reward function of each agent; The SAC algorithm is used to train the control strategies corresponding to each rational level of each agent step by step: the level 0 strategy of each agent is known, and when training the level k strategy of agent i, any other agent j∈-i in the environment is trained according to probability. Sample your own rationality level and execute your strategy Initialize the k-level strategy of agent i to randomly select actions and use the SAC algorithm for training. The obtained strategy is to obey the rational level The optimal strategy of the opponent agent After training all agents’ k-level strategies, you can increase the rationality level and traverse the agents for training until the maximum rationality level K is reached; is the proportion of the opponent agents whose rationality level is k that the opponent agents have a rationality level of κ or the proportion of all opponent agents with a rationality level of κ, κ∈{0,1,…,k-1}, -i is the set of opponent agents.

6. The human-machine interaction security control method based on bounded rationality game and reachability analysis according to claim 1, characterized in that The real-time inference of the rationality level of the opponent agent by initializing the belief probability of the opponent agent and dynamically updating the belief probability based on the difference between the observed action and the multi-level rationality prediction action, combined with normalization processing, specifically includes: Use T (K=k) [j](t) represents the belief probability of modeling the opponent agent j as level k, which is initialized as a probability function with a uniform distribution. The belief of the opponent agent is updated at each moment: The update rule is to observe the action taken by the opponent agent j at the current time t Obtain the current state observation value from the perspective of the opponent agent j Predict the actions that the opponent agent will take when executing strategies with different rationality levels Calculate the difference between the predicted actions at each rationality level and the actual action taken by the opponent agent at the current time, and obtain the most likely rationality level κ of the opponent agent at the current time * : For κ * the belief increases by ΔP and is normalized: ΔP represents the belief increment about the most likely rational level of the opponent agent, which represents the updated belief that the rational level of the opponent agent is κ * and then updates the belief at the current moment by normalization.

7. The human-computer interaction security control method based on bounded rationality game and reachability analysis according to claim 1, characterized in that The obtaining of the backward reachable tube and safety control strategy of the low-conservative human-robot collaborative system through Hamilton-Jacobi reachability analysis specifically includes: Successively establish the relative system model between the controlled agent and a certain interacting agent; Divide the sampling values of the relative system state vector according to the set grid size and discretize the continuous state space of the relative system; Use a bounded, Lipschitz continuous signed distance function l(x rel ) of the joint state of the system to represent the set of unsafe states of the system x rel is the relative state vector of the agents in the relative system; Set the rules for pruning the action space of the opponent agent; Solve the backward reachable tube and safety control strategy through the level set method; Successively calculate the backward reachable tubes of all relative systems and take the union as the backward reachable tube of the human-robot collaborative system.

8. The human-computer interaction security control method based on bounded rationality game and reachability analysis according to claim 7, characterized in that The successive establishment of the relative system model between the controlled agent and a certain interacting agent specifically includes: Consider the interaction between the controlled agent i and any other agent j∈-i, where -i is the set of opponent agents. Then the dynamic equation of the relative system composed of these two agents is: where a i and a j are the control inputs of agent i and agent j, respectively. Assume that for a fixed a i , a j is uniformly continuous, bounded, and Lipschitz continuous.

9. The human - machine interaction security control method based on bounded - rationality game and reachability analysis according to claim 7, characterized in that, The setting of the rules for pruning the action space of the opponent agent specifically includes: The rule of pruning is to ignore the actions with probabilities lower than the threshold ∈, and obtain the pruned action space A of the opponent agent j under the current state j : A j (x rel ,a i ):={a j :P(a j |x rel ,a i )>∈}; a i and j are the control inputs of agent i and agent j respectively. P(a j |x rel , a i ) represents the probability that the opponent agent j selects the action a rel under the current state x j and the action of agent i; the action space of the opponent agent after clipping is represented by the instantaneous action boundary , where is the upper bound of the action, τ ∈ {0, …, T}, and T represents the prediction horizon of the backward reachable tube A j (x rel , a i ) = min A j (x rel τ , a i τ ) is the lower bound of the action.

10. The human-computer interaction security control method based on bounded rationality game and reachability analysis according to claim 7, characterized in that The solution of the backward reachable tube and safety control strategy through the level set method specifically includes: Use the level set method to determine the range of the backward reachable tube by solving the viscosity solution of the time-varying Hamilton-Jacobi-Isaacs partial differential equation: The set of states for which the value function v(x rel ,t) is zero with respect to the grid points of the discrete state space of the system is the zero level set, and the zero level set corresponds to the safety state boundary of the relative system. The negative level set is the backward reachable tube of the relative system. The Hamilton-Jacobi-Isaacs partial differential equation is expressed as: v(x rel ,0) = l(x rel ), t ∈ [-T, 0]; wherein, is the spatial derivative of the value function, is the Hamiltonian operator; x rel is the relative state vector, and -T represents the duration of the backward prediction of the backward reachable tube; The backward reachable tube V(t) is represented by the zero-lower subset of the value function: Meanwhile, the optimal safe control action of agent i is obtained:

Citation Information

Cited By

  • Man-machine dynamic game control method driven by double reinforcement learning

    CN122035054A