Hierarchical reinforcement learning method for crowd adaptive behavior simulation
By constructing statically planned paths and dynamically adjusting speed weights through a hierarchical reinforcement learning method, the problem of global path optimization in multi-directional high-density crowd environments is solved, achieving efficient navigation without causing local collisions and improving navigation success rate and learning efficiency.
Patent Information
- Application Number
- CN202511018669.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-07
Smart Images

Figure CN120909106A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent agents, and in particular to a hierarchical reinforcement learning method for crowd adaptive behavior simulation. BACKGROUND
[0002] In the field of crowd motion simulation, self-organizing behavior is a phenomenon of individuals spontaneously forming and orderly structured through local interaction. For example, in one-way, two-way or multi-directional crowds, individuals will spontaneously form specific channel structures by avoiding each other. This self-organizing behavior is particularly important in high-density and dynamically changing crowd environments. Therefore, effectively simulating this self-organizing behavior in crowds not only helps to study the dynamics of crowd movement, but also has important practical application value in optimizing crowd flow, preventing potential accidents in large events, managing emergency evacuation and designing public spaces. Therefore, reinforcement learning (RL) provides a potential solution for crowd self-organizing behavior simulation. Reinforcement learning is a machine learning method that allows an agent to autonomously explore and learn the optimal policy through interaction with the environment, and autonomously explores and learns the best decision-making policy.
[0003] In a dynamic crowd environment, individual behavior is mainly driven by local information, and decisions are made based on the current environment. The reinforcement learning mechanism is highly consistent with the self-organizing behavior of the crowd, which allows individuals (agents) to continuously optimize their behavior strategies through trial and error without relying on fixed rules, thereby adapting to environmental changes. This adaptability allows reinforcement learning to dynamically adjust individual behavior in real-time environments, simulating the process of individuals autonomously avoiding collisions in crowds.
[0004] However, current reinforcement learning methods for crowd adaptive behavior simulation still have some limitations in highly complex crowd scenarios, especially in multi-directional high-density crowd environments. Reinforcement learning methods have difficulty in effectively balancing individual global path planning and local collision avoidance tasks, lack of refined control of global and local strategies, leading to difficulty in achieving global path optimization without causing local collisions, which results in high collision rates and low navigation success rates in high-density environments.
[0005] Therefore, how to achieve global path optimization without causing local collisions has become a technical problem that needs to be solved in the application field of reinforcement learning in crowd adaptive behavior simulation. SUMMARY
[0006] The technical problem to be solved by the present application is to provide a hierarchical reinforcement learning method for crowd adaptive behavior simulation that can achieve global path optimization without causing local collisions.
[0007] The technical scheme adopted by the present application to solve the above technical problems is: a hierarchical reinforcement learning method for crowd adaptive behavior simulation, characterized in that it comprises the following steps:
[0008] Step S1, based on the current position of the agent in the crowd, the target point and the environment structure, a static planning optimal path of the agent in a static closed environment is constructed;
[0009] Step S2, based on the static planning optimal path of the agent constructed and based on the unit tangent segmentation of the current position point of the agent, a target trend speed of the agent moving from the current position point to the target point is generated;
[0010] Step S3, a nonlinear mapping of laser radar observation data to collision avoidance speed is established, and dynamic obstacle avoidance guidance is provided based on the environment around the agent;
[0011] Step S4, the weights of the target trend speed and the collision avoidance speed are dynamically adjusted according to the personnel distribution information around the agent, to obtain an adjusted target trend speed weight value and an adjusted collision avoidance speed weight value;
[0012] Step S5, the current target trend speed, the adjusted target trend speed weight value, the target trend speed weight adjustment value and the adjusted collision avoidance speed weight value are coupled by speed weighting, to obtain an adaptive walking speed of the agent.
[0013] Improvements are made in the hierarchical reinforcement learning method for crowd adaptive behavior simulation, in step S1, a path search algorithm based on heuristic function is used, and based on the current position of the agent in the crowd, the target point and the environment structure, a static planning optimal path of the agent is constructed.
[0014] Further, in the hierarchical reinforcement learning method for crowd adaptive behavior simulation, the target trend speed is generated as follows:
[0015] v goal =(v x ,v y );
[0016]
[0017] Wherein, v goal is the target trend speed of the agent, v x is the horizontal component of the target trend speed v goal of the agent, ||v x || is the modulus of the component speed v x , and v y is the target trend speed v goalthe speed component in the vertical direction y ||is the modulus of the speed component v y M is the static planning optimal path of the agent obtained by the path search algorithm based on the heuristic function, is the unit tangent vector generated by the agent at the position point (p x , p y ) on the static planning optimal path M.
[0018] Further improvement, in the invention, the hierarchical reinforcement learning method for crowd adaptive behavior simulation further comprises:
[0019] Based on the state of other agents in the environment around the agent, the collision risk value of the agent at the current time is obtained;
[0020] According to the obtained agent collision risk value of the agent, the weight of the target trend speed and the weight of the collision avoidance speed of the agent are adjusted respectively, so as to realize the weighted coupling of the target trend speed and the collision avoidance speed of the agent.
[0021] Further, in the hierarchical reinforcement learning method for crowd adaptive behavior simulation, the collision risk value of the agent at the current time comprises the following steps a1-a3:
[0022] Step a1, obtaining the observable state information of other agents around the agent; wherein the observable state information is represented as follows:
[0023]
[0024] Wherein, represents the observable state of the i-th other agent around the agent, and Grid Sensor() represents the perception model for perceiving the state of the environment around the agent; Grid Sensor() is a perception model based on ray detection designed in Unity3D engine to support reinforcement learning model, which is used to represent the perception model for perceiving the state of the environment around the agent; n is the total number of other agents around the agent;
[0025] Step a2, encoding the position and trajectory of the agent according to the obtained observable state information of the agent; wherein:
[0026]
[0027] Wherein, represents the value of the speed vector of the j-th agent at time step t, represents the encoded position of the i-th agent around the agent, is the hidden state of the gating recurrent unit H-GRU, and || is the splicing operation; WP For embedding functions The parameter W P It is a weight matrix or weight parameters; Let be the embedding vector of the i-th agent at time step t. Through function From input state Mapping yields W G is a function The weight parameters are the weight matrix in the neural network, used to weight the input state. Mapped to embedding vector For the weight parameter W G Initialization embedding function; W H Hidden state The weights;
[0028] Step a3: Calculate the agent's collision risk value at the current moment; whereby the agent's collision risk value at the current moment is denoted as DANGER. gs :
[0029]
[0030]
[0031] Among them, DANGER gs Let g be the collision risk value of the gs-th agent, and softmax(·) represent the normalized exponential function characterizing the relative danger level between agents; N is the total number of agents in the scene. It is a hyperparameter; [gs-1] represents the danger value of the extracted gs-th agent;
[0032] is the query vector of the gs-th agent at time step t, used in the attention mechanism to calculate the correlation with other agents; W is the key vector of the ts-th agent at time step t, used in the attention mechanism to evaluate the interaction with the gs-th agent; QE It is a function The weight matrix is used to weight the velocity. Mapped to query vector This represents the velocity vector of the gs-th agent at time step t. It is the target's approach speed or final speed;
[0033] W QE The initialization embedding function, W KE It is a function The weight matrix is used to weight the velocity. Mapping as a key vector is by W KE initialized embedding function, denotes the velocity vector of the ts-th agent at time step t, which velocity vector is the target trend velocity or final velocity.
[0034] Further improvement, in the hierarchical reinforcement learning method for crowd adaptive behavior simulation, in step S5, the adaptive walking speed of the agent is calculated as follows:
[0035] v output = TY1 x v goal + DANGER gs x TY2 x v collision ;
[0036] Wherein, v output is the adaptive walking speed of the agent on the final planned path, TY1 is the first hyperparameter, v goal is the target trend velocity of the agent, DANGER gs is the collision risk value of the gs-th agent, TY2 is the second hyperparameter, v collision is the collision avoidance speed for the agent.
[0037] Further, in the invention, the hierarchical reinforcement learning method for crowd adaptive behavior simulation further comprises: based on a preset comprehensive reward function, encouraging the agent to reach the designated target location as soon as possible without mutual collision.
[0038] Further, in the hierarchical reinforcement learning method for crowd adaptive behavior simulation, the preset comprehensive reward function is set as follows:
[0039]
[0040]
[0041]
[0042] Further, in the hierarchical reinforcement learning method for crowd adaptive behavior simulation, the r warning1 is the VO obstacle avoidance area calculated by the RVO obstacle avoidance algorithm; wherein:
[0043] When the agent agent A is less than 0.5m away from any other agent agent B, the reward signal given to punish the impending collision, its value r warning1 is set to -0.5;
[0044] When the agent agent A is less than 1m away from any other agent agent B, the given future collision penalty, its value r warning2 Is set to -0.25.
[0045] Improved in the application, the hierarchical reinforcement learning method for crowd adaptive behavior simulation further comprises:
[0046] Divide the reinforcement learning of the agent into a low-level structure and a high-level structure;
[0047] Let the steps S1-S4 be executed in the low-level structure, and let the step S5 be executed in the high-level structure;
[0048] And, by means of a discount reward, the behavior of the agent in the crowd is optimized; wherein the discount reward is marked as G t :
[0049]
[0050] Wherein G global Is the discount reward value introduced in the high-level, G local Is the discount reward value introduced in the low-level; Indicates the k-th power of the discount factor γ global Of the global reward, that is, the discount weight of the global reward at time step t+k; R global Indicates the local collision avoidance reward item introduced in the high-level; Indicates the k-th power of the discount factor γ local Of the global reward, that is, the discount weight of the global reward at time step t+k; R local Indicates the local collision avoidance reward item introduced in the low-level, and k indicates the offset of the time step.
[0051] Compared with the prior art, the application has the following advantages:
[0052] Firstly, the hierarchical reinforcement learning method for crowd adaptive behavior simulation of the application constructs a static planning optimal path of the agent in a static closed environment based on the current position of the agent in the crowd, the target point and the environment structure, generates a target trend speed of the agent moving from the current position point to the target point based on the unit tangent segmentation of the current position point of the agent on the basis of the static planning optimal path, establishes a nonlinear mapping of laser radar observation data to collision avoidance speed and provides dynamic obstacle avoidance guidance based on the environment around the agent, dynamically adjusts the weight of the target trend speed and the weight of the collision avoidance speed according to the personnel distribution information around the agent, and finally couples the speed according to the current target trend speed of the agent, the adjusted target trend speed weight value, the target trend speed weight adjustment value and the adjusted collision avoidance speed weight value, to obtain the adaptive walking speed of the agent. In this way, both the overall flow direction factor of the crowd when the agent moves in the crowd and the local obstacle avoidance demand of the agent in the real-time environment are considered, the adaptive path adjustment needs of the agent in the high-density crowd environment are ensured, and the global path optimization under the condition that the agent does not cause local collision is realized.
[0053] Secondly, the hierarchical reinforcement learning method for crowd adaptive behavior simulation of the application also introduces a comprehensive reward function to enable the agent to reach the execution target point as soon as possible without mutual collision, thereby improving the learning efficiency of the entire hierarchical reinforcement learning.
[0054] Finally, the application divides the reinforcement learning of the agent into a low-level structure and a high-level structure, so that steps S1-S4 are executed in the low-level structure and step S5 is executed in the high-level structure, thereby realizing the cooperative cooperation of the low-level structure and the high-level structure in the learning process to ensure the balance between global path planning and local obstacle avoidance of the agent, and realizing the self-organizing behavior simulation of the agent more naturally. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1 The figure is a flowchart of the hierarchical reinforcement learning method for crowd adaptive behavior simulation in the embodiment of the application. DETAILED DESCRIPTION
[0056] The application will be further described in detail below with reference to the embodiments of the drawings.
[0057] The embodiment provides a hierarchical reinforcement learning method for crowd adaptive behavior simulation. Specifically, referring to FIG. 1, Figure 1 The hierarchical reinforcement learning method for crowd adaptive behavior simulation of the embodiment includes the following steps S1-S5:
[0058] Step S1, based on the current position of the agent in the crowd, the target point and the environment structure, a static planning optimal path of the agent in the static closed environment is constructed;
[0059] For example, in this embodiment, a path search algorithm based on heuristic function (i.e. A* algorithm) is used, and based on the current position of the agent in the crowd, the target point and the environment structure, a static planning optimal path of the agent is constructed;
[0060] In this embodiment, the "environment structure" here refers to the static layout and physical characteristics in the crowd environment where the agent is located, which is used to support the construction of the static planning optimal path; the environment structure usually includes the distribution of obstacles, the layout of passages, walls or boundaries and other fixed elements in the static closed environment, which determine the limitations and possibilities of the feasible path of the agent;
[0061] Step S2, based on the static planning optimal path of the agent which has been constructed, and based on the unit tangent segmentation of the current position point of the agent, a target trend speed of the agent moving from the current position point to the target point is generated;
[0062] Step S3, a nonlinear mapping of laser radar observation data to collision avoidance speed is established, and dynamic obstacle avoidance guidance is provided based on the environment around the agent; wherein, this step S3 can be realized by using mature existing technologies;
[0063] Step S4, the weights of the target trend speed and the weights of the collision avoidance speed are dynamically adjusted according to the personnel distribution information around the agent, to obtain the adjusted weight value of the target trend speed and the adjusted weight value of the collision avoidance speed; wherein, in this embodiment, the grid sensor can be used to perceive the information of other agents around the agent in real time;
[0064] For example, in this embodiment, the specific adjustment measures of dynamically adjusting the weights of the target trend speed and the weights of the collision avoidance speed according to the personnel distribution information around the agent are as follows:
[0065] 1. Evaluate the density of personnel: use laser radar data or Grid Sensor() to obtain the distribution state of agents around the agent Calculate the local personnel density;
[0066]
[0067] Wherein, ρ i is the local personnel density of the grid unit i, w j is the distance weight of the jth neighbor agent in the observation range, A i is the area of the grid unit i, and n is the total number of agents in the grid;
[0068] 2. Dynamic weight adjustment rule:
[0069] When the personnel density ρ i is lower than the preset density threshold ρ thresh , the target trend speed is given priority, the target trend speed weight w target = 0.8, and the collision avoidance speed weight w avoid = 0.2, so as to ensure that the agent moves efficiently towards the target direction;
[0070] When the personnel density ρ i is higher than the preset density threshold ρ thresh , the weight of the collision avoidance speed is increased, adjusted to w target = 0.3, and the collision avoidance speed weight w avoid = 0.7, so as to enhance the obstacle avoidance ability of the agent and reduce the risk of collision;
[0071] The weight adjustment adopts a smooth transition function, for example, w target = 1-δ(ρ i -ρ thresh ), w avoid = δ(ρ i -ρ thresh ); wherein δ is a sigmoid function, ensuring that the weight is within the range [0, 1] and the sum is 1;
[0072] 3. Real-time update mechanism: the hidden state m' i of the H-GRU model (see step a2) tracks the dynamic changes of personnel distribution, and the personnel density ρ i is recalculated every fixed time step (such as 0.1 seconds) and the target trend speed weight w target and the collision avoidance speed weight w avoid are updated to adapt to the changing environment;
[0073] Step S5, according to the current target trend speed of the agent, the adjusted target trend speed weight value, the target trend speed weight adjustment value, and the adjusted collision avoidance speed weight value, the speed is weighted and coupled to obtain the adaptive walking speed of the agent.
[0074] Specifically, in the above step S2 of this embodiment, the target trend speed generation method is as follows:
[0075] v goal = (v x ,v y );
[0076]
[0077] wherein v goal is the target trend speed of the agent, vx is the target tendency velocity of the agent v goal is the horizontal component of the velocity v x is the vertical component of the velocity v x is the modulus (i.e., scalar value) of the velocity v y is the target tendency velocity of the agent v goal is the vertical component of the velocity v y is the vertical component of the velocity v y is the modulus (i.e., scalar value) of the velocity v is the position point (p x , p y ) of the agent on the static planning optimal path M
[0078] As an improved way of the hierarchical reinforcement learning method in this embodiment, the hierarchical reinforcement learning method for crowd adaptive behavior simulation of this embodiment further comprises:
[0079] calculating the collision risk value of the agent at the current time based on the states of other agents in the surrounding environment of the agent;
[0080] adjusting the weights of the target tendency velocity and the collision avoidance velocity of the agent respectively according to the collision risk value of the agent, so as to realize the weighted coupling of the target tendency velocity and the collision avoidance velocity of the agent.
[0081] For example, in this embodiment, the collision risk value calculation method of the agent at the current time comprises the following steps a1-a3:
[0082] Step a1, obtaining the observable state information of other agents around the agent; wherein, the observable state information is represented as follows:
[0083]
[0084] wherein, represents the observable state of the i-th other agent around the agent, and Grid Sensor() represents a perception model for perceiving the state of the surrounding environment of the agent; Grid Sensor() is a perception model based on ray detection designed in Unity3D engine to support reinforcement learning model, which is used to represent the perception model for perceiving the state of the surrounding environment of the agent; n is the total number of other agents around the agent.
[0085] As is well known to those skilled in the art, Grid Sensor() is a model or algorithm based on a grid sensor used to sense the environmental state around an agent. It detects the agent's position, motion state, or other relevant information by using a gridded environment. Grid Sensor() is integrated into the Unity3D engine.
[0086] Step a2: Based on the acquired observable state information of the agent, perform position and trajectory encoding; where:
[0087]
[0088] in, This represents the value of the velocity vector of the j-th agent at time step t. This represents the encoded position of the i-th agent surrounding the current agent. For the hidden state of the gated loop unit H-GRU, || represents the splicing operation; W P For embedding functions The parameter W P It is a weight matrix or weight parameters; Let be the embedding vector of the i-th agent at time step t. Through function From input state Mapping yields W G is a function The weight parameters are the weight matrix in the neural network, used to weight the input state. Mapped to embedding vector For the weight parameter W G Initialization embedding function; W H Hidden state The weights;
[0089] The computation method (through the H-GRU model) implicitly contains the meaning of trajectory encoding; H-GRU (Gate-Controlled Recurrent Unit) is a recurrent neural network structure specifically designed for processing temporal data, and its hidden states... Capable of capturing the location of the agent Dynamic information that changes over time, i.e., motion trajectory characteristics;
[0090] Indicates hidden state It was in a hidden state at the previous moment. Based on this, combined with the current input (from encoding) The process is essentially encoding the temporal information of the agent's position into trajectory representation; therefore, the "trajectory encoding" is mainly reflected in the H-GRU model's encoding of the agent's position The modeling and updating vary with time step t;
[0091] Step a3, calculate the agent collision risk value of the resulting agent at the current time; wherein the agent collision risk value of the agent at the current time is marked as DANGER gs :
[0092]
[0093]
[0094]
[0095] wherein DANGER gs is the collision risk value of the gs-th agent, softmax(·) represents a normalized exponential function representing the relative danger degree between agents; N is the total number of agents in the scene, is a hyperparameter; [gs-1] represents the danger value of the extracted gs-th agent;
[0096] is the query vector of the gs-th agent at time step t, used in the attention mechanism to calculate the correlation between the gs-th agent and other agents;
[0097] is the key vector of the ts-th agent at time step t, used in the attention mechanism to evaluate the interaction with the gs-th agent; W QE is the weight matrix of the function , used to map the velocity to the query vector represents the velocity vector of the gs-th agent at time step t, which is the target trend velocity or final velocity;
[0098] is the embedding function initialized by W QE , W KE is the weight matrix of the function , used to map the velocity to the key vector is the embedding function initialized by W KE , represents the velocity vector of the ts-th agent at time step t, which is the target trend velocity or final velocity.
[0099] In addition, it needs to be explained that in the above step S5 of this embodiment, the adaptive walking speed of the agent is calculated as follows:
[0100] v output = TY1 x v goal + DANGER gs x TY2 x v collision ;
[0101] wherein v output is the adaptive walking speed of the agent on the final planned path, TY1 is the first hyperparameter, v goal is the target approaching speed of the agent, DANGER gs is the collision risk value of the agent, TY2 is the second hyperparameter, and v collision is the collision avoidance speed of the agent.
[0102] In this embodiment, the hierarchical reinforcement learning method for crowd adaptive behavior simulation further comprises: encouraging the agent to reach the designated target location as soon as possible on the premise that no mutual collision occurs based on a preset comprehensive reward function. Specifically, the preset comprehensive reward function is set as follows:
[0103]
[0104]
[0105]
[0106] wherein in this embodiment, r warning1 is the VO obstacle avoidance area calculated by the RVO obstacle avoidance algorithm; wherein:
[0107] When the agent agentA is less than 0.5m away from any other agent agentB, the reward signal given to punish the impending collision is set to have a value r warning1 of -0.5;
[0108] When the agent agentA is less than 1m away from any other agent agentB, the reward signal given to punish the impending collision is set to have a value r warning2 of -0.25.
[0109] According to the actual needs of hierarchical reinforcement learning, the hierarchical reinforcement learning method for crowd adaptive behavior simulation of this embodiment further comprises:
[0110] dividing the reinforcement learning of the agent into a low-level structure and a high-level structure;
[0111] letting steps S1-S4 be executed in the low-level structure, and letting step S5 be executed in the high-level structure;
[0112] And, the behavior of the agent in the crowd is optimized by means of a discount reward, wherein the discount reward is marked as G t :
[0113]
[0114] wherein G global is the discount reward value introduced in the high layer, G local is the discount reward value introduced in the low layer; represents the k-th power of the discount factor γ global of the global reward, that is, the discount weight of the global reward at time step t+k; R global represents the local collision avoidance reward item introduced in the high layer; represents the k-th power of the discount factor γ local of the global reward, that is, the discount weight of the global reward at time step t+k; R local represents the local collision avoidance reward item introduced in the low layer, and k represents the offset of the time step.
[0115] Although the preferred embodiments of the present application are described in detail above, it should be clearly understood that various modifications and changes can be made to the present application by those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A hierarchical reinforcement learning method for crowd adaptive behavior simulation, characterized in that, Comprising the following steps: Step S1, based on the current position of the agent in the crowd, the target point and the environment structure, a static planning optimal path of the agent in the static closed environment is constructed; Step S2, based on the static planning optimal path of the agent constructed, and based on the current position point of the agent, a unit tangent segmentation is made, and a target trend speed of the agent moving from the current position point to the target point is generated; Step S3, a nonlinear mapping of laser radar observation data to collision avoidance speed is established, and dynamic obstacle avoidance guidance is provided based on the environment around the agent; Step S4, the weight of the target trend speed and the weight of the collision avoidance speed are dynamically adjusted according to the personnel distribution information around the agent, to obtain the adjusted target trend speed weight value and the adjusted collision avoidance speed weight value; Step S5, the adaptive walking speed of the agent is obtained by coupling the current target trend speed, the adjusted target trend speed weight value, the target trend speed weight adjustment value and the adjusted collision avoidance speed weight value.
2. The hierarchical reinforcement learning method for crowd adaptive behavior simulation according to claim 1, wherein, In step S1, a path search algorithm based on heuristic function is used, and based on the current position of the agent in the crowd, the target point and the environment structure, the static planning optimal path of the agent is constructed.
3. The hierarchical reinforcement learning method for crowd adaptive behavior simulation according to claim 2, wherein, The target trend speed generation method is as follows: v goal = (v x , v y ); wherein v goal is the target trend velocity of the agent, v x is the target trend velocity v goal of the agent in the horizontal direction, ||v x || is the modulus of the component velocity v x , v y is the target trend velocity v goal of the agent in the vertical direction, ||v y || is the modulus of the component velocity v y , M is the static planning optimal path of the agent obtained by the path search algorithm based on the heuristic function, and M' (px,py) is the unit tangent vector generated by the agent at the position point (p x , p y ) on the static planning optimal path M.
4. The hierarchical reinforcement learning method for crowd adaptive behavior simulation according to claim 3, wherein, Further comprising: Based on the state of other agents in the environment around the agent, the agent collision risk value at the current time is obtained; According to the agent collision risk value of the agent obtained, the weight of the target trend speed and the weight of the collision avoidance speed of the agent are adjusted respectively, so as to realize the weighted coupling of the target trend speed and the collision avoidance speed of the agent.
5. The hierarchical reinforcement learning method for crowd adaptive behavior simulation according to claim 4, wherein, The agent collision risk value at the current time is calculated in the following steps: Step a1, the observable state information of other agents around the agent is obtained; wherein, the observable state information is represented as follows: wherein, Grid Sensor() represents the observable state of the i-th other agent around the agent, Grid Sensor() represents the perception model for perceiving the state of the environment around the agent; Grid Sensor() is a perception model based on ray detection designed in the Unity3D engine to support reinforcement learning models, which is used to represent the perception model for perceiving the state of the environment around the agent; n is the total number of other agents around the agent; Step a2, according to the obtained observable state information of the agent, the position and trajectory coding are performed; wherein: in, This represents the value of the velocity vector of the j-th agent at time step t. This represents the encoded position of the i-th agent surrounding the current agent. For the hidden state of the gated loop unit H-GRU, || represents the splicing operation; W P For embedding functions The parameter W P It is a weight matrix or weight parameters; Let be the embedding vector of the i-th agent at time step t. Through function From input state Mapping yields W G is a function The weight parameters are the weight matrix in the neural network, used to weight the input state. Mapped to embedding vector For the weight parameter W G Initialization embedding function; W H Hidden state The weights; Step a3, the agent collision risk value of the calculated agent at the current time is calculated; wherein the agent collision risk value of the agent at the current time is marked as DANGER gs : where DANGER gs is the collision risk value of the gs-th agent, softmax(·) represents a normalized exponential function representing the relative degree of danger between agents; N is the total number of agents in the scene, is a hyperparameter; [gs-1] represents the danger value of the gs-th agent extracted; is the query vector of the gs-th agent at time step t, used in the attention mechanism to compute the relevance to other agents; is the key vector of the ts-th agent at time step t, used in the attention mechanism to evaluate the interaction with the gs-th agent; W QE is the function is the weight matrix of the function mapping the velocity to the query vector is the target velocity or final velocity; is initialized by W QE is an embedding function initialized by W KE is a weight matrix of the function is used to map the velocity into a key vector is initialized by W KE is an embedding function initialized by W denotes the velocity vector of the ts-th agent at time step t, which velocity vector is the target velocity or final velocity.
6. The hierarchical reinforcement learning method for crowd adaptive behavior simulation according to claim 5, wherein, In step S5, the adaptive walking speed of the agent is calculated as follows: v output = TY1 x v goal + DANGER gs x TY2 x v collision ; wherein v output is the adaptive walking speed of the agent on the final planned path, TY1 is a first hyperparameter, v goal is the target tendency speed of the agent, DANGER gs is the collision risk value of the gs-th agent, TY2 is a second hyperparameter, v collision is the collision avoidance speed for the agent.
7. The hierarchical reinforcement learning method for crowd adaptive behavior simulation according to claim 6, wherein, Further comprising: Based on the preset comprehensive reward function, the agent is encouraged to reach the specified target location as soon as possible without mutual collision.
8. The hierarchical reinforcement learning method for crowd adaptive behavior simulation according to claim 7, wherein, The preset comprehensive reward function is set as follows: