Method for generating crowd evacuation and escape actions in three-dimensional scene
By constructing a three-dimensional real-time evacuation simulation framework and a personalized gait converter, the problems of realism and personalization in existing crowd evacuation simulations have been solved. This has enabled realistic, personalized, and real-time crowd evacuation simulations in three-dimensional scenes, improving obstacle avoidance capabilities and personalized speed decision-making.
Patent Information
- Application Number
- CN202511083189.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies cannot achieve realistic, online, personalized, and dynamic 3D perception of crowd evacuation and escape simulation in large-scale scenarios, and falls and collisions are prone to occur in densely populated areas.
A three-dimensional real-time evacuation simulation framework is constructed, including a perception module, a decision-making module, and an action module. The escape speed is calculated using a three-dimensional adaptive social force model, and personalized escape actions are generated through a personalized gait converter, combined with a physics engine for simulation.
It achieves realistic and personalized real-time crowd evacuation simulation in a 3D scene, solves the problem of the gap between 2D simulation and reality, and improves obstacle avoidance ability and personalized speed decision-making in a 3D environment.
Smart Images

Figure CN120976493A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of three-dimensional vision, and relates to a crowd evacuation escape action generation method in a three-dimensional scene. BACKGROUND
[0002] Evacuation simulation is a valuable tool for assessing the likelihood of crowd and stampede events, estimating evacuation time, and evaluating virtual reality escape training. However, existing methods cannot simulate the individualized and spatially-aware three-dimensional motion of hundreds of people online. The present application aims to develop a real-time online, individualized, dynamically-aware, multi-agent parallel three-dimensional crowd evacuation simulation framework. The present application framework can achieve online physically realistic simulation, avoid collisions within a suitable range, achieve individualized speed and gait, and be applicable to any scene and hundreds of people.
[0003] Traditional crowd simulation methods (Ding N, Zhu Y, Liu X, et al. A modified social force model for crowd evacuation considering collision predicting behaviors [J]. Applied Mathematics and Computation, 2024, 466: 128448., Integration of smoke effect and blind evacuation strategy (SEBES) within fire evacuation simulation) explore various strategies under different scenarios. However, as these methods represent people as two-dimensional points, they cannot combine three-dimensional motion with realistic behavior, resulting in physically unconvincing motion. For example, at exits or corridors with high-density crowds, traditional methods usually assume that the crowd will eventually evacuate smoothly after deceleration, whereas in real-world scenarios, individual falls and even stampedes can occur. Moreover, most crowd simulation methods are primarily designed to model regular crowd flow behavior rather than evacuation processes. Such methods also cannot adapt to scenarios with diverse terrains. Physics engine-based character control methods (Coros S, Beaudoin P, Van de Panne M. Generalized biped walking control [J]. ACM Transactions On Graphics (TOG), 2010, 29(4): 1-9., Peng X B, Guo Y, Halper L, et al. ASE: Large-scale reusable adversarial skill embeddings for physically simulated characters [J]. ACM Transactions On Graphics (TOG), 2022, 41(4): 1-17., Tessler C, Kasten Y, Guo Y, et al. CALM: Conditional adversarial latent models for directable virtual characters [C] / / ACM SIGGRAPH 2023 Conference Proceedings. 2023: 1-9.) enable autonomous 3D motion. However, these methods face issues such as a lack of personalized motion and can result in falls and collisions in densely populated scenarios.Some diffusion model-based motion generation methods (Guo C, Zou S, Zuo X, et al. Generating diverse and natural 3d human motions from text [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022:5152-5161., Shafir Y, Tevet G, Kapon R, et al. Human motion diffusion as a generative prior [J]. arXiv preprint arXiv:2303.01418, 2023.) can generate various 3D motions, but they still face challenges in controllability and physical realism. All the above methods cannot produce realistic, online, personalized evacuation motions in real time, while supporting multi-agent parallel simulation.
[0004] To solve the above problems, the present application proposes a method for generating evacuation motions of crowds in a three-dimensional scene. SUMMARY
[0005] The present application aims to provide a method for generating evacuation motions of crowds in a three-dimensional scene to solve the problem that existing methods cannot produce realistic, online, personalized, dynamic three-dimensional perception of crowd evacuation simulation results in large scenes.
[0006] To achieve the above purpose, the present application adopts the following technical solutions:
[0007] A method for generating evacuation motions of crowds in a three-dimensional scene, comprising the following steps:
[0008] S1, a three-dimensional real-time evacuation simulation framework simulating the brain perception-decision-action process of the crowd is constructed, and the framework is composed of a perception module, a decision module and an action module;
[0009] S2, in each time step, the perception module obtains the state of each agent, the state of other observable agents and the state of the environment;
[0010] S3, based on the state information obtained by the perception module in S2, the evacuation speed of each agent is calculated online using a three-dimensional adaptive social force model, wherein the social force model introduces individualized coefficients in combination with agent attributes;
[0011] S4, the escape speed beneficial to the calculation of S3 is to sample a local path for each agent, and a non-individualized escape action frame corresponding to the speed is obtained using an existing reinforcement learning-based pedestrian action control method with the local path as a condition;
[0012] S5, based on the non-individualized escape action frame obtained in S4, an individualized gait converter is used to generate an individualized action matched with the attributes of each agent, the action is simulated in a physics engine, and the part-level stress is visualized, and the simulation result is used as the input of the S2 perception module in the next time step.
[0013] Preferably, S1 abstracts each individual in the evacuation process as an online running agent, and the running at each time step goes through a perception module, a decision module, and an action module. The three modules are cyclically repeated in sequence, and the simulation result of the action module can be used for the perception of the perception module in the next time step.
[0014] Preferably, S2 specifically includes the following contents:
[0015] The perception module of each agent can obtain three parts of state information: its own state s self , part of the observable state s other of other agents, and the environment state s envir . Specifically, at time t, the own state s of agent i includes its position, velocity, and humanoid pose parameter h i,t ; the state s of other agents contains the root position information of surrounding agents relative to it; and the environment state s provides the position of the target path point relative to its root node.
[0016] Preferably, S3 specifically includes the following contents:
[0017] S301, first, the key path points of the escape path are determined by the A* search algorithm, then the social force is used to drive each agent to move towards the target path point and avoid collision, and finally the speed is calculated.
[0018] S302, the social force model used contains driving force and repulsive force. The driving force directs the agent to the desired target, and according to the difference between the desired speed and the current position, a weighted interpolation is performed combined with a relaxation parameter. The repulsive force maintains a safe inter-agent distance, and the repulsive force is proportional to the proximity of neighboring agents. Parameters such as repulsion coefficient, spatial decay coefficient, and interaction radius are introduced in the calculation.
[0019] S303、Unlike in two dimensions, in three dimensions, modeling the avoidance between agents only by repulsive forces can cause congestion or tripping. Therefore, on the basis of the original social force model, an evasive force is added, and by calculating the vertical direction towards the target path and considering the neighboring agents around the agent, intelligent detours can be achieved. The evasive force of the ith agent is defined as follows:
[0020] F evasive =A·sgn(o i ·p i )·p i
[0021] Where sgn is the sign function, p i represents the vertical vector of the desired path, o i represents the average observation direction of the neighboring agents, and A is the proportional coefficient.
[0022] S304、To reflect the differences in individual escape ability, individualized social force coefficients are assigned to different individuals. The specific method is to set v real as the actual escape speed obtained from the literature, and set v sim as the speed simulated by the physical engine, and optimize the input speed v setting so that the simulation speed v sim approximates the real value v real . The agents in the present application are divided into five categories: teenagers, middle-aged people, old people, patients and disabled people. Through this process, the present application can realize personalized coefficient fitting, and the error is controlled within 0.005 m / s.
[0023] S305、The final target speed is calculated by the following formula:
[0024]
[0025] Where Δt represents the time step, v i is the current speed, F drive , F repulsive and F evasive correspond to the driving force, the repulsive force and the evasive force respectively.
[0026] Preferably, the S4 is specifically a system that uses a robust controller PACER based on physical simulation as a basic component to follow the path, and the control strategy π PACER generates a two-dimensional trajectory τ i,t by sampling the desired speed and follows the trajectory. Under the given state S (including the agent position, the humanoid state and the environment state), the action a PACER is calculated by π i,t .
[0027] Preferably, S5 is specifically as follows:
[0028] S501, based on the classical gait cycle theory, the distance between the two ankle joints is analyzed, and the characteristic waveform of the class sine function can be obtained, and the peak, zero point and zero point of the waveform correspond to the four core phases in the gait cycle. Assign 0, 0.3, 0.5, 0.75 reference values to the gait frames in the core phase, and generate the remaining frames by linear interpolation. In the 100 STYLE data set, the action of the Neutral style is taken as the non-personalized gait frame, and the action of the non-Neutral style is taken as the personalized gait frame, and the non-personalized gait frame and the personalized gait frame with the same phase value are matched. When there are multiple candidate frames, the nearest neighbor matching strategy of joint angle is adopted. In the training stage, a data enhancement mechanism is introduced, and random rotation disturbance is applied to the root joints of the paired frames synchronously.
[0029] S502, the application designs a personalized gait converter based on the CAMDM network, which is based on the paired non-personalized / individualized action frames in S501, so as to realize the conversion of the non-personalized action frame a i,t to the individualized action frame . The converter is a probability diffusion model, which inputs a noise sample diffusion step t and individualized gait label c in each denoising step, and then learns to predict the original clean
[0030] S503, in the simulation of the physical engine, the upper body action of the non-personalized action frame a i,t is replaced by the upper body action of the individualized action frame . The simulated action in the physical engine can be perceived by the perception module for decision-making and action in the next time step.
[0031] S504, the method integrates force sensors in each body part of the role, and adopts gradual color coding: the lighter the color, the smaller the force, and the darker the color, the greater the force.
[0032] Compared with the prior art, the application provides a crowd evacuation escape action generation method in a three-dimensional scene, which has the following beneficial effects:
[0033] (1) The application provides a crowd evacuation escape action generation method in a three-dimensional scene, which simulates the framework of crowd evacuation escape in an emergency scene, and can realize realistic, personalized and real-time crowd evacuation escape simulation through the initial position of the scene and the crowd;
[0034] (2) The application solves the problem of idealization and actual difference in two-dimensional simulation by simulating the human perception-decision-action process.
[0035] (3) The application designs a three-dimensional adaptive social force simulation to make up for the three-dimensional congestion problem and realize personalized speed decision-making; meanwhile, through the design of a personalized gait controller, personalized escape gait control is realized. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 An escape simulation framework diagram for the embodiment 1 of the application is provided.
[0037] Figure 2 An escape process result schematic diagram for the embodiment 2 of the application is provided.
[0038] Figure 3 An escape process detail schematic diagram for the embodiment 2 of the application is provided. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the application will be apparently and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the application.
[0040] The application provides a crowd evacuation action generation method in a three-dimensional scene, which gives a three-dimensional scene model, an exit position and an initial position of the crowd. The application can realize real-time online, personalized, dynamic perception, multi-agent parallel three-dimensional crowd evacuation simulation. It is worth noting that it is suitable for evacuation tasks by combining the human-like paradigm, while ensuring personalized decision-making and action. The application provides a perception-decision-action joint framework to solve the problem of poor controllability and physical rationality. The perception module provides self-awareness, perception of others and perception of the exit. The decision module uses a three-dimensional adaptive social force model decision mechanism, which enhances the decision-making ability of individuals and the obstacle avoidance ability in a three-dimensional environment. Personalization of spatial perception is realized through the decision and action modules: the decision module has personalized optimization of social force coefficients, and calculates the speed customized for individual attributes, while the action module has a personalized gait controller to generate movements with personalized gait, solving the problem of lack of personalized movement. The application also introduces part-level force visualization to assist analysis. The crowd evacuation action generation method in a three-dimensional scene provided by the application will be described below in combination with the drawings and examples, which specifically includes the following contents.
[0041] Embodiment 1
[0042] Please refer to Figure 1 The application provides a crowd evacuation action generation method in a three-dimensional scene, which includes the following steps:
[0043] S1, constructing a three-dimensional real-time evacuation simulation framework simulating the perception-decision-action process of human brain, the framework is composed of a perception module (S2), a decision module (S3) and an action module (S4, S5), S1 specifically abstracts each individual in the evacuation process as an online running agent, and the running at each time step will go through the perception module, the decision module and the action module. The three modules circulate in order and the simulation results of the action module can be used for the perception of the next time step.
[0044] S2, at each time step, the perception module of each agent can obtain three parts of state information: its own state s self , part of the observable state s other of other agents and the environment state s envir . Specifically, at time t, the state of the agent i itself s includes its position, velocity and humanoid posture parameters h i,t ; the state of other agents s contains the relative position information of surrounding agents to its root position; and the environment state s provides the position of the target path point relative to its root node.
[0045] S3, based on the state information obtained by the S2 perception module, the escape speed of each agent is calculated online using a three-dimensional adaptive social force model, wherein the social force model introduces individualized coefficients in combination with agent attributes, specifically as follows:
[0046] S301, first, the A* search algorithm (Hart P E, Nilsson N J, Raphael B. A formal basis for the heuristic determination of minimum cost paths [J]. IEEE transactions on Systems Science and Cybernetics, 1968, 4(2): 100-107.) is used to determine the key path points of the escape path, then the social force is used to drive each agent to move towards the target path point and avoid collision, and finally the speed is calculated.
[0047] S302, the social force model used contains driving force and repulsive force. The driving force directs the agent to the desired target, and according to the difference between the desired speed and the current position, a weighted interpolation is performed in combination with a relaxation parameter. The repulsive force maintains a safe inter-agent distance, and the repulsive force is proportional to the proximity of neighboring agents, and parameters such as repulsion coefficient, spatial decay coefficient and interaction radius are introduced in the calculation.
[0048] S303、Unlike in two-dimensional point simulation, in three-dimensional case, only repulsive force to model the avoidance between agents may cause congestion or stumble. Therefore, on the basis of the original social force model, the avoidance force is added, by calculating the vertical direction towards the target path, and considering the surrounding agents of the agent, the intelligent detour can be realized. The avoidance force of the ith agent is defined as follows:
[0049] F evasive i i i
[0050] Where sgn is the sign function, p i represents the vertical vector of the desired path, o i represents the average observation direction of the surrounding agents, and A is the proportional coefficient.
[0051] S304、To reflect the differences in individual escape ability, individualized social force coefficients are assigned to different individuals. The specific method is to set v real as the actual escape speed obtained from the literature, and set v sim as the speed simulated by the physical engine, and optimize the input speed v setting so that the simulation speed v sim approximates the real value v real . The agents are divided into five categories: teenagers, middle-aged people, old people, patients and disabled people. Through this process, the invention can realize the fitting of individualized coefficients, and the error is controlled within 0.005 m / s.
[0052] S305、The final target speed is calculated by the following formula:
[0053]
[0054] Where Δt represents the time step, v i is the current speed, F drive , F repulsive and F evasive correspond to the driving force, repulsive force and avoidance force respectively.
[0055] S4, the escape speed calculated in S3 is to sample a local path for each agent, and use the existing reinforcement learning-based pedestrian action control method to obtain non-personalized escape action frames corresponding to the speed based on the local path. Specifically, the system uses the robust controller PACER (Rempe D, Luo Z, Bin Peng X, et al. Trace and pace: Controllable pedestrian animation via guided trajectory diffusion [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023: 13756-13766.) based on physical simulation as a basic component to follow the path, and its control strategy π PACER The two-dimensional trajectory τ generated by sampling the desired speed i,t And follow the trajectory. Under the condition of a given state S (including the agent's position, the humanoid state, and the environment state), the action a PACER is calculated by π i,t .
[0056] S5, based on the non-personalized escape action frames obtained in S4, a personalized gait converter is used to generate personalized actions that match the attributes of each agent. The action simulation is performed in the physics engine, and the part-level stress is visualized. The simulation results are used as inputs for the S2 perception module in the next time step. Specifically:
[0057] S501, based on the classical gait cycle theory, the distance between the two ankle joints is analyzed to obtain a characteristic waveform similar to a cosine function. The wave crest, zero point, wave trough, and zero point position of the waveform correspond to the four core phases in the gait cycle. Assign 0, 0.3, 0.5, and 0.75 as reference values to the gait frames in the core phases, and generate the remaining frames through linear interpolation. In the 100STYLE dataset, the Neutral style action is used as the non-personalized gait frame, and the non-Neutral style action is used as the personalized gait frame. Match the non-personalized gait frames and personalized gait frames with the same phase value. When there are multiple candidate frames, use the nearest neighbor matching strategy for joint angles. In the training stage, introduce a data augmentation mechanism to apply random rotation disturbances to the root joints of the paired frames.
[0058] S502、The application designs a personalized gait converter on the basis of CAMDM network (Chen R, Shi M, Huang S, et al. Taming diffusion probabilistic models for character control [C] / / ACM SIGGRAPH 2024 Conference Papers. 2024: 1-10.), which is based on the paired non-personalized / personalized action frames in S501, so as to realize the conversion of non-personalized action frames a i,t to personalized action frames . The converter is a probabilistic diffusion model, which inputs a noise sample diffusion step t and personalized gait label c at each denoising step, and then learns to predict the original clean
[0059] S503、In the simulation of the physical engine, the application replaces the upper body action of the non-personalized action frame a i,t with the upper body action of the personalized action frame . The simulated action in the physical engine can be perceived by the perception module for decision-making and action in the next time step.
[0060] S504、The method integrates force sensors in each body part of the character, and adopts gradual color coding: the lighter the color, the smaller the force, and the darker the color, the greater the force.
[0061] Embodiment 2:
[0062] Please refer to Figures 1-3 , which is based on embodiment 1 but has some differences. The application combines specific examples to explain a method for generating crowd evacuation actions in a three-dimensional scene:
[0063] (I) Scene and crowd initialization:
[0064] Input the three-dimensional grid of the scene, the number of escapees, and the location of the scene exit. First, generate a scene map using the navmesh algorithm, which marks the walkable areas in the scene, and randomly initialize the position of the people on the scene map. Use the A* algorithm to find the shortest path point to the scene exit on the scene map.
[0065] (II) Perception and decision-making process:
[0066] In each time step, the perception module of each agent can obtain three parts of state information: its own state s self , part of the observable state s other of other agents, and the environment state senvir Set v real Let v be the actual escape speed obtained from the literature statistics. sim Optimize the input velocity v to match the velocity simulated by the physics engine. setting This makes the simulated velocity v sim Approximating the true value v real v sim Personalized social force coefficients are assigned to different individuals. Using a three-dimensional adaptive social force model, including driving force, repulsive force, and the bypass force designed in this invention, the escape speed of each agent is calculated.
[0067] (III) Personalized escape gait control:
[0068] Based on the escape speed of the intelligent agent calculated by the decision-making module, a robust controller PACER (Rempe D, Luo Z, Bin Peng X, et al. Trace and pace: Controllable pedestrian animation via guided trajectory diffusion[C] / / Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition.2023:13756-13766.) based on physical simulation is used as the basic component to follow the path. Its control strategy π PACER The two-dimensional trajectory τ generated by sampling the desired velocity i,t It also generates escape actions that follow the trajectory.
[0069] This invention designs a personalized gait converter based on the CAMDM network (Chen R, Shi M, Huang S, et al. Taming diffusion probabilistic models for character control [C] / / ACM SIGGRAPH 2024 Conference Papers. 2024: 1-10.), and enables the implementation of non-personalized / personalized action frames based on the paired non-personalized / personalized action frames in S501. i,t To personalized motion frame a~ i,t The conversion. In the simulation of the physics engine, this invention will convert non-personalized action frames a i,t Replace the upper body motion with personalized motion frames. The upper body movements. Simulated movements in the physics engine can be perceived by the perception module in order to make decisions and take actions in the next time step.
[0070] In addition, the method integrates force sensors in each body part of the role, and adopts gradual color coding: the lighter the color, the smaller the force, and the darker the color, the greater the force.
[0071] (Four) Comparison of escape simulation:
[0072] The present application simulates in Isaac Gym, simulates in small and large scenes of escape.
[0073] The method of embodiment 1 is compared with the current most advanced motion generation method OmniControl (Xie Y, Jampani V, Zhong L, et al. OmniControl: Control any joint at any time for human motion generation [J]. arXiv preprint arXiv: 2310.08580, 2023.) and the role control method MaskedMimic (Tessler C, Guo Y, Nabati O, et al. MaskedMimic: Unified physics-based character control through masked motion inpainting [J]. ACM Transactions on Graphics (TOG), 2024, 43(6): 1-21.) in the qualitative results of the scene. The specific comparison results can be referred to Figure 2 and Figure 3 .
[0074] Referring to Figure 2 , OmniControl path confusion and mid-stop problems. MaskedMimic lacks obstacle avoidance mechanism, is easy to cause collision in straight line motion and cause group accumulation, and speed individualization is almost invisible. In contrast, the present method can simulate a more reasonable evacuation process. Referring to Figure 3 , OmniControl will cause motion distortion due to the inability to generate long-distance trajectories with strong constraints. The behavior of all agents in MaskedMimic is almost the same, and individual differences cannot be reflected. In contrast, the present method can generate actions consistent with individual attributes.
[0075] The method of embodiment 1 is compared with the current most advanced motion generation method OmniControl and the role control method MaskedMimic in the quantitative results of the scene. The specific comparison results can be referred to Table 1:
[0076] Table 1 Comparison of quantitative results of the method of the present application with OmniControl and MaskedMimic
[0077] Method Average escape success rate Average number of falls OmniControl 0.48 — MaskedMimic 0.60 18.55 The method of this example 0.84 12.26
[0078] From Table 1, it can be seen that there are two evaluation indexes of quantitative results, which are average escape success rate and average fall number, respectively. The first index value is higher the better, and the second index value is lower the better. Among them, the average escape success rate is the average value of the proportion of the number of successful escape in each run to the total number of people, which measures the completion of the escape task. The average fall number is the average value of the total number of falls of all agents in the escape process, which measures the real rationality of escape. Since OmniControl has no physical reality, its results will not produce falls, so the average fall number will not be calculated. The method of the present application is better than OmniControl and MaskedMimic in all scenarios, with higher success rate and fewer falls. OmniControl has insufficient controllability for long-distance movement, which leads to the deviation of the agent from the trajectory and premature stop. Compared with MaskedMinic, the agent in the method of the present application shows better pathfinding, mutual avoidance and balancing ability. Therefore, the simulation of the present application achieves the best results.
[0079] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any skilled person in the art can make equivalent replacement or change according to the technical solution and improvement concept of the present application within the technical range disclosed by the present application, which should be covered in the protection scope of the present application.
Claims
1. A method for generating crowd evacuation escape action in a three-dimensional scene, characterized in that, The method comprises the following steps: S1, constructing a three-dimensional real-time evacuation simulation framework simulating the perception-decision-action process of human brain, which comprises a perception module, a decision module and an action module; S2, in each time step, using the perception module to obtain the state of each agent, the state of other observable agents and the state of the environment; S3, based on the state information obtained by the perception module in S2, using a three-dimensional adaptive social force model to calculate the escape speed of each agent online, wherein the social force model introduces a personalized coefficient in combination with the attributes of the agent; S4, using the escape speed calculated in S3 to sample a local path for each agent, using a pedestrian action control method based on reinforcement learning to obtain a non-personalized escape action frame corresponding to the escape speed based on the local path; S5, based on the non-personalized escape action frame obtained in S4, using a personalized gait converter to generate personalized actions matching the attributes of each agent, executing action simulation in a physics engine, visualizing the force at the part level, and using the simulation results as input for the perception module in the next time step.
2. The method of claim 1, wherein, The three-dimensional real-time evacuation simulation framework in S1 abstracts each individual in the evacuation process as an online running agent, and the running in each time step sequentially passes through the perception module, the decision module and the action module; the three modules of the perception module, the decision module and the action module are cyclically repeated in sequence and the simulation results of the action module are used for the perception of the perception module in the next time step.
3. The method of claim 2, wherein, The agents are divided into: teenagers, middle-aged people, old people, patients and disabled people.
4. The method of claim 3, wherein, The agent's own state in S2 is denoted s self , the observable other agent states are denoted s other , and the environment state is denoted s envir ; at time t, the agent's own state includes its position, velocity, and humanoid pose parameters h i,t ; the other agent states contain the positions of the surrounding agents relative to its root node; and the environment state provides the position of the goal waypoint relative to its root node.
5. The method of claim 4, wherein, S3 specifically includes the following contents: S301, determining the key path points of the escape path through an A* search algorithm, using a three-dimensional adaptive social force model to drive each agent to move towards the target path point and avoid collision, and finally calculating its speed; S302, the social force model used includes a driving force and a repulsive force, wherein the driving force is used to guide the agent to the desired target, and the difference between the desired speed and the current position is weighted and interpolated in combination with a relaxation parameter; the repulsive force is used to maintain a safe distance between agents, and the size of the repulsive force is proportional to the proximity of the adjacent agent, and a repulsion coefficient, a spatial decay coefficient and an interaction radius are introduced in the calculation; S303, adding an avoidance force based on the original social force model, calculating the perpendicular direction towards the target path, and considering the adjacent agents around the agent to achieve intelligent detour; the avoidance force of the i-th agent is defined as follows: F evasive = A - sgn(o i · p i ) - p i where sgn is the sign function; p i represents the perpendicular vector of the desired path; o i represents the average observation direction of neighboring agents; A is a proportional coefficient; S304、for different individuals to allocate personalized social force coefficient to reflect the difference in individual escape ability, the specific method is: set v real for the actual escape speed obtained by statistics, set v sim for the speed simulated by the physical engine, optimize the input speed v setting , so that the simulation speed v sim approximates the real value v real , and then realize the fitting of personalized coefficient; S305, final target speed is calculated by the following equation: where Δt represents a time step; v i represents a current velocity; F drive , F repulsive , and F evasive respectively correspond to a driving force, a repulsive force, and an evasive force.
6. The method of claim 5, wherein, The method for controlling pedestrian action based on reinforcement learning in S4 is used to obtain non-personalized escape action frames corresponding to the escape speed under the condition of a local path. Specifically, a robust controller PACER based on physical simulation is used as a basic component to follow the path, and its control strategy π PACER A two-dimensional trajectory τ generated by sampling the expected speed i,t And follow the trajectory; under the condition of a given state S, the control strategy π PACER Calculate the action a i,t ; wherein the state S includes the position of the agent, the state of the humanoid, and the state of the environment.
7. The method of claim 6, wherein, S5 specifically includes the following contents: S501, based on the classical gait cycle theory, analyzing the distance between the two ankle joints to obtain a characteristic waveform similar to a sine function, and the wave peak, zero point, wave trough and zero point position of the waveform correspond to the four core phases in the gait cycle; assigning reference values of 0, 0.3, 0.5 and 0.75 to the gait frames in the core phases, and generating the remaining frames by linear interpolation; In the 100 STYLE dataset, the actions of Neutral style are taken as non-personalized gait frames, and the actions of non-Neutral style are taken as personalized gait frames, and the non-personalized gait frames with the same phase value are matched with the personalized gait frames; When there are multiple candidate frames, a joint angle nearest neighbor matching strategy is adopted; In the training stage, a data enhancement mechanism is introduced, and random rotation disturbance is applied to the root joints of the paired frames synchronously. S502, a personalized gait converter is designed on the basis of the CAMDM network, based on the paired non-personalized / personalized action frames in S501, so that the non-personalized action frames a i,t are converted into personalized action frames ; the personalized gait converter is a probability diffusion model, in each denoising step, the model inputs a noise sample diffusion step t, and a personalized gait label c, and then learns to predict the original clean S503, in the simulation of the physical engine, replace the upper body action of the non-personalized action frame a i,t with the upper body action of the personalized action frame ; the simulated action in the physical engine is perceived by the perception module to make decisions and actions in the next time step; S504, integrate force sensors in each body part of the role, and adopt a gradual color coding: the lighter the color, the smaller the force, and the darker the color, the greater the force.