Evacuation path planning method, computer equipment and storage medium
By using deep reinforcement learning models in rail transit stations, combining video streaming data and floor drawings, dynamically planning evacuation paths is solved, and the problem of poor rationality and effectiveness of evacuation path planning in the existing technology is solved, achieving safer and more efficient evacuation.
Patent Information
- Application Number
- CN202510233157.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, there are problems with poor rationality and effectiveness in evacuation path planning of rail transit stations, especially in the dynamic environment, where dynamic adjustments cannot be made.
By obtaining video stream data and floor drawings, pedestrian density and distribution information are determined, and a deep reinforcement learning model is constructed. Based on this model, the personnel path mapping information is determined and the evacuation path is dynamically planned.
It improves the accuracy and effectiveness of evacuation paths, and can dynamically adjust evacuation paths according to real-time pedestrian distribution and dangerous situations, improving the safety and efficiency of evacuation of personnel.
Smart Images

Figure CN120146344A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of crowd evacuation path planning. Specifically, it relates to an evacuation path planning method, a computer device, and a storage medium. Background Art
[0002] Rail transit is an important part of urban public transportation. As a place with a high density of people, once an emergency such as a fire occurs in a rail transit station, the evacuation efficiency of people is directly related to the lives of passengers and the normal operation of the rail transit system.
[0003] Currently, mainly relying on fire broadcasts and fixed indicator lights to guide passengers to evacuate, the evacuation guidance is fixed and single. For example, the buried lights only simply indicate the direction of the exit, without considering factors such as the real-time pedestrian distribution, obstacle positions, and environmental changes, which easily leads to congestion at the exit and even causes secondary accidents. Although existing detection means such as electronic turnstiles, ticketing systems, and infrared counters can roughly count the number of people, they are not accurate enough in detecting the crowd density in local areas. And traditional path planning algorithms can only find the optimal path in a static environment, but perform poorly in dynamic scenarios (such as the constantly changing environment at the fire scene), and cannot be dynamically adjusted according to the dangerous situation and the crowd evacuation situation, so the rationality and effectiveness of the evacuation path are relatively poor.
[0004] Therefore, there are certain limitations in the evacuation path planning of rail transit stations in the prior art. Summary of the Invention
[0005] The purpose of the present application is to provide an evacuation path planning method, a computer device, and a storage medium for the deficiencies in the above-mentioned prior art, so as to solve the problem of the poor rationality and effectiveness of the evacuation path in the prior art.
[0006] To achieve the above purpose, the technical solutions adopted in the embodiments of the present application are as follows:
[0007] In a first aspect, an embodiment of the present application provides an evacuation path planning method, and the method includes:
[0008] Obtain the video stream data sent by each video source and the floor plan of the target venue;
[0009] Determine pedestrian information according to the video stream data, where the pedestrian information includes pedestrian density information and pedestrian distribution information;
[0010] Construct a deep reinforcement learning model according to the floor plan of the target venue and the multiple historical pedestrian distribution information of the target venue, and determine at least one personnel path mapping information of the target venue through the deep reinforcement learning model. Each personnel path mapping information is used to indicate the mapping relationship between a pedestrian density information and a candidate evacuation path;
[0011] Determine a target evacuation path according to the pedestrian information, the deep reinforcement learning model and the at least one personnel path mapping information, so as to evacuate pedestrians according to the target evacuation path.
[0012] As an alternative implementation, the determining pedestrian information according to the video stream data includes:
[0013] Extract each video frame according to the video stream data;
[0014] Determine the pedestrian information according to each video frame by a pre-constructed pedestrian recognition model.
[0015] As an alternative implementation, the pedestrian recognition model is a multi-scale convolutional neural network model;
[0016] The determining the pedestrian information according to each video frame by a pre-constructed pedestrian recognition model includes:
[0017] Input each video frame into the multi-scale convolutional neural network model, and the multi-scale convolutional neural network model detects the number of pedestrians in the corresponding area of each video frame through a target detection algorithm;
[0018] Determine the pedestrian density information according to the number of pedestrians in the corresponding area of each video frame;
[0019] Determine the planar coordinates of each pedestrian in each video frame according to the pixel coordinates of each pedestrian in each video frame and a preset perspective transformation matrix, and use the planar coordinates of each pedestrian in each video frame as the pedestrian distribution information.
[0020] As an alternative implementation, the constructing a deep reinforcement learning model according to the floor plan of the target venue and the multiple historical pedestrian distribution information of the target venue, and determining at least one personnel path mapping information of the target venue through the deep reinforcement learning model includes:
[0021] Build a simulation environment of the target venue according to the floor plan of the target venue;
[0022] Determine the motion parameters of pedestrians in the target venue according to a preset pedestrian motion model and the simulation environment, where the motion parameters include a motion direction and a motion speed;
[0023] Construct the deep reinforcement learning model according to the motion parameters, the historical pedestrian distribution information, a pre-constructed reward function, and an initial deep reinforcement learning model, and determine at least one personnel path mapping information of the target venue through the deep reinforcement learning model, where the reward function is used to characterize the relationship among the distance between a pedestrian and an exit, the staying time of the pedestrian in the target venue, and the congestion degree of the target venue, and the maximum value of the reward function is related to the area of the target venue.
[0024] As an optional implementation manner, the pedestrian motion model is a mechanical model, including: a pedestrian target driving force model, a repulsive force model between pedestrians, and a repulsive force model between a wall and a pedestrian.
[0025] As an optional implementation manner, the constructing the deep reinforcement learning model according to the motion parameters, the historical pedestrian distribution information, a pre-constructed reward function, and an initial deep reinforcement learning model, and determining at least one personnel path mapping information of the target venue through the deep reinforcement learning model includes:
[0026] Input the motion parameters and the historical pedestrian distribution information into the initial deep reinforcement learning model, and the initial deep reinforcement learning model determines the reward value of the reward function;
[0027] The initial deep reinforcement learning model iteratively corrects the model parameters according to the reward value of the reward function, and obtains the deep reinforcement learning model when the iteration ends;
[0028] Determine the at least one personnel path mapping information according to the deep reinforcement learning model.
[0029] As an optional implementation manner, the determining the at least one personnel path mapping information according to the deep reinforcement learning model includes:
[0030] Obtain a plurality of reference pedestrian distribution information;
[0031] Input the plurality of reference pedestrian distribution information into the deep reinforcement learning model, and the deep reinforcement learning model predicts and generates candidate evacuation paths corresponding to each reference pedestrian density information, where the reference pedestrian density information is the pedestrian density information corresponding to the reference pedestrian distribution information;
[0032] Establish a mapping relationship between each reference pedestrian density information and the corresponding candidate evacuation path respectively, and obtain the at least one personnel path mapping information.
[0033] As an alternative implementation manner, determining a target evacuation path according to the pedestrian information, the deep reinforcement learning model, and the at least one personnel path mapping information includes:
[0034] Determine whether an emergency event has occurred;
[0035] If so, the deep reinforcement learning model predicts and generates the target evacuation path according to the pedestrian distribution information;
[0036] Otherwise, determine an alternative evacuation path that matches the pedestrian density information in the pedestrian information according to the at least one personnel path mapping information, and use the alternative evacuation path that matches the pedestrian density information in the pedestrian information as the target evacuation path.
[0037] In a second aspect, an embodiment of the present application provides an evacuation path planning device, and the device includes:
[0038] An acquisition module, configured to acquire video stream data sent by each video source and a floor plan of a target venue;
[0039] A determination module, configured to determine pedestrian information according to the video stream data, where the pedestrian information includes pedestrian density information and pedestrian distribution information;
[0040] A construction module, configured to construct a deep reinforcement learning model according to the floor plan of the target venue and multiple historical pedestrian distribution information of the target venue, and determine at least one personnel path mapping information of the target venue through the deep reinforcement learning model, and each personnel path mapping information is respectively used to indicate a mapping relationship between a pedestrian density information and an alternative evacuation path;
[0041] The determination module is further configured to determine a target evacuation path according to the pedestrian information, the deep reinforcement learning model, and the at least one personnel path mapping information, so as to evacuate pedestrians according to the target evacuation path.
[0042] In a third aspect, an embodiment of the present application provides a computer device, including: a processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus, and the processor executes the machine-readable instructions to perform the steps of the evacuation path planning method described in the first aspect above.
[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it performs the steps of the evacuation path planning method described in the first aspect above.
[0044] The beneficial effects of the present application are as follows:
[0045] The present application provides an evacuation path planning method, a computer device, and a storage medium. Video stream data sent by each video source and a floor plan of a target venue are obtained, and pedestrian detection is performed based on the video stream data to determine pedestrian information, including pedestrian density information and pedestrian distribution information. A deep reinforcement learning model is constructed based on the floor plan and historical pedestrian distribution information. The deep reinforcement learning model is an evacuation path model based on the deep reinforcement learning algorithm. At least one personnel path mapping information of the target venue is determined through the deep reinforcement learning model to indicate a mapping relationship between a pedestrian density information and an alternative evacuation path, and the at least one personnel path mapping information provides a decision basis for the safe evacuation of the target venue. According to the pedestrian density information and pedestrian distribution information in the pedestrian information, the trained deep reinforcement learning model, and the at least one personnel path mapping information, a target evacuation path is determined to guide pedestrians to evacuate along the target evacuation path. The pedestrian detection based on the video stream data improves the accuracy of pedestrian information. By constructing an evacuation path model based on the deep reinforcement learning algorithm, the safety and effectiveness of evacuation path planning are improved. The at least one personnel path mapping information of the target venue determined through the deep reinforcement learning model facilitates quickly finding the corresponding alternative evacuation path from the personnel path mapping information according to the pedestrian density information, thereby improving the evacuation efficiency. Brief Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a flowchart of the evacuation path planning method provided by the embodiment of the present application;
[0048] Figure 2 It is a flowchart of determining pedestrian information of the evacuation path planning method provided by the embodiment of the present application;
[0049] Figure 3 It is another flowchart of determining pedestrian information of the evacuation path planning method provided by the embodiment of the present application;
[0050] Figure 4 It is a flowchart of constructing a deep reinforcement learning model of the evacuation path planning method provided by the embodiment of the present application;
[0051] Figure 5Another process schematic diagram for constructing a deep reinforcement learning model of the evacuation path planning method provided by the embodiments of the present application;
[0052] Figure 6 A process schematic diagram for determining the personnel path mapping information of the evacuation path planning method provided by the embodiments of the present application;
[0053] Figure 7 A process schematic diagram for determining the target evacuation path of the evacuation path planning method provided by the embodiments of the present application;
[0054] Figure 8 A module structure diagram of the evacuation path planning device provided by the embodiments of the present application;
[0055] Figure 9 A schematic structural diagram of a computer device provided by the embodiments of the present application. Detailed implementation manners
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn in actual proportions. The flowcharts used in the present application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and the steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0057] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts belong to the protection scope of the present application.
[0058] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated hereinafter, but does not exclude adding other features.
[0059] In places with a high density of people such as rail transit stations, the evacuation efficiency of people in case of emergencies such as fires is crucial. At present, passengers are mainly guided to evacuate by fire broadcasts and fixed indicator lights, and the evacuation guidance is fixed and single. Moreover, the existing means of detecting the number of people have low accuracy in detecting the flow density in local areas, and traditional path planning algorithms cannot be dynamically adjusted according to dangerous situations and the evacuation situation of the crowd. The rationality and effectiveness of the evacuation path are poor, and there are certain limitations.
[0060] Based on the above problems, the embodiments of the present application propose an evacuation path planning method. Through video recognition technology, accurate and real-time detection of the flow density information of people in each area of the target place is realized, and the number and location distribution of people to be evacuated in the target place are clarified. By establishing a deep reinforcement learning model, rapid decision-making and dynamic planning of the optimal evacuation path are realized to ensure the rapid and safe evacuation of pedestrians.
[0061] Figure 1 It is a schematic flowchart of the evacuation path planning method provided by the embodiments of the present application. The execution subject of this method can be any computer device with computing and processing capabilities. As Figure 1 shown, this method includes:
[0062] S101. Obtain the video stream data sent by each video source and the floor plan of the target place.
[0063] Optionally, install multiple cameras in different areas of the target place to form a monitoring system of the target place to achieve full-round monitoring of the target place. Among them, these cameras are the sources of the video stream data, that is, each camera in the monitoring system is used as a video source. The continuous pictures captured by the cameras are processed through encoding and other processes to form video stream data, which contains the dynamic information occurring in real time in the target place, such as the activities of people, the movement of objects, etc.
[0064] The computer device obtains the video stream data sent by each video source to understand the flow of people in the target place in real time, such as the states of pedestrians and obstacles in different time periods and different areas of the target place. Among them, the target place is a specific area to be monitored or concerned, and different target places have different layout and structural characteristics. Exemplarily, the target place can be a rail transit station.
[0065] Obtain the floor plan of the target place. The floor plan is a two-dimensional presentation of the spatial structure of the target place, which details the building structure, room layout, channel direction, door and window positions, etc. in the target place, so that the computer device can determine the spatial structure of the target place based on the floor plan of the target place.
[0066] S102. Determine pedestrian information according to the video stream data. The pedestrian information includes pedestrian density information and pedestrian distribution information.
[0067] Optionally, according to the video stream data, pedestrian detection is performed to extract pedestrian information from the video stream data, including pedestrian density information and pedestrian distribution information. Among them, the target venue covered by the video stream is divided into multiple small areas, and the number of pedestrians in each small area is calculated through pedestrian detection. Then, based on the number of pedestrians in each small area, the pedestrian density information is determined. When a pedestrian is detected, the position information of each pedestrian is recorded, and the distribution of pedestrians in the target venue covered by the video stream is reflected through the position information of the pedestrians, that is, the pedestrian distribution information.
[0068] Exemplarily, a 100-square-meter target area is divided into 10 small areas of 10 square meters each. If there are 20 pedestrians in a certain small area, the pedestrian density of this small area is 2 pedestrians per square meter. By statistically analyzing the pedestrian densities of all small areas, the pedestrian density information of the entire target venue covered by the video stream can be obtained.
[0069] S103. According to the floor plan of the target venue and multiple historical pedestrian distribution information of the target venue, construct a deep reinforcement learning model, and determine at least one personnel path mapping information of the target venue through the deep reinforcement learning model. Each personnel path mapping information is used to indicate the mapping relationship between a kind of pedestrian density information and a candidate evacuation path.
[0070] Optionally, the floor plan of the target venue provides a spatial reference for constructing the deep reinforcement learning model. For example, through the floor plan, it can be clarified which areas are the main channels and which areas are relatively narrow, so as to comprehensively consider spatial factors when considering the evacuation path.
[0071] The historical pedestrian distribution information of the target venue is obtained through the monitoring system of the target venue. The historical pedestrian distribution information records the distribution of pedestrians in the target venue at different times and under different circumstances in the past. For example, through long-term analysis of the monitoring videos of rail transit stations, the number of pedestrians and their distribution locations in each area of the rail transit station during different time periods (such as weekday periods, weekend peak periods, etc.) are obtained. The historical pedestrian distribution information can reflect the laws and characteristics of pedestrian activities in the target venue and provide behavioral samples for the construction of the deep reinforcement learning model.
[0072] Based on the obtained plane drawings and historical pedestrian distribution information, a deep reinforcement learning model is constructed. Among them, the deep reinforcement learning model is an evacuation path model based on the deep reinforcement learning algorithm, and the deep reinforcement learning algorithm combines the deep learning algorithm and the reinforcement learning algorithm. Through learning and model training, the deep reinforcement learning model can obtain candidate evacuation paths corresponding to various pedestrian density information, so as to determine at least one personnel path mapping information of the target venue. According to the real-time pedestrian density information, a suitable evacuation path can be quickly selected through at least one personnel path mapping information, improving the evacuation efficiency and safety.
[0073] Among them, the personnel path mapping information refers to the mapping relationship between a pedestrian density information and a candidate evacuation path, providing a decision-making basis for the subsequent safe evacuation of the target venue. For example, when the pedestrian density in a certain area of the target venue reaches a certain level (such as 2 people per square meter), the corresponding candidate evacuation path may be to evacuate from a certain passage near this area to the safety exit.
[0074] It should be noted that the computer device can store at least one personnel path mapping information of the target venue determined by the deep reinforcement learning model in the database, so as to directly call the database subsequently and find the corresponding candidate evacuation path through the pedestrian density information.
[0075] S104. Determine the target evacuation path according to the pedestrian information, the deep reinforcement learning model and at least one personnel path mapping information, so as to evacuate pedestrians according to the target evacuation path.
[0076] Optionally, determine the target evacuation path according to the pedestrian density information and pedestrian distribution information in the pedestrian information, the trained deep reinforcement learning model and at least one personnel path mapping information stored in the database, so that the staff can guide pedestrians to evacuate according to the target evacuation path by means of broadcasting or signpost guiding, etc.
[0077] Among them, the pedestrian density information can reflect the degree of personnel concentration in different areas of the target venue. For example, some areas may be overcrowded, while other areas are relatively loose. The pedestrian distribution information indicates the specific location of pedestrians in the venue. For example, pedestrians may be concentrated on certain floors or passages. Using the pedestrian information as the basis for evacuation path decision-making can avoid guiding pedestrians to areas that are already crowded. In the crowd evacuation scenario, the optimal target evacuation path is obtained according to the pedestrian information, the deep reinforcement learning model and at least one personnel path mapping information.
[0078] Specifically, the personnel path mapping information is an empirical summary obtained based on historical pedestrian distribution information and deep reinforcement learning model training, providing a direct reference basis for determining the target evacuation path. In practical applications, according to the current real-time pedestrian density information, the corresponding alternative evacuation paths can be quickly found from the personnel path mapping information, narrowing the scope of evacuation path selection and improving the decision-making efficiency.
[0079] In this embodiment, video stream data sent by each video source and the floor plan of the target venue are obtained, and pedestrian detection is performed based on the video stream data to determine pedestrian information, including pedestrian density information and pedestrian distribution information. Based on the floor plan and historical pedestrian distribution information, a deep reinforcement learning model is constructed. The deep reinforcement learning model is an evacuation path model based on the deep reinforcement learning algorithm. At least one personnel path mapping information of the target venue is determined through the deep reinforcement learning model to indicate the mapping relationship between a kind of pedestrian density information and an alternative evacuation path. The at least one personnel path mapping information provides a decision-making basis for the safe evacuation of the target venue. According to the pedestrian density information and pedestrian distribution information in the pedestrian information, the trained deep reinforcement learning model, and the at least one personnel path mapping information, the target evacuation path is determined to guide pedestrians to evacuate along the target evacuation path. The pedestrian detection based on the video stream data improves the accuracy of pedestrian information. By constructing an evacuation path model based on the deep reinforcement learning algorithm, the safety and effectiveness of evacuation path planning are improved. The at least one personnel path mapping information of the target venue determined by the deep reinforcement learning model facilitates quickly finding the corresponding alternative evacuation path from the personnel path mapping information according to the pedestrian density information, improving the evacuation efficiency.
[0080] Hereinafter, the process of determining pedestrian information based on video stream data will be described in detail.
[0081] Figure 2 It is a schematic flow chart of determining pedestrian information for the evacuation path planning method provided by the embodiment of the present application. As Figure 2 shown, in the above step S102, determining pedestrian information according to the video stream data includes:
[0082] S201. Extract each video frame according to the video stream data.
[0083] Optionally, the obtained video stream data is continuous video data in time series. To perform pedestrian recognition and detection, the video stream data needs to be separated into individual video frames.
[0084] Exemplarily, the command-line tool of FFmpeg is used to extract each video frame from the video stream data. Specifically, a preset frame rate is set, and the video stream data is extracted as static image frames, that is, video frames, at preset time intervals for subsequent pedestrian recognition and detection.
[0085] S202. Determine pedestrian information based on each video frame by using a pre-constructed pedestrian recognition model.
[0086] Optionally, perform data augmentation on each video frame, including mosaic augmentation, random rotation (±15°), and HSV color perturbation. Among them, mosaic augmentation randomly stitches four different video frames into a new image, and at the same time adjusts the position and size of the annotation box to enhance the diversity of the image. Randomly rotating the video frame can simulate the postures of pedestrians at different angles and improve the recognition and detection ability of the pedestrian recognition model for pedestrians at different perspectives. HSV color perturbation randomly adjusts the hue (H), saturation (S), and value (V) of the video frame to change the color attributes of the video frame and simulate image changes under different lighting conditions, making the pedestrian recognition model more robust to lighting changes.
[0087] Input the augmented video frames into the pre-constructed pedestrian recognition model, and the pedestrian recognition model predicts and outputs pedestrian information based on the input video frames. Among them, the pedestrian recognition model is a pre-constructed and trained model. The COCO-Person dataset is used during the model training process. The COCO-Person dataset contains a large number of pedestrian images and annotation information. Using the COCO-Person dataset can reduce the workload of manual annotation and directly use these annotations to train the pedestrian recognition model.
[0088] In this embodiment, each static video frame is extracted from the video stream data, and data augmentation is performed on each video frame. The augmented video frames are input into the pre-constructed pedestrian recognition model, and the pedestrian recognition model predicts and outputs pedestrian information based on the input video frames. The accuracy of pedestrian recognition and detection is improved, and thus the accuracy of pedestrian information is enhanced.
[0089] As an optional implementation manner, the pedestrian recognition model is a multi-scale convolutional neural network model.
[0090] Optionally, in a pedestrian recognition scenario, due to the different distances between pedestrians and the camera, pedestrians may occupy different pixel areas in the video frame, that is, they appear in different scales. Multi-scale means considering the feature representations of pedestrians in a video frame at multiple different sizes and resolutions. A Convolutional Neural Network (CNN) model is a deep learning model for processing data with a grid structure (such as images), which extracts features from the video frame. Among them, the multi-scale detection heads of the multi-scale CNN model can be set to grids of three scales: 80×80, 40×40, and 20×20, corresponding to the long-distance area (>15m), the middle-distance area (5-15m), and the short-distance area (<5m) respectively, to improve the accuracy of pedestrian recognition.
[0091] Next, the process of determining pedestrian information based on each video frame by a pre-constructed pedestrian recognition model will be described in detail.
[0092] Figure 3 Another schematic diagram of the process for determining pedestrian information in the evacuation path planning method provided by the embodiments of this application is shown in Figure 3 As shown, in step S202 above, determining pedestrian information based on each video frame by a pre-constructed pedestrian recognition model includes:
[0093] S301. Input each video frame into a multi-scale convolutional neural network model, and the multi-scale convolutional neural network model detects the number of pedestrians in the corresponding area of each video frame through an object detection algorithm.
[0094] Optionally, when each video frame is input into the multi-scale CNN model, the multi-scale CNN model performs a series of convolution, pooling, and activation operations on the input video frame, extracts features of the video frame from different scales, and performs object (pedestrian) localization and classification on the feature map through an object detection algorithm, and finally outputs the number of pedestrians in the corresponding area of each video frame.
[0095] Among them, the expression of the loss function in the object detection algorithm is as follows, including coordinate loss, confidence loss, and classification loss:
[0096] Among them, λ coord is a preset first weight, λ noobj is a preset second weight, S is the number of grids, B is the number of bounding boxes predicted for each grid, indicates whether the j-th bounding box in the i-th grid is responsible for detecting the target, is the true coordinate of the center of the bounding box, is the predicted coordinate of the center of the bounding box, is the width and height of the true bounding box, is the predicted width and height, Indicates that the j-th bounding box in the i-th grid is not responsible for detecting the target, is the true confidence, is the predicted confidence. is the true label that the target in the j-th bounding box in the i-th grid belongs to class c, is the predicted label that the target in the j-th bounding box in the i-th grid belongs to class c.
[0097] S302. Determine the pedestrian density information according to the number of pedestrians in the corresponding area of each video frame.
[0098] Optionally, the pedestrian density information is the number of pedestrians per unit area. After obtaining the number of pedestrians in the corresponding area of each video frame and combining it with the area of this area, the pedestrian density can be calculated. For example, if the area of the area is A and the number of pedestrians in this area is N, then the pedestrian density D is the quotient of the number of pedestrians N in this area and the area A. Among them, the area of each area can be determined through a plane drawing.
[0099] S303. Determine the planar coordinates of each pedestrian in each video frame according to the pixel coordinates of each pedestrian in each video frame and a preset perspective transformation matrix, and use the planar coordinates of each pedestrian in each video frame as the pedestrian distribution information.
[0100] Optionally, in the video frame, the position of the pedestrian is represented by pixel coordinates, and in the actual scene, the position of the pedestrian is represented by planar coordinates. Through the preset perspective transformation matrix M, the pixel coordinates of each pedestrian in the video frame can be projected and transformed into planar coordinates.
[0101] Specifically, establish a perspective transformation matrix M from pixel coordinates to planar coordinates. The expression of the perspective transformation matrix M is as follows:
[0102]
[0103] For a point (x, y) on the pixel coordinate system, the perspective transformation maps it to a point (x ′ , y ′ ) on the planar coordinate system. The transformation formula is:
[0104]
[0105] After expansion, it is obtained:
[0106]
[0107] For each video frame in the monitored position area, a set of source points and target points are determined. Through the perspective transformation matrix M, the source points can be mapped to the target points by M, and the planar coordinates of each pedestrian in each video frame are obtained. The planar coordinates of each pedestrian in each video frame are used as pedestrian distribution information. Among them, the source point is the point (x, y) in the pixel coordinate system of the video frame, and the target point is the point (x′, y′) of the source point in the actual physical coordinate system, that is, the planar coordinate system.
[0108] It should be noted that for cameras that can detect the same area, by comparing their respective effective monitoring areas, the effective monitoring images are calculated inversely through the perspective transformation matrix M. For such monitoring cameras, only the pedestrians in the effective monitoring images are identified, rather than global pedestrian identification, which improves the accuracy of identification.
[0109] In this embodiment, each video frame is input into a multi-scale convolutional neural network model. The multi-scale convolutional neural network model extracts the features of the video frame from different scales according to the input video frames, and performs pedestrian localization and classification on the feature map through a target detection algorithm to obtain the number of pedestrians in the corresponding area of each video frame. According to the number of pedestrians in the corresponding area of each video frame and the area of each area, the pedestrian density information is calculated. Through a preset perspective transformation matrix, the pixel coordinates of each pedestrian in the video frame are converted into planar coordinates, and the planar coordinates of each pedestrian in each video frame are used as pedestrian distribution information. This further improves the detection accuracy of pedestrian recognition and the accuracy of pedestrian information.
[0110] Next, the process of constructing a deep reinforcement learning model based on the floor plan of the target site and multiple historical pedestrian distribution information of the target site, and determining at least one personnel path mapping information of the target site through the deep reinforcement learning model will be described in detail.
[0111] Figure 4 This is a schematic flowchart of the process of constructing a deep reinforcement learning model for the evacuation path planning method provided by the embodiment of the present application. As Figure 4 shown, in the above step S103, constructing a deep reinforcement learning model based on the floor plan of the target site and multiple historical pedestrian distribution information of the target site, and determining at least one personnel path mapping information of the target site through the deep reinforcement learning model includes:
[0112] S401. Build a simulation environment of the target site according to the floor plan of the target site.
[0113] Optionally, the floor plan of the target venue contains detailed layout information of the target venue. Using this layout information, a virtual simulation environment consistent with the structure of the real target venue can be built using simulation software. In this environment, the physical properties and spatial relationships of various elements correspond to those of the actual target venue, providing a virtual interactive environment that approximates the real continuous space for subsequent pedestrian movement simulation.
[0114] S402. Determine the movement parameters of pedestrians in the target venue according to the preset pedestrian movement model and the simulation environment. The movement parameters include the movement direction and movement speed.
[0115] Optionally, the preset pedestrian movement model is a mathematical model established based on pedestrian behavior and dynamics principles, used to describe the movement laws of pedestrians in different scenarios. Applying the pedestrian movement model to the built simulation environment and combining the geometric information of the environment and the movement state of pedestrians, the movement parameters of pedestrians in the target venue can be calculated, including the movement direction and movement speed.
[0116] For example, the movement state of pedestrians can be that pedestrians will try to avoid collisions between pedestrian groups during movement and will also receive feedback when colliding with positions such as walls.
[0117] S403. Construct a deep reinforcement learning model according to the movement parameters, historical pedestrian distribution information, pre-constructed reward function, and initial deep reinforcement learning model, and determine at least one personnel path mapping information of the target venue through the deep reinforcement learning model. The reward function is used to characterize the relationship between the distance of pedestrians from the exit, the stay time of pedestrians in the target venue, and the degree of crowding in the target venue, and the maximum value of the reward function is related to the area of the target venue.
[0118] Optionally, the movement parameters provide dynamic information of pedestrians, reflecting the real-time state of pedestrians in the target venue. The historical pedestrian distribution information contains the distribution of pedestrians in the target venue at different past time points, which can reflect the laws and trends of pedestrian flow in the target venue.
[0119] The pre-constructed reward function is used to measure the reward value obtained by the agent for choosing each evacuation path, characterizing the relationship between the distance of pedestrians from the exit, the stay time of pedestrians in the target venue, and the degree of crowding in the target venue, and the maximum value of the reward function is related to the area of the target venue. Among them, the stay time of pedestrians in the target venue can include the movement time and residence time of pedestrians in the target venue.
[0120] Iteratively train the initial deep reinforcement learning model according to the motion parameters, historical pedestrian distribution information, and a pre-constructed reward function to construct a deep reinforcement learning model, and determine at least one personnel path mapping information of the target venue through the trained deep reinforcement learning model. Among them, the trained deep reinforcement learning model can output corresponding optimal evacuation paths under different pedestrian distribution information, and the personnel path mapping information is a mapping relationship formed by associating different pedestrian density distributions with the corresponding optimal evacuation paths.
[0121] In this embodiment, according to the floor plan of the target venue, build a simulation environment of the target venue to provide a virtual interaction environment approaching a real continuous space. Apply the pedestrian motion model to the built simulation environment, and combine the geometric information of the environment and the motion state of the pedestrians to determine the motion parameters of the pedestrians in the target venue, including the motion direction and the motion speed. Iteratively train the initial deep reinforcement learning model according to the motion parameters, historical pedestrian distribution information, and a pre-constructed reward function to construct a deep reinforcement learning model, and determine at least one personnel path mapping information of the target venue through the trained deep reinforcement learning model. By learning a large amount of historical pedestrian distribution information, the deep reinforcement learning model can avoid pedestrians choosing crowded or unreasonable paths during evacuation, reduce the evacuation time, and improve the evacuation efficiency. The reward function comprehensively considers various factors such as the distance between pedestrians and exits, the staying time, and the degree of crowding, and the maximum value is related to the area of the target venue, enabling the deep reinforcement learning model to adapt to target venues of different scales and layouts, as well as various complex pedestrian flow situations, and improving the adaptability of the deep reinforcement learning model to complex evacuation scenarios.
[0122] As an optional implementation manner, the pedestrian motion model is a mechanical model, including: a pedestrian target driving force model, an inter-pedestrian repulsive force model, and a wall repulsive force model.
[0123] Optionally, the pedestrian motion model adopts a mechanical model, in which the pedestrian target driving force model prompts the pedestrian to move towards the exit direction, the inter-pedestrian repulsive force model enables pedestrians to avoid colliding with each other, and the wall repulsive force model prevents pedestrians from passing through the wall in the simulation environment and takes actions after colliding with the wall.
[0124] Specifically, the expression of the pedestrian target driving force model is as follows:
[0125]
[0126] Among them, m is the mass of the pedestrian, v 0 is the expected walking speed, represents the unit direction vector of the action-taking direction, is the current actual speed, and τ is the time constant for speed adjustment.
[0127] The expression of the repulsive force model between pedestrians is as follows:
[0128]
[0129] where A is the psychological repulsion intensity, E is the action range coefficient, r ab -d ab represents the distance between pedestrians, θ represents the Heaviside step function, is the normal vector between pedestrians, represents the unit vector from pedestrian a to pedestrian b, and k is the elastic coefficient.
[0130] The expression of the repulsive force model of the wall is as follows:
[0131]
[0132] where μ w is the coefficient related to the wall repulsive force intensity, d w is the shortest distance from the pedestrian to the wall, is the normal vector of the wall at the position of the pedestrian, and σ is the attenuation length parameter.
[0133] In this embodiment, the pedestrian motion model is a mechanical model, including a pedestrian target driving force model, a repulsive force model between pedestrians, and a repulsive force model of the wall. The pedestrian target driving force model prompts the pedestrian to move towards the exit direction, the repulsive force model between pedestrians prevents pedestrians from colliding with each other, and the repulsive force model of the wall prevents pedestrians from passing through the wall in the simulation environment and takes actions after colliding with the wall. Through the constraints of multiple mechanical models, the accuracy of the motion parameters of pedestrians in the target place is improved.
[0134] Next, the process of constructing a deep reinforcement learning model based on motion parameters, historical pedestrian distribution information, a pre-constructed reward function, and an initial deep reinforcement learning model, and determining at least one personnel path mapping information of the target place through the deep reinforcement learning model will be described in detail.
[0135] Figure 5 is another process schematic diagram of constructing a deep reinforcement learning model for the evacuation path planning method provided by the embodiment of the present application. As Figure 5 shown, in the above step S403, constructing a deep reinforcement learning model based on motion parameters, historical pedestrian distribution information, a pre-constructed reward function, and an initial deep reinforcement learning model, and determining at least one personnel path mapping information of the target place through the deep reinforcement learning model includes:
[0136] S501. Input the motion parameters and historical pedestrian distribution information into the initial deep reinforcement learning model, and the initial deep reinforcement learning model determines the reward value of the reward function.
[0137] Optionally, the motion parameters describe the dynamic characteristics of pedestrians in the target venue, including the motion direction and motion speed. The historical pedestrian distribution information records the distribution information of pedestrians in the target venue over a past period of time, which can reflect the laws and trends of pedestrian flow. These two types of information are input into the initial deep reinforcement learning model as the basis for model training and decision-making.
[0138] The reward function is used in deep reinforcement learning to evaluate the effect of an agent taking a certain action. In this process, the initial deep reinforcement learning model selects an evacuation path based on the input motion parameters and historical pedestrian distribution information, and calculates the reward value in the current state in combination with the rules of the reward function. Among them, the reward function comprehensively considers multiple factors, including the distance between pedestrians and the exit, the staying time of pedestrians in the target venue, and the congestion degree of the target venue.
[0139] S502. The initial deep reinforcement learning model iteratively corrects the model parameters according to the reward value of the reward function, and obtains the deep reinforcement learning model at the end of the iteration.
[0140] Optionally, the initial deep reinforcement learning model determines whether the current evacuation path decision is correct according to the obtained reward value. If the reward value is high, it means that the current evacuation path decision is effective, and the initial deep reinforcement learning model tends to strengthen this strategy; if the reward value is low, the initial deep reinforcement learning model will try to adjust the strategy. By continuously interacting with the environment, the initial deep reinforcement learning model uses the Proximal Policy Optimization (PPO) algorithm to iteratively correct its own parameters.
[0141] As the number of iterations increases, the initial deep reinforcement learning model gradually learns how to make decisions that can obtain the maximum reward value under different motion parameters and historical pedestrian distribution information. When the preset number of iterations is reached or certain convergence conditions are met, the iteration process ends. At this time, the obtained model is the trained deep reinforcement learning model, and the deep reinforcement learning model can accurately make the optimal evacuation path decision according to the input information.
[0142] S503. Determine at least one personnel path mapping information according to the deep reinforcement learning model.
[0143] Optionally, the trained deep reinforcement learning model can learn the optimal evacuation paths under various pedestrian distribution information. The deep reinforcement learning model associates different pedestrian distribution information with the corresponding optimal evacuation paths, forms a mapping relationship, and records all the mapping relationships to obtain the personnel path mapping information.
[0144] In this embodiment, the motion parameters and historical pedestrian distribution information are input into the initial deep reinforcement learning model as the basis for model training and decision-making. The initial deep reinforcement learning model selects an evacuation path according to the input motion parameters and historical pedestrian distribution information, and determines the reward value of the reward function in combination with the rules of the reward function. The initial deep reinforcement learning model iteratively corrects the model parameters according to the reward value of the reward function and obtains the deep reinforcement learning model at the end of the iteration, so that the deep reinforcement learning model can make an optimal evacuation path decision relatively accurately according to the input information. The trained deep reinforcement learning model learns the optimal evacuation paths under various pedestrian distribution information, associates different pedestrian distribution information with the corresponding optimal evacuation paths, forms a mapping relationship, and records all the mapping relationships to obtain the personnel path mapping information. The use of historical pedestrian distribution information enables the deep reinforcement learning model to learn the long-term laws and trends of pedestrian flow in the target venue, so that when facing new pedestrian distribution information, it can still output reasonable evacuation paths according to the learned laws, enhancing the generalization ability of the deep reinforcement learning model.
[0145] Hereinafter, the process of determining at least one personnel path mapping information according to the deep reinforcement learning model will be described in detail.
[0146] Figure 6 FIG. is a schematic flow chart of determining the personnel path mapping information for the evacuation path planning method provided by the embodiment of the present application. As Figure 6 shown, in step S503 above, determining at least one personnel path mapping information according to the deep reinforcement learning model includes:
[0147] S601. Obtain multiple reference pedestrian distribution information.
[0148] Optionally, the reference pedestrian distribution information is data describing the distribution of pedestrians in different regions at different times in the target venue. Among them, the pedestrian distribution information obtained by pedestrian recognition in the foregoing embodiment can be used as the reference pedestrian distribution information.
[0149] Alternatively, in the simulation environment, according to factors such as the layout of the target venue and the expected pedestrian flow, reference pedestrian distribution information in different scenarios is simulated and generated, including randomly generating different numbers of pedestrians, dispersing them in different regions, and obtaining the distribution information of pedestrians in the simulation environment as the reference pedestrian distribution information.
[0150] S602. Input multiple reference pedestrian distribution information into a deep reinforcement learning model, and the deep reinforcement learning model predicts and generates candidate evacuation paths corresponding to each reference pedestrian density information, where the reference pedestrian density information is the pedestrian density information corresponding to the reference pedestrian distribution information.
[0151] Optionally, the deep reinforcement learning model is a trained intelligent model that can learn to make optimal evacuation path decisions in different environmental states to maximize the reward value. After inputting multiple reference pedestrian distribution information into the deep reinforcement learning model, the deep reinforcement learning model will analyze the scenarios corresponding to each reference pedestrian distribution information according to the strategies and rules learned internally, and combine the decision-making mechanism of the model itself to predict the pedestrian density information corresponding to various reference pedestrian distribution information, that is, the candidate evacuation paths corresponding to the reference pedestrian density information.
[0152] Specifically, the deep reinforcement learning model evaluates the advantages and disadvantages of different evacuation paths according to the input reference pedestrian distribution information. The deep reinforcement learning model will consider various factors, such as the length of the evacuation path, the distribution of obstacles, the evacuation time, etc., to determine the path that can evacuate pedestrians to the safe area fastest and most safely under the current reference pedestrian distribution.
[0153] Among them, the reference pedestrian density information can be obtained according to the reference pedestrian distribution information to reflect the pedestrian density in different areas within the target venue.
[0154] S603. Establish a mapping relationship between each reference pedestrian density information and the corresponding candidate evacuation path respectively to obtain at least one personnel-path mapping information.
[0155] Optionally, associate each reference pedestrian density information with the corresponding candidate evacuation path predicted and generated by the deep reinforcement learning model to form a mapping relationship. By establishing multiple such mapping relationships, personnel-path mapping information is obtained.
[0156] Among them, the personnel-path mapping information can be stored in a table, database or other data structures for convenient query and use in actual applications. When the pedestrian density information in the actual scenario is determined, the corresponding candidate evacuation path can be quickly found according to the personnel-path mapping information, providing data support for pedestrian evacuation.
[0157] In this embodiment, multiple reference pedestrian distribution information is obtained and input into a trained deep reinforcement learning model. The deep reinforcement learning model analyzes the scenarios corresponding to each reference pedestrian distribution information, and combines its own decision-making mechanism to predict the pedestrian density information corresponding to various reference pedestrian distribution information, that is, the candidate evacuation paths corresponding to the reference pedestrian density information. Each reference pedestrian density information is associated with the corresponding candidate evacuation path predicted and generated by the deep reinforcement learning model to form a mapping relationship. By establishing multiple such mapping relationships, personnel path mapping information is obtained, and the personnel path mapping information is stored so as to directly determine the corresponding candidate evacuation path according to the pedestrian density information in the actual scenario, improving the evacuation efficiency and the effectiveness of the evacuation path.
[0158] Hereinafter, the process of determining the target evacuation path according to the pedestrian information, the deep reinforcement learning model, and at least one personnel path mapping information will be described in detail.
[0159] Figure 7 It is a schematic flowchart of determining the target evacuation path of the evacuation path planning method provided by the embodiment of the present application. As Figure 7 shown, in the above step S104, determining the target evacuation path according to the pedestrian information, the deep reinforcement learning model, and at least one personnel path mapping information includes:
[0160] S701. Determine whether an emergency event occurs.
[0161] Optionally, an emergency event refers to a sudden situation that may pose a threat to the lives and property of people, such as a fire, an earthquake, etc. In the target venue, various monitoring means are required to determine whether an emergency event has occurred.
[0162] Specifically, it can be achieved by means of a variety of sensors and monitoring systems. For example, install smoke alarms and temperature sensors to detect whether a fire has occurred; use seismic monitoring equipment to sense the occurrence of an earthquake; judge whether there are signs of violent behavior through multiple cameras and face recognition in the monitoring system. Once relevant abnormal signals are detected, it is determined that an emergency event has occurred.
[0163] S702. If so, the deep reinforcement learning model predicts and generates a target evacuation path according to the pedestrian distribution information.
[0164] Optionally, when it is determined that an emergency event occurs, the deep reinforcement learning model will analyze and make decisions according to the current real-time pedestrian distribution information. The pedestrian distribution information includes the specific location and distribution of pedestrians in the venue.
[0165] As an intelligent model trained with a large amount of data, the deep reinforcement learning model can comprehensively consider various factors such as the layout of the venue, the location of obstacles, the location of exits, and the dynamic distribution of pedestrians, and predict and generate an optimal target evacuation path. The target evacuation path enables pedestrians to evacuate to a safe area in the shortest time with the highest safety. For example, in the event of a fire, the model will avoid the areas where the fire is spreading and the passages with high smoke concentration, and guide pedestrians to choose a safe evacuation route.
[0166] S703. Otherwise, determine an alternative evacuation path that matches the pedestrian density information in the pedestrian information according to at least one personnel path mapping information, and use the alternative evacuation path that matches the pedestrian density information in the pedestrian information as the target evacuation path.
[0167] Optionally, the personnel path mapping information is established by associating a large amount of reference pedestrian density information with the corresponding alternative evacuation paths before. Under normal circumstances without an emergency event, the computer device searches for an alternative evacuation path that matches the current pedestrian density information from the stored at least one personnel path mapping information, and uses it as the target evacuation path.
[0168] For example, if the pedestrian density in a certain area is 2 people per square meter currently, the system will find the alternative evacuation path corresponding to a pedestrian density of 2 people per square meter in the personnel path mapping information and determine it as the current target evacuation path.
[0169] In this embodiment, it is monitored in real time whether an emergency event occurs in the target venue. If so, the deep reinforcement learning model analyzes and makes decisions according to the current real-time pedestrian distribution information, comprehensively considers various factors such as the layout of the venue, the location of obstacles, the location of exits, and the dynamic distribution of pedestrians, and predicts and generates an optimal target evacuation path. Otherwise, the computer device searches for an alternative evacuation path that matches the current pedestrian density information from the stored at least one personnel path mapping information, and uses it as the target evacuation path. The emergency response ability is improved, and the real-time planning and dynamic adjustment of the evacuation path are realized through the deep learning reinforcement model. And under normal circumstances without an emergency event, according to the pre-stored personnel path mapping information, the determination efficiency and accuracy of the target evacuation path are improved. Different evacuation path planning methods are adopted according to whether an emergency event occurs, which can not only make accurate decisions quickly in case of emergency, but also make reasonable plans using the existing personnel path mapping information under normal circumstances, thus enhancing the adaptability and reliability of planning the target evacuation path in different scenarios.
[0170] Based on the same inventive concept, an evacuation route planning device corresponding to the evacuation route planning method is further provided in the embodiments of the present application. Since the principle of solving problems by the device in the embodiments of the present application is similar to that of the above-mentioned evacuation route planning method in the embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0171] Figure 8 It is a module structure diagram of the evacuation route planning device provided in the embodiments of the present application. As Figure 8 shown, the device includes:
[0172] An acquisition module 801, configured to acquire video stream data sent by each video source and a floor plan of the target venue.
[0173] A determination module 802, configured to determine pedestrian information according to the video stream data, where the pedestrian information includes pedestrian density information and pedestrian distribution information.
[0174] A construction module 803, configured to construct a deep reinforcement learning model according to the floor plan of the target venue and multiple historical pedestrian distribution information of the target venue, and determine at least one personnel path mapping information of the target venue through the deep reinforcement learning model. Each personnel path mapping information is respectively used to indicate the mapping relationship between a kind of pedestrian density information and an alternative evacuation path.
[0175] The determination module 802 is further configured to determine a target evacuation path according to the pedestrian information, the deep reinforcement learning model, and at least one personnel path mapping information, so as to evacuate pedestrians according to the target evacuation path.
[0176] As an alternative implementation manner, the determination module 802 is specifically configured to:
[0177] Extract each video frame according to the video stream data.
[0178] Determine pedestrian information according to each video frame by a pre-constructed pedestrian recognition model.
[0179] As an alternative implementation manner, the determination module 802 is specifically configured to:
[0180] Input each video frame into a multi-scale convolutional neural network model, and the multi-scale convolutional neural network model detects the number of pedestrians in the corresponding area of each video frame through an object detection algorithm.
[0181] Determine pedestrian density information according to the number of pedestrians in the corresponding area of each video frame.
[0182] Determine the plane coordinates of each pedestrian in each video frame according to the pixel coordinates of each pedestrian in each video frame and a preset perspective transformation matrix, and use the plane coordinates of each pedestrian in each video frame as pedestrian distribution information.
[0183] As an alternative implementation, the construction module 803 is specifically configured to:
[0184] Build a simulation environment of the target venue according to the floor plan of the target venue.
[0185] Determine the movement parameters of pedestrians in the target venue according to the preset pedestrian movement model and the simulation environment, where the movement parameters include movement direction and movement speed.
[0186] Build a deep reinforcement learning model according to the movement parameters, historical pedestrian distribution information, pre-constructed reward function, and initial deep reinforcement learning model, and determine at least one personnel path mapping information of the target venue through the deep reinforcement learning model. Among them, the reward function is used to characterize the relationship between the distance between pedestrians and the exit, the stay time of pedestrians in the target venue, and the degree of crowding in the target venue, and the maximum value of the reward function is related to the area of the target venue.
[0187] As an alternative implementation, the pedestrian movement model is a mechanical model, including: pedestrian target driving force model, pedestrian repulsive force model, and wall repulsive force model.
[0188] As an alternative implementation, the construction module 803 is specifically configured to:
[0189] Input the movement parameters and historical pedestrian distribution information into the initial deep reinforcement learning model, and the initial deep reinforcement learning model determines the reward value of the reward function.
[0190] The initial deep reinforcement learning model iteratively corrects the model parameters according to the reward value of the reward function, and obtains the deep reinforcement learning model at the end of the iteration.
[0191] Determine at least one personnel path mapping information according to the deep reinforcement learning model.
[0192] As an alternative implementation, the construction module 803 is specifically configured to:
[0193] Obtain multiple reference pedestrian distribution information.
[0194] Input the multiple reference pedestrian distribution information into the deep reinforcement learning model, and the deep reinforcement learning model predicts and generates candidate evacuation paths corresponding to each reference pedestrian density information, where the reference pedestrian density information is the pedestrian density information corresponding to the reference pedestrian distribution information.
[0195] Establish the mapping relationship between each reference pedestrian density information and the corresponding candidate evacuation path respectively, and obtain at least one personnel path mapping information.
[0196] As an alternative implementation, the determination module 802 is specifically configured to:
[0197] Determine whether an emergency event has occurred.
[0198] If so, the deep reinforcement learning model predicts and generates a target evacuation path based on the pedestrian distribution information.
[0199] Otherwise, determine an alternative evacuation path that matches the pedestrian density information in the pedestrian information according to at least one personnel path mapping information, and use the alternative evacuation path that matches the pedestrian density information in the pedestrian information as the target evacuation path.
[0200] An embodiment of the present application also provides a computer device, as Figure 9 shown, which is a schematic structural diagram of the computer device provided by the embodiment of the present application, including: a processor 91, a memory 92, and a bus 93. The memory 92 stores machine-readable instructions executable by the processor 91 (for example, Figure 8 the execution instructions corresponding to the acquisition module 801, the determination module 802, and the construction module 803 in the device in
[0201] When the computer device runs, the processor 91 communicates with the memory 92 through the bus 93. When the machine-readable instructions are executed by the processor 91, the steps of the evacuation path planning method in the above embodiment are executed.
[0202] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the method embodiments, which will not be elaborated in the present application. In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces. The indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical, or other form.
[0203] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0204] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application.
Claims
1. A method for evacuation path planning, characterized in that: include: Obtain the video stream data sent by each video source and the floor plan of the target location; Determining pedestrian information according to the video stream data, the pedestrian information including pedestrian density information and pedestrian distribution information; According to the plan drawing of the target place and multiple historical pedestrian distribution information of the target place, a deep reinforcement learning model is constructed, and at least one personnel path mapping information of the target place is determined by the deep reinforcement learning model, each of the personnel path mapping information is used to indicate a mapping relationship between a type of pedestrian density information and a to-be-selected evacuation path; A target evacuation path is determined according to the pedestrian information, the deep reinforcement learning model, and the at least one personnel path mapping information, so as to evacuate the pedestrians according to the target evacuation path.
2. The method according to claim 1, characterized in that The determining pedestrian information according to the video stream data includes: Extracting each video frame according to the video stream data; The pedestrian information is determined according to each of the video frames using a pre-built pedestrian recognition model.
3. The method according to claim 2, characterized in that The pedestrian recognition model is a multi-scale convolutional neural network model; Determining the pedestrian information according to each of the video frames by using a pre-built pedestrian recognition model includes: Inputting each of the video frames into the multi-scale convolutional neural network model, and using the multi-scale convolutional neural network model to detect the number of pedestrians in the corresponding area of each of the video frames through a target detection algorithm; Determining pedestrian density information according to the number of pedestrians in the area corresponding to each of the video frames; According to the pixel coordinates of each pedestrian in each video frame and a preset perspective transformation matrix, the plane coordinates of each pedestrian in each video frame are determined, and the plane coordinates of each pedestrian in each video frame are used as pedestrian distribution information.
4. The method according to claim 1, characterized in that: The step of constructing a deep reinforcement learning model according to the plan drawing of the target place and a plurality of historical pedestrian distribution information of the target place, and determining at least one personnel path mapping information of the target place by using the deep reinforcement learning model includes: According to the plan drawing of the target place, build a simulation environment of the target place; Determining the motion parameters of the pedestrian in the target location according to the preset pedestrian motion model and the simulation environment, wherein the motion parameters include motion direction and motion speed; The deep reinforcement learning model is constructed according to the motion parameters, the historical pedestrian distribution information, a pre-constructed reward function and an initial deep reinforcement learning model, and at least one personnel path mapping information of the target place is determined by the deep reinforcement learning model, wherein the reward function is used to characterize the relationship between the distance between the pedestrian and the exit, the pedestrian's stay time in the target place and the degree of congestion in the target place, and the maximum value of the reward function is related to the area of the target place.
5. The method according to claim 4, characterized in that The pedestrian motion model is a mechanical model, including: a pedestrian target driving force model, an inter-pedestrian repulsion force model and a wall repulsion force model.
6. The method according to claim 4, characterized in that The step of constructing the deep reinforcement learning model according to the motion parameters, the historical pedestrian distribution information, the pre-constructed reward function, and the initial deep reinforcement learning model, and determining at least one personnel path mapping information of the target place through the deep reinforcement learning model includes: Inputting the motion parameters and the historical pedestrian distribution information into the initial deep reinforcement learning model, and determining the reward value of the reward function by the initial deep reinforcement learning model; The initial deep reinforcement learning model iteratively modifies model parameters according to the reward value of the reward function, and obtains the deep reinforcement learning model at the end of the iteration; Determine at least one personnel path mapping information according to the deep reinforcement learning model.
7. The method according to claim 6, characterized in that The determining, according to the deep reinforcement learning model, at least one personnel path mapping information comprises: Obtain multiple reference pedestrian distribution information; Inputting the plurality of reference pedestrian distribution information into the deep reinforcement learning model, and using the deep reinforcement learning model to predict and generate a candidate evacuation path corresponding to each reference pedestrian density information, wherein the reference pedestrian density information is pedestrian density information corresponding to the reference pedestrian distribution information; A mapping relationship between each reference pedestrian density information and a corresponding evacuation path to be selected is established respectively to obtain the at least one personnel path mapping information.
8. The method according to claim 1, characterized in that The determining a target evacuation path according to the pedestrian information, the deep reinforcement learning model, and the at least one personnel path mapping information includes: Determine if an emergency has occurred; If yes, the deep reinforcement learning model predicts and generates the target evacuation path according to the pedestrian distribution information; Otherwise, a candidate evacuation path matching the pedestrian density information in the pedestrian information is determined according to the at least one personnel path mapping information, and the candidate evacuation path matching the pedestrian density information in the pedestrian information is used as the target evacuation path.
9. A computer device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the steps of the evacuation path planning method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the evacuation route planning method according to any one of claims 1 to 8 are executed.