Escape path determination method and device, equipment, medium and product
By acquiring user location and environmental information within a building, candidate escape routes are generated and scored, solving the problem of non-dynamic path adjustment in existing technologies. This enables safe and reliable escape route planning in high-rise buildings and other locations, reducing the risk of casualties.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-03-10
AI Technical Summary
Existing fire prevention and control systems cannot adjust path priorities according to real-time environment, which may lead people to rush towards fire or congested areas in densely populated places such as high-rise buildings. This reduces the reliability of path planning in emergency scenarios and increases the risk of casualties.
By acquiring user location and environmental information within the building, the status of sub-areas is determined, multiple candidate escape routes are generated, and scores are applied based on safety indicator data. The route with the highest score is selected as the target escape route, and the escape plan is dynamically adjusted to adapt to environmental changes.
Ensuring that escape plans remain optimal in dynamic environments significantly reduces the risk of injury or death due to path selection errors and enhances the safety assurance capability of path planning in emergency scenarios.
Smart Images

Figure CN121638607A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of emergency management technology, and in particular relates to a method, device, equipment, medium and product for determining an escape route. Background Technology
[0002] In emergency management scenarios involving high-rise buildings and other densely populated areas, fires can spread rapidly, and the indoor environment can change in an instant. Providing trapped individuals with personalized escape guidance during the escape process, helping them avoid crowded areas and fire hazard points, will greatly increase the success rate of their escape.
[0003] Today, large public spaces and high-rise buildings are generally equipped with fire prevention and control systems. Common measures include escape warning voice prompt systems and lighting guidance systems, which can guide people to escape to a certain extent. However, while current fire prevention and control systems can provide voice or lighting guidance, they cannot adjust path priorities according to the real-time environment. This leads to a "one-size-fits-all" approach that may guide people towards fire or congested areas, resulting in blind spots in escape routes that appear safe but are actually dangerous. This limitation reduces the reliability of path planning in emergency scenarios and, because it cannot dynamically adapt to changes in the environment and the needs of people, indirectly increases the risk of casualties in fires and other emergencies. Summary of the Invention
[0004] This application provides a method, apparatus, equipment, medium, and product for determining escape routes, which can significantly improve the safety assurance capability of route planning in emergency scenarios and effectively reduce the risk of casualties in emergency events such as fires.
[0005] In a first aspect, embodiments of this application provide a method for determining an escape route, including: In the event of an emergency within a building, obtain the location information of each user within the building and the environmental information within the building; Based on the location and environmental information of each user, determine the regional status information of each sub-area within the building; Based on the regional status information of each sub-region, multiple candidate escape paths for the target user are determined; Obtain safety indicator data corresponding to candidate escape routes; The candidate escape routes are scored based on safety index data to obtain the route score for each candidate escape route. The candidate escape path corresponding to the highest path score is determined as the target escape path.
[0006] Based on the same inventive concept, in a second aspect, embodiments of this application also provide an escape route determination device, comprising: The acquisition module is used to acquire the location information of each user in the building and the environmental information of the building in the event of an emergency. The determination module is used to determine the regional status information of each sub-area within the building based on the location and environmental information of each user; The determination module is also used to determine multiple candidate escape paths for the target user based on the regional status information of each sub-region; The acquisition module is also used to acquire safety indicator data corresponding to candidate escape paths; The scoring module is used to score the candidate escape routes based on the safety index data corresponding to the candidate escape routes, and obtain the path score of each candidate escape route. The determination module is also used to determine the candidate escape path corresponding to the highest path score as the target escape path.
[0007] Based on the same inventive concept, in a third aspect, embodiments of this application also provide an escape path determination device, the device including a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the first aspect, or the escape path determination method in any embodiment of the first aspect.
[0008] Based on the same inventive concept, in a fourth aspect, embodiments of this application also provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the first aspect, or the method for determining an escape path in any embodiment of the first aspect.
[0009] Based on the same inventive concept, in a fifth aspect, embodiments of this application also provide a computer program product, wherein instructions in the computer program product, when executed by a processor of a device, enable the device to execute the method for determining an escape path in the first aspect or any embodiment of the first aspect.
[0010] According to the escape route determination method, apparatus, equipment, medium, and product provided in this application embodiment, in the event of an emergency within a building, the location information of each user within the building and the environmental information within the building are obtained. Then, based on the location information of each user and the environmental information, the regional status information of each sub-area within the building can be determined. Next, based on the regional status information of each sub-area, multiple candidate escape routes can be planned for the target user. Then, based on the safety index data corresponding to the candidate escape routes, the candidate escape routes are scored to obtain the path score of each candidate escape route. Thus, the candidate escape route corresponding to the highest path score can be determined as the target escape route, so that the target user can escape according to the target escape route. This application embodiment, through multiple path alternatives and a scientific path scoring mechanism, can ensure that the escape plan always remains optimal in a dynamic environment, significantly reducing the risk of injury or death caused by path selection errors, ensuring the optimality and reliability of the escape plan, significantly improving the safety assurance capability of path planning in emergency scenarios, and effectively reducing the risk of personnel injury or death in emergencies such as fires. Attached Figure Description
[0011] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings, in which the same or similar reference numerals denote the same or similar features, and the drawings are not drawn to scale.
[0012] Figure 1 This is a flowchart illustrating a method for determining an escape route provided in an embodiment of this application; Figure 2 This is a schematic diagram of the system architecture of the escape path determination system in the escape path determination method provided in the embodiments of this application; Figure 3 This is another flowchart illustrating the method for determining an escape route provided in the embodiments of this application; Figure 4 This is another flowchart illustrating the method for determining an escape route provided in the embodiments of this application; Figure 5 This is another flowchart illustrating the method for determining an escape route provided in the embodiments of this application; Figure 6 This is another flowchart illustrating the method for determining an escape route provided in the embodiments of this application; Figure 7 This is another flowchart illustrating the method for determining an escape route provided in the embodiments of this application; Figure 8 This is another flowchart illustrating the method for determining an escape route provided in the embodiments of this application; Figure 9This is another flowchart illustrating the method for determining an escape route provided in the embodiments of this application; Figure 10 This is another flowchart illustrating the method for determining an escape route provided in the embodiments of this application; Figure 11 This is another flowchart illustrating the method for determining an escape route provided in the embodiments of this application; Figure 12 This is another flowchart illustrating the method for determining an escape route provided in the embodiments of this application; Figure 13 This is a schematic diagram of a device for determining an escape path provided in an embodiment of this application; Figure 14 This is a schematic diagram of a device for determining an escape path provided in an embodiment of this application. Detailed Implementation
[0013] The features and exemplary embodiments of various aspects of this application will now be described in detail. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only configured to explain this application and are not configured to limit this application. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples of this application.
[0014] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0015] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0016] Various modifications and variations can be made to this application without departing from its spirit or scope, which will be apparent to those skilled in the art. Therefore, this application is intended to cover modifications and variations falling within the scope of the corresponding claims (the claimed technical solutions) and their equivalents. It should be noted that the implementation methods provided in the embodiments of this application can be combined with each other without contradiction.
[0017] Before describing the technical solutions provided in the embodiments of this application, in order to facilitate understanding of the embodiments of this application, this application first specifically explains the problems existing in the related technologies: In emergency management scenarios involving high-rise buildings and other densely populated areas, fires spread rapidly, and the indoor environment changes in an instant. During rescue operations, real-time information on the distribution and movement of people inside the building allows for the rational allocation of rescue resources. Providing trapped individuals with personalized escape guidance during the escape process, helping them avoid crowded areas and fire hotspots, will greatly increase the success rate of their escape.
[0018] Today, large public spaces and high-rise buildings are generally equipped with fire prevention and control systems. Common measures include escape warning voice prompt systems and lighting guidance systems, which can guide people to escape to a certain extent. However, while current fire prevention and control systems can provide voice or lighting guidance, they cannot adjust path priorities according to the real-time environment. This leads to a "one-size-fits-all" approach that may guide people towards fire or congested areas, resulting in blind spots in escape routes that appear safe but are actually dangerous. This limitation reduces the reliability of path planning in emergency scenarios and, because it cannot dynamically adapt to changes in the environment and the needs of people, indirectly increases the risk of casualties in fires and other emergencies.
[0019] Based on this, the embodiments of this application provide a method, device, equipment, medium and program product for determining an escape route. Through multiple path alternatives and a scientific path scoring mechanism, it can ensure that the escape plan always remains optimal in a dynamic environment, significantly reduce the risk of casualties caused by path selection errors, ensure the optimality and reliability of the escape plan, significantly improve the safety assurance capability of path planning in emergency scenarios, and effectively reduce the risk of casualties in emergency events such as fires.
[0020] The method for determining the escape route provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0021] Figure 1 This is a schematic flowchart of a method for determining an escape route provided in an embodiment of this application, as shown below. Figure 1 As shown, the method may include steps S110 to S160.
[0022] S110, in the event of an emergency within a building, acquires the location information of each user within the building and the environmental information within the building.
[0023] Emergency events refer to events that require users to evacuate quickly and safely, such as fires and earthquakes.
[0024] Environmental information within a building can include environmental information such as temperature and smoke concentration in various sub-areas of the building.
[0025] Specifically, in the event of an emergency within a building, real-time location information of each user within the building can be obtained through positioning devices, and environmental information such as temperature and smoke concentration in various sub-areas within the building can be obtained through sensors.
[0026] S120 determines the regional status information of each sub-area within the building based on the location and environmental information of each user.
[0027] The sub-region's status information can include its location, the presence or absence of obstacles, its temperature, smoke concentration, and congestion level. The sub-region's location is constant and can therefore be pre-entered into the system.
[0028] Specifically, in the event of an emergency within a building, it is necessary to determine the real-time status information of each sub-area within the building to facilitate the planning of escape routes. For example, based on the real-time location information of each user and the real-time environmental information, the status information of each sub-area within the building (whether there are obstacles in the area, the temperature information in the area, the smoke concentration in the area, the degree of congestion in the area, etc.) can be determined for subsequent planning of escape routes.
[0029] S130: Based on the regional status information of each sub-region, determine multiple candidate escape paths for the target user.
[0030] The regional status information of each sub-region can reflect whether the sub-region is safe. Therefore, based on the regional status information of each sub-region, the candidate escape routes for the target user can be determined.
[0031] Specifically, multiple path generation rules can be set for candidate escape paths, thereby using multiple path generation rules to determine multiple candidate escape paths for the target user based on the regional status information of each sub-region.
[0032] S140, Obtain the safety indicator data corresponding to the candidate escape routes.
[0033] Specifically, after identifying multiple candidate escape routes for the target user, it is necessary to select the optimal escape route from these routes. This requires first obtaining the safety indicator data corresponding to the candidate escape routes. For example, safety indicator data may include escape time, path length, number of times obstacles are avoided, number of times fire points are avoided, number of times congestion points are passed, and path smoothness, etc.
[0034] S150: Based on safety index data, the candidate escape routes are scored to obtain the path score for each candidate escape route.
[0035] Specifically, candidate escape routes can be scored based on safety indicators such as escape time, route length, number of times obstacles are avoided, number of times fire points are avoided, number of times congestion points are passed, and route smoothness, thus obtaining a route score for each candidate escape route.
[0036] S160, determine the candidate escape path corresponding to the highest path score as the target escape path.
[0037] Specifically, after scoring each candidate escape route, the candidate escape route with the highest score can be selected as the target escape route for the target user.
[0038] It should be noted that after the target escape route is determined, S130~S160 can be repeatedly executed at certain time intervals during the target user's escape process according to the target escape route, or S130~S160 can be repeatedly executed when certain triggering events occur (for example, the degree of crowding, temperature, or smoke concentration in a certain key node / key area of the route exceeds a certain threshold), so as to replan the target user's escape route and ensure that the escape plan always remains optimal in the dynamic environment.
[0039] According to the escape route determination method provided in this application embodiment, in the event of an emergency within a building, the location information of each user within the building and the environmental information within the building are obtained. Then, based on the location information of each user and the environmental information, the regional status information of each sub-area within the building can be determined. Next, based on the regional status information of each sub-area, multiple candidate escape routes can be planned for the target user. Then, based on the safety index data corresponding to the candidate escape routes, the candidate escape routes are scored to obtain the path score of each candidate escape route. Thus, the candidate escape route corresponding to the highest path score can be determined as the target escape route, so that the target user can escape according to the target escape route. This application embodiment, through multiple path alternatives and a scientific path scoring mechanism, can ensure that the escape plan always remains optimal in a dynamic environment, significantly reducing the risk of injury or death caused by path selection errors, ensuring the optimality and reliability of the escape plan, significantly improving the safety assurance capability of path planning in emergency scenarios, and effectively reducing the risk of casualties in emergencies such as fires.
[0040] Taking the emergency escape system for fires in park buildings as an example, such as Figure 2 As shown, the system includes Bluetooth base station equipment, smartphones, remote servers, and fire environment monitoring equipment. By integrating data from various sensors, including temperature, smoke, and Bluetooth signals, the system can achieve comprehensive environmental monitoring and personnel location.
[0041] For example, the emergency escape system supports integration with existing fire alarm systems and intelligent building management systems, improving system compatibility and scalability. For instance, through real-time data collection and personnel distribution information, the rescue command center can dynamically adjust the allocation of rescue resources based on the actual situation, ensuring that resources are rationally and efficiently distributed to where they are most needed. Precise personnel location and environmental monitoring data can help the command center quickly identify high-risk areas and densely populated areas, optimizing rescue strategies.
[0042] The remote server can determine the location information of each user in the building through the user's smartphone and Bluetooth base station equipment. The remote server can also collect environmental information in the building through fire environmental monitoring equipment.
[0043] Low-power Bluetooth base station devices are deployed in the buildings of the park to construct a virtual space map. Buildings are divided into three categories: independent rooms, enclosed corridors, and open spaces, with Bluetooth base station devices deployed in designated areas. Each area is at least 1 square meter, and the Bluetooth base station devices are deployed in open, unobstructed areas to create a virtual space map of the building's Bluetooth base station device distribution. The virtual space map also includes security areas, such as first-floor emergency exits and fire escape floors.
[0044] When a fire occurs, trapped individuals open a smartphone app to collect signals from nearby Bluetooth Low Energy base stations and send them to a remote server. Environmental data such as temperature and smoke concentration are transmitted to the remote server via fire environmental monitoring equipment. The remote server calculates the user's location using the following algorithm, thereby generating a real-time distribution map of people in the building.
[0045] When the Bluetooth base station device with the strongest Bluetooth signal is located in an independent room or enclosed corridor, and the user's location is in that area (x, y), the user's location positioning accuracy is the location within that area.
[0046] When the device with the strongest Bluetooth signal acquired by the mobile phone is in an open space, an optimized triangulation centroid algorithm is used for position calculation: Assuming the coordinates of the three Bluetooth base station devices with the strongest signals in the nearby open space are (x1, y1), (x2, y2), and (x3, y3), then the coordinates of the phone's location Z(x, y) are: The distance di is calculated as follows: Where P is the received signal strength; P( ) is the reference value of signal strength at a distance of 1m; n is the signal attenuation coefficient, which is generally 2; Xσ is a Gaussian distribution with variance σ.
[0047] w represents the weight coefficient of each Bluetooth base station device, and the weight wi is calculated using the following formula: in, It is a small positive number to prevent the weight from being too large when the distance di is too small.
[0048] It should be noted that, according to the inventors' research, mobile terminals (such as smartphones) in related technologies convert Bluetooth signal strength into distance information with nearby Bluetooth base station devices based on an indoor signal attenuation model using Received Signal Strength Indication (RSSI). The three nearest nearby Bluetooth base station devices are selected, and the mobile terminal (such as a smartphone) is positioned within a triangle formed by these three devices. Compared to related technologies, this application's embodiment performs Kalman filtering and smoothing on the received signal strength before calculating the Bluetooth base station device weights to reduce the impact of signal noise. Furthermore, a weighted average of the signal strength is introduced in the weight calculation, and for each Bluetooth base station device, its weight coefficient is dynamically adjusted using historical data. This improves the accuracy and stability of the triangular centroid algorithm, providing more accurate positioning results for mobile phone Bluetooth positioning (user).
[0049] It should be noted that, according to the inventors' research, while indoor personnel location can be achieved using Wi-Fi signal strength and fingerprint positioning technology—Wi-Fi fingerprint positioning determines the user's location by matching a pre-collected Wi-Fi signal strength fingerprint database with real-time signal strength data—the accuracy of Wi-Fi positioning is generally lower than that of low-power Bluetooth base station devices, especially in complex indoor environments where it is susceptible to signal interference and multipath effects. Furthermore, establishing and maintaining a Wi-Fi signal fingerprint database requires significant initial deployment and subsequent updates. In the event of a power outage due to a fire, Wi-Fi devices will be inoperable.
[0050] This application embodiment acquires the real-time distribution of escaping personnel through a low-power Bluetooth base station device and utilizes a high-precision positioning calculation method to ensure that the location and distribution of people within a building can be quickly and accurately determined during a fire. Real-time personnel distribution information helps rescue personnel more accurately locate trapped individuals, improving rescue efficiency.
[0051] The path scoring mechanism in the escape path determination method provided in the embodiments of this application is described below: In some embodiments, such as Figure 3 As shown, the safety indicator data includes multiple sub-safety indicator data; step S150 scores the candidate escape routes based on the safety indicator data to obtain the path score for each candidate escape route, which may include step S151: S151, multiply each sub-safety indicator data by its corresponding preset weight and then sum them to obtain the path score of the candidate escape path.
[0052] In one example, the sub-safety metric data includes at least one of the following: escape time, path length, number of times obstacles were avoided, number of times fire points were avoided, number of times congestion points were passed, and path smoothness.
[0053] In yet another example, the formula for calculating the path score is as follows: Where V is the path score, T is the escape time, L is the path length, O is the number of times obstacles and fire points are avoided, C is the number of times congestion points are passed, S is the path smoothness, and W is the path rating. T W L W O W C W S Preset weights.
[0054] For example, the sub-safety index data for candidate escape route 1 are: escape time Tcurrent = 350 seconds, path length Lcurrent = 600 meters, number of times to avoid obstacles and fire points Ocurrent = 8, number of times to pass through congested points Ccurrent = 5 times, time 100 seconds, and path smoothness Scurrent = 0.7.
[0055] The sub-safety index data for candidate escape route 2 are as follows: escape time Topt = 300 seconds, path length Lopt = 550 meters, number of times to avoid obstacles and fire points Oopt = 10, number of times to pass through congested points Copt = 2, time 50 seconds, and path smoothness Sopt = 0.9.
[0056] The weights of each indicator are set as follows: WT=0.4, WL=0.3, WO=0.1, WC=0.1, WS=0.1.
[0057] (4) Calculate path score: The path score for candidate escape route 1 is: The path score for candidate escape route 2 is: Compare the path scores and select the path with the higher score as the final escape route to ensure that the final route is safer, more efficient and secure.
[0058] This application provides a method that combines multi-objective optimization technology to comprehensively evaluate escape routes and return high-scoring routes to the user's mobile phone, thereby improving the safety and effectiveness of escape routes.
[0059] The following describes the specific process of generating multiple candidate escape paths in the escape path determination method provided in the embodiments of this application.
[0060] In some embodiments, such as Figure 4 As shown, the candidate escape paths include a first candidate escape path and a second candidate escape path; step S130 determines multiple candidate escape paths for the target user based on the area status information of each sub-area, and may include steps S131~S136: S131, Determine the location of the target user as the starting node; S132, determine the location of the safe zone as the termination node; S133, Based on the regional status information of each sub-region, determine the first intermediate node group between the starting node and the ending node; S134, Based on the starting node, the first intermediate node group and the ending node, generate the first candidate escape path for the target user; S135, Based on the regional status information of each sub-region, determine the second intermediate node group between the starting node and the ending node; S136, based on the starting node, the second intermediate node group, and the ending node, generate the second candidate escape path for the target user.
[0061] Specifically, the starting node of each candidate escape path is the target user's current location, and the ending node is the location within the safe zone. By using different escape path generation rules, multiple intermediate node groups can be obtained. For example, using two different escape path generation rules can yield a first intermediate node group and a second intermediate node group. Therefore, based on the starting node, the first intermediate node group, and the ending node, a first candidate escape path for the target user can be generated, and based on the starting node, the second intermediate node group, and the ending node, a second candidate escape path can be generated. Then, the path scores for the first and second candidate escape paths are calculated, and the candidate escape path with the higher path score is determined as the target user's target escape path.
[0062] For example, the escape path has two generation rules: Escape path generation rule one: Find the intermediate node with the lowest replacement value through the cost function f(n) to obtain the first candidate escape path.
[0063] Escape path generation rule two: Calculate the action score Q using the model, find the intermediate node with the largest Q, and obtain the second candidate escape path.
[0064] This application embodiment defines the target user's location as the starting node and the safe zone as the ending node, uses different escape path generation rules to determine the first and second intermediate node groups, and then generates the first and second candidate escape paths. The path score is then calculated to select the optimal target escape path. This approach ensures the consistency of the start and end points of multiple candidate paths and generates diverse intermediate node groups using different rules. It can comprehensively consider multiple factors and provide target users with more reasonable and better escape path selection in different scenarios, thereby improving the escape success rate and safety.
[0065] The following describes the specific implementation process of the escape path generation rule one (finding the intermediate node with the lowest cost through the cost function f(n) to determine the first intermediate node group).
[0066] In some embodiments, such as Figure 5As shown, the first intermediate node group includes a first intermediate node and a second intermediate node; the area status information includes area location, area temperature, area smoke concentration, and area population density; step S133 determines the first intermediate node group between the starting node and the ending node based on the area status information of each sub-area, which may include steps S1331~S1333: S1331, Given that the first intermediate node has been determined, determine multiple candidate second intermediate nodes that are adjacent to the first intermediate node based on the regional location; S1332, determine the cost value corresponding to the candidate second intermediate node based on the regional location, regional temperature, regional smoke concentration and regional population density; S1333, among multiple candidate second intermediate nodes, determine the candidate second intermediate node corresponding to the minimum cost value as the second intermediate node.
[0067] The first intermediate node refers to the i-th intermediate node in the escape path.
[0068] The second intermediate node refers to the (i+1)th intermediate node in the escape path, which is the node following the i-th intermediate node.
[0069] Specifically, the rule for determining the next intermediate node (second intermediate node) after the first intermediate node in the escape path is as follows: Given the first intermediate node of the escape path, multiple candidate second intermediate nodes adjacent to the first intermediate node can be found based on the regional location. For example, any node within 50 meters of the first intermediate node can be considered a candidate second intermediate node. After finding multiple candidate second intermediate nodes, the safest node needs to be selected as the second intermediate node of the escape path (i.e., the node next to the first intermediate node in the escape path). Specifically, the cost value corresponding to each candidate second intermediate node can be calculated based on the regional location, regional temperature, regional smoke concentration, and regional personnel density of the sub-region where the candidate second intermediate node is located. Then, among the multiple candidate second intermediate nodes, the candidate second intermediate node with the minimum cost value is selected as the second intermediate node.
[0070] This application's embodiments determine candidate second intermediate nodes adjacent to the first intermediate node based on regional location, and then calculate the cost value of the candidate second intermediate nodes by comprehensively considering regional location, temperature, smoke concentration, and personnel density. Finally, the node with the lowest cost value is selected as the second intermediate node. This approach can accurately consider multiple environmental factors, providing a reliable basis for generating safe and reasonable escape routes, and effectively improving the scientific and practical nature of escape route planning.
[0071] In some embodiments, such as Figure 6As shown, step S1332 determines the cost value corresponding to the candidate second intermediate node based on the area location, area temperature, area smoke concentration, and area population density, and may include S13321~S13323: S13321, Calculate the first distance between the candidate second intermediate node and the first intermediate node based on the region location. S13322, Calculate the second distance between the candidate second intermediate node and the termination node based on the region location; S13323, multiply the first distance, the second distance, the regional temperature, the regional smoke concentration, and the regional population density by the corresponding preset influence factors to obtain the cost value corresponding to the candidate second intermediate node.
[0072] For example, the A* algorithm is used as the cost function f(n): Where f(n) is the cost value, g(n) is the first distance, h(n) is the second distance, α, β, and γ are preset influence factors, and ensity(n), temperature(n), and smoke(n) represent population density, temperature, and smoke concentration, respectively. It should be noted that, according to the inventors' research, related technologies calculate evacuation routes from within a building based on the user's location information and the location and extent of the fire, using a shortest path algorithm, and then send these routes to the mobile terminal. In contrast, the embodiments of this application use the A* algorithm to generate initial escape routes, incorporating parameters such as personnel density, temperature, and smoke concentration into the cost function. This results in more effective escape routes; that is, the cost function integrates multiple factors such as personnel density, temperature, and smoke concentration, thus improving the accuracy of the routes.
[0073] In calculating the cost value of a candidate second intermediate node, this embodiment first calculates its distance to the first intermediate node and the termination node, then multiplies the distance by the area temperature, smoke concentration, personnel density, and corresponding preset influencing factors to obtain the cost value. Furthermore, it employs the A* algorithm as the cost function, incorporating a comprehensive consideration of multiple parameters. Compared to related technologies that only calculate evacuation routes based on the shortest path algorithm, this embodiment more comprehensively integrates various environmental factors, resulting in more effective and accurate escape routes and providing a more reliable guarantee for escape.
[0074] In some embodiments, such as Figure 7 As shown, step S13321 calculates the first distance between the candidate second intermediate node and the first intermediate node based on the region location, which may include steps S133211~S133213: S133211, Based on the regional location, determine the search area between the candidate second intermediate node and the first intermediate node; S133212, In the case that there are no obstacles in the search area, the straight-line distance between the candidate second intermediate node and the first intermediate node is determined as the first distance; S133213, In the case of obstacles in the search area, determine the shortest distance between the candidate second intermediate node and the first intermediate node to bypass the obstacles as the first distance.
[0075] For example, a rectangular or circular area can be defined between two nodes as the search region.
[0076] Simply put, if there are no obstacles, walk in a straight line and the first distance is the straight-line distance; if there are obstacles, go around them and find the first distance, which is the shortest detour.
[0077] In calculating the first distance between the candidate second intermediate node and the first intermediate node, this embodiment first determines the search area based on the location of the area, and then determines the straight-line distance or the shortest detour distance as the first distance in two cases: with or without obstacles. This processing method takes into account all aspects and can flexibly and accurately calculate the distance according to the actual scenario, laying a solid foundation for generating reasonable and accurate escape routes in the future, and effectively improving the reliability and practicality of escape route planning.
[0078] In one example, such as Figure 8 As shown, the specific process of determining the next intermediate node (second intermediate node) of the first intermediate node in the escape path may include steps S01 to S11.
[0079] S01, Construct a building grid map.
[0080] The entire building is divided into multiple grids (sub-areas) based on its structure.
[0081] S02, Key area marking: Determine the start node and end node and mark the start node as X_s, the end node as X_e, and the intermediate node as X_k.
[0082] S03, determine an intermediate node in the escape route.
[0083] For example, based on the location of the area, multiple candidate first intermediate nodes adjacent to the starting node are determined, and then the candidate first intermediate node with the lowest cost value is selected as the first intermediate node of the escape path from the multiple candidate first intermediate nodes.
[0084] S04. Given that the first intermediate node has been determined, determine multiple candidate second intermediate nodes that are adjacent to the first intermediate node based on the regional location.
[0085] S05, Based on the regional location, determine the search area between the candidate second intermediate node and the first intermediate node.
[0086] S06, Determine whether there are obstacles in the search area.
[0087] S07, if there are no obstacles in the search area, determine the straight-line distance between the candidate second intermediate node and the first intermediate node as the first distance.
[0088] S08, if there are obstacles in the search area, determine the shortest distance between the candidate second intermediate node and the first intermediate node to bypass the obstacles as the first distance.
[0089] S09, Calculate the second distance between the candidate second intermediate node and the termination node based on the region location.
[0090] S10, multiply the first distance, the second distance, the regional temperature, the regional smoke concentration, and the regional population density by the corresponding preset influence factors to obtain the cost value corresponding to the candidate second intermediate node.
[0091] S11, among multiple candidate second intermediate nodes, determine the candidate second intermediate node corresponding to the minimum cost value as the second intermediate node.
[0092] The embodiments of this application comprehensively consider building structure, obstacles and various environmental factors, and can accurately plan safe and reasonable escape routes, effectively improving the success rate and safety of escape.
[0093] The following describes the specific implementation of the second escape path generation rule (calculating the action score Q through the model, finding the intermediate node with the largest Q, and obtaining the second candidate escape path).
[0094] In some embodiments, such as Figure 9 As shown, the second intermediate node group includes the first intermediate node and the second intermediate node; the area status information includes the area location, whether there are obstacles in the area, the area temperature, the area smoke concentration, and the area personnel density; step S135 determines the second intermediate node group between the starting node and the ending node based on the area status information of each sub-area, which may include steps S1351~S1353: S1351, Given that the first intermediate node has been determined, determine multiple candidate second intermediate nodes that are adjacent to the first intermediate node based on the regional location. S1352, using the pre-trained action scoring model, determine the action score value (denoted as Q) of action a corresponding to the first intermediate node to the candidate second intermediate node based on the area location, whether there are obstacles in the area, the area temperature, the area smoke concentration, and the area personnel density. S1353, among multiple candidate second intermediate nodes, determine the candidate second intermediate node corresponding to the maximum action score value Qmax as the second intermediate node.
[0095] Specifically, the regional state information of the first intermediate node and the regional state information of each candidate second intermediate node (such as regional location, presence of obstacles, regional temperature, regional smoke concentration, and regional personnel density) can be input into the trained action scoring model. The action scoring model will then output the action score values corresponding to the actions from the first intermediate node to each candidate second intermediate node. This allows us to determine the candidate second intermediate node corresponding to the maximum action score value as the second intermediate node. Then, we use this second intermediate node as the new first intermediate node and continue to determine the second intermediate node corresponding to the first intermediate node. This process continues until the second intermediate node becomes the termination node, thus obtaining the second candidate escape path.
[0096] For example, the action scoring model is a deep Q-network containing multiple convolutional neural networks (CNNs) to extract features from grid state information (i.e., region state information). The input layer of the model is grid state information (region state information); the hidden layers include several convolutional layers and fully connected layers; the output layer is the action score value Q corresponding to each action a.
[0097] The trained action scoring model in this embodiment can comprehensively consider complex information from multiple dimensions such as location, presence of obstacles, temperature, smoke concentration, and personnel density to scientifically evaluate the actions from the first intermediate node to each candidate second intermediate node and output a score value. By selecting the node with the highest score value, the second candidate escape path is gradually constructed, which can accurately and quickly generate a safer and more reasonable escape path.
[0098] The following describes the training process of the action scoring model. The purpose of model training is to make the model's output (action score value) more accurate.
[0099] In some embodiments, such as Figure 10 As shown, before step S1352, which uses the trained action scoring model to determine the action score value of the action from the first intermediate node to the candidate second intermediate node based on the area location, whether there are obstacles in the area, the area temperature, the area smoke concentration, and the area personnel density, the method may also include S171 to S174.
[0100] S171, Establish an emergency scenario simulation environment for the building. The emergency scenario simulation environment includes the simulated area status information of multiple sub-areas within the building (such as area location, whether there are obstacles in the area, area temperature, area smoke concentration, area personnel density, etc.).
[0101] For example, the process of establishing a fire scenario simulation environment is as follows: 1) The fire scenario simulation environment includes building structure, fire point, congestion point, and good areas. The environment can simulate different fire scenarios and personnel distribution.
[0102] 2) Gridded Environment: The environment is divided into grids, with each grid representing a state. Each state includes information such as the current grid's position, whether it contains an obstacle, the level of congestion, temperature, and smoke concentration.
[0103] 3) State: The state information of each grid (area state information), including location, obstacles, crowding level, temperature, smoke concentration, etc.
[0104] For example, the environment is divided into grids, and the state of each grid (region state) is represented as follows: .
[0105] 4) Action: Actions that the agent can take, such as moving up, down, left, or right.
[0106] S172, given that the first historical intermediate node has been determined, the initial action score value of the action corresponding to each historical candidate second intermediate node is calculated using the initial action scoring model based on the state information of the simulated area.
[0107] S173, obtain the reference action score value of the action corresponding to each historical candidate second intermediate node from the first historical intermediate node.
[0108] The reference action score is the expected output of the model.
[0109] S174. Based on the scoring error between the reference action score and the initial action score, the initial action scoring model is trained to obtain a trained action scoring model. The scoring error of the trained action scoring model is less than the preset error value.
[0110] Specifically, the initial action score values calculated using the initial action score model for the actions corresponding to the first historical intermediate node and each historical candidate second intermediate node are not accurate enough. Therefore, it is necessary to obtain reference action score values and use them as the expected output of the model to optimize the parameters in the model until the model output is close to the expected output.
[0111] This application embodiment constructs a simulated emergency building scenario environment containing rich simulated area state information, meticulously divides the grid and clarifies the state and actions, calculates the initial action score value using an initial action scoring model, and then introduces a reference action score value. The model is trained based on the error between the two, which can effectively improve the accuracy of the action scoring model. This allows the model to more scientifically and accurately consider multiple factors such as area location, obstacles, and temperature when determining the action score value from the first intermediate node to the candidate second intermediate node, providing strong support for generating a safe and reasonable escape route.
[0112] In some embodiments, such as Figure 11 As shown, step S173 obtains the reference action score value of the action corresponding to each historical candidate second intermediate node from the first historical intermediate node, which may include steps S1731 and S1732.
[0113] S1731, when the second intermediate node of the historical candidate is the termination node, obtain the instant reward value corresponding to the second intermediate node of the historical candidate; determine the instant reward value as the reference action score value of the action corresponding to the action from the first historical intermediate node to the second intermediate node of the historical candidate.
[0114] The instant reward value (which can be denoted as R) refers to the reward value of a node. For example, the node reward value of a historical candidate second intermediate node is the instant reward value of the historical candidate second intermediate node.
[0115] For example, the immediate reward function (node reward function) is: Where R is the instant reward value; a high reward of R is given for reaching the safe exit. exit Entering the fire site will result in a high penalty. fire Appropriate penalties will be imposed for entering congested areas. crowded Appropriate rewards will be given for entering the excellent area. safe Entering the normal area will grant a small reward (R). normal .
[0116] S1732, if the second intermediate node of the historical candidate is not the termination node, obtain the immediate reward value and future reward value corresponding to the second intermediate node of the historical candidate; based on the immediate reward value and future reward value, determine the reference action score value of the action corresponding to the action from the first historical intermediate node to the second intermediate node of the historical candidate.
[0117] For example, the future reward value corresponding to the historical candidate second intermediate node can be obtained based on the node reward value of the next node of the historical candidate second intermediate node.
[0118] This application embodiment, by comprehensively considering the safety of the current node and subsequent paths, enables the trained action scoring model to be more scientific and reasonable, thereby providing accurate basis for generating safer and more efficient escape routes.
[0119] In some embodiments, determining the reference action score value for the action corresponding to the first historical intermediate node to the historical candidate second intermediate node based on the immediate reward value and the future reward value in step S1732 may include: Among them, Q ref r is the reference action score for the action corresponding to the action from the first historical intermediate node to the second historical candidate intermediate node. i+1 r represents the instantaneous reward value corresponding to the second intermediate node in the historical candidate list. i+2 max r is the instantaneous reward value corresponding to the third intermediate node of the historical candidate that is adjacent to the second intermediate node of the historical candidate. i+2 For future reward values (i.e., from multiple r) j+2 Find the largest rj+2 as the future reward value. A discount factor (0≤γ≤1) is used to weigh the importance of immediate rewards against future rewards.
[0120] For example, the reference action score Q for the action corresponding to the action from the first historical intermediate node to the second historical candidate intermediate node. ref The calculation formula is: In the case where the second intermediate node of the historical candidate is the termination node, the instantaneous reward value r corresponding to the second intermediate node of the historical candidate is obtained. i+1 Determine the instant reward value r i+1 The reference action score Q is the action score corresponding to the action from the first historical intermediate node to the second historical candidate intermediate node. ref .
[0121] If the second intermediate node in the historical candidate is not the termination node, obtain the instantaneous reward value r corresponding to the second intermediate node in the historical candidate. i+1 and future reward value r i+2 Based on the instant reward value r i+1 and future reward value r i+2 Determine the reference action score Q for the action corresponding to the action from the first historical intermediate node to the second historical candidate intermediate node. ref .
[0122] In this embodiment, when obtaining the reference action score, the system handles the historical candidate second intermediate node differently depending on whether it is a termination node. If it is a termination node, the score is determined directly by the immediate reward value. If it is not a termination node, the score combines the immediate reward value with the future reward value obtained based on subsequent nodes, and a discount factor is introduced to weigh the importance of both. This meticulous consideration of current and future benefits allows the reference action score to more comprehensively and realistically reflect the value of the action, thereby making the trained action scoring model more scientific and reasonable, and providing accurate and powerful support for generating safe and efficient escape paths.
[0123] In some embodiments, such as Figure 12 As shown, obtaining the instant reward value corresponding to the second intermediate node of the historical candidate may include steps S181 to S183.
[0124] S181, Based on the regional status information corresponding to the historical candidate second intermediate node, determine the target access feasibility level corresponding to the historical candidate second intermediate node.
[0125] For example, areas with normal monitoring data and people passing by within the last minute are considered excellent areas (good passability) with a high level of passability; areas with no monitoring data or no one passing by recently are considered ordinary areas (average passability) with a relatively high level of passability; areas with dense population distribution and low mobility are considered crowded areas (poor passability) with a low level of passability; and areas with excessively high temperatures or high smoke concentrations are considered obstacle / fire areas (impassable) with a low level of passability.
[0126] S182, in the preset correspondence between accessibility levels and instant reward values, query the target instant reward value corresponding to the target accessibility level.
[0127] For example, the preset correspondence between the accessibility level and the instant reward value is as follows: reaching the safe exit (very high accessibility level) grants a high reward R. exit Entering a fire site (with a low feasibility level) will result in a high penalty (R). fire Entering congested areas (where the feasibility level is low) will result in a moderate penalty. crowded Entering a superior area (high feasibility level) will be given an appropriate reward. safe Entering a normal area (with a relatively high feasibility level) will grant a small reward of R. normal .
[0128] S183, determine the target instant reward value as the instant reward value corresponding to the second intermediate node of the historical candidate.
[0129] Specifically, the process of determining the instant reward value (the reward value of a single node) is as follows: When obtaining the instant reward value corresponding to the historical candidate second intermediate node, firstly, based on the monitoring data, personnel distribution, temperature, and smoke concentration of the area where the node is located, it is divided into different accessibility levels, such as excellent, normal, congested, and obstacle / fire areas, corresponding to different levels of accessibility; then, in the preset correspondence between accessibility levels and instant reward values, the target instant reward value corresponding to the target accessibility level of the node is found; finally, this target instant reward value is determined as the instant reward value corresponding to the historical candidate second intermediate node, thereby reasonably quantifying the node value.
[0130] For example, regions can be labeled after they are categorized: Fire point (denoted as X_d): Areas with excessively high temperature or high smoke concentration as shown by environmental monitoring data are marked as fire points and are not allowed to pass through.
[0131] Crowded areas (denoted as X_c): Areas with dense population distribution and low mobility are marked as crowded areas, indicating poor passability.
[0132] Normal area (denoted as X_n): Areas with no monitoring data or no one passing through recently are marked as normal areas, with average passability.
[0133] Excellent area (denoted as X_g): Areas with normal monitoring data and where someone has passed through within the last minute are marked as excellent areas, indicating good passability.
[0134] This application embodiment determines the feasibility level of historical candidate second intermediate nodes by comprehensively considering regional status information, and then obtains the instant reward value based on the preset correspondence. This can accurately and meticulously quantify the value of nodes, so that the instant reward value fully reflects the safety and accessibility of nodes in the escape path, providing a reliable and reasonable basis for the training of subsequent action scoring models and the generation of safe and efficient escape paths.
[0135] In one example, the training process for the action scoring model is as follows: 1) Obtain the initial action scoring model. The output of the initial action scoring model is denoted as Q(s, a; θ), where θ is the parameter of the initial model.
[0136] 2) Obtain the expected output of the action scoring model, denoted as Q′(s, a; θ-), where θ- represents the parameters of the target action scoring model. Q′ is the known expected output, but θ- is unknown; therefore, the parameters θ- of the target action scoring model need to be obtained through continuous training.
[0137] 3) Training parameter θ.
[0138] At each time step t, according to A greedy strategy selects action at, i.e., with probability. Randomly select an action with a probability of 1 Select the action with the largest Q-value in the current Q-network. Execute the action at and observe the immediate reward rt and the next node state st+1. Store the sample (st, at, rt, st+1) in buffer D. Using the experience replay technique, randomly sample mini-batch samples from the experience replay buffer for training (sj, aj, rj, sj+1), where rj determines the reference action score Qref (see Equation 8).
[0139] 4) The parameters θ- of the target action scoring model are obtained by training by minimizing the loss function L(θ), that is, the target action scoring model is trained.
[0140] For example, using the Adam optimizer to minimize the loss function L(θ), the loss function L(θ) is: If the loss value of the loss function L(θ) is less than a certain threshold, it indicates that the model training is complete.
[0141] For example, model parameters can be updated periodically to stabilize the training process and produce a stable model.
[0142] In the embodiments of this application, model training is performed using... The greedy action selection strategy effectively balances the relationship between exploration and exploitation, avoids getting trapped in local optima, and can accurately find the optimal model parameters by continuously minimizing the loss function. This results in a high-performance, stable, and reliable action scoring model, which provides strong support for subsequent decisions such as generating high-quality escape paths.
[0143] In one example, a motion scoring model can be used to optimize the first candidate escape path to obtain a second candidate escape path. For instance, key path nodes can be extracted from the first candidate escape path as reference points for generating the second candidate escape path. By using a trained motion scoring model to optimize the existing first candidate escape path, a more efficient and safer optimal escape path can be generated based on the original path.
[0144] It should be noted that the escape route determination system can be applied to various high-rise buildings, commercial complexes, industrial parks, public facilities, and other large public places. It is particularly suitable for places requiring high security, such as schools, hospitals, and office buildings, meeting various emergency management needs. Through real-time positioning, personalized escape guidance, dynamic resource allocation, and user feedback mechanisms (for example, during the escape process, user feedback is collected through a mobile application, recording escape time, path deviation, and other data. The server continuously optimizes and adjusts the path planning algorithm based on user feedback, improving the system's usability and safety), it comprehensively improves the efficiency of fire escape and rescue, meeting the market's urgent demand for efficient emergency management systems.
[0145] Based on the same inventive concept, embodiments of this application also provide an escape route determination device, such as... Figure 11 As shown, the device 1100 may include an acquisition module 1110, a determination module 1120, and a scoring module 1130: The acquisition module 1110 is used to acquire the location information of each user in the building and the environmental information of the building in the event of an emergency. The determination module 1120 is used to determine the regional status information of each sub-area within the building based on the location information and environmental information of each user; The determination module 1120 is also used to determine multiple candidate escape paths for the target user based on the regional status information of each sub-region; The acquisition module 1110 is also used to acquire safety indicator data corresponding to the candidate escape paths; The scoring module 1130 is used to score the candidate escape routes based on the safety index data corresponding to the candidate escape routes, and obtain the path score of each candidate escape route. The determination module 1120 is also used to determine the candidate escape path corresponding to the highest path score as the target escape path.
[0146] In some embodiments, the safety indicator data includes multiple sub-safety indicator data; the scoring module is used to score the candidate escape routes based on the safety indicator data, obtaining a path score for each candidate escape route, specifically for: The path score of the candidate escape route is obtained by multiplying each sub-safety index data by its corresponding preset weight and then summing them.
[0147] In some embodiments, the sub-safety indicator data includes at least one of escape time, path length, number of times obstacles are avoided, number of times fire points are avoided, number of times congestion points are passed, and path smoothness.
[0148] In some embodiments, the candidate escape paths include a first candidate escape path and a second candidate escape path; the determining module is used to determine multiple candidate escape paths for the target user based on the area status information of each sub-area, specifically for: The starting node is determined by the location of the target user. The location of the safe zone is determined as the termination node; Based on the regional status information of each sub-region, determine the first intermediate node group between the start node and the end node; Based on the starting node, the first intermediate node group, and the ending node, generate the first candidate escape path for the target user. Based on the regional status information of each sub-region, determine the second intermediate node group between the start node and the end node; Based on the starting node, the second intermediate node group, and the ending node, a second candidate escape path for the target user is generated.
[0149] In some embodiments, the first intermediate node group includes a first intermediate node and a second intermediate node; the area status information includes area location, area temperature, area smoke concentration, and area population density; the determining module is used to determine the first intermediate node group between the starting node and the ending node based on the area status information of each sub-area, specifically for: Given that the first intermediate node has been determined, multiple candidate second intermediate nodes adjacent to the first intermediate node are determined based on the regional location. The cost value corresponding to the candidate second intermediate node is determined based on the regional location, regional temperature, regional smoke concentration, and regional population density. Among multiple candidate second intermediate nodes, the candidate second intermediate node corresponding to the minimum cost value is determined as the second intermediate node.
[0150] In some embodiments, the determining module is used to determine the cost value corresponding to the candidate second intermediate node based on the area location, area temperature, area smoke concentration, and area population density, specifically for: Based on the region location, calculate the first distance between the candidate second intermediate node and the first intermediate node. Based on the region location, calculate the second distance between the candidate second intermediate node and the termination node; Multiply the first distance, the second distance, the regional temperature, the regional smoke concentration, and the regional population density by the corresponding preset influence factors to obtain the cost value corresponding to the candidate second intermediate node.
[0151] In some embodiments, the determining module is configured to calculate a first distance between the candidate second intermediate node and the first intermediate node based on the region location, specifically for: Based on the location of the region, determine the search area between the candidate second intermediate node and the first intermediate node; If there are no obstacles in the search area, the straight-line distance between the candidate second intermediate node and the first intermediate node is determined as the first distance; If there are obstacles in the search area, the shortest distance between the candidate second intermediate node and the first intermediate node to bypass the obstacles is determined as the first distance.
[0152] In some embodiments, the second intermediate node group includes a first intermediate node and a second intermediate node; the area status information includes area location, presence or absence of obstacles within the area, area temperature, area smoke concentration, and area population density; the determining module is used to determine the second intermediate node group between the starting node and the ending node based on the area status information of each sub-area, specifically for: Given that the first intermediate node has been determined, multiple candidate second intermediate nodes adjacent to the first intermediate node are determined based on the regional location. Using a pre-trained action scoring model, the action score value of the action corresponding to the first intermediate node to the candidate second intermediate node is determined based on the area location, whether there are obstacles in the area, the area temperature, the area smoke concentration, and the area population density. Among multiple candidate second intermediate nodes, the candidate second intermediate node corresponding to the maximum action score value is determined as the second intermediate node.
[0153] In some embodiments, before the determining module uses a trained action scoring model to determine the action score value of the action corresponding to the action from the first intermediate node to the candidate second intermediate node based on the area location, whether there are obstacles in the area, the area temperature, the area smoke concentration, and the area population density, the device further includes an establishment module, a calculation module, and a training module. The module is used to create an emergency scenario simulation environment for a building. The emergency scenario simulation environment includes the state information of multiple sub-areas within the building. The calculation module is used to determine the first timeline. history In the case of intermediate nodes, the initial action score value of the action corresponding to each historical candidate second intermediate node is calculated based on the simulated area state information using the initial action score model. The acquisition module is also used to acquire reference action score values for actions corresponding to each historical candidate second intermediate node from the first historical intermediate node. The training module is used to train the initial action scoring model based on the scoring error between the reference action score and the initial action score, so as to obtain a trained action scoring model. The scoring error of the trained action scoring model is less than the preset error value.
[0154] In some embodiments, the acquisition module is used to acquire reference action score values for actions corresponding to each historical candidate second intermediate node from the first historical intermediate node, specifically for: If the second intermediate node in the historical candidate is the termination node, obtain the instant reward value corresponding to the second intermediate node in the historical candidate; determine the instant reward value as the reference action score value for the action from the first intermediate node in the historical candidate to the second intermediate node in the historical candidate; If the second intermediate node in the historical candidate is not the termination node, obtain the immediate reward value and future reward value corresponding to the second intermediate node in the historical candidate; based on the immediate reward value and future reward value, determine the reference action score value of the action corresponding to the action from the first historical intermediate node to the second intermediate node in the historical candidate.
[0155] In some embodiments, the determining module is used to determine a reference action score value for the action corresponding to the first historical intermediate node to the historical candidate second intermediate node based on the immediate reward value and the future reward value, specifically for: Where Qref is the reference action score for the action corresponding to the first historical intermediate node to the second historical candidate intermediate node, r i+1 This represents the immediate reward value corresponding to the second intermediate node in the historical candidate list. r is the discount factor. i+2 max r is the instantaneous reward value corresponding to the third intermediate node of the historical candidate that is adjacent to the second intermediate node of the historical candidate. i+2 This is for future reward value.
[0156] In some embodiments, the acquisition module is used to acquire the instant reward value corresponding to the historical candidate second intermediate node, specifically for: Based on the regional status information corresponding to the historical candidate second intermediate nodes, determine the target access feasibility level corresponding to the historical candidate second intermediate nodes; Within the preset correspondence between feasibility levels and instant reward values, check... Inquiry The immediate reward value corresponding to the target's accessibility level; The target instant reward value is determined to be the instant reward value corresponding to the second intermediate node of the historical candidates.
[0157] The various modules in the escape path determination device provided in this application embodiment can achieve... Figures 1 to 10 The functions of each step in the method for determining the escape route, and the corresponding technical effects they achieve, will not be elaborated here for the sake of brevity.
[0158] Figure 12 A schematic diagram of the hardware structure of the escape path determination device provided in an embodiment of this application is shown.
[0159] The device for determining the escape route may include a processor 1201 and a memory 1202 storing computer program instructions.
[0160] Specifically, the processor 1201 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0161] Memory 1202 may include mass storage for data or instructions. For example, and not limitingly, memory 1202 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 1202 may include removable or non-removable (or fixed) media. Where appropriate, memory 1202 may be internal or external to a device defining an escape path. In a particular embodiment, memory 1202 is a non-volatile solid-state memory.
[0162] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.
[0163] The processor 1201 reads and executes computer program instructions stored in the memory 1202 to implement any of the escape path determination methods in the above embodiments.
[0164] In one example, the escape route determination device may further include a communication interface 1203 and a bus 1204. For example, Figure 12 As shown, the processor 1201, memory 1202, and communication interface 1203 are connected through bus 1204 and complete communication with each other.
[0165] The communication interface 1203 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0166] This device can execute the escape path determination method in the embodiments of this application based on each unit / component in the escape path determination device, thereby achieving a combination of Figures 1 to 10 The method for determining the escape route is described.
[0167] Furthermore, in conjunction with the escape path determination method in the above embodiments, this application embodiment can provide a computer storage medium for implementation. This computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the escape path determination methods in the above embodiments.
[0168] This application also provides a computer program product in which the instructions, when executed by a processor of an electronic device, cause the electronic device to perform various processes implementing any of the above-described escape path determination method embodiments.
[0169] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0170] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (ROM), flash memory, erasable read-only memory (EROM), floppy disks, compact disc read-only memory (CD-ROM), optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0171] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0172] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0173] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method of determining an escape route, characterized by, The method comprises: In the case of an emergency occurring in a building, obtaining location information of each user in the building and environmental information in the building; According to the location information of each user and the environmental information, determine the area state information of each sub-area in the building; According to the area state information of each sub-area, determine a plurality of candidate escape paths of a target user; Obtain the safety index data corresponding to the candidate escape path; According to the safety index data, score the candidate escape path, and obtain the path score of each candidate escape path; Determine the candidate escape path corresponding to the highest path score as the target escape path.
2. The method of claim 1, wherein, The safety index data includes a plurality of sub-safety index data; according to the safety index data, score the candidate escape path, and obtain the path score of each candidate escape path, comprising: Multiply each sub-safety index data by the corresponding preset weight and add it to obtain the path score of the candidate escape path.
3. The method of claim 2, wherein, The sub-safety index data includes at least one of the escape time, path length, number of times to avoid obstacles, number of times to avoid fire points, number of times to pass through crowded points, and path smoothness.
4. The method of claim 1, wherein, The candidate escape path includes a first candidate escape path and a second candidate escape path; according to the area state information of each sub-area, determine a plurality of candidate escape paths of a target user, comprising: Determine the location of the target user as the starting node; Determine the location of the safe area as the terminal node; According to the area state information of each sub-area, determine a first intermediate node group between the starting node and the terminal node; Based on the starting node, the first intermediate node group, and the terminal node, generate the first candidate escape path of the target user; According to the area state information of each sub-area, determine a second intermediate node group between the starting node and the terminal node; Based on the starting node, the second intermediate node group, and the terminal node, generate the second candidate escape path of the target user.
5. The method of claim 4, wherein, The first intermediate node group includes a first intermediate node and a second intermediate node; the area state information includes area location, area temperature, area smoke concentration, and area personnel density; According to the area state information of each sub-area, determine a first intermediate node group between the starting node and the terminal node, comprising: After determining the first intermediate node, determine a plurality of candidate second intermediate nodes adjacent to the first intermediate node according to the area location; According to the area location, the area temperature, the area smoke concentration, and the area personnel density, determine the generation value corresponding to the candidate second intermediate node; Among a plurality of candidate second intermediate nodes, determine the candidate second intermediate node corresponding to the minimum generation value as the second intermediate node.
6. The method of claim 5, wherein, According to the area location, the area temperature, the area smoke concentration, and the area personnel density, determine the generation value corresponding to the candidate second intermediate node, comprising: According to the region position, a first distance between the candidate second intermediate node and the first intermediate node is calculated, According to the region position, a second distance between the candidate second intermediate node and the terminal node is calculated; The first distance, the second distance, the region temperature, the region smoke concentration and the region personnel density are respectively multiplied by corresponding preset influence factors to obtain a generation value corresponding to the candidate second intermediate node.
7. The method of claim 6, wherein, The first distance between the candidate second intermediate node and the first intermediate node is calculated according to the region position, comprising: According to the region position, a search region between the candidate second intermediate node and the first intermediate node is determined; In the case that there is no obstacle in the search region, a straight-line distance between the candidate second intermediate node and the first intermediate node is determined as the first distance; In the case that there is an obstacle in the search region, a shortest distance of detouring around the obstacle between the candidate second intermediate node and the first intermediate node is determined as the first distance.
8. The method of claim 4, wherein, The second intermediate node group includes a first intermediate node and a second intermediate node; and the region state information includes a region position, whether there is an obstacle in the region, a region temperature, a region smoke concentration and a region personnel density; The second intermediate node group between the starting node and the terminal node is determined according to the region state information of each sub-region, comprising: In the case that the first intermediate node has been determined, a plurality of candidate second intermediate nodes adjacent to the first intermediate node are determined according to the region position; An action score model is trained, and an action score value of an action corresponding to the first intermediate node to the candidate second intermediate node is determined according to the region position, whether there is an obstacle in the region, the region temperature, the region smoke concentration and the region personnel density; Among the plurality of candidate second intermediate nodes, the candidate second intermediate node corresponding to the maximum action score value is determined as the second intermediate node.
9. The method of claim 8, wherein, Before the action score model is trained, and the action score value of the action corresponding to the first intermediate node to the candidate second intermediate node is determined according to the region position, whether there is an obstacle in the region, the region temperature, the region smoke concentration and the region personnel density, the method further comprises: An emergency scene simulation environment of the building is established, and the emergency scene simulation environment includes simulation region state information of a plurality of sub-regions in the building; In a case where the first historical intermediate node has been determined History In a case where the first historical intermediate node has been determined In a case where the first historical intermediate node has been determined Reference action score values of actions corresponding to the first historical intermediate node to each historical candidate second intermediate node are obtained; According to a score error between the reference action score value and the initial action score value, the initial action score model is trained to obtain a trained action score model, and a score error corresponding to the trained action score model is less than a preset error value.
10. The method of claim 9, wherein, The reference action score values of the actions corresponding to the first historical intermediate node to each historical candidate second intermediate node are obtained, comprising: In a case where the historical candidate second intermediate node is a terminal node, an immediate reward value corresponding to the historical candidate second intermediate node is obtained; and a reference action score value of an action corresponding to the first historical intermediate node to the historical candidate second intermediate node is determined according to the immediate reward value. In a case where the historical candidate second intermediate node is not a terminal node, an immediate reward value and a future reward value corresponding to the historical candidate second intermediate node are obtained; and a reference action score value of an action corresponding to the first historical intermediate node to the historical candidate second intermediate node is determined according to the immediate reward value and the future reward value.
11. The method of claim 10, wherein, The determining of the reference action score value of the action corresponding to the first historical intermediate node to the historical candidate second intermediate node according to the immediate reward value and the future reward value comprises: wherein, the Qref is a reference action score value of an action corresponding to the first historical intermediate node to the historical candidate second intermediate node pair, r i+1 is an immediate reward value corresponding to the historical candidate second intermediate node, the is a discount factor, the r i+2 is an immediate reward value corresponding to a historical candidate third intermediate node adjacent to the historical candidate second intermediate node, the max r i+2 is the future reward value.
12. The method of claim 10, wherein, The obtaining of the immediate reward value corresponding to the historical candidate second intermediate node comprises: According to the region state information corresponding to the historical candidate second intermediate node, a target traffic feasibility level corresponding to the historical candidate second intermediate node is determined. In the preset correspondence between the pass feasibility level and the instant reward value, the target instant reward value corresponding to the target pass feasibility level is searched The determining of the target immediate reward value as the immediate reward value corresponding to the historical candidate second intermediate node is performed. the target instant reward value corresponding to the target pass feasibility level Comprise:
13. An escape route determination apparatus characterized by comprising: The obtaining module is configured to, in a case where an emergency event occurs in a building, obtain position information of each user in the building and environment information in the building; The determining module is configured to determine region state information of each sub-region in the building according to the position information of each user and the environment information. The determining module is further configured to determine a plurality of candidate escape paths of a target user according to the region state information of each sub-region. The obtaining module is further configured to obtain safety index data corresponding to the candidate escape paths. The scoring module is configured to score the candidate escape paths according to the safety index data corresponding to the candidate escape paths, to obtain path scores of the candidate escape paths. The determining module is further configured to determine that a candidate escape path corresponding to the highest path score is a target escape path. The device comprises a processor and a memory storing computer program instructions; and the processor implements the method for determining an escape path according to any one of claims 1 to 12 when executing the computer program instructions.
14. An escape route determination device, characterized by, The computer readable storage medium stores computer program instructions; and the computer program instructions are executed by a processor to implement the method for determining an escape path according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, The instructions in the computer program product are executed by a processor of a device, so that the device can execute the method for determining an escape path according to any one of claims 1 to 12.
16. A computer program product, characterised in that,