Method and device for determining urban land evacuation identification information

Through the collaborative work of the strategy network and value network of the identification agent, urban land evacuation signs are determined in real time, which solves the shortcomings of manual design and traditional reinforcement learning in existing technologies and realizes efficient and convenient evacuation sign design and optimization.

CN120372258BActive Publication Date: 2025-09-23TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510872952.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-23
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In existing technologies, the design of urban land evacuation signs relies on manual design and lacks quantitative indicators. Traditional reinforcement learning methods are limited to indoor evacuation. Numerical simulation software has large computational complexity and is time-consuming. It cannot quickly respond to changes in actual evacuation scenarios and has a high threshold for use.

Method used

The identification agent is used to work together through the strategy network and the value network. Based on the plot information and initial evacuation parameters, the evacuation sign layout and status are determined in real time. The identification agent is obtained through reinforcement learning training to avoid dependence on mathematical models and algorithms.

Benefits of technology

It realizes the real-time determination of urban land evacuation sign information, improves the performance and effect of sign intelligence, simplifies the convenience of operation, improves the universality of evacuation signs and the universality of design, and reduces computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372258B_ABST
    Figure CN120372258B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for determining urban land parcel evacuation identification information, which can be applied to the fields of artificial intelligence and emergency evacuation technology. The method includes: an identification agent obtains an action for laying out evacuation identification based on land parcel information and initial evacuation parameters; the action is executed in an environment to determine the layout and evacuation status of the evacuation identification; an initial strategy network obtains sample actions, sample evacuation identification layout results, sample evacuation status results, and sample distribution parameters based on sample land parcel information and sample initial evacuation parameters; an initial value network evaluates the sample actions, sample evacuation identification layout results, and sample evacuation status results based on an evaluation strategy to obtain an evaluation result; and based on the evaluation result, sample distribution parameters, and a loss function, the initial parameters of the initial strategy network and the initial value network are updated to obtain an identification agent. This method can reduce reliance on mathematical models and algorithms and reduce computing resource consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and emergency evacuation technology, and in particular to a method and device for determining evacuation identification information of urban plots. Background Art

[0002] Optimizing the design of evacuation signs can provide clearer and more accurate directional guidance for evacuees in urban emergency situations, allowing them to quickly determine evacuation routes and thus speed up evacuation. In related technologies, urban evacuation sign design optimization methods mainly rely on manual design or rule-based optimization algorithms. For example, evacuation data is classified, the time when the instructions are issued is compared, the locations to be evacuated are determined, different emergency sign devices are used to remind people at each location to evacuate, or Legion simulation software is combined with real-time simulation and analysis of the flow of people in subway stations to provide the optimal evacuation path; leader evacuation path planning based on reinforcement learning, collaborative guidance strategies based on evacuation signs and leaders, and simulation test verification are used to improve crowd evacuation efficiency; clustering algorithms are used to rationally classify newly discovered factors to simplify complex data and obtain evacuation evaluation results.

[0003] However, manual design in related technologies relies on people's subjective feelings and lacks quantitative indicators; traditional reinforcement learning methods are limited to indoor evacuation sign design and cannot be applied to emergency evacuation in urban areas; numerical simulation software or tools rely on complex mathematical models and algorithms to simulate crowd behavior, which requires large amounts of calculations and takes a long time for each simulation, and cannot quickly respond to changes in actual evacuation scenarios. In addition, simulation tools usually require professional knowledge and skills to set up and run, which increases the threshold for use. Summary of the Invention

[0004] In view of the above problems, the present invention provides a method, apparatus, device, medium and program product for determining evacuation identification information of urban plots.

[0005] According to a first aspect of the present invention, a method for determining evacuation identification information of urban plots is provided, comprising: an identification agent obtains an action for laying out evacuation identification based on acquired plot information of the plot to be identified and initial evacuation parameters corresponding to the plot information, wherein the initial evacuation parameters are determined by building information and object information in the plot information; an action is performed in an environment corresponding to the plot to be identified to determine a layout result and an evacuation status result of the evacuation identification corresponding to the plot to be identified; wherein the identification agent is trained in the following manner: an initial strategy network obtains sample actions, layout results of sample evacuation identification, sample evacuation status results and sample distribution parameters based on sample plot information of a sample plot and sample initial evacuation parameters corresponding to the sample plot information.

[0006] Preferably, the sample initial evacuation parameters are determined by the building information and object information in the plot information; the initial value network evaluates the sample action, the layout results of the sample evacuation identification and the sample evacuation status results based on the evaluation strategy to obtain the evaluation results; based on the evaluation results, the sample distribution parameters and the loss function, the initial parameters of the initial strategy network and the initial value network are updated to obtain the identification intelligent agent.

[0007] The second aspect of the present invention provides a device for determining the evacuation identification information of urban plots, including: an action determination module, which is used to identify the action of an intelligent agent to obtain the layout evacuation identification based on the acquired plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, wherein the initial evacuation parameters are determined by the building information and object information in the plot information; a result determination module, which is used to perform actions in an environment corresponding to the plot to be identified, and determine the layout result and evacuation status result of the evacuation identification corresponding to the plot to be identified.

[0008] Preferably, the identification agent is trained in the following manner: the initial strategy network obtains sample actions, layout results of sample evacuation identification, sample evacuation status results and sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, wherein the sample initial evacuation parameters are determined by the building information and object information in the plot information; the initial value network evaluates the sample actions, layout results of sample evacuation identification and sample evacuation status results based on the evaluation strategy to obtain an evaluation result; based on the evaluation result, the sample distribution parameter and the loss function, the initial parameters of the initial strategy network and the initial value network are updated to obtain the identification agent.

[0009] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0010] The fourth aspect of the present invention further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0011] The fifth aspect of the present invention further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.

[0012] According to an embodiment of the present invention, the action that the identification agent is currently performing can be determined in real time through the plot information and initial evacuation parameters of the plot to be identified, so as to flexibly obtain the layout result and evacuation status result of the evacuation sign after executing the action. Since the identification agent is obtained through a closed-loop update process formed by the collaborative work of the strategy network and the value network, the value network can feedback the evacuation status of the urban plot in real time, and the strategy network can adjust the sign layout in time according to the feedback information, so as to realize the real-time interaction of the evacuation layout result, the evacuation status result and the corresponding evaluation result, thereby improving the overall performance and effect of the identification agent. By automatically determining the evacuation sign information of the urban plot through the identification agent, the dependence on mathematical models and algorithms is avoided, and the consumption of computing resources is reduced. The identification agent automatically determines the evacuation sign information, without the user having to spend a lot of time to learn and master professional knowledge and skills. The operation is simple and convenient, further improving the universality of evacuation sign design and optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings.

[0014] Figure 1 A diagram illustrating an application scenario of a method, apparatus, device, medium, and program product for determining urban land parcel evacuation identification information according to an embodiment of the present invention is shown.

[0015] Figure 2 A flow chart of a method for determining urban plot evacuation identification information according to an embodiment of the present invention is shown.

[0016] Figure 3 An example schematic diagram of an optimization update process of an identification agent implemented by a proximal policy optimization algorithm based on a policy-value framework according to an embodiment of the present invention is shown.

[0017] Figure 4 A structural block diagram of a device for determining urban land parcel evacuation identification information according to an embodiment of the present invention is shown.

[0018] Figure 5 A block diagram of an electronic device suitable for implementing a method for determining urban land parcel evacuation identification information according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0019] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0020] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0022] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0023] In the process of conceiving the present invention, the inventors found that in the related technologies, manual design relies on people's subjective feelings and lacks quantitative indicators; the related reinforcement learning methods are limited to the design of indoor evacuation signs and cannot be used for emergency evacuation of urban plots; numerical simulation software or tools rely on complex mathematical models and algorithms to simulate crowd behavior, which has a large amount of calculations and each simulation takes a long time, and cannot quickly respond to changes in actual evacuation scenarios; and simulation tools usually require professional knowledge and skills to set up and run, which increases the threshold for use.

[0024] In view of the above technical problems, the present invention provides a method for determining the evacuation identification information of urban plots, comprising: an identification agent obtains an action for laying out an evacuation identification based on the acquired plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, wherein the initial evacuation parameters are determined by the building information and object information in the plot information; an action is performed in an environment corresponding to the plot to be identified to determine the layout result and evacuation status result of the evacuation identification corresponding to the plot to be identified; wherein the identification agent is trained in the following manner: an initial strategy network obtains sample actions, layout results of sample evacuation identification, sample evacuation status results and sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information; wherein the sample initial evacuation parameters are determined by the building information and object information in the plot information; an initial value network evaluates the sample actions, layout results and sample evacuation status results of the sample evacuation identification based on the evaluation strategy to obtain an evaluation result; based on the evaluation result, the sample distribution parameter and the loss function, the initial parameters of the initial strategy network and the initial value network are updated to obtain the identification agent.

[0025] According to an embodiment of the present invention, the action that the identification agent is currently performing can be determined in real time through the plot information and initial evacuation parameters of the plot to be identified, so as to flexibly obtain the layout result and evacuation status result of the evacuation sign after executing the action. Since the identification agent is obtained through a closed-loop update process formed by the collaborative work of the strategy network and the value network, the value network can feedback the evacuation status of the urban plot in real time, and the strategy network can adjust the sign layout in time according to the feedback information, so as to realize the real-time interaction of the evacuation layout result, the evacuation status result and the corresponding evaluation result, thereby improving the overall performance and effect of the identification agent. By automatically determining the evacuation sign information of the urban plot through the identification agent, the dependence on mathematical models and algorithms is avoided, and the consumption of computing resources is reduced. The identification agent automatically determines the evacuation sign information, without the user having to spend a lot of time to learn and master professional knowledge and skills. The operation is simple and convenient, further improving the universality of evacuation sign design and optimization.

[0026] Figure 1 A diagram illustrating an application scenario of a method, apparatus, device, medium, and program product for determining urban land parcel evacuation identification information according to an embodiment of the present invention is shown.

[0027] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0028] A user can use a terminal device 101 to interact with a server 103 via a network 102 to receive or send messages, etc. The terminal device 101 can be any electronic device with a display screen and web browsing support, including but not limited to a smartphone, tablet computer, laptop computer, desktop computer, etc. For example, a user can use the terminal device 101 to send land parcel information and related parameters of a land parcel to be identified to an agent.

[0029] Server 103 can be a server that provides various services. For example, server 103 receives various data from terminal devices 101 and stores it in a database. This data includes real-time environmental data, personnel distribution, device status, and user feedback, providing data support for subsequent analysis and decision-making. For example, server 103 runs an agent-based evacuation identification information determination system. The agent determines evacuation identification information for urban areas based on the received data.

[0030] It should be noted that the method for determining the evacuation identification information of urban plots provided in the embodiment of the present invention can generally be executed by the server 103. Accordingly, the device for determining the evacuation identification information of urban plots provided in the embodiment of the present invention can generally be set in the server 103. The method for determining the evacuation identification information of urban plots provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103. Accordingly, the device for determining the evacuation identification information of urban plots provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103.

[0031] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0032] Figure 2 A flow chart of a method for determining urban plot evacuation identification information according to an embodiment of the present invention is shown.

[0033] like Figure 2 As shown, the method for determining urban land parcel evacuation identification information in this embodiment includes operations S210 to S220.

[0034] In operation S210, the identification agent obtains an action of layout evacuation identification based on the acquired plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, wherein the initial evacuation parameters are determined by the building information and object information in the plot information.

[0035] In an embodiment of the present invention, the identification agent may include a strategy agent and a value agent. Based on real-time information acquired, it can provide evacuation route guidance, real-time information exchange, and evacuation identification design and optimization for the plots to be identified. It is the entity that executes decisions within the reinforcement learning framework. Plot information may include plot boundary information, evacuation location information, building information within the plots, and object information of the objects to be evacuated. Building information may include building layout information, building boundary information, and building entrance and exit location information. Object information may include the total number of people and population density corresponding to the plot to be identified or each building.

[0036] The action of laying out evacuation signs may include adding new evacuation signs to a plot of land or updating or adjusting existing evacuation signs. The initial evacuation parameters may be initial evacuation state parameters under the initial state, and may include the initial evacuation duration, the initial evacuation completion time corresponding to each building in the urban area, and the initial location information and initial number of evacuation concentration points.

[0037] For example, by extracting features from plot information and initial evacuation parameters, different features are obtained, and then the different features are spliced ​​together to obtain splicing results. Then, based on reinforcement learning and the splicing results, multiple initial actions are obtained, and appropriate actions are selected from the multiple initial actions as the actions for laying out evacuation signs.

[0038] In operation S220 , an action is performed in an environment corresponding to the land parcel to be identified, and a layout result and an evacuation status result of the evacuation sign corresponding to the land parcel to be identified are determined.

[0039] In an embodiment of the present invention, the environment can be obtained by processing the acquired plot information and initial evacuation parameters. The layout result of the evacuation sign can characterize the geometric parameters of multiple evacuation signs in the plot to be identified, and can include the center point coordinates, size information, height information and direction information of the evacuation sign. The evacuation status result can characterize the evacuation status parameters achieved in the urban area at the current moment according to the layout result of the current evacuation sign, and can include the evacuation duration, the evacuation completion time corresponding to each building in the urban area, and the point information and number of evacuation dense points.

[0040] For example, the plot information and initial evacuation parameters obtained at the current moment are combined to obtain the current environment of the plot to be marked, and the action of laying out evacuation signs is executed in the current environment to obtain the layout result of the current evacuation signs and the current evacuation status result.

[0041] Preferably, the identification agent is trained in the following manner, which may include operations S221 to S223.

[0042] In operation S221, the initial strategy network obtains sample actions, layout results of sample evacuation identifications, sample evacuation status results, and sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, wherein the sample initial evacuation parameters are determined by the building information and object information in the sample plot information.

[0043] In operation S222 , the initial value network evaluates the sample actions, the layout results of the sample evacuation identifiers, and the sample evacuation status results based on the evaluation strategy to obtain an evaluation result.

[0044] In operation S223, based on the evaluation results, the sample distribution parameters and the loss function, the initial parameters of the initial policy network and the initial value network are updated to obtain an identified intelligent agent.

[0045] In an embodiment of the present invention, the identification agent can be obtained by training an initial identification agent, which can include an initial policy agent and an initial value agent. The initial policy network can be a reinforcement learning model corresponding to the initial policy agent, which can be used to output an action policy with the goal of optimizing policy parameters to maximize the cumulative reward at a later time. The initial value network can be a reinforcement learning model corresponding to the initial value agent, which can guide the initial policy agent to output a policy based on the potential reward of the current state by learning the relationship between the environment and rewards. The environment can be a system for urban building layout, sign layout, and evacuation simulation.

[0046] For example, the urban evacuation optimization environment can be initialized, and the initial state is set to a plot containing only building layouts (no sign layout, i.e., a plot to be marked), and evacuation signs can be placed one by one through reinforcement learning. The parameters of the evacuation sign can include the center point position coordinates (x si ,y si ), sign length, sign width and sign height from the ground (h). The position constraint of the evacuation sign can be located within the outer contour line of the building and within a meter offset from the outer contour line. The role of the evacuation sign can be to indicate the evacuation direction of the location and to speed up the evacuation of personnel.

[0047] The initial plot can be simulated by inputting the population density distribution of each building to obtain the initial evacuation state; the evacuation state can include the total evacuation time, the time for all personnel to evacuate each building, and the coordinates and total number of all evacuation congestion points during the evacuation process; the initial plot and evacuation state are then used as the environment input for reinforcement learning to establish an urban evacuation optimization environment.

[0048] For example, an action set may be defined, which may include the change in coordinates of the center point of each marker, the change in length and width, the change in height from the ground, and the marker directivity.

[0049] In embodiments of the present invention, a reinforcement learning model can be constructed using a policy optimization algorithm to train an identity agent. This algorithm uses a neural network to predict action strategies and implement dynamic identity placement and adjustment. Combined with a policy-value framework, the policy network is responsible for outputting action strategies, aiming to optimize policy parameters to maximize future cumulative rewards. The value network, by learning the relationship between the environment and rewards, assesses the potential rewards of the current state and guides the policy network in outputting these strategies.

[0050] For example, based on reinforcement learning, multiple sets of random logo layouts can be generated as training data sets; the training data sets can then be used to drive the reinforcement learning model (which may include a policy network and a value network) to interact with the environmental modeling system and the evacuation simulation system to optimize the model strategy; the interaction process is repeated until the termination conditions are met and the training is completed, resulting in a trained reinforcement learning model.

[0051] In an embodiment of the present invention, the sample plot may be a plot similar to the plot to be identified within a historical period. The evaluation strategy may be a strategy that predicts the state value of the current state and calculates the action advantage of the current action based on the reward function of the initial identification agent. The evaluation result may include the state value and the action advantage. The loss function may include loss functions corresponding to the initial policy network and the initial value network respectively. The initial parameters may include initial policy parameters and initial value parameters. It is understandable that the information corresponding to the sample plot information, sample initial evacuation parameters, sample actions, etc. in the training phase is the same or similar to the information corresponding to the application phase, and the details will not be repeated here.

[0052] For example, feature extraction and splicing are performed on the sample plot information and the sample initial evacuation parameters to obtain sample actions and sample distribution parameters; then, based on the sample distribution parameters, the sample actions are randomly sampled to obtain the action information of the sample actions, and then the sampling actions are executed in the sample environment to obtain the layout results of the sample evacuation signs and the sample evacuation status results.

[0053] According to an embodiment of the present invention, the action that the identification agent is currently performing can be determined in real time through the plot information and initial evacuation parameters of the plot to be identified, so as to flexibly obtain the layout result and evacuation status result of the evacuation sign after executing the action. Since the identification agent is obtained through a closed-loop update process formed by the collaborative work of the strategy network and the value network, the value network can feedback the evacuation status of the urban plot in real time, and the strategy network can adjust the sign layout in time according to the feedback information, so as to realize the real-time interaction of the evacuation layout result, the evacuation status result and the corresponding evaluation result, thereby improving the overall performance and effect of the identification agent. By automatically determining the evacuation sign information of the urban plot through the identification agent, the dependence on mathematical models and algorithms is avoided, and the consumption of computing resources is reduced. The identification agent automatically determines the evacuation sign information, without the user having to spend a lot of time to learn and master professional knowledge and skills. The operation is simple and convenient, further improving the universality of evacuation sign design and optimization.

[0054] According to an embodiment of the present invention, the building information includes plane information and evacuation information; the method further comprises: determining the evacuation duration and evacuation target point information in the initial evacuation parameters based on the plane information, evacuation information and object information.

[0055] In embodiments of the present invention, planar information may include building layout information and building boundary information. Evacuation information may include building entrance and exit location information and building evacuation path information. Evacuation duration information may include the total time required to complete the evacuation of the identified plot (the plot to be evacuated) and the time required to complete the evacuation of each individual building within the plot. Evacuation target point information may include the location and number of target points. Target points may be locations where the density of all personnel during an emergency evacuation is greater than or equal to a preset density threshold.

[0056] For example, the evacuation simulation can be performed on the initial plot based on the input personnel density distribution of each building. The initial evacuation state, including the total evacuation time, can be calculated through the evacuation simulation tool. t , the time t for all personnel in each building to evacuate m , The density of people during evacuation exceeds 2 people / m 2 The position coordinates (x ci ,y ci ) and the total number of points c.

[0057] For example, each floor of a building can be divided into multiple areas, and a state network can be constructed. Each area can be abstracted as a state space, and the stairs can be abstracted as a state path. By analyzing the transfer of people between different areas and channels, the analysis model can be used to dynamically analyze the evacuation process of people, thereby determining potential congestion points and evacuation time.

[0058] According to an embodiment of the present invention, a dynamic optimization model is established based on state network theory. By assuming regional division and introducing concepts such as crowd density, the functional relationship between crowd density and travel speed can be determined, and the calculation formulas for the state transfer quantities of horizontal passages, stairs and doors are calculated. The evacuation path is optimized and adjusted in combination with the shortest path algorithm, which meets the requirements for effective identification and processing of complex situations such as traffic congestion during the evacuation process, thereby improving the rationality and feasibility of the evacuation plan.

[0059] According to an embodiment of the present invention, an action is performed in an environment corresponding to a land parcel to be identified to determine a layout result and an evacuation status result of an evacuation sign corresponding to the land parcel to be identified, including: constructing a current environment based on building information, current identification information at a current moment, and current evacuation information; performing a position identification action, a size identification action, and a function identification action in the current environment to obtain a layout result and a current evacuation status result of the current evacuation sign; when the layout result and the current evacuation status result of the current evacuation sign meet preset conditions, stopping the action, determining the layout result and the current evacuation status result of the current evacuation sign as the layout result and the evacuation status result of the evacuation sign, respectively, and outputting them.

[0060] In the embodiment of the present invention, the position identification action may include the coordinates of the center point of the evacuation mark (x si ,y si )、Height from ground (h i ) is marked. The size marking action may include the length of the evacuation mark (l i )、width(w i ) to mark the action. Functional marking action can represent the directionality of the evacuation sign (θ i ) to perform the marking action. The preset condition can be a condition defined by the user according to actual needs. For example, if the total number of evacuation signs arranged in the area exceeds a threshold, the action will be stopped.

[0061] For example, the building layout information and evacuation status information of the plot to be identified are obtained, and the trained identification agent is loaded; the building layout information and evacuation status information are used as input to the identification agent at the current moment to optimize the identification layout of the urban area, output the optimal identification layout of the area at the current moment, and update the building layout information (the building layout information after the identification action has been executed) and the current evacuation status result; and then, when the total number of evacuation signs in the entire area is greater than the number threshold (for example, 100), stop executing the action and output the evacuation sign layout result and evacuation status result.

[0062] According to an embodiment of the present invention, the identification agent obtains the action of layout evacuation identification based on the acquired plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, including: extracting features of the plot information and the initial evacuation parameters respectively to obtain plot features and state features; splicing the plot features and the state features to obtain a splicing vector; the identification agent determines multiple initial actions and multiple distribution parameters corresponding to the multiple initial actions based on the splicing vector; and determines an action from the multiple initial actions based on the multiple distribution parameters and preset screening rules.

[0063] In embodiments of the present invention, plot features and state features may be plot vectors and state vectors obtained by extracting features from plot information and initial evacuation parameters, satisfying the requirements of the identification agent. A concatenated vector may be a vector obtained by concatenating, combining, or fusing the plot vector and state vector. Distribution parameters may be action distribution parameters obtained by executing multiple initial actions in the plot to be identified. Preset screening rules may include either screening based on probability or screening based on advantage.

[0064] For example, feature extraction is performed on the acquired building layout and evacuation status information, and it is converted into plot features and status features that can be processed by the identification agent; the plot features and status features are then combined into a splicing vector, which is used as input information for the identification agent and input into the identification agent; the identification agent outputs multiple initial actions, action distribution parameters, and probability values ​​corresponding to the multiple initial actions; and then the initial action corresponding to the highest probability value is selected as the final action to be executed.

[0065] According to an embodiment of the present invention, the identification intelligent agent can dynamically determine and optimize the identification action strategy based on the real-time acquired plot information and initial evacuation parameters, and guide the evacuees to the optimal path in real time, making the evacuation path more direct and efficient, thereby shortening the overall evacuation time.

[0066] In a feasible embodiment, the method for determining urban plot evacuation identification information based on an identification agent may include operations S310 to S340.

[0067] In operation S310 , building information, object information, and initial evacuation parameters of an unidentified initial land parcel (a land parcel to be identified) are acquired.

[0068] In operation S320, the current building information, object information and initial evacuation parameters are used as input information of the identification agent, and the trained identification agent is used to optimize and update the identification layout of the target urban area, and the optimal evacuation sign layout result and evacuation status result of the area at the current moment are output.

[0069] In operation S330 , the operation S320 is repeatedly performed until the total number of evacuation signs in the target urban area is greater than or equal to the number threshold, and the marking operation is stopped.

[0070] In operation S340 , the final evacuation sign layout result and evacuation status result are output.

[0071] It can be understood that the above has described the method of applying the identification agent to determine the evacuation identification information of urban plots. The following will describe the process of how to train the identification agent.

[0072] The sample identification agent obtains an identification agent for determining the evacuation identification information of urban plots through interaction with the sample environment, the update process of the initial policy network and the initial value network, and the process of the policy optimization algorithm.

[0073] According to an embodiment of the present invention, the initial strategy network obtains sample actions, layout results of sample evacuation identification, sample evacuation status results and sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, including: processing the sample plot information and the sample initial evacuation parameters to obtain sample actions and sample distribution parameters; sampling from the sample actions based on the sample distribution parameters to obtain action information of the sample actions; executing the sampling action in the sample environment corresponding to the sample plot to obtain the layout results of the sample evacuation identification and the sample evacuation status results.

[0074] In embodiments of the present invention, the sample distribution parameters may be the action distribution parameters output by the initial policy agent. Sampled actions may be obtained through random sampling or empirical sampling of sample actions. It should be noted that the sample actions, sample evacuation sign layout results, sample evacuation status results, sample environment, and other related content in the training phase can refer to the corresponding content in the application phase. Similarly, the relevant content in the application phase can also refer to the corresponding content in the training phase. The details are not detailed here.

[0075] According to an embodiment of the present invention, the method for determining evacuation identification information further includes: determining the current state value of the sampling action at the current moment based on the layout result of the sample evacuation identification and the sample evacuation state result to perform result evaluation.

[0076] In an embodiment of the present invention, the current state value may be a state value obtained by inputting the layout result of the sample evacuation identification and the sample evacuation state result at the current moment into the value network.

[0077] For example, a sample environment may receive a sample action a t (For example, new identification parameters [x, y, l, w, h, θ]), output state result s t(including the layout results of the sample evacuation signs, the sample evacuation status results) and the instant reward value r t , and generate the next state result s t+1 .

[0078] Furthermore, based on the trajectory data, the current strategy is used Empirical data collected from the environment (s t , a t ,r t , s t+1 ) is used for subsequent calculations. t As a result of the current state, the environment is input to the policy network and the value network to predict actions and values.

[0079] On this basis, the initial policy network in the identification agent receives the state result s t , you can output sample actions , that is, the initial strategy network places or adjusts parameters according to the current layout prediction mark; the initial value network receives the current state result s t , output the current state value Vφ(s t ), the state value can represent the estimated future cumulative rewards of the current state results.

[0080] According to embodiments of the present invention, the initial value network's assessment of the current state provides more accurate feedback to the initial policy network, enabling it to output personalized identification strategies tailored to different areas and population densities. For example, for densely populated areas, the density and visibility of evacuation instructions can be increased; for areas farther from safe exits, clearer guidance directions and distance information can be provided, enabling intelligent and precise evacuation guidance. This optimized policy update approach makes policy updates within the policy network more stable, avoiding confusion in identification instructions caused by frequent and significant policy changes.

[0081] According to an embodiment of the present invention, the initial value network evaluates the sample action, the layout result of the sample evacuation identification and the sample evacuation state result based on the evaluation strategy to obtain the evaluation result, including: using the reward function in the evaluation strategy to process the sample action, the layout result of the sample evacuation identification and the sample evacuation state result, and determining the reward value after performing the sampling action in the sample environment, as well as the subsequent identification layout result, the subsequent evacuation state result and the subsequent state value at the subsequent moment, wherein the subsequent state value is determined by the subsequent identification layout result and the subsequent evacuation state result; and determining the evaluation result based on the reward value, the current state value, the subsequent state value and the weight corresponding to the subsequent state value.

[0082] In an embodiment of the present invention, the reward function may be an instantaneous reward function corresponding to a time step. The weights corresponding to the state values ​​at different moments may be determined based on actual conditions or empirical values, and are not specifically limited herein. The evaluation results may include the state value V obtained by the value network evaluation. φ (st) and the advantage result A of the sample action t .

[0083] Figure 3 An example schematic diagram of an optimization update process of an identification agent implemented by a proximal policy optimization algorithm based on a policy-value framework according to an embodiment of the present invention is shown.

[0084] like Figure 3 As shown, the identification agent 301 may include a policy network 301a and a value network 301b. The process of implementing the update optimization of the policy network and the value network may include: randomly initializing the parameters θ of the policy network 301a and the parameters φ of the value network 301b; the identification agent 301 may interact with the environment 303 based on the action space 302 of the identification agent, and obtain the initial state s from the environment 303. t , through the policy network, we can calculate the current policy πθ(a t ∣s t ) Select target action a from action space 302 t , and then perform action a in the environment t , get the current reward value r t and the updated status result s t+1 The action space 302 is also called an action set.

[0085] Based on the trajectory data 304, the current strategy πθ can be used old Empirical data collected from the environment (s t , a t , r t ,s t+1 ) is used for subsequent calculations. t As a result of the current state, the environment input is sent to the policy network and the value network to predict the action and state results. The value network 301b can be used to predict the current state s t The value of V ϕ (s t ) and predict the next state result s t+1 The corresponding value V ϕ (s t+1 ), and calculate the time series difference error, advantage function A t The timing difference error δ can be directly used t Estimated.

[0086] The value function is updated by updating the strategy, calculating the probability ratio 305 and the strategy loss 306, and calculating the value loss 307. The strategy loss 306 can be determined based on the probability ratio 305 and the advantage function 308. Thus, the total loss value can be obtained based on the strategy loss 306 and the value loss 307. Then, the Adam optimizer is used to update the parameters θ and φ of the strategy network and the value network based on the total loss L. The above steps are repeated until the preset number of training rounds or convergence conditions are reached.

[0087] For example, a reinforcement learning model for a sign agent can be constructed as follows. First, the state space of the sign agent is defined. The state space can be a dynamic vector S that changes with the placed evacuation signs and includes the geometric parameters of the currently placed signs and the evacuation state calculated through simulation.

[0088] The geometric parameters of the markers may include: the coordinates of the center point of each marker (x si ,y si ) length (l i )、width(w i ), height from the ground (h i ) and marker directivity (θ i ), if n markers have been placed, the geometric parameter part can be: [x1, y1, l1, w1 ,h1,θ1 ,…x n , y n , l n , w n , h n ,θ n The evacuation state parameters can be calculated by the evacuation simulation system, where the total evacuation time is the time when the last individual to be evacuated in the urban area arrives at the refuge location. t The evacuation time of all personnel in each building is the time when the last individual to be evacuated leaves the building. If there are m buildings, there are m building evacuation times, recorded as t m During the evacuation process, the density of all personnel exceeds the preset density threshold (for example, 2 people / m 2 ) point coordinates (x ci ,y ci ) and the number of points c.

[0089] It is understandable that the specific type of the evacuation simulation system can be determined according to actual needs, with the standard of meeting the simulated evacuation of the land to be identified, and the specific type is not limited here.

[0090] Furthermore, an action set A of the identification agent can be defined, and the action set A can include the placement of a new evacuation sign. The new evacuation sign placement can specify parameters for the pth sign: [x sp,y sp , l p , w p , h p ,θ p ], where x sp ,y sp It can be the center point coordinates, limited to the boundaries of the plot, and within the range of n meters offset from the building outline. n can be determined according to the specific urban area type and building size.

[0091] In an embodiment of the present invention, the parameters of the initial policy network and the initial value network can be updated through policy gradient and value loss.

[0092] For example, the reward function of the identity agent can be determined by the following formula (1):

[0093] (1);

[0094] in, Can be used to minimize the total evacuation time, Time t can be the total time for the last individual to be evacuated in the urban area to reach the refuge location; β.b can be the number of minimized identifiers, and b can be the number of identifiers; γ.c can be the number of minimized crowd congestion points, and c can be the number of congestion points; You can reward items for marking proximity to congestion points; You can create a bonus item for marking proximity to a "slow evacuation building"; Each is the weight of different rewards. and The determination method is as shown in the following formulas (2)-(3):

[0095] (2);

[0096] (3);

[0097] Among them, t i Can be the current time step, T can be the total time step, (x mi ,y mi ) can be the coordinates of the center point of the building, (x si ,y si ) can be used to identify the center point coordinates, (x ci ,y ci ) can be the coordinates of the congestion point, σ 2 To adjust the parameters, the attenuation degree of the exponential function exp can be controlled, which can be obtained based on experience or experiments.

[0098] For example, the method of updating the initial policy network parameters based on the policy gradient may include: first using the value network to calculate the current state value V φ (s t ) and the value V in the subsequent state φ (s t+1 ), the advantage is calculated by generalized advantage estimation, as shown in the following formulas (4)-(5):

[0099] (4);

[0100] (5);

[0101] Among them, A t It can represent the advantage function, indicating the result s in the current state t Next, select sample action a t Compared to the average level; k can represent the number of future time steps considered from time step t, which is used to calculate multi-step returns, δ t It can represent the difference between the predicted value and the actual value at time step t.

[0102] δ t+K It can be the temporal difference error, which represents the difference between the predicted value and the actual value at time step t+k; r t V can be the immediate reward at time step t (i.e. the reward function designed previously); ϕ (s t+1 ) can be the state value estimate at the later time; γ' can be V φ (s t+1 ), also called the discount factor, is taken as 0.99, for example, and λ can be a generalized objective estimation parameter, for example, 0.95.

[0103] It can be understood that the reward function can comprehensively consider the number of evacuation signs, the total evacuation time, and the time for all personnel to evacuate each building, minimize the number of evacuation congestion locations, encourage the optimization of evacuation capacity and meet the constraints.

[0104] According to an embodiment of the present invention, the proximal strategy optimization algorithm can utilize the collected samples more efficiently in each update, and improve the utilization rate of samples by updating the strategy through multiple iterations, thereby accelerating the convergence process and achieving a better balance between exploring new strategies and utilizing existing knowledge, which not only ensures sufficient exploration but also can effectively utilize existing information to accelerate convergence.

[0105] According to an embodiment of the present invention, the initial parameters include initial design parameters and initial evaluation parameters; based on the evaluation results, sample distribution parameters and loss function, the initial parameters of the initial policy network and the initial value network are updated to obtain an identification intelligent agent, including: determining the ratio between the current distribution parameters at the current moment and the previous distribution parameters at the previous moment in the sample distribution parameters as the strategy loss probability ratio; determining the loss result based on the loss probability ratio, the evaluation results and the loss function; updating the initial design parameters and the initial evaluation parameters using the updated strategy and the loss result, determining the target design parameters and the target evaluation parameters, so as to update the initial policy network using the target design parameters and update the initial value network using the target evaluation parameters to obtain an identification intelligent agent.

[0106] In an embodiment of the present invention, the strategy loss probability ratio may be the probability ratio between sample action strategies at different times. The loss result may include the strategy loss result corresponding to the strategy network and the value loss result corresponding to the value network. The update strategy may be an update algorithm based on an adaptive learning rate and momentum.

[0107] For example, the strategy loss probability ratio can be calculated using the following formula (6):

[0108] (6);

[0109] in, It can represent the action strategy output at the current moment, It can represent the action strategy output at the previous moment, θ can represent the parameters of the policy network, and πθ can represent the policy function.

[0110] According to an embodiment of the present invention, the loss function includes a strategy loss function and a value loss function; based on the loss probability ratio, the evaluation result and the loss function, the loss result is determined, including: based on the loss probability ratio, the evaluation result, the strategy loss function and the value loss function, the strategy loss value and the value loss value are determined; and the loss weights corresponding to the strategy loss value and the value loss value are used to perform weighted summation on the strategy loss value and the value loss value to obtain the loss result.

[0111] In an embodiment of the present invention, the policy loss value may be a loss value obtained by clipping or limiting the range of the data, such as a clip function. The value loss value may be used to evaluate the difference between the predicted state result and the actual state result.

[0112] In policy optimization, the policy network can be used to output action policies , the goal of the policy gradient is to optimize the policy parameters θ to maximize the future cumulative reward. The policy loss value can be determined by optimizing the objective function, as shown in the following formula (7):

[0113] (7);

[0114] in, The objective function that can characterize the policy gradient, It can express expectations, It can be used to limit the update range of the strategy, for example, the value is 0.02; the clip function can be used to t (θ) is limited to Within the range, r t (θ) can be the strategy loss probability ratio.

[0115] According to an embodiment of the present invention, the policy optimization algorithm limits the policy update range by introducing a clip function, thereby ensuring the stability of the policy update and achieving effective optimization of the policy.

[0116] For example, updating the initial value network parameters based on value loss can be achieved by minimizing the error between the predicted value and the actual return, as shown in the following formulas (8)-(9):

[0117] (8);

[0118] (9);

[0119] Among them, L VF Can represent the value loss value, Can express expectations, R t can be the cumulative return, and Vφ(st) can be the value predicted by the value network.

[0120] Furthermore, the policy loss value and the value loss value can be weighted and summed to obtain the total loss value (loss result); thereby determining the gradient of the total loss with respect to the policy network parameters and the value network parameters φ, and then updating the policy network parameters θ and the value network parameters φ based on the update algorithm based on the adaptive learning rate and momentum according to the determined gradient.

[0121] In a feasible embodiment, the method of identifying the intelligent agent training includes operations S401 to S405.

[0122] In operation S401, based on the state space of the initial identification agent and the current layout information of the site, the output action distribution parameters of the identification agent are generated through the strategy network to obtain the current strategy .

[0123] In operation S402, based on the reward function of the identified agent, the state value V is predicted through the value network. φ (st) and calculate the advantage result A of the action t .

[0124] In operation S403, based on the advantage result A t 、Status value V φ (st) and strategy , the total loss L is calculated through the loss function.

[0125] In operation S404 , according to the total loss L, the policy network parameters and the value network parameters are updated using an update algorithm with adaptive learning rate and momentum.

[0126] In operation S405 , the training step is repeated to iteratively update the policy network parameters and the value network parameters until a preset termination condition is reached, thereby obtaining the final parameters of the policy network and the final parameters of the value network of the identification agent.

[0127] Based on the above-mentioned method for determining the evacuation identification information of urban land parcels, the present invention also provides a device for determining the evacuation identification information of urban land parcels. Figure 4 The device is described in detail.

[0128] Figure 4 A structural block diagram of a device for determining urban land parcel evacuation identification information according to an embodiment of the present invention is shown.

[0129] like Figure 4 As shown, the apparatus 400 for determining urban land parcel evacuation identification information in this embodiment includes an action determination module 410 and a result determination module 420 .

[0130] Action determination module 410 is configured to determine the action of the identification agent to obtain a layout evacuation identification based on the acquired parcel information of the to-be-identified parcel and the initial evacuation parameters corresponding to the parcel information. The initial evacuation parameters are determined by the building and object information in the parcel information. In one embodiment, action determination module 410 can be used to perform operation S210 described above and will not be further described here.

[0131] The result determination module 420 is configured to execute an action in the environment corresponding to the land parcel to be identified, and determine the layout result and evacuation status result of the evacuation sign corresponding to the land parcel to be identified. In one embodiment, the result determination module 420 can be configured to execute the operation S220 described above, which will not be described in detail here.

[0132] Preferably, the identification agent is trained in the following manner: the initial strategy network obtains sample actions, layout results of sample evacuation identification, sample evacuation status results and sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, wherein the sample initial evacuation parameters are determined by the building information and object information in the plot information; the initial value network evaluates the sample actions, layout results of sample evacuation identification and sample evacuation status results based on the evaluation strategy to obtain an evaluation result; based on the evaluation result, the sample distribution parameter and the loss function, the initial parameters of the initial strategy network and the initial value network are updated to obtain the identification agent.

[0133] According to an embodiment of the present invention, through the action determination module 410 and the result determination module 420 in the device 400 for determining the evacuation identification information of urban plots, the action to be currently performed by the identification agent can be determined in real time based on the plot information of the plot to be identified and the initial evacuation parameters, thereby flexibly obtaining the layout result and evacuation status result of the evacuation identification after executing the action. Since the identification agent is obtained through a closed-loop update process formed by the collaborative work of the strategy network and the value network, the value network can feedback the evacuation status of the urban plot in real time, and the strategy network can adjust the identification layout in time according to the feedback information, realizing real-time interaction between the evacuation layout result, the evacuation status result and the corresponding evaluation result, thereby improving the overall performance and effect of the identification agent. By automatically determining the evacuation identification information of the urban plot through the identification agent, the dependence on mathematical models and algorithms is avoided, and the consumption of computing resources is reduced. The method of automatically determining the evacuation identification information by the identification agent does not require the user to spend a lot of time learning and mastering professional knowledge and skills. The operation is simple and convenient, further improving the universality of evacuation sign design and optimization.

[0134] According to an embodiment of the present invention, the building information includes plane information and evacuation information; the device also includes: an information determination module for determining the evacuation duration and evacuation target point information in the initial evacuation parameters based on the plane information, evacuation information and object information.

[0135] According to an embodiment of the present invention, the result determination module 420 includes: an environment construction submodule, an action execution submodule, and an action stop submodule. The environment construction submodule is used to construct the current environment based on building information, current identification information at the current moment, and current evacuation information; the action execution submodule is used to execute the position identification action, size identification action, and function identification action in the current environment to obtain the layout result of the current evacuation sign and the current evacuation status result; the action stop submodule is used to stop the action when the layout result of the current evacuation sign and the current evacuation status result meet preset conditions, determine the layout result of the current evacuation sign and the current evacuation status result as the evacuation sign layout result and the evacuation status result, and output them respectively.

[0136] According to an embodiment of the present invention, the action determination module 410 includes: an extraction submodule, a splicing submodule, a parameter determination submodule, and an action determination submodule. The extraction submodule is configured to extract features from the plot information and initial evacuation parameters, respectively, to obtain plot features and state features; the splicing submodule is configured to splice the plot features and state features to obtain a splicing vector; the parameter determination submodule is configured to identify the intelligent agent and determine multiple initial actions and multiple distribution parameters corresponding to the multiple initial actions based on the splicing vector; and the action determination submodule is configured to determine an action from the multiple initial actions based on the multiple distribution parameters and preset screening rules.

[0137] According to an embodiment of the present invention, the initial strategy network obtains sample actions, layout results of sample evacuation identification, sample evacuation status results and sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, including: processing the sample plot information and the sample initial evacuation parameters to obtain sample actions and sample distribution parameters; sampling from the sample actions based on the sample distribution parameters to obtain action information of the sample actions; executing the sampling action in the sample environment corresponding to the sample plot to obtain the layout results of the sample evacuation identification and the sample evacuation status results.

[0138] According to an embodiment of the present invention, the apparatus further comprises: a value determination module for determining the current state value of the sampling action at the current moment based on the layout result of the sample evacuation identification and the sample evacuation state result, so as to perform result evaluation.

[0139] According to an embodiment of the present invention, the initial value network evaluates the sample action, the layout result of the sample evacuation identification and the sample evacuation state result based on the evaluation strategy to obtain the evaluation result, including: using the reward function in the evaluation strategy to process the sample action, the layout result of the sample evacuation identification and the sample evacuation state result, and determining the reward value after performing the sampling action in the sample environment, as well as the subsequent identification layout result, the subsequent evacuation state result and the subsequent state value at the subsequent moment, wherein the subsequent state value is determined by the subsequent identification layout result and the subsequent evacuation state result; and determining the evaluation result based on the reward value, the current state value, the subsequent state value and the weight corresponding to the subsequent state value.

[0140] According to an embodiment of the present invention, the initial parameters include initial design parameters and initial evaluation parameters; based on the evaluation results, sample distribution parameters and loss function, the initial parameters of the initial policy network and the initial value network are updated to obtain an identification intelligent agent, including: determining the ratio between the current distribution parameters at the current moment and the previous distribution parameters at the previous moment in the sample distribution parameters as the strategy loss probability ratio; determining the loss result based on the loss probability ratio, the evaluation results and the loss function; updating the initial design parameters and the initial evaluation parameters using the updated strategy and the loss result, determining the target design parameters and the target evaluation parameters, so as to update the initial policy network using the target design parameters and update the initial value network using the target evaluation parameters to obtain an identification intelligent agent.

[0141] According to an embodiment of the present invention, the loss function includes a strategy loss function and a value loss function; based on the loss probability ratio, the evaluation result and the loss function, the loss result is determined, including: based on the loss probability ratio, the evaluation result, the strategy loss function and the value loss function, the strategy loss value and the value loss value are determined; and the loss weights corresponding to the strategy loss value and the value loss value are used to perform weighted summation on the strategy loss value and the value loss value to obtain the loss result.

[0142] According to embodiments of the present invention, any multiple modules in action determination module 410 and result determination module 420 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present invention, at least one of action determination module 410 and result determination module 420 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of software, hardware, and firmware, or any appropriate combination of these. Alternatively, at least one of action determination module 410 and result determination module 420 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.

[0143] Figure 5 A block diagram of an electronic device suitable for implementing a method for determining urban land parcel evacuation identification information according to an embodiment of the present invention is shown.

[0144] like Figure 5As shown, an electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 502 or programs loaded from a storage unit 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0145] Various programs and data required for the operation of the electronic device 500 are stored in the RAM 503. The processor 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The processor 501 executes the programs in the ROM 502 and / or RAM 503 to perform various operations according to the method flow of the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 502 and RAM 503. The processor 501 may also execute the programs stored in the one or more memories to perform various operations according to the method flow of the embodiment of the present invention.

[0146] According to an embodiment of the present invention, electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to bus 504. Electronic device 500 may also include one or more of the following components connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 508 including a hard disk; and a communication section 509 including a network interface card such as a LAN card or modem. Communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. Removable media 511, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 510 as needed, so that computer programs read from the removable media can be installed into storage section 508 as needed.

[0147] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0148] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 502 and / or RAM 503 described above, and / or one or more memories other than ROM 502 and RAM 503.

[0149] Embodiments of the present invention also include a computer program product comprising a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code is configured to cause the computer system to implement the method for determining urban land parcel evacuation identification information provided in an embodiment of the present invention.

[0150] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when the computer program is executed by the processor 501. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0151] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 509, and / or installed from a removable medium 511. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0152] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509 and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above-described functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0153] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0155] It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or couplings fall within the scope of the present invention.

[0156] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A method for determining urban land evacuation identification information, characterized in that: The method comprises: The identification agent obtains an action of laying out evacuation identification based on the acquired plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, wherein the initial evacuation parameters are determined by the building information and object information in the plot information, the plot information includes the plot boundary information, evacuation location information, building information in the plot, and object information of the object to be evacuated of different plots in the target urban area, the building information includes the building layout information, building boundary information, and entrance and exit location information of the building, the object information includes the total number of people and population density corresponding to the plot to be identified or each building, the action of laying out evacuation identification includes the action of adding a new evacuation identification to the plot to be identified and the action of updating and adjusting the existing evacuation identification, the initial evacuation parameters are initial evacuation state parameters in the initial state, including the initial evacuation duration, the initial evacuation completion time corresponding to the buildings in the target urban area, and the initial point information and initial number of evacuation dense points; Executing the action in an environment corresponding to the land parcel to be identified, and determining a layout result and an evacuation status result of the evacuation sign corresponding to the land parcel to be identified, wherein the environment is obtained by processing the acquired land parcel information and initial evacuation parameters, the layout result of the evacuation sign represents the geometric parameters of multiple evacuation signs in the land parcel to be identified, including the center point coordinates, size information, height information, and orientation information of the evacuation sign, and the evacuation status result represents the evacuation status parameters achieved in the target urban area at the current moment according to the current layout result of the evacuation sign, including the evacuation duration, the evacuation completion time corresponding to the buildings in the target urban area, and the location information and number of evacuation dense points; The identification agent is trained in the following way: The initial strategy network obtains sample actions, sample evacuation identification layout results, sample evacuation state results, and sample distribution parameters based on sample plot information of the sample plot and sample initial evacuation parameters corresponding to the sample plot information, wherein the sample initial evacuation parameters are determined by building information and object information in the sample plot information; The initial value network evaluates the sample action, the layout result of the sample evacuation mark, and the sample evacuation state result based on the evaluation strategy to obtain an evaluation result; Based on the evaluation result, the sample distribution parameter and the loss function, the initial parameters of the initial policy network and the initial value network are updated to obtain the identification agent.

2. The method according to claim 1, characterized in that The building information includes plan information and evacuation information; the method further includes: Based on the plane information, the evacuation information and the object information, the evacuation duration and the evacuation target point information in the initial evacuation parameters are determined.

3. The method according to claim 1, characterized in that Executing the action in an environment corresponding to the land parcel to be identified, and determining a layout result and an evacuation status result of the evacuation sign corresponding to the land parcel to be identified, including: constructing a current environment based on the building information, current identification information at a current moment, and current evacuation information; Executing the position identification action, the size identification action, and the function identification action in the actions in the current environment to obtain a layout result of the current evacuation sign and a current evacuation state result; When the layout result of the current evacuation sign and the current evacuation status result meet the preset conditions, the action is stopped, and the layout result of the current evacuation sign and the current evacuation status result are respectively determined as the layout result of the evacuation sign and the evacuation status result and output.

4. The method according to claim 1, wherein The identification agent obtains an action of arranging evacuation identification based on the acquired land parcel information of the land parcel to be identified and the initial evacuation parameters corresponding to the land parcel information, including: Performing feature extraction on the plot information and the initial evacuation parameters respectively to obtain plot features and state features; Splicing the plot feature and the state feature to obtain a splicing vector; The identification agent determines a plurality of initial actions and a plurality of distribution parameters corresponding to the plurality of initial actions based on the splicing vector; The action is determined from the plurality of initial actions based on the plurality of distribution parameters and preset screening rules.

5. The method according to claim 1, wherein The initial strategy network obtains sample actions, sample evacuation identification layout results, sample evacuation state results, and sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, including: Processing the sample plot information and the sample initial evacuation parameters to obtain the sample action and the sample distribution parameters; Sampling the sample action based on the sample distribution parameter to obtain action information of the sample action; The sampling action is performed in a sample environment corresponding to the sample plot to obtain a layout result of the sample evacuation mark and a sample evacuation status result.

6. The method according to claim 5, characterized in that The method further comprises: Based on the layout result of the sample evacuation mark and the sample evacuation state result, the current state value of the sampling action at the current moment is determined to perform result evaluation.

7. The method according to claim 6, characterized in that The initial value network evaluates the sample action, the layout result of the sample evacuation mark, and the sample evacuation state result based on the evaluation strategy to obtain an evaluation result, including: The sample action, the sample evacuation sign layout result, and the sample evacuation state result are processed using the reward function in the evaluation strategy to determine a reward value after executing the sample action in the sample environment, as well as a subsequent sign layout result, a subsequent evacuation state result, and a subsequent state value at a subsequent time, wherein the subsequent state value is determined by the subsequent sign layout result and the subsequent evacuation state result; The evaluation result is determined based on the reward value, the current state value, the subsequent state value, and a weight corresponding to the subsequent state value.

8. The method according to claim 7, characterized in that The initial parameters include initial design parameters and initial evaluation parameters; Based on the evaluation result, the sample distribution parameter and the loss function, updating the initial parameters of the initial policy network and the initial value network to obtain the identification agent, including: Determine the ratio between the current distribution parameter at the current moment and the previous distribution parameter at the previous moment in the sample distribution parameters as the strategy loss probability ratio; determining a loss result based on the loss probability ratio, the evaluation result, and the loss function; The initial design parameters and the initial evaluation parameters are updated using the update strategy and the loss results to obtain target design parameters and target evaluation parameters, and the initial strategy network is updated using the target design parameters and the initial value network is updated using the target evaluation parameters to obtain the identification agent.

9. The method according to claim 8, characterized in that The loss function includes a strategy loss function and a value loss function; Determining a loss result based on the loss probability ratio, the evaluation result, and the loss function includes: Determining a strategy loss value and a value loss value based on the loss probability ratio, the evaluation result, the strategy loss function, and the value loss function; The loss result is obtained by weighting and summing the strategy loss value and the value loss value respectively using the loss weights corresponding to the strategy loss value and the value loss value.

10. A device for determining urban land evacuation identification information, characterized in that: The device comprises: an action determination module for identifying an intelligent agent, based on the acquired plot information of a plot to be identified and initial evacuation parameters corresponding to the plot information, to obtain an action for arranging evacuation identification, wherein the initial evacuation parameters are determined by the building information and object information in the plot information, the plot information including the plot boundary information, evacuation location information, building information in the plot, and object information of the object to be evacuated of different plots in the target urban area, the building information including the building layout information, building boundary information, and entrance and exit location information of the building, the object information including the total number of people and population density corresponding to the plot to be identified or each building, the action for arranging evacuation identification including the action of adding a new evacuation identification to the plot to be identified and the action of updating and adjusting the existing evacuation identification, the initial evacuation parameters being initial evacuation state parameters in the initial state, including the initial evacuation duration, the initial evacuation completion time corresponding to the buildings in the target urban area, and the initial point information and initial number of evacuation density points; a result determination module for executing the action in an environment corresponding to the land parcel to be identified, and determining a layout result and an evacuation status result of the evacuation sign corresponding to the land parcel to be identified, wherein the environment is obtained by processing the acquired land parcel information and initial evacuation parameters; the layout result of the evacuation sign represents the geometric parameters of multiple evacuation signs in the land parcel to be identified, including the center point coordinates, size information, height information, and orientation information of the evacuation sign; and the evacuation status result represents the evacuation status parameters achieved in the target urban area at the current moment according to the current layout result of the evacuation sign, including the evacuation duration, the evacuation completion time corresponding to the buildings in the target urban area, and the location information and number of evacuation dense points; The identification agent is trained in the following way: The initial strategy network obtains sample actions, sample evacuation identification layout results, sample evacuation state results, and sample distribution parameters based on sample plot information of the sample plot and sample initial evacuation parameters corresponding to the sample plot information, wherein the sample initial evacuation parameters are determined by building information and object information in the sample plot information; The initial value network evaluates the sample action, the layout result of the sample evacuation mark, and the sample evacuation state result based on the evaluation strategy to obtain an evaluation result; Based on the evaluation result, the sample distribution parameter and the loss function, the initial parameters of the initial policy network and the initial value network are updated to obtain the identification agent.

Citation Information

Patent Citations

  • Crowd evacuation method and system based on multi-carrier intelligent guidance

    CN111767789A

  • Evaluation method and system for design rationality of emergency evacuation indication sign system

    CN115688386A