Urban parcel evacuation identification information determination method and device

Through the coordinated work of the policy network and value network of the identification agent, the evacuation identification information of urban plots is determined in real time, which solves the problem of high consumption of manual and computing resources in the existing technology, and realizes efficient and convenient evacuation identification design.

CN120372258AActive Publication Date: 2025-07-25TIANJIN UNIV

Patent Information

Application Number
CN202510872952.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In the prior art, urban plot evacuation mark design relies on manual design to lack quantitative indicators, and traditional reinforcement learning methods are limited to indoor evacuation. Numerical simulation software has a large amount of calculation and takes a long time, so it cannot quickly respond to changes in actual evacuation scenarios, and the threshold for use is high.

Method used

The identification agent is used to work collaboratively through the policy network and the value network, and based on the plot information and initial evacuation parameters, the layout and status results of the evacuation identifier are determined in real time, and the identification agent is obtained through reinforcement learning training to avoid dependence on mathematical models and algorithms.

Benefits of technology

Real-time determination of urban plot evacuation sign information is realized, computing resource consumption is reduced, operation is simple and convenient, and the universality of evacuation sign design and optimization is improved, and the dependence on professional knowledge is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372258A_ABST
    Figure CN120372258A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for determining urban plot evacuation identification information, which can be applied to the technical field of artificial intelligence and emergency evacuation. The method comprises the steps that an identification agent obtains an action of layout evacuation identification based on land parcel information and initial evacuation parameters; executing an action in the environment, and determining a layout result of the evacuation identifier and an evacuation state result; the initial strategy network obtains a sample action, a layout result of a sample evacuation identifier, a sample evacuation state result and a sample distribution parameter based on the sample plot information and the sample initial evacuation parameter; the initial value network evaluates a sample action based on an evaluation strategy, a layout result of a sample evacuation identifier and a sample evacuation state result to obtain an evaluation result; and based on the evaluation result, the sample distribution parameter and the loss function, updating respective initial parameters of the initial policy network and the initial value network to obtain an identified agent. Through the method, dependence on a mathematical model and an algorithm can be reduced, and consumption of computing resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and emergency evacuation, and particularly relates to a method and device for determining evacuation identification information of urban plots. Background Art

[0002] Optimizing the design of evacuation identification can provide clearer and more accurate direction guidance for evacuees in the urban emergency state, enabling them to quickly determine the evacuation route and thus accelerating the evacuation speed. In related technologies, the methods for optimizing the design of urban evacuation identification mainly rely on manual design or rule-based optimization algorithms. For example, classifying evacuation data, comparing the instruction issuance time, determining the positions to be evacuated, using different emergency identification devices to remind people at each position to evacuate, or combining Legion simulation software to conduct real-time simulation and analysis of the crowd flow in the subway station to provide the optimal evacuation route; planning the evacuation route of leaders based on reinforcement learning, improving the crowd evacuation efficiency through the collaborative guidance strategy and simulation test verification of evacuation identification and leaders; adopting a clustering algorithm to reasonably classify newly discovered factors to simplify the complex data and obtain the evacuation evaluation result.

[0003] However, in related technologies, manual design relies on people's subjective feelings and lacks quantitative indicators; traditional reinforcement learning methods are limited to the design of indoor evacuation identification and cannot be applied to the emergency evacuation of urban plots; numerical simulation software or tools rely on complex mathematical models and algorithms to simulate crowd behavior, with a large amount of calculation and long time consumption for each simulation, unable to quickly respond to changes in the actual evacuation scenario, and the simulation tools usually require professional knowledge and skills for setting and running, increasing the usage threshold. Summary of the Invention

[0004] In view of the above problems, the present invention provides a method, device, equipment, medium and program product for determining evacuation identification information of urban plots.

[0005] According to the first aspect of the present invention, a method for determining evacuation identification information of urban plots is provided, including: the identification agent obtains the action of arranging evacuation identification based on the plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, where the initial evacuation parameters are determined by the building information and object information in the plot information; execute the action in the environment corresponding to the plot to be identified to determine the layout result and evacuation status result of the evacuation identification corresponding to the plot to be identified; wherein, the identification agent is trained in the following manner: the initial policy network obtains the sample action, the layout result of the sample evacuation identification, the sample evacuation status result and the sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information.

[0006] Preferably, the initial evacuation parameters of the sample are determined by the building information and object information in the plot information; the initial value network evaluates the sample actions, the layout results of the sample evacuation signs, and the sample evacuation status results based on the evaluation strategy to obtain the evaluation results; based on the evaluation results, the sample distribution parameters, and the loss function, the initial parameters of the initial policy network and the initial value network are updated respectively to obtain the sign agent.

[0007] The second aspect of the present invention provides a device for determining urban plot evacuation sign information, including: an action determination module, configured to enable the sign agent to obtain the action of arranging evacuation signs based on the plot information of the plot to be signed and the corresponding initial evacuation parameters, wherein the initial evacuation parameters are determined by the building information and object information in the plot information; a result determination module, configured to execute the action in the environment corresponding to the plot to be signed and determine the layout result and the evacuation status result of the evacuation signs corresponding to the plot to be signed.

[0008] Preferably, the sign agent is trained in the following manner: the initial policy network obtains the sample actions, the layout results of the sample evacuation signs, the sample evacuation status results, and the sample distribution parameters based on the sample plot information of the sample plot and the corresponding sample initial evacuation parameters, wherein the sample initial evacuation parameters are determined by the building information and object information in the plot information; the initial value network evaluates the sample actions, the layout results of the sample evacuation signs, and the sample evacuation status results based on the evaluation strategy to obtain the evaluation results; based on the evaluation results, the sample distribution parameters, and the loss function, the initial parameters of the initial policy network and the initial value network are updated respectively to obtain the sign agent.

[0009] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory, configured to store one or more computer programs, wherein the above-mentioned one or more processors execute the above-mentioned one or more computer programs to implement the steps of the above-mentioned method.

[0010] The fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the above-mentioned computer program or instruction is executed by a processor, the steps of the above-mentioned method are implemented.

[0011] The fifth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the above-mentioned computer program or instruction is executed by a processor, the steps of the above-mentioned method are implemented.

[0012] According to an embodiment of the present invention, the actions that the identification agent needs to perform currently can be determined in real time through the plot information of the plot to be identified and the initial evacuation parameters, so as to flexibly obtain the layout result and evacuation status result of the evacuation identification after performing the actions. Since the identification agent is obtained through a closed-loop update process formed by the collaborative work of the policy network and the value network, the value network can provide real-time feedback on the evacuation status of urban plots, and the policy network can adjust the identification layout in a timely manner according to the feedback information, realizing real-time interaction among the evacuation layout result, the evacuation status result, and the corresponding evaluation result, improving the overall performance and effect of the identification agent. By means of the identification agent automatically determining the evacuation identification information of urban plots, the dependence on mathematical models and algorithms is avoided, and the consumption of computing resources is reduced. The way for the identification agent to automatically determine the evacuation identification information does not require users to spend a lot of time learning and mastering professional knowledge and skills, and the operation is simple and convenient, further improving the universality of the evacuation identification design and optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present invention will become more clear.

[0014] Figure 1 The application scenario diagram of the method, device, equipment, medium, and program product for determining the evacuation identification information of urban plots according to an embodiment of the present invention is shown.

[0015] Figure 2 The flowchart of the method for determining the evacuation identification information of urban plots according to an embodiment of the present invention is shown.

[0016] Figure 3 The example schematic diagram of the proximal policy optimization algorithm based on the policy-value framework for realizing the optimization and update process of the identification agent according to an embodiment of the present invention is shown.

[0017] Figure 4 The structural block diagram of the device for determining the evacuation identification information of urban plots according to an embodiment of the present invention is shown.

[0018] Figure 5 The block diagram of the electronic device suitable for implementing the method for determining the evacuation identification information of urban plots according to an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.

[0020] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "comprising", "including" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0022] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning that those skilled in the art usually understand such expressions (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0023] In the process of conceiving the present invention, the inventors found that in the related art, manual design relies on human subjective feelings and lacks quantitative indicators; the related reinforcement learning methods are limited to the design of indoor evacuation signs and cannot be used for the emergency evacuation of urban plots; numerical simulation software or tools rely on complex mathematical models and algorithms to simulate crowd behavior, with a large amount of calculation and a long time-consuming for each simulation, and cannot quickly respond to changes in actual evacuation scenarios; and simulation tools usually require professional knowledge and skills to set up and run, increasing the usage threshold.

[0024] In view of the above technical problems, the present invention provides a method for determining evacuation identification information of urban plots, including: an identification agent obtains an action of arranging evacuation identifications based on the plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, where the initial evacuation parameters are determined by the building information and object information in the plot information; performing the action in the environment corresponding to the plot to be identified to determine the layout result and evacuation status result of the evacuation identifications corresponding to the plot to be identified; wherein, the identification agent is trained in the following manner: an initial policy network obtains a sample action, a layout result of sample evacuation identifications, a sample evacuation status result, and sample distribution parameters based on the sample plot information of a sample plot and the sample initial evacuation parameters corresponding to the sample plot information, where the sample initial evacuation parameters are determined by the building information and object information in the plot information; an initial value network evaluates the sample action, the layout result of sample evacuation identifications, and the sample evacuation status result based on an evaluation policy to obtain an evaluation result; based on the evaluation result, the sample distribution parameters, and a loss function, the initial parameters of the initial policy network and the initial value network are updated respectively to obtain the identification agent.

[0025] According to an embodiment of the present invention, through the plot information of the plot to be identified and the initial evacuation parameters, the action that the identification agent currently needs to perform can be determined in real time, so as to flexibly obtain the layout result and evacuation status result of the evacuation identifications after performing the action. Since the identification agent is obtained through a closed-loop update process formed by the collaborative work of the policy network and the value network, the value network can provide real-time feedback on the evacuation status of the urban plot, and the policy network can adjust the identification layout in a timely manner according to the feedback information, realizing real-time interaction between the evacuation layout result, the evacuation status result, and the corresponding evaluation result, improving the overall performance and effect of the identification agent. By means of the identification agent automatically determining the evacuation identification information of urban plots, the dependence on mathematical models and algorithms is avoided, and the consumption of computing resources is reduced. The way that the identification agent automatically determines the evacuation identification information does not require users to spend a lot of time learning and mastering professional knowledge and skills, and the operation is simple and convenient, further improving the universality of evacuation identification design and optimization.

[0026] Figure 1 The application scenario diagram of the method, device, equipment, medium, and program product for determining evacuation identification information of urban plots according to an embodiment of the present invention is shown.

[0027] As Figure 1 shown, the application scenario 100 according to this embodiment may include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0028] A user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. The terminal device 101 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, and so on. For example, the user can send the plot information and related parameters of the plot to be identified to the agent through the terminal device 101.

[0029] The server 103 can be a server providing various services. For example, the server 103 receives various types of data from the terminal device 101 and stores them in the database. These data include real-time environmental data, personnel distribution, device status, and user feedback, etc., providing data support for subsequent analysis and decision-making. For example, the server 103 runs an evacuation identification information determination system based on the agent. The agent determines the evacuation identification information of the urban plot according to the received data.

[0030] It should be noted that the method for determining the evacuation identification information of the urban plot provided by the embodiments of the present invention can generally be executed by the server 103. Correspondingly, the device for determining the evacuation identification information of the urban plot provided by the embodiments of the present invention can generally be set in the server 103. The method for determining the evacuation identification information of the urban plot provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 103 and capable of communicating with the terminal device 101 and / or the server 103. Correspondingly, the device for determining the evacuation identification information of the urban plot provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 103 and capable of communicating with the terminal device 101 and / or the server 103.

[0031] It should be understood that Figure 1 the numbers of the terminal device, the network, and the server in

[0032] Figure 2 are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0032] Figure 2 shows a flowchart of the method for determining the evacuation identification information of the urban plot according to the embodiments of the present invention.

[0033] As Figure 2 shown, the method for determining the evacuation identification information of the urban plot in this embodiment includes operation S210 to operation S220.

[0034] In operation S210, the identification agent obtains the action of arranging the evacuation identification based on the obtained plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, where the initial evacuation parameters are determined by the building information and object information in the plot information.

[0035] In an embodiment of the present invention, the identification agent may include a policy agent and a value agent, which can perform evacuation path guidance, real-time information interaction, and evacuation sign design and optimization for the to-be-identified plot according to the information obtained in real time, and is an entity that executes decisions in the reinforcement learning framework. The plot information may include the plot boundary information, evacuation location information, building information in the plot, and object information of the to-be-evacuated objects of different plots in the target urban area. The building information may include building layout information, building boundary information, and the location information of the entrances and exits of the building. The object information may include the total number of people and population density corresponding to the to-be-identified plot or each building.

[0036] The actions of arranging evacuation signs may include the action of adding new evacuation signs in the plot and the action of updating and adjusting the existing evacuation signs. The initial evacuation parameters may be the initial evacuation state parameters in the initial state, and may include the initial evacuation duration, the initial evacuation completion time corresponding to each building in the urban area, the initial point position information and the initial number of evacuation dense points.

[0037] For example, by extracting features from the plot information and the initial evacuation parameters, different features are obtained, and then the different features are spliced to obtain a splicing result. Furthermore, based on the reinforcement learning and the splicing result, multiple initial actions are obtained, and a suitable action is selected from the multiple initial actions as the action of arranging evacuation signs.

[0038] In operation S220, perform an action in the environment corresponding to the to-be-identified plot, and determine the layout result of the evacuation signs and the evacuation state result corresponding to the to-be-identified plot.

[0039] In an embodiment of the present invention, the environment may be obtained by processing the acquired plot information and the initial evacuation parameters. The layout result of the evacuation signs may characterize the geometric parameters of multiple evacuation signs in the to-be-identified plot, and may include the center point coordinates, size information, height information, and pointing information of the evacuation signs. The evacuation state result may characterize the evacuation state parameters achieved in the urban area at the current moment according to the current layout result of the evacuation signs, and may include the evacuation duration, the evacuation completion time corresponding to each building in the urban area, the point position information and the number of evacuation dense points.

[0040] For example, splice the plot information and the initial evacuation parameters obtained at the current moment to obtain the current environment of the to-be-identified plot, and perform the action of arranging evacuation signs in the current environment to obtain the layout result of the current evacuation signs and the current evacuation state result.

[0041] Preferably, the identification agent is trained in the following manner, which may include operation S221 to operation S223.

[0042] In operation S221, the initial policy network obtains a sample action, a layout result of sample evacuation identifiers, a sample evacuation status result, and sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, where the sample initial evacuation parameters are determined by the building information and object information in the sample plot information.

[0043] In operation S222, the initial value network evaluates the sample action, the layout result of the sample evacuation identifiers, and the sample evacuation status result based on the evaluation policy to obtain an evaluation result.

[0044] In operation S223, based on the evaluation result, the sample distribution parameters, and the loss function, the initial parameters of the initial policy network and the initial value network are updated respectively to obtain an identifier agent.

[0045] In an embodiment of the present invention, the identifier agent can be obtained by training an initial identifier agent, and the initial identifier agent can include an initial policy agent and an initial value agent. The initial policy network can be a reinforcement learning model corresponding to the initial policy agent, which can be used to output an action policy, and the goal is to optimize the policy parameters to maximize the cumulative reward at a later time; the initial value network can be a reinforcement learning model corresponding to the initial value agent, which can learn the relationship between the environment and the reward, and guide the initial policy agent to perform policy output according to the potential reward of the current state. The environment can be a system for urban building layout, identifier layout, and evacuation simulation.

[0046] For example, the urban evacuation optimization environment can be initialized, and the initial state can be set as a plot containing only the building layout (without identifier arrangement, i.e., the plot to be identified), and evacuation identifiers can be placed one by one through reinforcement learning. The parameters of the evacuation identifier can include the center point position coordinates (x si , y si ), the identifier length, the identifier width, and the height from the ground of the identifier (h). The position constraint condition of the evacuation identifier can be located within a range of a meters offset from the outer contour line of the building; the function of the evacuation identifier can be to indicate the evacuable direction of this position and increase the evacuation speed of personnel.

[0047] Evacuation simulation can be performed on the initial plot through the input of the personnel density distribution of each building to obtain the initial evacuation state; the evacuation state can include the total evacuation time, the time for all personnel in each building to evacuate, the coordinates and total number of all evacuation congestion points during the evacuation process; and then the initial plot and the evacuation state are used as the environmental input of reinforcement learning to establish an urban evacuation optimization environment.

[0048] For example, an action set can be defined, and the action set can include the change amount of the center point coordinates of each identifier, the change amount of the length and width, the change amount of the height from the ground, and the identifier directivity.

[0049] In an embodiment of the present invention, a reinforcement learning model can be constructed through a policy optimization algorithm to train an identification agent. The policy optimization algorithm can predict an action policy through a neural network to achieve the placement and adjustment of dynamic identification. Combining the policy-value framework, the policy network is responsible for outputting the action policy, and the goal is to optimize the policy parameters to maximize the future cumulative reward. The value network can evaluate the potential reward of the current state by learning the relationship between the environment and the reward, and guide the policy network to perform policy output.

[0050] For example, multiple sets of random identification layouts can be generated based on reinforcement learning as a training data set; thus, the training data set is used to drive the interaction between the reinforcement learning model (which can include a policy network and a value network), an environment modeling system, and an evacuation simulation system to optimize the model policy; the interaction process is repeated until the termination condition is met to complete the training and obtain a trained reinforcement learning model.

[0051] In an embodiment of the present invention, the sample plot can be a plot similar to the plot to be identified during a historical period. The evaluation policy can be a policy that predicts the state value of the current state and calculates the action advantage of the current action based on the reward function of the initial identification agent. The evaluation result can include the state value and the action advantage. The loss function can include the loss functions corresponding to the initial policy network and the initial value network respectively. The initial parameters can include the initial policy parameters and the initial value parameters. It can be understood that the information corresponding to the sample plot information, the sample initial evacuation parameters, and the sample actions, etc. during the training stage is the same or similar to the information corresponding to the application stage, and will not be elaborated here specifically.

[0052] For example, feature extraction and splicing are performed on the sample plot information and the sample initial evacuation parameters to obtain the sample action and the sample distribution parameters; thus, the action information of the sampled action is obtained by randomly sampling the sample action based on the sample distribution parameters, and then the sampled action is executed in the sample environment to obtain the layout result of the sample evacuation identification and the sample evacuation state result.

[0053] According to an embodiment of the present invention, the actions that the identification agent currently needs to perform can be determined in real time through the plot information of the plot to be identified and the initial evacuation parameters, so as to flexibly obtain the layout result and evacuation status result of the evacuation identification after executing the actions. Since the identification agent is obtained through a closed-loop update process formed by the collaborative work of the policy network and the value network, the value network can provide real-time feedback on the evacuation status of urban plots, and the policy network can adjust the identification layout in a timely manner according to the feedback information, realizing real-time interaction among the evacuation layout result, the evacuation status result, and the corresponding evaluation result, improving the overall performance and effect of the identification agent. By means of the identification agent automatically determining the evacuation identification information of urban plots, the dependence on mathematical models and algorithms is avoided, and the consumption of computing resources is reduced. The way that the identification agent automatically determines the evacuation identification information does not require users to spend a lot of time learning and mastering professional knowledge and skills, and the operation is simple and convenient, further improving the universality of the evacuation identification design and optimization.

[0054] According to an embodiment of the present invention, the building information includes plane information and evacuation information; the method further includes: determining the evacuation duration and evacuation target point position information in the initial evacuation parameters based on the plane information, the evacuation information, and the object information.

[0055] In an embodiment of the present invention, the plane information may include building layout information and building boundary information. The evacuation information may include the position information of the entrances and exits of the building and the building evacuation path information. The evacuation duration information may include the total duration for the plot to be identified (plot to be evacuated) to complete the evacuation and the duration for each individual building in the plot to complete the evacuation respectively. The evacuation target point position information may include the position information and quantity information of the target points. The target points may be the points where the population density is greater than or equal to the preset density threshold during the emergency evacuation process.

[0056] For example, evacuation simulation can be performed on the initial plot based on the input population density distribution of each building, and through the evacuation simulation tool, the initial evacuation status can be calculated, including the total evacuation time Time t , the time t for all the people in each building to evacuate m , the position coordinates (x 2 , y ci ) where the population density exceeds 2 people / m ci during the evacuation process, and the total number c of the points.

[0057] For example, each floor of the building can be divided into multiple areas, a state network can be constructed, each area can be abstracted into a state space, and the stairs can be abstracted into state paths. By analyzing the transfer situation of people between different areas and channels, the analysis model can be used to dynamically analyze the personnel evacuation process, so as to determine potential congestion points and evacuation duration.

[0058] According to an embodiment of the present invention, a dynamic optimization model is established based on the state network theory. By assuming area division and introducing concepts such as pedestrian flow density, the functional relationship between pedestrian flow density and walking speed can be determined. The calculation formulas for the state transition amounts of horizontal channels, stairs, and doors are calculated, and the evacuation path is optimized and adjusted in combination with the shortest path algorithm, meeting the effective identification and handling of complex situations such as traffic congestion during the evacuation process and improving the rationality and feasibility of the evacuation plan.

[0059] According to an embodiment of the present invention, actions are performed in the environment corresponding to the plot to be marked, and the layout result and evacuation state result of the evacuation mark corresponding to the plot to be marked are determined, including: constructing the current environment based on building information, the current mark information at the current moment, and the current evacuation information; performing position marking actions, size marking actions, and function marking actions in the current environment to obtain the layout result of the current evacuation mark and the current evacuation state result; when the layout result of the current evacuation mark and the current evacuation state result meet the preset conditions, stop the actions, and respectively determine the layout result of the current evacuation mark and the current evacuation state result as the layout result and the evacuation state result of the evacuation mark and output them.

[0060] In an embodiment of the present invention, the position marking action may include an action of marking the central point coordinates (x si , y si ) and the height from the ground of the mark (h i ) of the evacuation mark. The size marking action may include an action of marking the length (l i ) and width (w i ) of the evacuation mark. The function marking action may represent an action of marking the directivity (θ i ) of the evacuation mark. The preset condition may be a condition defined by the user according to actual needs. For example, when the total number of evacuation marks laid out in this area exceeds the quantity threshold, stop executing the actions.

[0061] For example, obtain the building layout information and evacuation state information of the plot to be marked, and load the trained mark intelligent agent; thus, use the building layout information and evacuation state information as the input of the mark intelligent agent at the current moment to optimize the mark layout of the urban area, output the optimal mark layout of the area at the current moment, and update the building layout information (the building layout information after the mark action has been executed) and the current evacuation state result; furthermore, when the total number of evacuation marks in the entire area is greater than the quantity threshold (for example, 100), stop executing the actions and output the evacuation mark layout result and the evacuation state result.

[0062] According to an embodiment of the present invention, the identification agent obtains the action of arranging evacuation signs based on the obtained plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, including: respectively extracting features from the plot information and the initial evacuation parameters to obtain plot features and state features; concatenating the plot features and the state features to obtain a concatenated vector; the identification agent determines a plurality of initial actions and a plurality of distribution parameters corresponding to the plurality of initial actions based on the concatenated vector; and determines an action from the plurality of initial actions based on the plurality of distribution parameters and a preset screening rule.

[0063] In an embodiment of the present invention, the plot features and the state features may be a plot vector and a state vector obtained by extracting features from the plot information and the initial evacuation parameters and satisfying the processing of the identification agent. The concatenated vector may be a vector obtained by concatenating, combining or fusing the plot vector and the state vector. The distribution parameter may be an action distribution parameter obtained after performing a plurality of initial actions in the plot to be identified. The preset screening rule may include any one of screening according to a probability value and screening according to an advantage.

[0064] For example, extract features from the obtained building layout and evacuation state information, and convert them into plot features and state features that the identification agent can process; thereby combine the plot features and the state features into a concatenated vector as the input information of the identification agent and input it into the identification agent; the identification agent outputs a plurality of initial actions, action distribution parameters, and probability values corresponding to the plurality of initial actions; and then selects the initial action corresponding to the highest probability value as the final action to be executed.

[0065] According to an embodiment of the present invention, through the identification agent, the identification action strategy can be dynamically determined and optimized according to the real-time obtained plot information and initial evacuation parameters, so as to guide the optimal path for the evacuees in real time, make the evacuation path more direct and efficient, and thus shorten the overall evacuation time.

[0066] In a feasible embodiment, the method for determining the evacuation identification information of urban plots based on the identification agent may include operation S310 to operation S340.

[0067] In operation S310, obtain the building information, object information, and initial evacuation parameters of the unmarked initial plot (plot to be identified).

[0068] In operation S320, take the building information, object information, and initial evacuation parameters at the current moment as the input information of the identification agent, apply the trained identification agent to optimize and update the identification layout of the target urban area, and output the layout result and evacuation state result of the optimal evacuation identification in the area at the current moment.

[0069] In operation S330, the action of S320 is repeatedly executed until the total number of evacuation signs in the target urban area is greater than or equal to the quantity threshold, and then the sign action is stopped.

[0070] In operation S340, the final layout result of the evacuation signs and the evacuation status result are output.

[0071] It can be understood that the method of applying the sign agent to determine the evacuation sign information of urban plots has been described above. Next, the process of training the sign agent will be described.

[0072] The sample sign agent obtains the sign agent for determining the evacuation sign information of urban plots through the interaction with the sample environment, the update process of the initial policy network and the initial value network, and the process of the policy optimization algorithm.

[0073] According to an embodiment of the present invention, the initial policy network obtains a sample action, a layout result of sample evacuation signs, a sample evacuation status result, and a sample distribution parameter based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, including: processing the sample plot information and the sample initial evacuation parameters to obtain a sample action and a sample distribution parameter; sampling from the sample actions based on the sample distribution parameter to obtain the action information of the sampled action; executing the sampled action in the sample environment corresponding to the sample plot to obtain the layout result of the sample evacuation signs and the sample evacuation status result.

[0074] In an embodiment of the present invention, the sample distribution parameter may be the action distribution parameter output by the initial policy agent. The sampled action can be obtained by random sampling or sampling the sample actions based on experience. It should be noted that the relevant contents such as the sample actions, the layout results of the sample evacuation signs, the sample evacuation status results, and the sample environment in the training stage can refer to the corresponding contents in the application stage; similarly, the relevant contents in the application stage can also refer to the corresponding contents in the training stage, which will not be elaborated here.

[0075] According to an embodiment of the present invention, the method for determining evacuation sign information further includes: determining the current state value of the sampled action at the current moment based on the layout result of the sample evacuation signs and the sample evacuation status result for result evaluation.

[0076] In an embodiment of the present invention, the current state value may be the state value obtained by inputting the layout result of the sample evacuation signs and the sample evacuation status result at the current moment into the value network.

[0077] For example, the sample environment can receive the sample action a t (such as the new sign parameter [x, y, l, w, h, θ]), and output the state result s t(including the layout result of the sample evacuation signs and the sample evacuation status result) and the immediate reward value r t , and generate the next state result s t+1 .

[0078] Furthermore, based on the trajectory data, use the current policy The experience data collected from the environment (s t , a t , r t , s t+1 ) is used for subsequent calculations. s t As the current state result, it is input from the environment to the policy network and the value network for predicting actions and values.

[0079] On this basis, the initial policy network in the identification agent can output sample actions t by receiving the state result s , that is, the initial policy network predicts the sign placement or adjusts the parameters according to the current layout; the initial value network outputs the current state value Vφ(s t ) based on the received current state result s t . The state value can represent the future cumulative reward estimate of the current state result.

[0080] According to the embodiments of the present invention, the value evaluation of the current state by the initial value network can provide more accurate feedback to the initial policy network, enabling the initial policy network to output personalized sign strategies according to different regions and different personnel density situations. For example, for areas with a high personnel density, increase the density and prominence of evacuation instructions; for areas far from the safety exit, provide more explicit guidance directions and distance information to achieve intelligent and accurate evacuation guidance. Based on the policy optimization update method, the policy update of the policy network is more stable, avoiding sign indication chaos caused by frequent and large fluctuations in the policy.

[0081] According to the embodiments of the present invention, the initial value network evaluates the sample actions, the layout result of the sample evacuation signs, and the sample evacuation status result based on the evaluation policy to obtain an evaluation result, including: using the reward function in the evaluation policy to process the sample actions, the layout result of the sample evacuation signs, and the sample evacuation status result to determine the reward value after performing the sampling action in the sample environment, as well as the subsequent sign layout result, the subsequent evacuation status result, and the subsequent state value at a later time, where the subsequent state value is determined by the subsequent sign layout result and the subsequent evacuation status result; determining the evaluation result based on the reward value, the current state value, the subsequent state value, and the weight corresponding to the subsequent state value.

[0082] In an embodiment of the present invention, the reward function may be an immediate reward function corresponding to a time step. The weights corresponding to the state values at different times may be determined according to actual situations or empirical values, and are not specifically limited herein. The evaluation result may include the state value V φ (st) obtained by evaluating with the value network and the advantage result A of the sample action t .

[0083] Figure 3 FIG. shows an exemplary schematic diagram of implementing the optimization and update process of an identification agent based on the proximal policy optimization algorithm of the policy-value framework according to an embodiment of the present invention

[0084] As Figure 3 shown, the identification agent 301 may include a policy network 301a and a value network 301b. The process of implementing the update and optimization of the policy network and the value network may include: randomly initializing the parameters θ of the policy network 301a and the parameters φ of the value network 301b; the identification agent 301 may interact with the environment 303 based on the action space 302 of the identification agent to obtain an initial state s t , and through the policy network, according to the current policy πθ(a t ∣s t ), select a target action a from the action space 302 t , and then execute the action a in the environment t , obtain the current reward value r t and the updated state result s t+1 . Among them, the action space 302 is also called the action set

[0085] Based on the trajectory data 304, the experience data (s old , a t , r t , s t ) collected from the environment using the current policy πθ t+1 can be used for subsequent calculations. s t As the current state result, it is input from the environment to the policy network and the value network to predict the action and the state result. The value network 301b can be used to predict the value V t of the current state s ϕ (s t ) and the value V t+1 corresponding to the predicted next state result s ϕ (s t+1 ), and calculate the temporal difference error. The advantage function A t can be directly estimated by the temporal difference error δ t .

[0086] The update of the value function is achieved by updating the policy, calculating the probability ratio 305 and the policy loss 306, and calculating the value loss 307. The policy loss 306 can be determined based on the probability ratio 305 and the advantage function 308. Thus, the total loss value can be obtained based on the policy loss 306 and the value loss 307. Furthermore, the Adam optimizer is used to update the parameters θ and φ of the policy network and the value network according to the total loss L. The above steps are repeated until the preset number of training rounds or the convergence condition is reached.

[0087] For example, the reinforcement learning model of the identification agent can be constructed in the following way. First, the state space of the identification agent is defined. The state space can be a dynamic vector S, which changes with the placed evacuation signs and includes the geometric parameters of the currently placed signs and the evacuation state obtained through simulation calculation.

[0088] The geometric parameters of the signs can include: the coordinates of the center point of each sign (x si , y si ), length (l i ), width (w i ), height from the ground (h i ), and sign directivity (θ i ). If n signs have been placed, the geometric parameter part can be: [x1, y1, l1, w1, h1, θ1,…x n , y n , l n , w n , h n , θ n . The evacuation state parameters can be obtained through simulation calculation by the evacuation simulation system. The total evacuation time is the time Time when the last individual to be evacuated in the urban area reaches the shelter location t ; the evacuation time of all personnel in each building is the time when the last individual to be evacuated in each building leaves the building. If there are m buildings, there are m building evacuation times, denoted as t m ; the coordinates (x 2 ), y ci of the points where the density of all personnel exceeds the preset density threshold (e.g., 2 people / m ci ) during the evacuation process and the number of points c.

[0089] It can be understood that the specific type of the evacuation simulation system can be determined according to actual needs, with the standard of meeting the simulated evacuation of the plot to be marked. The specific type is not limited here.

[0090] Furthermore, the action set A of the identification agent can be defined. The action set A can include placing new evacuation signs. Placing new evacuation signs can specify parameters for the p-th sign: [x sp, y sp , l p , w p , h p , θ p , where x sp , y sp can be the coordinates of the center point, restricted within the plot boundary and within the range of n meters offset from the building outline. n can be determined according to the specific urban area type and building size.

[0091] In the embodiments of the present invention, the parameter updates of the initial policy network and the initial value network can be achieved through policy gradients and value losses.

[0092] For example, the reward function of the identification agent can be determined by the following formula (1):

[0093] (1);

[0094] Where can be to minimize the total evacuation time, Time t can be the total time for the last individual to be evacuated to the shelter location in the urban area; β·b can be to minimize the number of identifications, b can be the number of identifications; γ·c can be to minimize the number of crowd congestion points, c can be the number of congestion points; can be the reward term for the identification being close to the congestion point; can be the reward term for the identification being close to the "slow evacuation building"; are the weights of different rewards. Where and are determined as shown in the following formulas (2)-(3):

[0095] (2);

[0096] (3);

[0097] Where t i can be the current time step, T can be the total time step, (x mi , y mi ) can be the coordinates of the building center point, (x si , y si ) can be the coordinates of the identification center point, (x ci , y ci ) can be the coordinates of the congestion point, σ 2 is a tuning parameter that can control the decay degree of the exponential function exp and can be obtained based on experience or experiments.

[0098] For example, the method for updating the initial policy network parameters based on policy gradients may include: First, use the value network to calculate the current state value V φ (s t ) and the subsequent state value V φ (s t+1 ), and calculate the advantage through Generalized Advantage Estimation (GAE), as shown in the following formulas (4)-(5):

[0099] (4);

[0100] (5);

[0101] Where A t can represent the advantage function, indicating how much higher or lower the selected sample action a t is compared to the average level under the current state result s t ; k can represent the number of future time steps considered starting from time step t, used to calculate the multi-step return, and δ t can represent the difference between the predicted value and the actual value at time step t.

[0102] δ t+K can be the temporal difference error, representing the difference between the predicted value and the actual value at time step t + k; r t can be the immediate reward at time step t (i.e., the previously designed reward function); V ϕ (s t+1 ) can be the state value estimate at a subsequent time; γ' can be the weight of V φ (s t+1 ), also known as the discount factor, for example, taking 0.99, and λ can be the GAE parameter, for example, taking 0.95.

[0103] It can be understood that the reward function can comprehensively consider the number of evacuation signs, the total evacuation time, the time for all personnel in each building to evacuate, minimize the number of evacuation congestion locations, encourage optimizing the evacuation ability and meet the constraint conditions.

[0104] According to the embodiments of the present invention, the Proximal Policy Optimization (PPO) algorithm can more efficiently utilize the collected samples during each update, update the policy through multiple iterations, improve the utilization rate of samples, thereby accelerating the convergence process, achieving a better balance between exploring new policies and exploiting existing knowledge, ensuring both sufficient exploration and effectively using existing information to accelerate convergence.

[0105] According to an embodiment of the present invention, the initial parameters include initial design parameters and initial evaluation parameters; based on the evaluation result, the sample distribution parameter, and the loss function, update the initial parameters of the initial policy network and the initial value network respectively to obtain an identification agent, including: determining the ratio between the current distribution parameter at the current moment and the previous distribution parameter at the previous moment in the sample distribution parameter as the policy loss probability ratio; determining the loss result based on the policy loss probability ratio, the evaluation result, and the loss function; using the updated policy and the loss result to update the initial design parameters and the initial evaluation parameters to determine the target design parameters and the target evaluation parameters, so as to update the initial policy network using the target design parameters and update the initial value network using the target evaluation parameters to obtain an identification agent.

[0106] In an embodiment of the present invention, the policy loss probability ratio can be the probability ratio between the sample action policies at different moments. The loss result can include a policy loss result corresponding to the policy network and a value loss result corresponding to the value network. The updated policy can be an update algorithm based on an adaptive learning rate and momentum.

[0107] For example, the policy loss probability ratio can be calculated by the following formula (6):

[0108] (6);

[0109] Wherein, can represent the action policy output at the current moment, can represent the action policy output at the previous moment, θ can represent the parameters of the policy network, and πθ can represent the policy function.

[0110] According to an embodiment of the present invention, the loss function includes a policy loss function and a value loss function; determining the loss result based on the policy loss probability ratio, the evaluation result, and the loss function includes: determining a policy loss value and a value loss value based on the policy loss probability ratio, the evaluation result, the policy loss function, and the value loss function; using the loss weights corresponding to the policy loss value and the value loss value respectively to perform weighted summation on the policy loss value and the value loss value to obtain the loss result.

[0111] In an embodiment of the present invention, the policy loss value can be a loss value obtained by operating on the data such as clipping or limiting the range, such as the clip function. The value loss value can be used to evaluate the difference between the predicted state result and the actual state result.

[0112] In policy optimization, the policy network can be used to output an action policy , and the goal of the policy gradient is to optimize the policy parameter θ to maximize the future cumulative reward. The policy loss value can be determined by optimizing the objective function, as shown in the following formula (7):

[0113] (7);

[0114] Among them, The objective function that can characterize the policy gradient Can represent the expectation Can represent restricting the policy update amplitude, for example, taking the value 0.02; the clip function can clip r t (θ) within the range, r t (θ) can be the policy loss probability ratio.

[0115] According to the embodiments of the present invention, the policy optimization algorithm restricts the policy update range by introducing a truncation function (clip function), which not only ensures the stability of the policy update but also realizes the effective optimization of the policy.

[0116] For example, the method of updating the initial value network parameters based on the value loss can be achieved by minimizing the error between the predicted value and the actual return for update and optimization, as shown in the following formulas (8)-(9):

[0117] (8);

[0118] (9);

[0119] Among them, L VF Can characterize the value loss value Can represent the expectation, R t Can be the cumulative return, and Vφ(st) can be the value predicted by the value network.

[0120] Furthermore, the policy loss value and the value loss value can be weighted and summed to obtain the total loss value (loss result); thereby determining the gradients of the total loss with respect to the policy network parameter θ and the value network parameter φ, and then updating the policy network parameter θ and the value network parameter φ based on the update algorithm of the adaptive learning rate and momentum according to the determined gradients.

[0121] In a feasible embodiment, the ways to identify the agent training include operations S401~S405.

[0122] In operation S401, based on the state space of the initial identification agent and the site current layout information, the output action distribution parameters of the identification agent can be generated through the policy network to obtain the current policy .

[0123] In operation S402, based on the reward function of the identification agent, predict the state value V φ (st) through the value network and calculate the advantage result A t .

[0124] In operation S403, based on the dominant result A t , the state value V φ (st), and the policy , calculate the total loss L through the loss function.

[0125] In operation S404, according to the total loss L, use an update algorithm with an adaptive learning rate and momentum to update the policy network parameters and the value network parameters.

[0126] In operation S405, repeat the training steps, iteratively update the policy network parameters and the value network parameters until a preset termination condition is reached, thereby obtaining the final parameters of the policy network and the final parameters of the value network that identify the agent.

[0127] Based on the above method for determining the evacuation identification information of urban plots, the present invention also provides a device for determining the evacuation identification information of urban plots. The following will be combined with Figure 4 Describe this device in detail.

[0128] Figure 4 Shows a structural block diagram of a device for determining evacuation identification information of urban plots according to an embodiment of the present invention.

[0129] As Figure 4 Shown, the device 400 for determining the evacuation identification information of urban plots in this embodiment includes an action determination module 410 and a result determination module 420.

[0130] The action determination module 410 is used for the identification agent to obtain the action of arranging evacuation identification based on the plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, where the initial evacuation parameters are determined by the building information and object information in the plot information. In one embodiment, the action determination module 410 can be used to execute the operation S210 described above, which will not be elaborated here.

[0131] The result determination module 420 is used to execute the action in the environment corresponding to the plot to be identified, and determine the layout result and evacuation status result of the evacuation identification corresponding to the plot to be identified. In one embodiment, the result determination module 420 can be used to execute the operation S220 described above, which will not be elaborated here.

[0132] Preferably, the identification agent is trained as follows: The initial policy network obtains sample actions, layout results of sample evacuation identifiers, sample evacuation state results, and sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, where the sample initial evacuation parameters are determined by the building information and object information in the plot information; the initial value network evaluates the sample actions, layout results of sample evacuation identifiers, and sample evacuation state results based on the evaluation policy to obtain an evaluation result; based on the evaluation result, sample distribution parameters, and loss function, the initial parameters of the initial policy network and the initial value network are updated to obtain the identification agent.

[0133] According to an embodiment of the present invention, through the action determination module 410 and the result determination module 420 in the urban plot evacuation identification information determination device 400, the action that the identification agent currently needs to execute can be determined in real time through the plot information of the plot to be identified and the initial evacuation parameters, so as to flexibly obtain the layout result of the evacuation identifier and the evacuation state result after executing the action. Since the identification agent is obtained through the closed-loop update process formed by the collaborative work of the policy network and the value network, the value network can feedback the evacuation state of the urban plot in real time, and the policy network can adjust the identification layout in time according to the feedback information, realizing the real-time interaction between the evacuation layout result, the evacuation state result, and the corresponding evaluation result, improving the overall performance and effect of the identification agent. By the way that the identification agent automatically determines the urban plot evacuation identification information, the dependence on mathematical models and algorithms is avoided, and the consumption of computing resources is reduced. The way that the identification agent automatically determines the evacuation identification information does not require the user to spend a lot of time learning and mastering professional knowledge and skills, and the operation is simple and convenient, further improving the universality of the evacuation identification design and optimization.

[0134] According to an embodiment of the present invention, the building information includes plane information and evacuation information; the device further includes: an information determination module, configured to determine the evacuation duration and evacuation target point information in the initial evacuation parameters based on the plane information, evacuation information, and object information.

[0135] According to an embodiment of the present invention, the result determination module 420 includes: an environment construction sub-module, an action execution sub-module, and an action stop sub-module. The environment construction sub-module is configured to construct the current environment based on the building information, the current identification information at the current moment, and the current evacuation information; the action execution sub-module is configured to execute the position identification action, size identification action, and function identification action in the action in the current environment to obtain the layout result of the current evacuation identifier and the current evacuation state result; the action stop sub-module is configured to stop the action when the layout result of the current evacuation identifier and the current evacuation state result meet the preset conditions, and determine the layout result of the current evacuation identifier and the current evacuation state result as the evacuation identifier layout result and the evacuation state result respectively and output them.

[0136] According to an embodiment of the present invention, the action determination module 410 includes: an extraction sub-module, a splicing sub-module, a parameter determination sub-module, and an action determination sub-module. The extraction sub-module is configured to perform feature extraction on the plot information and the initial evacuation parameters respectively to obtain a plot feature and a state feature; the splicing sub-module is configured to splice the plot feature and the state feature to obtain a spliced vector; the parameter determination sub-module is configured to identify an agent to determine a plurality of initial actions and a plurality of distribution parameters corresponding to the plurality of initial actions based on the spliced vector; the action determination sub-module is configured to determine an action from the plurality of initial actions based on the plurality of distribution parameters and a preset screening rule.

[0137] According to an embodiment of the present invention, the initial policy network obtains a sample action, a layout result of sample evacuation identifiers, a sample evacuation state result, and sample distribution parameters based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, including: processing the sample plot information and the sample initial evacuation parameters to obtain a sample action and sample distribution parameters; sampling from the sample actions based on the sample distribution parameters to obtain the action information of the sampled action; executing the sampled action in the sample environment corresponding to the sample plot to obtain the layout result of the sample evacuation identifiers and the sample evacuation state result.

[0138] According to an embodiment of the present invention, the device further includes: a value determination module, configured to determine the current state value of the sampled action at the current moment based on the layout result of the sample evacuation identifiers and the sample evacuation state result for result evaluation.

[0139] According to an embodiment of the present invention, the initial value network evaluates the sample action, the layout result of the sample evacuation identifiers, and the sample evacuation state result based on the evaluation policy to obtain an evaluation result, including: processing the sample action, the layout result of the sample evacuation identifiers, and the sample evacuation state result by using the reward function in the evaluation policy to determine the reward value after executing the sampled action in the sample environment, as well as the subsequent identifier layout result, the subsequent evacuation state result, and the subsequent state value at a subsequent moment, where the subsequent state value is determined from the subsequent identifier layout result and the subsequent evacuation state result; determining the evaluation result based on the reward value, the current state value, the subsequent state value, and the weight corresponding to the subsequent state value.

[0140] According to an embodiment of the present invention, the initial parameters include initial design parameters and initial evaluation parameters; based on the evaluation result, the sample distribution parameter, and the loss function, the initial parameters of the initial policy network and the initial value network are updated respectively to obtain an identification agent, including: determining the ratio between the current distribution parameter at the current moment and the previous distribution parameter at the previous moment in the sample distribution parameter as the policy loss probability ratio; determining the loss result based on the policy loss probability ratio, the evaluation result, and the loss function; using the updated policy and the loss result to update the initial design parameters and the initial evaluation parameters to determine the target design parameters and the target evaluation parameters, so as to update the initial policy network using the target design parameters and update the initial value network using the target evaluation parameters to obtain an identification agent.

[0141] According to an embodiment of the present invention, the loss function includes a policy loss function and a value loss function; determining the loss result based on the policy loss probability ratio, the evaluation result, and the loss function includes: determining a policy loss value and a value loss value based on the policy loss probability ratio, the evaluation result, the policy loss function, and the value loss function; performing weighted summation on the policy loss value and the value loss value respectively using the loss weights corresponding to the policy loss value and the value loss value to obtain the loss result.

[0142] According to an embodiment of the present invention, any multiple modules among the action determination module 410 and the result determination module 420 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the action determination module 410 and the result determination module 420 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits and other hardware or firmware, or can be implemented in any one of the three implementation manners of software, hardware, and firmware or in any appropriate combination of several of them. Or, at least one of the action determination module 410 and the result determination module 420 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.

[0143] Figure 5 The block diagram of an electronic device suitable for implementing the method for determining evacuation identification information of urban plots according to an embodiment of the present invention is shown.

[0144] As Figure 5As shown, the electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage section 508 into a random access memory (RAM) 503. The processor 501 can include, for example, a general-purpose microprocessor (e.g., CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), etc. The processor 501 can also include on-board memory for caching purposes. The processor 501 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0145] In the RAM 503, various programs and data required for the operation of the electronic device 500 are stored. The processor 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. The processor 501 performs various operations of the method flow according to an embodiment of the present invention by executing the programs in the ROM 502 and / or the RAM 503. It should be noted that the program can also be stored in one or more memories other than the ROM 502 and the RAM 503. The processor 501 can also perform various operations of the method flow according to an embodiment of the present invention by executing the programs stored in the one or more memories.

[0146] According to an embodiment of the present invention, the electronic device 500 can further include an input / output (I / O) interface 505, and the input / output (I / O) interface 505 is also connected to the bus 504. The electronic device 500 can further include one or more of the following components connected to the input / output (I / O) interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output (I / O) interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read therefrom can be installed into the storage section 508 as needed.

[0147] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0148] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 502 and / or RAM 503 and / or one or more memories other than ROM 502 and RAM 503.

[0149] An embodiment of the present invention further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the method for determining urban plot evacuation identification information provided by the embodiments of the present invention.

[0150] When the computer program is executed by the processor 501, the above functions defined in the system / apparatus of the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0151] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium and be downloaded and installed through the communication part 509, and / or be installed from the removable medium 511. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0152] In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above functions defined in the system of the embodiments of the present invention are executed. According to the embodiments of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0153] According to the embodiments of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0155] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0156] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A method for determining evacuation sign information of urban plots, characterized in that The method includes: The identification agent obtains the action of arranging evacuation signs based on the plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, where the initial evacuation parameters are determined by the building information and object information in the plot information; Execute the action in the environment corresponding to the plot to be identified, and determine the layout result and evacuation status result of the evacuation signs corresponding to the plot to be identified; Among them, the identification agent is trained in the following way: The initial policy network obtains a sample action, a layout result of sample evacuation signs, a sample evacuation status result, and a sample distribution parameter based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, where the sample initial evacuation parameters are determined by the building information and object information in the sample plot information; The initial value network evaluates the sample action, the layout result of the sample evacuation signs, and the sample evacuation status result based on the evaluation policy to obtain an evaluation result; Based on the evaluation result, the sample distribution parameter, and the loss function, update the initial parameters of the initial policy network and the initial value network respectively to obtain the identification agent.

2. The method according to claim 1, characterized in that, The building information includes plane information and evacuation information; the method further includes: Based on the plane information, the evacuation information, and the object information, determine the evacuation duration and evacuation target point information in the initial evacuation parameters.

3. The method according to claim 1, characterized in that Executing the action in the environment corresponding to the plot to be identified and determining the layout result and evacuation status result of the evacuation signs corresponding to the plot to be identified includes: Construct the current environment based on the building information, the current identification information at the current moment, and the current evacuation information; Execute the position identification action, size identification action, and function identification action in the action in the current environment to obtain the layout result of the current evacuation signs and the current evacuation status result; When the layout result of the current evacuation signs and the current evacuation status result meet the preset conditions, stop the action, and determine the layout result of the current evacuation signs and the current evacuation status result as the layout result and the evacuation status result of the evacuation signs respectively and output them.

4. The method according to claim 1, characterized in that The identification agent obtains the action of arranging evacuation signs based on the plot information of the plot to be identified and the initial evacuation parameters corresponding to the plot information, including: Extract features from the plot information and the initial evacuation parameters respectively to obtain plot features and state features; Concatenate the plot features and the state features to obtain a concatenated vector; The identification agent determines a plurality of initial actions and a plurality of distribution parameters corresponding to the plurality of initial actions based on the concatenated vector; Based on the plurality of distribution parameters and the preset screening rules, determine the action from the plurality of initial actions.

5. The method according to claim 1, characterized in that The initial policy network obtains a sample action, a layout result of sample evacuation signs, a sample evacuation status result, and a sample distribution parameter based on the sample plot information of the sample plot and the sample initial evacuation parameters corresponding to the sample plot information, including: Process the sample plot information and the sample initial evacuation parameters to obtain the sample actions and the sample distribution parameters; Sample from the sample actions based on the sample distribution parameters to obtain the action information of the sampled actions; Execute the sampled actions in the sample environment corresponding to the sample plot to obtain the layout result of the sample evacuation identifiers and the sample evacuation status result.

6. The method according to claim 5, characterized in that, The method further includes: Based on the layout result of the sample evacuation identifiers and the sample evacuation status result, determine the current state value of the sampled actions at the current moment for result evaluation.

7. The method according to claim 6, characterized in that, The initial value network evaluates the sample actions, the layout result of the sample evacuation identifiers, and the sample evacuation status result based on the evaluation strategy to obtain an evaluation result, including: Use the reward function in the evaluation strategy to process the sample actions, the layout result of the sample evacuation identifiers, and the sample evacuation status result to determine the reward value after executing the sampled actions in the sample environment, as well as the subsequent identifier layout result, subsequent evacuation status result, and subsequent state value at a subsequent moment, where the subsequent state value is determined by the subsequent identifier layout result and the subsequent evacuation status result; Based on the reward value, the current state value, the subsequent state value, and the weight corresponding to the subsequent state value, determine the evaluation result.

8. The method according to claim 7, wherein The initial parameters include initial design parameters and initial evaluation parameters; Based on the evaluation result, the sample distribution parameters, and the loss function, update the initial parameters of the initial policy network and the initial value network respectively to obtain the identifier agent, including: Determine the ratio between the current distribution parameter at the current moment and the previous distribution parameter at the previous moment in the sample distribution parameters as the policy loss probability ratio; Based on the loss probability ratio, the evaluation result, and the loss function, determine the loss result; Use the update strategy and the loss result to update the initial design parameters and the initial evaluation parameters to obtain target design parameters and target evaluation parameters, and use the target design parameters to update the initial policy network and the target evaluation parameters to update the initial value network to obtain the identifier agent.

9. The method according to claim 8, wherein The loss function includes a policy loss function and a value loss function; Based on the loss probability ratio, the evaluation result, and the loss function, determine the loss result, including: Based on the loss probability ratio, the evaluation result, the policy loss function, and the value loss function, determine the policy loss value and the value loss value; Use the loss weights corresponding to the policy loss value and the value loss value respectively to perform weighted summation on the policy loss value and the value loss value to obtain the loss result.

10. An apparatus for determining evacuation identification information of urban land parcels, characterized in that, The device includes: An action determination module, configured to enable the identifier agent to obtain the actions of arranging evacuation identifiers based on the plot information of the plot to be identified obtained and the initial evacuation parameters corresponding to the plot information, where the initial evacuation parameters are determined by the building information and object information in the plot information; A result determination module, configured to execute the action in the environment corresponding to the to-be-identified plot, and determine a layout result and an evacuation status result of an evacuation sign corresponding to the to-be-identified plot; Wherein, the identification agent is trained in the following manner: An initial policy network obtains a sample action, a layout result of a sample evacuation sign, a sample evacuation status result, and a sample distribution parameter based on the sample plot information of a sample plot and sample initial evacuation parameters corresponding to the sample plot information, wherein the sample initial evacuation parameters are determined by building information and object information in the sample plot information; An initial value network evaluates the sample action, the layout result of the sample evacuation sign, and the sample evacuation status result based on an evaluation policy to obtain an evaluation result; Based on the evaluation result, the sample distribution parameter, and a loss function, update the initial parameters of the initial policy network and the initial value network respectively to obtain the identification agent.

Citation Information

Patent Citations

  • Crowd evacuation method and system based on multi-carrier intelligent guidance

    CN111767789A

  • Evaluation method and system for design rationality of emergency evacuation indication sign system

    CN115688386A

  • Emergency evacuation scheme intelligent generation method and system

    CN117556977A

  • Multi-agent joint layout model training method and device, equipment and medium

    CN119312760A

  • Multi-agent-based disaster prevention and risk avoiding greenbelt system layout optimization method and system

    CN119761059A

Cited By

  • Control network training method and device of power distribution network and control method of power distribution network

    CN120999641A