A Fire Monitoring and Prediction Method Based on Multi-Agent Reinforcement Learning

Through the multi-agent reinforcement learning method, combined with multi-source data and intelligent collaboration, the problems of small coverage, slow response and weak prediction capabilities of the fire monitoring system are solved, intelligent fire monitoring and prediction are realized, and fire prevention and control efficiency and emergency response capabilities are improved.

CN120013094BActive Publication Date: 2025-07-18NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510503081.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing fire monitoring system has limited coverage, slow response speed and weak prediction capabilities, making it difficult to deal with complex and changeable fire risks.

Method used

Multi-agent reinforcement learning method is adopted, and data processing and decision-making are coordinated by navigation and acquisition of agents, map construction of agents, fire feature extraction agents, fire spread prediction agents and comprehensive decision-making agents, combined with multi-source data to monitor and predict fire conditions, and data processing and decision-making optimization are used using DQN, Monte Carlo tree search and hierarchical analysis methods.

Benefits of technology

It realizes intelligent monitoring and prediction of fire situations, expands monitoring coverage, improves early warning accuracy and monitoring efficiency, supports the generation of personalized emergency plans, and has all-weather and all-round dynamic monitoring capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013094B_ABST
    Figure CN120013094B_ABST
Patent Text Reader

Abstract

The present invention discloses a fire monitoring and prediction method based on multi-agent reinforcement learning, belonging to the fields of artificial intelligence and disaster prevention and control. The method conducts early detection, real-time monitoring, accurate prediction and rapid response to fires based on intelligent means; includes the calculation and evaluation of multiple indicators such as fire risk degree, fire spread degree, fire hazard degree, environmental impact factors and emergency response degree; and uses reinforcement learning methods: deep Q learning, Monte Carlo tree search, policy gradient method, actor-critic framework and double DQN network to improve the coverage, response speed and prediction ability of fire monitoring. The method uses multi-agent cooperation to conduct intelligent risk assessment and spread prediction of fires, effectively improving the efficiency and accuracy of fire prevention and control, and having broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and disaster prevention and control, and particularly relates to a fire monitoring and prediction method based on multi-agent reinforcement learning. Background Art

[0002] Fire is a highly destructive disaster, posing a serious threat to the ecological environment, people's lives and property safety, and social stability. The current fire monitoring systems mainly rely on fixed sensor networks and manual inspections, suffering from problems such as limited coverage, poor real-time performance, and insufficient fire prediction ability. Due to the complex and changeable nature of fire situations, traditional monitoring systems are difficult to effectively cope with fire risks in large-scale and complex environments. The rapid spread and unpredictability of fires pose higher requirements for prevention and control work. In addition, the processing ability of a single agent is limited and it is difficult to comprehensively and quickly handle the multi-task requirements in fire monitoring.

[0003] As an important branch in the field of artificial intelligence, multi-agent reinforcement learning can optimize strategies through continuous interaction with the environment and solve decision-making problems in dynamic and complex environments. Applying multi-agent reinforcement learning to the field of fire monitoring and prediction can make full use of fire historical data, environmental information, and real-time sensing data to intelligently conduct fire risk assessment, spread prediction, and emergency response planning, and combine the collaborative capabilities of multiple agents, thereby improving the scientific nature and efficiency of fire prevention and control. Summary of the Invention

[0004] The object of the present invention is to overcome the defects in the above background art. The present invention provides a fire monitoring and prediction method based on multi-agent reinforcement learning, aiming to solve problems such as limited coverage, slow response speed, and weak prediction ability in the prior art for fire monitoring.

[0005] The present invention adopts the following technical solutions to solve the above technical problems:

[0006] A fire monitoring and prediction method based on multi-agent reinforcement learning, comprising the following steps:

[0007] Step S1: Based on a navigation acquisition agent, obtain multi-source data related to the fire; and preprocess the data;

[0008] Step S2: A map construction agent constructs a Top-Down map and a depth image according to the data provided by the navigation acquisition agent;

[0009] Step S3: A fire feature extraction agent is responsible for extracting fire-related features from the Top-Down map and the depth image;

[0010] Step S4: A fire spread prediction agent predicts the fire spread direction and speed based on the Top-Down map and the depth image;

[0011] Step S5: The navigation optimization agent is responsible for optimizing the path of the navigation agent to improve the efficiency of data collection and emergency response;

[0012] Step S6: The comprehensive decision-making agent fuses the analysis results of navigation data, Top-Down map, and depth image to generate fire prediction and emergency suggestions.

[0013] Furthermore, the multi-source data in step S1 includes historical fire data, geographical information, meteorological data, environmental data, and sensor data.

[0014] Furthermore, step S1 is specifically as follows:

[0015] The navigation acquisition agent includes a dynamic acquisition agent and a static acquisition agent;

[0016] Step S11: The dynamic acquisition agent performs circumferential navigation around the fire area to dynamically acquire visual and depth data;

[0017] Step S12: The static acquisition agent uses fixed sensors to collect environmental data, including temperature, humidity, wind direction and wind speed;

[0018] Step S13: The agents share the collected data through a communication network and transmit the collected raw data to other agents;

[0019] Step S14: Use DQN (Deep Q-Network) to train the navigation path to maximize the data coverage range and assist the map construction agent to establish a Top-Down map and a depth image of the fire area;

[0020] Step S15: Standardize the data to ensure that the dimensions of various types of data are consistent;

[0021] Step S16: Construct a state space matrix S(t) to convert various types of input data at different time points into a state representation that the model can process;

[0022] Step S17: Fill in or delete missing values and process abnormal data.

[0023] Furthermore, step S2 is specifically as follows:

[0024] Step S21: Stitch multi-view RGB images and infrared images into a top-down map, and mark the fire source location, combustion area, smoke distribution, and potential obstacles;

[0025] Step S22: Process lidar or depth camera data to generate a three-dimensional depth map of the fire area and extract key features from it;

[0026] Step S23: As the navigation agent moves, the map construction agent will update the map and depth image in real time.

[0027] Further, the specific steps of step S3 are as follows:

[0028] Step S31: Analyze the fire source location, combustion range, and smoke diffusion situation in the Top-Down map, and identify the fire points in the map;

[0029] Step S32: Extract the fire height, heat distribution, and the influence of obstacles on the fire from the depth image, and analyze the three-dimensional fire characteristics in the depth image;

[0030] Step S33: Convert the extracted features into a state representation that can be processed by the reinforcement learning model.

[0031] Further, the specific steps of step S4 are as follows:

[0032] Step S41: Use the fire source location and smoke distribution in the Top-Down map, combined with the fire height and obstacle information in the depth image, to calculate the spread probability FS in each direction;

[0033] Step S42: Use Monte Carlo tree search to simulate the diffusion path of the fire on the Top-Down map;

[0034] Step S43: Output the spread prediction result and mark the high-risk areas.

[0035] Further, the specific steps of step S5 are as follows:

[0036] Step S51: Dynamically adjust the navigation path according to the fire source location and high-risk areas in the Top-Down map;

[0037] Step S52: Use the reinforcement learning algorithm to train the navigation strategy, with the goal of minimizing the acquisition time and maximizing the coverage area;

[0038] Step S53: The agents share the analysis results to provide navigation support for emergency response.

[0039] Further, the specific steps of step S6 are as follows:

[0040] Step S61: Integrate the fire characteristics and spread prediction to calculate the fire risk degree FR, hazard degree FH, and emergency response degree ER;

[0041] Step S62: Use the analytic hierarchy process to determine the weights of each index and comprehensively evaluate the current fire situation:

[0042] TS = α1·FR + α2·FS + α3·FH + α4·ER

[0043] Among them, the weight coefficients α1 to α4 are determined by the analytic hierarchy process to measure the relative importance of different indicators.

[0044] Furthermore, the calculation formula for the spread probability is as follows:

[0045]

[0046] Among them, P i is the spread probability in the i-th direction, V i is the spread speed, W e is the environmental weight coefficient, H d is the fire height factor in the depth image, and n is the number of sampling times.

[0047] Furthermore, the calculation formulas for calculating the fire risk degree FR, the hazard degree FH, and the emergency response degree ER are as follows:

[0048] FR = Q(S(t), a FR ) = R(t) + γ FR ·max(Q(S(t + 1), a′ FR ))

[0049]

[0050] ER = Q(s, a ER ) = r + γ ER ·Q′(s′, argmax(Q(s′, a′ ER )))

[0051] Among them, Q(S(t), a FR ) represents the value function of taking action a FR at time t and state S, R(t) represents the immediate reward, that is, the feedback value immediately obtained by the system after taking an action at a certain time point, and γ FR is the fire risk degree discount factor; π(a FH ∣s) is the policy function, r t is the reward value at time t, is the attenuation coefficient at time t; Q(s, a ER ) is the main network, which is used to select the best action a ER to take in state s, r represents the immediate reward, γ ER is the emergency response degree discount factor, and Q(s′, a′ ER ) is the target network, which is used to estimate the state value after taking the next action.

[0052] Compared with the prior art, the present invention adopts the above technical solutions and has the following beneficial effects:

[0053] (1) A fire monitoring and prediction method based on multi-agent reinforcement learning provided by the present invention realizes intelligent monitoring and prediction of fires and improves the early warning accuracy rate.

[0054] (2) A fire monitoring and prediction method based on multi-agent reinforcement learning provided by the present invention expands the monitoring coverage and improves the monitoring efficiency through multi-agent cooperation.

[0055] (3) A fire monitoring and prediction method based on multi-agent reinforcement learning provided by the present invention can achieve all-weather and omni-directional dynamic monitoring.

[0056] (4) A fire monitoring and prediction method based on multi-agent reinforcement learning provided by the present invention supports the generation of personalized emergency plans, and the system can continuously learn and optimize, with strong adaptability.

[0057] (5) A fire monitoring and prediction method based on multi-agent reinforcement learning provided by the present invention improves the coordination and rapid response ability in fire response through a multi-agent cooperation mechanism. Brief Description of the Drawings

[0058] Figure 1 is the overall system architecture diagram of the present invention;

[0059] Figure 2 is the data flow diagram of fire monitoring;

[0060] Figure 3 is the Top-Down diagram and depth image in the present invention;

[0061] Figure 4 is the schematic diagram of the fire spread prediction model;

[0062] Figure 5 is the curve graph of the change of the fire spread factor FS with time under different environments in the present invention

[0063] Figure 6 is the schematic diagram of fire risk degree assessment;

[0064] Figure 7 is the curve graph of the change of four indicators, namely the fire risk degree FR, hazard degree FH, emergency response degree ER, and comprehensive fire situation TS, with time in the present invention;

[0065] Figure 8 is the emergency response flow chart;

[0066] Figure 9 is the schematic diagram of the fire fighting and rescue system adopting the present invention. Detailed Embodiment

[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0068] As Figure 1 , Figure 2 shown, the present invention provides a fire monitoring and prediction method based on multi-agent reinforcement learning, and the specific implementation process is as follows:

[0069] Step 1: The data collection agent collects data.

[0070] Distributed data collection: The dynamic and static data collection agents cooperate to collect multi-source data related to fires. The navigation collection agent conducts circumferential navigation around the fire area to dynamically collect visual and depth data, such as RGB images, infrared images, etc.; the static collection agent uses fixed sensors to collect environmental data such as temperature, humidity, wind direction, and wind speed; the agents share the collected data through a communication network and transmit the collected raw data to other agents (such as the map construction agent); use DQN to train the navigation path to maximize the data coverage range and assist the map construction agent to establish a top-down map and depth image of the fire area.

[0071] The collected data includes: historical fire data, fire records within the responsible area of each agent, including occurrence time, location, scale, etc.; meteorological data, the agents collect real-time temperature, humidity, wind speed, etc. through meteorological sensors; geographical information, static information such as the terrain, vegetation, and building distribution of the responsible area; sensor data, dynamic data obtained by the sensor network deployed by each agent.

[0072] The agents standardize the data through a data synchronization and sharing mechanism to ensure that the dimensions of various types of data are consistent, which is convenient for subsequent model processing. Construct a state space matrix S(t) to convert various types of input data at different time points into a state representation that the model can process. Fill or delete missing values and process abnormal data to ensure data quality.

[0073] Step 2: The map construction agent constructs a top-down map and a depth image.

[0074] The map construction agent generates a top-down map and a depth image based on the data of the navigation collection agent as Figure 3As shown in the figure. The agent stitches multi-view RGB images and infrared images into an overhead map, and marks the fire source location, burning area, smoke distribution, and potential obstacles; processes lidar or depth camera data to generate a three-dimensional depth map of the fire area, and updates the map and depth image in real time as the navigation agent moves, while sharing the map data with the fire risk assessment and spread prediction agents; during the map construction process, the agent will optimize the update frequency and accuracy of the map, and its corresponding state space is the current image data volume and coverage range, and the reward function is the map accuracy improvement value.

[0075] Step 3: The fire feature extraction agent extracts fire-related features:

[0076] Analyze the fire source location, burning range, and smoke diffusion situation in the Top-Down map. Extract the fire height, heat distribution, and the impact of obstacles on the fire from the depth image. Convert the extracted features into a state representation that can be processed by the reinforcement learning model.

[0077] Step 4: The fire spread prediction agent predicts the fire spread degree FS, as Figure 4 shown:

[0078] Step 4.1: Collaborative environment modeling:

[0079] Step 4.1.1: The multi-agent system constructs a global model by sharing environmental information.

[0080] Step 4.1.2: Divide the monitored area into multiple grids, each grid representing a different spatial area. The multi-agent system distributes the calculation of the flammability index of the grids, where the flammability index is determined according to data such as geographical information, vegetation conditions, and building distributions.

[0081] Step 4.1.3: The multi-agents collaborate to update the environmental model.

[0082] Step 4.2: Distributed spread prediction:

[0083] Step 4.2.1: Use the Monte Carlo tree search method to predict the spread direction and speed of the fire.

[0084] Step 4.2.2: Calculate the spread probability and speed: where, P i is the spread probability in the i-th direction, V i is the spread speed, W e is the environmental weight coefficient, H d is the fire height factor in the depth image, and n is the number of sampling times. As Figure 5As shown, it is used to characterize the dynamic evolution characteristics of forest fires in different scenarios. The horizontal axis represents time (unit: hour), ranging from 0 to 10 hours; the vertical axis represents the value of the fire spread factor FS, which is determined by the calculation formula proposed in the present invention. The figure shows the FS curves of five typical environmental scenarios (forest, grassland, city, mountain, and wetland), and each curve is distinguished by different colors and marking symbols: Forest (green, circular mark): represents a natural environment with high combustible density and volume. Grassland (yellow, square mark): represents an open area with medium combustible density and low volume. City (blue, diamond mark): represents an area with low combustible density but complex structure. Mountain (orange, triangular mark): represents a scenario with significant terrain factors. Wetland (purple, X-shaped mark): represents an environment with high humidity and low spread speed.

[0085] Step 4.2.3: The multi-agent system simulates multiple spread paths, predicts the spread trend of the fire in real time, and fuses the prediction results.

[0086] Step 5: The navigation optimization agent optimizes and updates the navigation path:

[0087] Step 5.1: Dynamically adjust the navigation path according to the fire source location and high-risk areas in the Top-Down map.

[0088] Step 5.2: Use the reinforcement learning algorithm to train the navigation strategy to minimize the acquisition time and maximize the coverage area.

[0089] The state space is the current position and the fire source distribution in the Top-Down map; the action space is the moving direction, speed adjustment, and change of the surrounding radius; the reward function is the degree of covering the fire source area and the path efficiency. Step 6: The comprehensive decision-making agent evaluates various indicators and generates a final decision:

[0090] Step 6.1: Evaluate the fire risk degree FR, as Figure 6 shown:

[0091] Step 6.1.1: Distributed risk assessment model: Evaluate the fire risk degree of this area based on the Deep Q-Network (DQN). Establish a risk information sharing mechanism among agents. Construct the state space S(t) and the action space a (for example: monitoring, warning, alarm, etc.).

[0092] Step 6.1.2, Collaborative Risk Degree Calculation: Use the Q-value function to evaluate the value of taking a certain action at a specific time and state: FR = Q(S(t), a) = R(t) + γ·max(Q(S(t + 1), a')), where Q(S(t), a) represents the value function of taking action a in state S at time t, R(t) represents the immediate reward, and γ is the discount factor. The system dynamically updates the fire risk assessment based on real-time sensor data, generates a fire risk map, and uses an attention mechanism to fuse the risk assessment results of multiple agents.

[0093] Step 6.2, Fire Hazard Degree (FH) Assessment

[0094] Step 6.2.1, Distributed Data Collection: The multi-agent system conducts hazard assessment based on the protection targets (such as residential areas, factories, forests, etc.) in the monitored area, analyzes the losses and impacts that the fire may cause to these targets. Jointly evaluate the difficulty of rescue and potential casualties and property losses.

[0095] Step 6.2.2, Collaborative Hazard Degree Calculation: Use the policy gradient method to evaluate the hazard degree of the fire. The policy gradient function is: FH = π(a∣s)·∑(r t ·γ t ), where π(a∣s) is the policy function, r t is the reward value at time t, and γ is the time decay coefficient.

[0096] Step 6.3, The comprehensive decision-making agent conducts the emergency response degree ER assessment

[0097] Step 6.3.1, Emergency Response Planning: The system pre-constructs an emergency plan library, which includes emergency response strategies, resource allocation plans, and evacuation routes under different fire conditions. The multi-agent jointly evaluates the availability and scheduling plans of emergency response resources (such as fire trucks, fire extinguishing equipment, rescue teams, etc.).

[0098] Step 6.3.2, Joint Response Degree Calculation: Use the double DQN network to calculate the emergency response degree. The multi-agent jointly evaluates the effectiveness of each emergency strategy. The formula is as follows: ER = Q(s, a) = r + γ·Q'(s', argmax(Q(s', a'))), where Q(s, a) is the main network, r represents the immediate reward, γ is the discount factor, and Q(s', a') is the target network.

[0099] Step 6.4, The comprehensive decision-making agent conducts comprehensive scoring and decision-making

[0100] Step 6.4.1, Multi-index comprehensive scoring: The system comprehensively evaluates the current fire situation according to indicators such as fire risk degree (FR), fire spread degree (FS), fire hazard degree (FH), and emergency response degree (ER), and integrates the evaluation results of multiple agents. As Figure 7 shown, the horizontal axis represents time (unit: hour), ranging from 0 to 10 hours; the vertical axis represents the values of each indicator, and the four curves in the figure are drawn using smooth interpolation. The analytic hierarchy process is used to determine the weights α1 to α4 of each indicator, and the comprehensive score is calculated by weighted summation: TS = α1·FR + α2·FS + α3·FH + α4·ER

[0101] Step 6.4.2, Intelligent decision generation: Multiple agents generate corresponding emergency response suggestions according to the comprehensive score, including decision-making information such as early warning, alarm, resource scheduling, fire extinguishing strategies, and evacuation routes for personnel; among them, the threshold τ for whether a fire occurs is obtained through learning from actual situations. As Figure 8 shown, the system can adjust the emergency response plan in real time according to the development trend of the fire situation to ensure rapid and effective control of the fire and reduce fire losses.

[0102] Step 6.5, Multi-agent system optimization and learning

[0103] Step 6.5.1, Distributed learning optimization: The system adopts a paradigm of centralized training and distributed execution. Through experience sharing and knowledge transfer among agents, as well as a collaborative policy optimization mechanism, it continuously optimizes its fire monitoring and prediction strategies. With continuous interaction with the environment, the system can learn new decision-making experiences during the handling of each fire and apply them to future fire monitoring and emergency responses.

[0104] Step 6.5.2, Adaptive model update: As the fire data and environmental data increase, based on the model update mechanism feedback from multiple agents, the system will regularly perform distributed parameter optimization, update the model parameters and strategies, so that it always maintains sensitivity to the latest fire data, and combined with collaborative verification and evaluation of multiple agents, further improve the accuracy of fire prediction and response efficiency.

[0105] As Figure 9 shown, in a complex fire scene, a multi-agent system composed of multiple fire-fighting intelligent vehicles needs to perform collaborative decision-making and precise actions. The multi-agent reinforcement learning algorithm proposed by the present invention can achieve the collaborative cooperation of the intelligent vehicle group. Specifically, it includes the following aspects:

[0106] Distributed Environment Perception and State Sharing: Multiple intelligent vehicles, as independent agents, collect environmental information distributedly through in-vehicle sensor networks, including real-time data such as temperature, humidity, and wind speed; an information sharing mechanism is established among the agents to exchange the environmental data and historical fire information collected by each; a global environmental state space is collaboratively constructed based on the shared information to provide complete situation awareness for decision-making.

[0107] Collaborative Action Planning (Actor Part): The agent group collaboratively plans the optimal action strategy based on the shared environmental information; considering the mutual influence among agents, it optimizes vehicle scheduling, fire extinguishing path planning, and resource allocation; through an improved multi-agent policy network, it coordinates specific operations such as the water spraying intensity and spraying angle of multiple vehicles.

[0108] Distributed Evaluation and Feedback (Critic Part): Design a value evaluation mechanism considering agent collaboration to evaluate the overall effect of group actions; based on the combination of global rewards and local rewards, guide agents to optimize collaborative decisions; accelerate the group learning process through experience sharing among agents.

[0109] Collaborative Learning and Optimization: Adopt a learning paradigm of centralized training and distributed execution; the agent group continuously accumulates and shares experience during the task execution process; through a knowledge transfer mechanism, improve the emergency response ability of the entire system.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A fire monitoring and prediction method based on multi-agent reinforcement learning, characterized in that, The method includes the following steps: Step S1: The navigation acquisition agent obtains multi-source data related to the fire based on navigation; and preprocesses the data. Step S2: The map construction agent constructs a Top-Down map and a depth image according to the data provided by the navigation acquisition agent. Step S3: The fire feature extraction agent is responsible for extracting fire-related features from the Top-Down map and the depth image. Step S4: The fire spread prediction agent predicts the fire spread direction and speed based on the Top-Down map and the depth image. Step S5: The navigation optimization agent is responsible for optimizing the path of the navigation agent to improve the efficiency of data collection and emergency response. Step S6: The comprehensive decision-making agent fuses the analysis results of navigation data, the Top-Down map, and the depth image to generate fire prediction and emergency suggestions. The specific content of step S1 is as follows: The navigation acquisition agent includes a dynamic acquisition agent and a static acquisition agent. Step S11: The dynamic acquisition agent conducts circumferential navigation around the fire area and dynamically acquires visual and depth data. Step S12: The static acquisition agent uses fixed sensors to collect environmental data, including temperature, humidity, wind direction, and wind speed. Step S13: The agents share the collected data through a communication network and transmit the collected raw data to other agents. Step S14: Use DQN (Deep Q-Network) to train the navigation path to maximize the data coverage range and assist the map construction agent in establishing a Top-Down map and a depth image of the fire area. Step S15: Standardize the data to ensure that the dimensions of various types of data are consistent. Step S16: Construct a state space matrix S(t) to convert various types of input data at different time points into a state representation that the model can process. Step S17: Fill or delete missing values and process abnormal data. The specific content of step S6 is as follows: Step S61: Integrate the fire characteristics and spread prediction to calculate the fire risk degree FR, the hazard degree FH, and the emergency response degree ER. Step S62: Use the analytic hierarchy process to determine the weights of each index and comprehensively evaluate the current fire situation: TS = α1·FR + α2·FS + α3·FH + α4·ER Among them, the weight coefficients α1 to α4 are determined by the analytic hierarchy process to measure the relative importance of different indexes. The calculation formulas for calculating the fire risk degree FR, the hazard degree FH, and the emergency response degree ER in step S61 are as follows: FR = Q(S(t), a FR ) = R(t) + γ FR ·max(Q(S(t + 1), a′ FR )) ER = Q(s,a ER ) = r + γ ER ·Q′(s′, argmax(Q(s′,a′ ER ))) Among them, Q(S(t), a FR ) represents the value function of taking action a at time t and state S FR , R(t) represents the immediate reward, that is, the feedback value immediately obtained by the system after taking an action at a certain time point, and γ FR is the fire risk degree discount factor; π(a FH |s) is the policy function, r t is the reward value at time t, is the attenuation coefficient at time t; Q(s, a ER ) is the main network, which is used to select the best action a to be taken in state s ER , r represents the immediate reward, and γ ER is the emergency response degree discount factor, and Q(s′, a′ ER ) is the target network, which is used to estimate the state value after taking the next action.

2. The fire monitoring and prediction method based on multi-agent reinforcement learning according to claim 1, characterized in that, The multi-source data in step S1 includes historical fire data, geographical information, meteorological data, environmental data, and sensor data.

3. A fire monitoring and prediction method based on multi-agent reinforcement learning according to claim 1, characterized in that The specific content of step S2 is as follows: Step S21: Stitch multi-view RGB images and infrared images into a top-down map, and mark the fire source location, the burning area, the smoke distribution, and potential obstacles. Step S22: Process lidar or depth camera data to generate a three-dimensional depth map of the fire area and extract key features from it. Step S23: As the navigation agent moves, the map construction agent will update the map and the depth image in real time.

4. A fire monitoring and prediction method based on multi-agent reinforcement learning according to claim 1, characterized in that, The specific content of step S3 is as follows: Step S31: Analyze the fire source location, combustion range, and smoke diffusion in the Top-Down map, and identify the fire points in the map; Step S32: Extract the fire height, heat distribution, and the impact of obstacles on the fire from the depth image, and analyze the three-dimensional fire characteristics in the depth image; Step S33: Convert the extracted features into a state representation that can be processed by the reinforcement learning model.

5. A fire monitoring and prediction method based on multi-agent reinforcement learning according to claim 1, characterized in that, The specific steps of step S4 are as follows: Step S41: Calculate the spread probability FS in each direction by using the fire source location and smoke distribution in the Top-Down map, combined with the fire height and obstacle information in the depth image; Step S42: Use Monte Carlo tree search to simulate the diffusion path of the fire on the Top-Down map; Step S43: Output the spread prediction result and mark the high-risk areas.

6. The method for fire monitoring and prediction based on multi-agent reinforcement learning according to claim 1, wherein The specific steps of step S5 are as follows: Step S51: Dynamically adjust the navigation path according to the fire source location and high-risk areas in the Top-Down map; Step S52: Use the reinforcement learning algorithm to train the navigation strategy, with the goal of minimizing the acquisition time and maximizing the coverage area; Step S53: The agent shares the analysis results to provide navigation support for emergency response.

7. A fire monitoring and prediction method based on multi-agent reinforcement learning according to claim 6, characterized in that, The calculation formula for the spread probability is: Among them, P i is the spread probability in the i-th direction, V i is the spread speed, W e is the environmental weight coefficient, H d is the fire height factor in the depth image, and n is the number of sampling times.

Citation Information

Patent Citations

  • Intelligent fire safety monitoring system

    CN117649130A

  • Study defense method and system based on digital twinborn and intelligent sensor

    CN119418462A