Image-fused end-side cloud collaborative intelligent fire-fighting fire monitoring system

Through multimodal image fusion and distributed edge computing, combined with end-edge cloud collaborative decision-making, the problems of accuracy and delayed response of fire field development trend prediction in forest fire monitoring are solved, and efficient and low-latency fire monitoring and early warning are achieved.

CN120356294AActive Publication Date: 2025-07-22HANGZHOU ZIPENG TECH CO LTD

Patent Information

Application Number
CN202510813525.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-22
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

In the forest fire monitoring, traditional single-area monitoring methods are difficult to accurately predict the development trend of the fire field, and centralized cloud processing has data transmission delays and computing power bottlenecks, which cannot meet the monitoring requirements of low latency and high reliability.

Method used

A multi-modal image fusion module is used to generate a dynamic scanning priority map, a dual-light fusion algorithm is used to distinguish natural heat sources from abnormal fire conditions, and a distributed edge computing module is used to perform spatiotemporal synchronization and dynamic resource allocation of cross-modal data. A federated learning-driven model sharing network is built through the end-edge cloud collaborative decision-making module to optimize the drone formation density.

Benefits of technology

It improves the accuracy of fire point positioning, reduces false alarm rate, ensures low-latency response, improves coverage efficiency and reduces energy consumption, and meets the real-time early warning needs of forest fire monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356294A_ABST
    Figure CN120356294A_ABST
Patent Text Reader

Abstract

The invention discloses an end-side cloud collaborative intelligent fire-fighting fire monitoring system based on image fusion, and relates to the technical field of intelligent fire-fighting, the system is composed of a plurality of functional modules, and the system comprises a multi-modal image fusion module which generates a dynamic scanning priority map based on prior data, distinguishes a natural heat source from an abnormal fire by using a dual-light fusion algorithm, and sends an image fusion result to a cloud server; a scanning area is divided according to the thermal risk grade, and the thermal imaging resolution is dynamically adjusted; the distributed edge computing module is used for carrying out space-time synchronization on cross-modal data through a multi-modal feature alignment network, and carrying out dynamic allocation on a CUDA core and CPU resources through adaptive computing scheduling; an improved artificial bee colony algorithm is adopted, the bandwidth of the multi-sensor data flow is dynamically allocated through a time-sharing multiplexing protocol, and three-dimensional path planning is carried out; and the end-side cloud collaborative decision module constructs a federated learning driven model sharing network, and each edge node trains a lightweight YOLOv5s pruning model based on local data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent fire protection, and specifically to an edge-cloud collaborative intelligent fire monitoring system with image fusion for fire detection. Background Art

[0002] In the forest fire prevention and control system, fire monitoring, as the core link, is not only important in terms of the accuracy of real-time fire perception, but also lies in the ability to deeply analyze the dynamic evolution law of the fire scene; especially in the scenario of large-scale mountain forest fires, the spatio-temporal evolution of the fire presents significant multi-dimensional coupling characteristics: on the one hand, the combustion intensity of local areas is directly controlled by physical conditions such as vegetation type, terrain slope, and fuel load; on the other hand, the spread of fire in adjacent areas forms complex non-linear correlations through mechanisms such as heat radiation transfer, fire whirl generation, and flying fire propagation; this dual driving mechanism makes it difficult for traditional single-region monitoring methods to accurately predict the development trend of the fire scene, and there is an urgent need to construct a dynamic monitoring model that integrates spatio-temporal correlation characteristics.

[0003] Existing technologies rely on data collected by devices such as satellite remote sensing, UAV inspections, and ground sensors, which have differences in spatio-temporal resolution, coordinate systems, and data formats. How to achieve efficient collaboration remains a technical difficulty; the evolution of the fire is affected by multiple factors such as vegetation type, terrain slope, wind speed and direction, and the mechanisms such as heat radiation transfer and flying fire diffusion between adjacent areas are highly non-linear, making it difficult for traditional statistical models to accurately describe their spatio-temporal correlations; existing systems mostly rely on centralized cloud processing, which has data transmission delays and computing power bottlenecks, and it is difficult to meet the strict requirements of low latency and high reliability for forest fire monitoring. Summary of the Invention

[0004] To achieve the above objectives, the present invention is realized through the following technical solutions: A multi-modal image fusion module, which generates a dynamic scanning priority map based on prior data, uses a dual-light fusion algorithm to distinguish natural heat sources from abnormal fire situations, divides the scanning area according to the heat risk level, and dynamically adjusts the thermal imaging resolution; A distributed edge computing module, which synchronizes spatio-temporal cross-modal data through a multi-modal feature alignment network, dynamically allocates CUDA cores and CPU resources through adaptive computing scheduling; uses an improved artificial bee colony algorithm to dynamically allocate the bandwidth of multi-sensor data streams through a time-division multiplexing protocol, and performs three-dimensional path planning; An edge-cloud collaborative decision-making module, which constructs a model sharing network driven by federated learning. Each edge node trains a lightweight YOLOv5s pruned model based on local data, generates a globally optimized model and synchronizes it backward, conducts dynamic networking, and adjusts the density of the UAV formation in real time.

[0005] Furthermore, the process of generating the dynamic scanning priority map is as follows: Based on the historical data of fire points, where the historical data includes fire data, environmental information, and geographical location information, obtain it as prior data, and classify the prior data into text information, image annotations, and sensor data; structure the prior data by OCR recognizing the text information, image annotations, and sensor data, and integrate the text, image, and sensor data using multimodal fusion technology; YOLOv5s outputs the bounding box coordinates to locate the fire point position, Mask R-CNN segments the burning area, quantifies the burning area and the fire line morphology, rotates and simulates the smoke of the fire satellite image, and at the same time cleans and aligns the sensor data; extract keywords from the report extracted by OCR, convert them into numerical vectors TF-IDF or word embeddings, use ResNet to extract the high-level features of the burning area, and calculate the burning intensity through the masked area.

[0006] Furthermore, the process of distinguishing natural heat sources from abnormal fire situations is as follows: Collect data from the visible light channel and the thermal imaging channel, and for the visible light channel, perform color space conversion through a defogging algorithm and image enhancement technology, use the YOLOv5s model to detect flame or smoke features, and extract morphological features; The thermal imaging channel extracts the temperature gradient and hot spot distribution through NSCT (Non-Subsampled Contourlet Transform); perform high-frequency feature fusion using a Pulse Coupled Neural Network, and process the low-frequency information using an adaptive fuzzy logic algorithm; Utilize dual-light fusion, dynamically adjust the weights according to the scene, adopt the DAF-Net (Domain Adaptive Dual-Branch Feature Decomposition Fusion Network) to reduce the distribution difference between infrared and visible light images through multi-kernel maximum mean discrepancy, and use the Transformer-CNN structure to retain cross-modal features.

[0007] Furthermore, the process of dividing the scanning area according to the heat risk level is as follows: Through random forest, perform a heat risk score on the area and divide the levels; capture the image of the area through high-definition drone photography, and extract the temperature gradient and hot spot distribution through NSCT; dynamically adjust the weights according to the scene, adopt the EMMA framework, and retain cross-modal features; exclude natural heat sources through temperature thresholds and morphological features, and dynamically update the classification model based on historical false alarms; Locally magnify the suspicious area through an AI algorithm; carry an NPU on the drone to analyze the thermal imaging data in real time. If a temperature anomaly is detected, automatically trigger the high-resolution mode, and the cloud calls the NSCT fusion model to process complex scenes and dynamically optimize the resolution strategy; otherwise, continue with the low-resolution mode for recognition; Divide the area into grids, extract the mean values of each feature within the grids, calculate the average humidity change rate in the past 7 days, and at the same time count the fire frequencies in the same period in the past 3 years. If a fire has occurred in the grid within 30 days, add a binary feature, otherwise it is 0; based on historical data, if a fire has occurred in the grid during the observation period, mark it as high-risk = 1, otherwise it is low-risk = 0; perform semi-supervised enhancement, for unlabeled areas, use K-Means clustering to generate pseudo-labels; use the XGBoost model to identify key risk drivers through the built-in feature importance ranking.

[0008] Further, the process of spatio-temporal synchronization of the cross-modal data is as follows: Through timestamp calibration, synchronize different sensors in time, and unify the coordinate systems of different sensors to the same reference system.

[0009] Further, the process of dynamic allocation of CUDA cores and CPU resources is as follows: Dynamically allocate GPU and CPU resources, prioritize the real-time image processing tasks of fire monitoring, adjust the resource allocation ratio according to the real-time task requirements, and dynamically adjust the scheduling strategy under the fluctuations of sensor data streams and network bandwidth changes; Dynamic resource allocation strategy: Real-time collect the load, temperature, and memory occupancy of the GPU and CPU through hardware performance counters; combine the task queue status to dynamically adjust the resource allocation.

[0010] Further, the process of dynamic bandwidth allocation for multi-sensor data streams is as follows: Collect the bandwidth occupancy, latency, and packet loss rate of each sensor data stream, monitor the network load between the edge node and the cloud, extract the resolution, frame rate of the thermal imaging video, and the characteristics and update frequency of the temperature and humidity sensor data stream; identify the thermal imaging data and environmental temperature data for fire monitoring; classify the sensor data streams based on the traffic service model; Dynamically allocate weights through confidence or task requirements, use the LSTM model to predict the data stream requirements of each sensor in the future for a period of time, combine the traffic prediction model to evaluate the network load trend; calculate the current bandwidth requirements of each sensor data stream.

[0011] Further, the process of three-dimensional path generation is as follows: Obtain fixed starting and ending points, and optimize the intermediate points through the ABC algorithm; randomly generate an initial path, select excellent paths according to the fitness value, and update the positions of neighborhood points; other drones perform local search according to the probability to select paths; if the path is not improved, randomly generate a new path; use spline interpolation to eliminate jagged paths; combine lidar or vision sensors to dynamically update the obstacle map and trigger path replanning; at the same time, optimize the coverage of high-risk areas in path planning.

[0012] Further, the process of constructing the federated learning-driven model sharing network is as follows: Obtain the data after spatio-temporal alignment of multi-modal data. After receiving the gradients of all edge nodes through the regional master node, use the federated averaging algorithm in federated learning for weighted aggregation to generate a global model; allocate weights according to task priorities, and adjust hyperparameters such as the learning rate and regularization coefficient of the global model according to the aggregation results.

[0013] Further, the process of real-time adjusting the formation density of unmanned aerial vehicles is as follows: Adopt a wireless Mesh network, with each unmanned aerial vehicle as a node, supporting multi-hop relay transmission; divide the unmanned aerial vehicles into multiple clusters based on a clustering algorithm, and specify a master node in each cluster to be responsible for channel allocation. According to the heat risk score, adjust the scanning strategy, and dynamically adjust the formation shape by setting a leader and virtual leader points; and according to environmental changes, real-time adjust the behavior rules of the unmanned aerial vehicles; at the same time, the cloud calls the NSCT fusion model to process complex scenarios and dynamically optimize the resolution strategy.

[0014] An edge-cloud collaborative intelligent fire monitoring system for image fusion provided by the present invention has the following beneficial effects: (1) By using multi-modal data such as historical fire data, environmental information, and geographical location information, the present invention generates a structured dynamic scanning priority map through technologies such as OCR recognition and Mask R-CNN segmentation of the combustion area; improves the accuracy of fire point positioning, and dynamically adjusts the risk level according to real-time meteorological data; marks high-risk areas as priority scanning areas, and reduces the scanning frequency in low-risk areas to save resources.

[0015] (2) Through data collection and processing of the visible light channel and the thermal imaging channel, combined with the DAF-Net domain adaptive double-branch feature decomposition fusion network, the present invention effectively distinguishes natural heat sources from abnormal fire situations; uses the Transformer-CNN structure to retain cross-modal features, and dynamically updates the classification model based on historical false alarm or missed alarm data, thereby ensuring the ability to capture micro fire sources and reducing the false alarm rate at the same time.

[0016] (3) The present invention realizes spatio-temporal synchronization of cross-modal data through a multi-modal feature alignment network, combined with the dynamic allocation of CUDA cores and CPU resources, preferentially processes real-time image processing tasks for fire monitoring, and ensures low-latency response; in scenarios such as sensor data stream fluctuations and network bandwidth changes, dynamically adjusts the scheduling strategy to avoid single-point resource bottlenecks; uses an improved artificial bee colony algorithm for three-dimensional path planning, comprehensively considering factors such as path length, terrain cost, and energy consumption, realizes the optimal path selection of the unmanned aerial vehicle formation, improves the coverage efficiency and reduces the energy consumption. Description of the Drawings

[0017] Figure 1 This is a schematic diagram of the system process of the present invention. Detailed implementation manners

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment Please refer to Figure 1 , the embodiment of the present application provides an edge-cloud collaborative intelligent fire monitoring system for image fusion, and the system includes: A multimodal image fusion module, which generates a dynamic scanning priority map based on prior data, uses a dual-light fusion algorithm to distinguish natural heat sources from abnormal fire situations, divides the scanning area according to the heat risk level, and dynamically adjusts the thermal imaging resolution; Prior data: Based on the historical data of fire points, the historical data includes fire data, environmental information, and geographical location information. Among them, the historical fire data, such as data on spatial distribution, combustion intensity, and spread speed, combined with vegetation types, environmental information such as the flammability of pine trees, and geographical location information including the heat retention characteristics of canyons or steep slopes, are obtained as prior data, and the prior data is further classified into text information, image annotations, and sensor data; a heat risk assessment model is constructed; Historical fire data: including fire point distribution, combustion frequency, fire spread speed, etc.; Environmental information: vegetation coverage, wind speed and direction, humidity, terrain elevation, etc.; Geographical location information: risk differences in different scenarios such as mountains, forests, industrial areas, etc., such as high-risk areas around transmission lines; The prior data can be obtained through the fire statistics data released by fire or emergency management departments in various countries; for example: the Fire and Rescue Bureau of the Ministry of Emergency Management of China regularly publishes fire data in residential places, including the causes of fires, regional distribution, casualty statistics, etc.; the National Interagency Fire Center in the United States provides records of annual fire area, number of fire occurrences, and allocation of rescue resources, covering forest wildfires and urban fires; and through academic research and open databases, etc.; Generate a dynamic scanning priority map: Structurize the prior data by OCR to recognize text information, image annotations, and sensor data, and use multimodal fusion technology to integrate text, images, and sensor data; Text information in scanned fire reports, historical archive pictures, and on-site photos, such as the fire starting time, coordinates, and burning area, etc.; Use Tesseract OCR to detect text regions in images and generate text coordinate boxes; Extract keyword fields through pre-trained models such as BERT and map them to structured fields; Correct OCR errors in combination with the fire domain knowledge base; YOLOv5s outputs bounding box coordinates to locate the fire point, Mask R-CNN segments the burning area, quantifies the burning area and the fire line morphology, rotates and simulates smoke for the fire satellite images to improve the model's robustness to low visibility scenarios; At the same time, clean and align the sensor data; Extract keywords from the reports extracted by OCR, such as "wind direction northeast" and "combustible density 0.5 kg / m²", and convert them into numerical vectors TF-IDF or word embeddings. Use ResNet to extract high-level features of the burning area, such as flame texture and smoke diffusion morphology, and calculate the burning intensity through the masked area; Unify the fire point coordinates, satellite images, and sensor positions into the same projection coordinate system; Align the time series of each modality data based on the fire occurrence time. The code is as follows: # Pseudo-code: Multimodal fusion based on attention mechanism class MultimodalFusion(nn.Module): def __init__(self): super().__init__() self.text_encoder = BertModel.from_pretrained('bert-base') # Text encoding self.image_encoder = ResNet50(pretrained=True) # Image encoding self.sensor_lstm = nn.LSTM(input_size=5, hidden_size=128) # Sensor time series encoding # Cross-modal attention layer self.cross_attn = nn.MultiheadAttention(embed_dim=256, num_heads=4) def forward(self, text, image, sensor): text_feat = self.text_encoder(text).last_hidden_state[:,0,:]# [CLS] token image_feat = self.image_encoder(image).flatten(1) sensor_feat, _ = self.sensor_lstm(sensor) # Concatenate features and calculate attention weights combined = torch.cat([text_feat, image_feat, sensor_feat],dim=1) attn_output, _ = self.cross_attn(combined, combined,combined) return attn_output Distinguish natural heat sources from abnormal fire situations: Collect data from the visible light channel and the thermal imaging channel respectively; for visible light, use a high-definition camera, such as with 30x optical zoom and fog penetration function, to capture details such as flame morphology, smoke contour, color changes, etc. During the day, it can identify ignition points of 1×3 square meters within 15 kilometers; for thermal imaging, use a non-cooled vanadium oxide detector with a resolution of 640×512 and a working band of 8 - 14 μm to detect temperature anomalies at the 0.1℃ level through smoke and haze. For example, in the dual-spectrum system of transmission lines, thermal imaging can locate a fire source of 1 square meter within 3 kilometers; And preprocess the collected visible light channel and thermal imaging channel: For the visible light channel, through dehazing algorithms and image enhancement techniques, convert the HSV→YUV color space to enhance the flame contrast; use the YOLOv5s model to detect flames / smoke and extract morphological features such as flame jitter frequency and smoke diffusion speed; For the thermal imaging channel, through NSCT (Non-Subsampled Contourlet Transform), extract temperature gradients and hot spot distributions; use Pulse Coupled Neural Network for high-frequency feature fusion and adaptive fuzzy logic algorithm to process low-frequency information; NSCT decomposition: Multi-scale and multi-directional feature extraction; NSCT is a multi-scale geometric analysis tool that realizes image decomposition through non-subsampled pyramids and non-subsampled directional filter banks; in thermal imaging processing, the role of NSCT is as follows: Perform thermal imaging preprocessing. By inputting the original image of the thermal imaging channel, perform normalization processing on the image, such as mapping the temperature value to the range [0,1]; eliminate noise through Gaussian filtering or median filtering; and perform multi-scale decomposition of NSCT to decompose the thermal imaging image into low-frequency subbands and high-frequency subbands according to scales; Low-frequency sub-band: Contains global temperature distribution information, such as the overall temperature gradient; High-frequency sub-band: Contains local detail information, such as the edges of hot spots and rapidly changing temperature regions; Perform multi-directional decomposition on each high-frequency sub-band, such as 4 directions, to capture the hot spot texture features in different directions; Calculate the temperature gradient through the low-frequency coefficients, such as the temperature difference between adjacent pixels, to reflect the overall temperature change trend of the thermal imaging; extract the hot spot distribution through the high-frequency coefficients, such as the boundaries, shapes, and densities of local high-temperature regions, and extract the temperature gradient and hot spot distribution; PCNN high-frequency feature fusion to enhance the hot spot details; input the high-frequency sub-bands after NSCT decomposition into PCNN; PCNN enhances the contrast of the hot spot edges and details through the pulse coupling between neurons; dynamically adjusts the activation threshold of neurons to highlight the high-frequency regions; fuses the high-frequency sub-bands in multiple directions to generate clearer hot spot edge information; outputs the enhanced high-frequency features, and the edges and textures of the hot spot regions are significantly enhanced for subsequent analysis; Feature extraction and fusion: Independently detect through visible light and thermal imaging. Visible light detection: The YOLOv5s model is trained based on more than 100,000 fire scene samples to identify flame forms, such as color, shape, dynamic jitter, and smoke characteristics; Thermal imaging detection: Distinguish fire points from other heat sources through the brightness temperature threshold method and the dual-band algorithm; Perform dual-light fusion and dynamically adjust the weights according to the scene. For example, visible light is mainly used during the day, and thermal imaging is mainly used at night. For instance, in the forest fire prevention system, the weight of visible light accounts for 70% during the day, and the weight of thermal imaging accounts for 90% at night; Adopt the DAF-Net domain adaptive dual-branch feature decomposition fusion network to reduce the distribution difference between infrared and visible light images through multi-kernel maximum mean discrepancy, and use the Transformer-CNN structure to retain cross-modal features, such as the consistency between the flame edge and the hot spot position; Divide the scanning area according to the thermal risk level: Through machine learning models, such as random forest or XGBoost, score the thermal risk of the area and divide it into levels, such as high risk, medium risk, and low risk. High-risk areas, such as areas where fires have occurred recently, are marked as priority scanning areas, and the scanning frequency of low-risk areas is reduced to save resources; Capture details such as the flame form and smoke contour of the area through the high-definition camera of the drone; adopt a non-cooled vanadium oxide detector with a working band of 8 - 14μm to penetrate smoke and haze and detect temperature anomalies at the 0.1℃ level. For example, in the transmission line dual-spectrum system, thermal imaging can locate a 1㎡ fire source within 3 kilometers; detect flames / smoke through the YOLOv5s model and extract morphological features, and extract the temperature gradient and hot spot distribution through NSCT; Dynamically adjust the weights according to the scenario. For example, in the daytime, visible light is the main component, and in the nighttime, thermal imaging is the main component, and then perform weighted fusion. Adopt the EMMA framework or the Transformer-CNN structure to retain cross-modal features. Exclude natural heat sources through temperature thresholds and morphological features, and dynamically update the classification model based on historical false alarm or missed alarm data: Use the highest resolution mode of thermal imaging to ensure the capture of micro fire sources; switch to the low resolution mode to reduce power consumption; locally magnify the suspicious area through the AI algorithm; Analyze thermal imaging data in real time through the NPU or GPU carried by the drone. If a temperature anomaly is detected, automatically trigger the high resolution mode, and the cloud calls the NSCT fusion model to process complex scenarios and dynamically optimize the resolution strategy; Overlay satellite remote sensing data and ground Internet of Things sensors through the GIS system to generate a three-dimensional thermal risk map, and update the risk level based on real-time meteorological data (wind speed, humidity). Allocate a higher scanning frequency to high-risk areas; conversely, continuously use the low resolution mode for recognition; Divide the area into 1km×1km grids, extract the mean value of each feature within the grid, calculate the average humidity change rate in the past 7 days, and at the same time count the fire frequency in the same period in the past 3 years. If a fire has occurred in the grid within 30 days, add a binary feature, otherwise it is 0; Based on historical data, if a fire has occurred in the grid during the observation period, mark it as high risk = 1, otherwise it is low risk = 0; Perform semi-supervised enhancement. For unlabeled areas, use K-Means clustering to generate pseudo-labels; Use the XGBoost model to identify key risk driving factors through the built-in feature importance ranking. Code example: import xgboost as xgb from sklearn.model_selection import TimeSeriesSplit # Time series cross-validation (to prevent data leakage) tscv = TimeSeriesSplit(n_splits=5) for train_idx, test_idx in tscv.split(X): X_train, X_test = X.iloc[train_idx], X.iloc[test_idx] y_train, y_test = y.iloc[train_idx], y.iloc[test_idx] # Training parameters (after optimization) params = { 'objective': 'binary:logistic', 'max_depth': 8, 'learning_rate': 0.05, 'subsample': 0.7, 'colsample_bytree': 0.6, 'scale_pos_weight': 10 # Handling class imbalance (few high-risk samples) } model = xgb.train(params, xgb.DMatrix(X_train, y_train), num_boost_round=500) # Validate AUC preds = model.predict(xgb.DMatrix(X_test)) auc_score = roc_auc_score(y_test, preds) By running in low-resolution mode for 80% of the time, the overall energy consumption is reduced by 65%. The false negative rate for high-risk area detection is < 3%. The time from anomaly trigger to high-resolution imaging completion is < 1 second, meeting the real-time warning requirements. Table 1: Dynamic scanning strategy;

[0020] The dynamic adjustment of thermal imaging resolution is achieved through the dynamic activation of hardware-level pixel subgroups, threshold-triggered local super-resolution, risk-level-driven resolution hierarchical mapping, deep learning super-resolution reconstruction, real-time enhancement of details, optical-electronic collaborative compensation, and energy consumption adaptive regulation. This technology enables the intelligent switching of resolution between 0.1MP and 2MP for high-dynamic range scenarios such as fire monitoring; Distributed edge computing module, through a multi-modal feature alignment network, performs spatio-temporal synchronization of cross-modal data; through adaptive computing scheduling, dynamically allocates CUDA cores and CPU resources; adopts an improved artificial bee colony algorithm, through a time-division multiplexing protocol, dynamically allocates the bandwidth of multi-sensor data streams for three-dimensional path planning; Three-dimensional path planning: Multi-modal feature alignment network: Through the multi-modal feature alignment network, spatio-temporal alignment of heterogeneous data from different sensors is performed to ensure the consistency of data in different modalities in terms of time stamps and spatial coordinates; through timestamp calibration, the time deviation between sensors is eliminated; Spatial mapping: Unify the coordinate systems of different sensors into the same reference system, such as geographical coordinates or device coordinate systems, to achieve physical consistency in data fusion; Dynamic allocation of CUDA cores and CPU resources: Dynamically allocate GPU and CPU resources to avoid resource idleness or overload; prioritize the real-time image processing tasks of fire monitoring to ensure low-latency response for fire monitoring tasks; adjust the resource allocation ratio according to real-time task requirements to avoid single-point resource bottlenecks; dynamically adjust the scheduling strategy in scenarios such as sensor data stream fluctuations and network bandwidth changes. In the distributed edge computing module, the dynamic allocation of CUDA cores and CPU resources is achieved by real-time monitoring of hardware loads, such as GPU utilization, CPU core occupancy, and task characteristics, computationally intensive / logically intensive, and combined with an improved artificial bee colony algorithm for multi-objective optimization to dynamically adjust the resource allocation strategy; computationally intensive tasks are preferentially allocated to CUDA cores for accelerated parallel processing; such as matrix operations in 3D path planning, while logically intensive tasks are the responsibility of the CPU for decision-making control, such as path optimization logic; through a time-division multiplexing protocol, the system dynamically divides the bandwidth according to the priority of the sensor data stream to ensure that high-real-time tasks, such as lidar point cloud processing, are preferentially transmitted; at the same time, asynchronous memory transfer and shared memory optimization are adopted to reduce the data copy latency between the GPU and the CPU; during the dynamic scheduling process, the reinforcement learning model provides real-time feedback on hardware performance metrics and adaptively adjusts the thread block division and task offloading strategy, such as migrating some tasks to the CPU when the GPU is overloaded, to achieve maximum resource utilization and energy consumption balance, and ultimately efficiently support complex tasks such as 3D path planning in scenarios such as autonomous driving and industrial robots. Task classification and priority division: For high-computation-load tasks such as image feature extraction and deep learning inference, they are scheduled to the GPU and accelerated using CUDA parallel computing capabilities, such as during the inference task of the YOLOv5s model; based on the real-time analysis of thermal imaging images and multi-modal data fusion. For low-latency tasks such as sensor data preprocessing and protocol parsing, they are scheduled to the CPU to handle lightweight tasks and ensure the real-time nature of the data stream, such as sensor data format conversion, timestamp calibration, and cross-modal alignment. Dynamic resource allocation strategy: Real-time collect metrics such as the load, temperature, and memory occupancy of the GPU and CPU through hardware performance counters; dynamically adjust resource allocation in combination with the task queue status, such as task priority and estimated execution time. First, task priority division is based on business requirements or real-time requirements. For example, high-priority tasks, such as emergency decision-making in 3D path planning, are preferentially allocated GPU cores for accelerated computing, while low-priority tasks, such as data log storage, are processed serially by the CPU. Secondly, execution time estimation uses a prediction model trained with historical task data, such as LSTM or ARIMA, to estimate the resource requirements and execution duration of tasks, and dynamically adjusts the resource allocation ratio. For example, if a task is expected to occupy 80% of the GPU's computing power, the system will reserve the corresponding resources and prevent other tasks from preempting them. On this basis, a dynamic scheduling algorithm, such as an improved artificial bee colony algorithm or priority queue scheduling combined with hardware performance metrics, GPU utilization, CPU load, and task queue status, generates a resource allocation plan in real time: high-priority tasks trigger a resource preemption mechanism to forcibly reclaim the GPU resources of low-priority tasks; tasks with longer execution times dynamically adjust the time slice allocation through a time-sharing multiplexing protocol to ensure maximum resource utilization; finally, the system continuously optimizes the scheduling strategy through a reinforcement learning model to adapt to task load fluctuations, such as sudden high-priority tasks, to achieve efficient resource utilization and service response stability. Improved artificial bee colony algorithm: To address the problem of low search efficiency of the standard ABC algorithm in discrete space, a neighborhood search strategy based on path local smoothness is introduced. For example, a point on the path is randomly selected, and new points are generated in its neighborhood and the path quality is evaluated through a probability density function; adaptive learning is carried out by dynamically adjusting parameters, such as the number of leading bees and the probability of scout bees, to quickly explore the global space in the initial stage of the search and finely search for local optimal solutions in the later stage; at the same time, a differential evolution algorithm is introduced in the scout stage to improve the convergence speed. Comprehensively consider path length, terrain cost, such as steepness, energy consumption, such as turning radius, threat avoidance, such as obstacle distance and high-risk area scanning density; Dynamic bandwidth allocation for multi-sensor data streams: Real-time collect indicators such as bandwidth occupancy, latency, and packet loss rate of each sensor data stream through tools; monitor the network load between edge nodes and the cloud, extract the resolution, frame rate of thermal imaging videos, and the characteristics and update frequencies of temperature and humidity sensor data streams; identify thermal imaging data and environmental temperature data for fire monitoring; classify sensor data streams based on a traffic service model: Assign high priority to thermal imaging videos and smoke detection signals; Assign medium priority to environmental data such as temperature, humidity, and air pressure; Assign low priority to device status data such as sensor health status and log information, and increase the priority of thermal imaging data according to real-time environmental conditions, such as high-temperature warnings; Dynamically allocate weights based on confidence or task requirements, such as increasing the weight of the IMU during high-speed movement; use the LSTM model to predict the data stream requirements of each sensor in the next period of time, and combine with the traffic prediction model to evaluate the network load trend; calculate the bandwidth requirements of the current data streams of each sensor and evaluate the total available bandwidth of the network link; By minimizing the latency of high-priority data streams and maximizing bandwidth utilization; generate an initial solution as a bandwidth allocation scheme, such as: Allocate 300 Mbps for thermal imaging data; Allocate 100 Mbps for temperature and humidity data; Allocate 100 Mbp for other data; By dynamically adjusting the bandwidth allocation ratio, combined with the time-division multiplexing protocol, divide the bandwidth into time slots and allocate the time slot lengths according to priorities; statistically perform time-division multiplexing dynamic allocation, dynamically allocate time slots according to real-time requirements, allocate more time slots for high-priority data streams, such as thermal imaging data occupying 60% of the time slots; low-priority data streams are only transmitted during idle periods, such as device status data being transmitted when thermal imaging is idle; use a buffer queue to temporarily store non-real-time data, such as temperature and humidity data, and transmit it in batches during low-load periods; control burst traffic through the token bucket algorithm to avoid congestion; apply the time-division multiplexing protocol, divide the total bandwidth into multiple equal-length time slots, allocate continuous time slots for high-priority data streams, and allocate fragmented time slots for low-priority data streams. Deploy a time-division multiplexer at the edge node to poll the data streams of each sensor according to time slots, and achieve low-latency time slot switching through hardware acceleration; Three-dimensional path generation: The path is represented by three-dimensional coordinate points (x, y, z), the starting point and the ending point are fixed, and the intermediate points are optimized by the ABC algorithm; randomly generate an initial path to ensure that flight constraints are met, such as the minimum turning radius and the maximum climbing angle; select excellent paths according to fitness values and update the positions of neighborhood points; other drones perform local searches by selecting paths according to probabilities; if the path is not improved, randomly generate a new path; use spline interpolation to eliminate jagged paths; combine lidar or vision sensors to dynamically update the obstacle map and trigger path replanning; at the same time, optimize covering high-risk areas and reducing energy consumption in path planning, for example, deploying dense formations in mountainous areas and sparse formations in flat areas; The edge-cloud collaborative decision-making module constructs a model sharing network driven by federated learning. Each edge node trains a lightweight YOLOv5s pruned model based on local data, generates a globally optimized model and synchronizes it backward, conducts dynamic networking, and adjusts the drone formation density in real time; Achieve cross-edge-node model collaborative training through federated learning, and adjust the drone formation density in real time through dynamic networking to optimize the fire monitoring efficiency; Construct a model sharing network driven by federated learning: Based on lightweight design, redundant convolutional layers are removed from the YOLOv5s model to reduce the computational load while maintaining detection accuracy; each drone and ground sensor node uses locally collected multimodal data such as thermal images, visible light images, and environmental parameters for model training; fire characteristics such as flames and smoke are detected and the risk level is output. After obtaining the spatio-temporally aligned multimodal data, after receiving the gradients of all edge nodes through the regional master node, the federated averaging algorithm in federated learning is used for weighted aggregation to generate a global model; weights are assigned according to the data volume or task priority of the edge nodes, and the nodes in the high fire risk area have higher weights; hyperparameters such as the learning rate and regularization coefficient of the global model are adjusted according to the aggregation result; the performance of the global model is evaluated using the reserved validation dataset.

[0021] Dynamic networking and adjustment of drone formation density: A wireless Mesh network is adopted, and each drone acts as a node, supporting multi-hop relay transmission; the drones are divided into multiple clusters based on the clustering algorithm, and a master node is designated in each cluster to be responsible for channel allocation to reduce conflicts; combined with Beidou / 5G fusion positioning to improve the positioning accuracy in complex environments; according to the heat risk score, a dense formation is deployed in high-risk areas and scanned once every 10 minutes; a sparse formation is adopted in low-risk areas and the scanning period is extended to 30 minutes. For example, in a forest fire prevention system, the drone swarm automatically encrypts the formation at the initial stage of the fire and covers a radius of 15 kilometers; the formation shape is dynamically adjusted by setting a leader and virtual leader points; and according to environmental changes, the behavior rules of the drones are adjusted in real time; at the same time, the cloud calls the NSCT fusion model to process complex scenarios and dynamically optimize the resolution strategy.

[0022] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution.

[0023] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0024] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. An edge-cloud collaborative intelligent fire monitoring system for image fusion, characterized in that, The system includes: A multi-modal image fusion module that generates a dynamic scanning priority map based on prior data, uses a dual-light fusion algorithm to distinguish natural heat sources from abnormal fire situations, divides the scanning area according to the heat risk level, and dynamically adjusts the thermal imaging resolution; A distributed edge computing module that synchronizes spatio-temporal cross-modal data through a multi-modal feature alignment network, dynamically allocates CUDA cores and CPU resources through adaptive computing scheduling; adopts an improved artificial bee colony algorithm to dynamically allocate the bandwidth of multi-sensor data streams through a time-division multiplexing protocol for three-dimensional path planning; An edge-cloud collaborative decision-making module that constructs a model sharing network driven by federated learning. Each edge node trains a lightweight YOLOv5s pruned model based on local data, generates a globally optimized model and synchronizes it backward for dynamic networking and real-time adjustment of the UAV formation density.

2. The edge-cloud collaborative intelligent fire monitoring system for image fusion according to claim 1, characterized in that, The process of generating the dynamic scanning priority map is as follows: Based on the historical data of fire points, which includes fire data, environmental information, and geographical location information, it is obtained as prior data and classified into text information, image annotations, and sensor data; the prior data is structured by OCR recognizing text information, image annotations, and sensor data, and multi-modal fusion technology is used to integrate text, images, and sensor data; YOLOv5s outputs the bounding box coordinates to locate the fire point position, Mask R-CNN segments the burning area, quantifies the burning area and the fire line morphology, rotates and simulates smoke for the fire satellite image, and at the same time cleans and aligns the sensor data; keywords are extracted from the reports extracted by OCR, converted into numerical vectors TF-IDF or word embeddings, high-level features of the burning area are extracted using ResNet, and the burning intensity is calculated through the masked area.

3. The edge-cloud collaborative intelligent fire monitoring system for image fusion according to claim 2, characterized in that, The process of distinguishing natural heat sources from abnormal fire situations is as follows: Data from the visible light channel and the thermal imaging channel are collected. For the visible light channel, through a dehazing algorithm and image enhancement technology, color space conversion is performed, and the YOLOv5s model is used to detect flame or smoke features and extract morphological features; The thermal imaging channel extracts the temperature gradient and hot spot distribution through NSCT (non-subsampled contourlet transform); a pulse-coupled neural network is used for high-frequency feature fusion, and an adaptive fuzzy logic algorithm is used to process low-frequency information; Using dual-light fusion, the weights are dynamically adjusted according to the scene. The DAF-Net (domain adaptive dual-branch feature decomposition fusion network) is adopted to reduce the distribution difference between infrared and visible light images through multi-core maximum mean discrepancy, and the Transformer-CNN structure is used to retain cross-modal features.

4. An edge-cloud collaborative intelligent fire monitoring system for image fusion according to claim 3, characterized in that, The process of dividing the scanning area according to the heat risk level is as follows: Through a random forest, the heat risk of the area is scored and classified; images of the area are captured by high-definition UAV cameras, and the temperature gradient and hot spot distribution are extracted through NSCT; the weights are dynamically adjusted according to the scene, and the EMMA framework is adopted to retain cross-modal features; Natural heat sources are excluded through temperature thresholds and morphological features, and the classification model is dynamically updated based on historical false alarms. The suspicious area is locally magnified through an AI algorithm; an NPU is carried by a drone to analyze thermal imaging data in real time. If an abnormal temperature is detected, the high-resolution mode is automatically triggered, and the NSCT fusion model is called by the cloud to process complex scenarios and dynamically optimize the resolution strategy; otherwise, continuous low-resolution mode recognition is performed. The area is divided into grids, the average value of each feature within the grid is extracted, and the average humidity change rate in the past 7 days is calculated. At the same time, the fire frequency in the same period in the past 3 years is statistically analyzed. If a fire has occurred in the grid within 30 days, a binary feature is added, otherwise it is 0; based on historical data, if a fire has occurred in the grid during the observation period, it is marked as high risk = 1, otherwise it is low risk = 0; semi-supervised enhancement is performed, and for unlabeled areas, K-Means clustering is used to generate pseudo-labels; the XGBoost model is used to identify key risk driving factors through the built-in feature importance ranking.

5. The edge-cloud collaborative intelligent fire monitoring system for image fusion according to claim 1, characterized in that, The process of spatio-temporal synchronization of the cross-modal data is as follows: Through timestamp calibration, different sensors are synchronized in time, and the coordinate systems of different sensors are unified to the same reference system.

6. The edge-cloud collaborative intelligent fire monitoring system for image fusion according to claim 5, characterized in that, The process of dynamic allocation of CUDA cores and CPU resources is as follows: GPU and CPU resources are dynamically allocated to preferentially process real-time image processing tasks for fire monitoring, and the resource allocation ratio is adjusted according to real-time task requirements. When the sensor data stream fluctuates and the network bandwidth changes, the scheduling strategy is dynamically adjusted. Dynamic resource allocation strategy: The load, temperature, and memory occupancy of the GPU and CPU are real-time collected through hardware performance counters; combined with the task queue status, the resource allocation is dynamically adjusted.

7. The edge-cloud collaborative intelligent fire monitoring system for image fusion according to claim 6, characterized in that The process of dynamic allocation of the bandwidth of multi-sensor data streams is as follows: The bandwidth occupancy, latency, and packet loss rate of each sensor data stream are collected, the network load between the edge node and the cloud is monitored, the resolution, frame rate of the thermal imaging video, and the characteristics of the temperature and humidity sensor data stream are extracted, and the update frequency is identified; the thermal imaging data and environmental temperature data for fire monitoring are identified; based on the traffic service model, the sensor data streams are classified. Weights are dynamically allocated through confidence or task requirements, and the LSTM model is used to predict the data stream requirements of each sensor in the future for a period of time. Combined with the traffic prediction model, the network load trend is evaluated. Calculate the current bandwidth requirements of each sensor data stream.

8. An edge-cloud collaborative intelligent fire monitoring system for image fusion according to claim 1, characterized in that, The process of three-dimensional path generation is as follows: The starting point and the ending point are fixed, and the intermediate points are optimized by the ABC algorithm; an initial path is randomly generated, excellent paths are selected according to the fitness value, and the positions of neighboring points are updated; other drones perform local searches according to the probability to select paths; if the path is not improved, a new path is randomly generated; spline interpolation is used to eliminate the jagged path; combined with lidar or vision sensors, the obstacle map is dynamically updated, and path replanning is triggered; at the same time, the coverage of high-risk areas is optimized during path planning.

9. The edge-cloud collaborative intelligent fire monitoring system for image fusion according to claim 1, wherein, The process of constructing a model sharing network driven by federated learning is as follows: Obtain the data after spatio-temporal alignment of multi-modal data. After receiving the gradients of all edge nodes through the regional master node, use the federated averaging algorithm in federated learning for weighted aggregation to generate a global model; allocate weights according to task priorities, and adjust the hyperparameters of the learning rate and regularization coefficient of the global model according to the aggregation results.

10. The edge-cloud collaborative intelligent fire monitoring system for image fusion according to claim 1, characterized in that, The process of real-time adjusting the density of the UAV formation is as follows: Adopt a wireless Mesh network, with each UAV as a node, supporting multi-hop relay transmission; divide the UAVs into multiple clusters based on a clustering algorithm, and specify a master node in each cluster to be responsible for channel allocation. According to the heat risk score, adjust the scanning strategy, and dynamically adjust the formation shape by setting leaders and virtual leading points; And according to environmental changes, real-time adjust the behavior rules of the UAVs; at the same time, the cloud calls the NSCT fusion model to process complex scenarios and dynamically optimize the resolution strategy.

Citation Information

Patent Citations

  • Fire instance segmentation method based on semi-supervised learning strategy

    CN114092798A

  • Neural network modeling method for training unmanned aerial vehicle cluster based on federated learning framework

    CN116847379A

  • Heterogeneous computing resource adaptive configuration and allocation method and device, and storage medium

    CN119363679A

  • Power transmission line forest fire risk assessment and early warning system and method

    CN119811051A

  • Power transmission line forest fire risk assessment method and system driven by flame combustion model

    CN119811052A

Cited By

  • Self-adaptive control system of hotspot acquisition equipment based on multi-modal data fusion

    CN120630726A

  • Hot spot acquisition device adaptive control system based on multi-modal data fusion

    CN120630726B

  • Fire-fighting data resource scheduling method based on shared Internet of Things

    CN120875435A

  • Intelligent factory multi-modal data processing system based on edge computing

    CN120951240A

  • An intelligent factory multi-modal data processing system based on edge computing

    CN120951240B