Small unmanned aerial vehicle airborne visual fire detection system based on chimeric cooperation model
Through the chimeric collaboration models A and B, the backbone and neck structure of YOLOv5n is optimized, combined with the lightweight LG Block and YOLOv8s, the problems of insufficient detection accuracy and excessive resource consumption of drone fire detection systems in urban environments are solved, efficient and accurate fire detection is achieved, and the timeliness and accuracy of detection is improved.
Patent Information
- Application Number
- CN202510297537.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-29
AI Technical Summary
The existing drone fire detection systems have problems such as insufficient detection accuracy and excessive resource consumption in urban environments, which are difficult to meet the needs of high accuracy and high real-time at the same time.
The chimeric collaboration model is adopted, including ultra-lightweight Model A and lightweight Model B. Model A is used for preliminary flame detection and Model B is used for multi-task detection. By optimizing the backbone and neck structure of YOLOv5n, it reduces the computational volume and memory usage, and combines the lightweight LG Block and YOLOv8s for collaborative work.
It realizes efficient and accurate fire detection in environments with limited drone resources, improves the timeliness and accuracy of detection, and provides strong support for fire rescue.
Smart Images

Figure CN120388303A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of UAV technology and computer vision, and specifically to a small UAV airborne vision fire detection system based on a chimeric cooperation model. The system aims to improve the accuracy and timeliness of urban fire detection, and is particularly suitable for high-speed flight UAV scenarios with limited resources. Background Art
[0002] Existing UAV fire detection systems mainly use lightweight object detection algorithms or image segmentation algorithms to detect fire sources and related targets. However, these solutions all have obvious drawbacks. Although the lightweight object detection algorithm has high real-time performance, its detection accuracy and effect are difficult to meet the high-precision requirements of urban fire detection. While the image segmentation algorithm has good detection effect, but it consumes a large amount of resources and has poor real-time performance, which does not meet the high real-time requirements of UAVs. Summary of the Invention
[0003] The present invention designs a deep learning-based chimeric cooperation model system for urban fire airborne edge computing scenarios. This model system is composed of two parts, model A and model B. The first part, model A, is an ultra-lightweight detection model with small computational requirements and very low memory occupancy. This model only focuses on the detection of flames and can achieve better detection effects compared to multi-task detection. The second part, model B, is a lightweight detection model with relatively rich model parameters. This model is responsible for multi-task detection, including the detection of flames, people, fire size, buildings, and floor windows. Its number of parameters is greater than that of model A and is more suitable for multi-task detection.
[0004] The first part, the ultra-lightweight model, is used for daily fire inspections, enabling the UAV to have high detection real-time performance and very low energy consumption under normal conditions, reasonably and effectively reducing the resource usage burden of the UAV. When the ultra-lightweight model detects a fire source, the second part of the model will be activated to obtain the feature information of the first stage and conduct a multi-task detailed and comprehensive detection of the fire scene to better judge the scene.
[0005] The structure of the ultra-lightweight model is as Figure 1 shown. Model A optimized the backbone based on YOLOv5n. Figure 1 The left image is input into the backbone. Based on YOLOv5n, we replaced the third, fourth, and fifth segments of the backbone in the figure with lightweight Conv&C3G (C3Ghost) modules. Among them, C3G×2 in the third stage represents two C3G modules, and the same applies to the fourth stage. And C3 was retained in the second stage to enhance the feature extraction ability of the shallow layer. Figure 1The middle position is the LG-PAN neck. In this neck structure, the C3 in the PAN neck structure of YOLOv5 is replaced by a self-developed lightweight LG block. As Figure 2 shown, we designed a lightweight LG block. The first channel of the LG block contains a Light Conv and a Ghostblock. As a lightweight feature extraction structure, the Ghost block is internally divided into two sub-channels. The first sub-channel sequentially includes a Ghost Conv, a DW Conv (depthwise separable convolution), and another Ghost Conv, which are used to capture detailed differential features. The second sub-channel consists of a DW Conv and a convolution module to enhance the diversity of features. The second channel uses a lightweight Light Conv module, which contains a Conv block without an activation function and a DW Conv to ensure the feature extraction ability. The output feature map sizes of both channels are half of the input, thereby reducing the computational amount and memory occupancy. Finally, through a concatenation operation (Concat) in the channel dimension, the outputs of the two channels are restored to the original model channel number.
[0006] The overall parameters of Model A are more lightweight. Compared with YOLOv5n, the number of parameters is reduced from 1.8M to 1.3M, and the computational amount is reduced from 4.1G to 2.9G, a decrease of 27.8% and 29.3% respectively.
[0007] Model B uses a lightweight YOLOv8s as the basic model. After Model 1 detects a fire, the input of the model camera video frame will switch to Model 2 for further confirmation of the fire. If it is confirmed by both models that a fire has occurred, Model 2 will continuously perform multi-task detection of the fire's flame, people, fire size, building, floor, and window, and return the detection results to the fire rescue command center through communication equipment. Description of the Drawings
[0008] Figure 1 Shows the overall structure diagram of the lightweight network of Model A, including the optimized backbone structure, the improved LG-PAN neck, and a simple schematic diagram of the self-developed LG Block.
[0009] Figure 2 Details the structure diagram of the LG Block, including its two internal sub-channels and their respective component modules.
[0010] Figure 3 Is the overall flowchart of the present invention, showing the conversion process of the system from the initial state to the safe state, the warning state, and the emergency state, as well as the cooperation mechanism between Model A and Model B.
[0011] Figure 4 Schematic diagram of the invention process Detailed implementation manners
[0012] The detailed implementation manners are as Figure 3 shown. The system includes two core models: Model A and Model B. The system takes real-time video frames as the original data input and sets the initial state as the "safe state".
[0013] First, the system starts Model A to perform preliminary fire detection on the input real-time video frames. If Model A fails to detect any signs of fire from the current video frame, it continues to continuously detect subsequent video frames; otherwise, once Model A detects a fire, the system immediately marks the current state as the "warning state" and triggers the intervention of Model B.
[0014] In the "warning state", the system transfers the video frames suspected of containing a fire to Model B for further verification and detection. If the detection result of Model B also fails to confirm the existence of a fire, the system regards this frame as a false alarm, resets the state to the "safe state", and Model A continues its normal fire detection process. However, if Model B also confirms the existence of a fire, the system state is upgraded to the "emergency state", indicating that the fire has been jointly confirmed by the two models.
[0015] In the "emergency state", the system activates the emergency response mechanism and commands Model B to continuously perform fire detection on the subsequently continuously input video frames to track the fire dynamics. If Model B fails to detect a fire source in the next five consecutive video frames, it is considered that the fire has been effectively controlled or eliminated, and the system state is adjusted back to the "safe state" again, ending the emergency response process.
[0016] Through the above dual-model cooperation mechanism, the present invention effectively improves the accuracy and timeliness of fire detection, providing strong support for taking timely fire extinguishing measures.
Claims
1. A small unmanned aerial vehicle (UAV) airborne vision fire detection system based on a chimeric collaboration model, characterized in that, Including: Model A, as an ultra-lightweight detection model, is specifically used for flame detection. Its structure is optimized based on YOLOv5n. Specifically, the third, fourth, and fifth segments of the backbone are replaced with lightweight Conv&C3G (C3Ghost) modules to reduce the computational load and memory occupancy; Model B, as a lightweight multi-task detection model, is used to perform multi-task detection on the fire scene, including flame, personnel, fire size, buildings, and floor windows, after Model A detects the fire source; The chimeric cooperation mechanism is used to trigger the intervention of Model B when Model A detects the fire source, realizing the combination of preliminary rapid fire detection and subsequent detailed multi-task detection.
2. The system according to claim 1, wherein The backbone optimization of Model A includes replacing the third, fourth, and fifth segments with lightweight Conv&C3G modules. The third and fourth stages respectively contain two C3G modules, and C3 is retained in the second stage to enhance the feature extraction ability of the shallow layer.
3. The system according to claim 1, wherein Model A also includes an LG-PAN neck structure, where C3 is replaced with a self-developed lightweight LG block. The LG block contains two sub-channels. The first sub-channel sequentially includes GhostConv, DW Conv, and another Ghost Conv to capture detailed differential features; the second sub-channel consists of DW Conv and a convolutional module to enhance the diversity of features.
4. The system according to claim 1, wherein The overall number of parameters of Model A is reduced compared to YOLOv5n. Specifically, it is reduced from 1.8M to 1.3M, and the computational load is reduced from 4.1G to 2.9G.
5. The system according to claim 1, wherein Model B uses lightweight YOLOv8s as the basic model. After Model A detects the fire, the video frames suspected of containing the fire are transmitted to Model B for further verification and detection.
6. The system according to claim 1, characterized in that, It also includes a state management mechanism for adjusting the system state according to the detection results of Model A and Model B, including "safe state", "warning state", and "emergency state", and performing corresponding operations according to the state.
7. The system according to claim 6, wherein In the "emergency state", the system activates the emergency response mechanism and commands Model B to continuously perform fire detection on the subsequently continuously input video frames to track the fire dynamics; If Model B does not detect the fire source in the next five consecutive video frames, it is considered that the fire has been effectively controlled or eliminated, and the system state is adjusted back to the "safe state".