Fire accurate rescue implementation method, system and equipment based on unmanned aerial vehicle and medium

By acquiring dual-light images of high-rise building fire areas using drones, removing smoke through downsampling and rolling guided filtering algorithms, and fusing the images with infrared images, an improved target detection model was used to detect trapped personnel. This solved the problem of difficult location in high-rise building fires and improved rescue efficiency and safety.

CN120953575APending Publication Date: 2025-11-14WUHAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510959823.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies are insufficient for effectively and quickly locating trapped individuals in high-rise building fire rescue operations. They are limited by the operating height of ground equipment, the range of water jets, the effects of dense smoke, and complex structures, resulting in low rescue efficiency and safety risks.

Method used

A precise fire rescue method based on drones is adopted. The drones acquire dual-light images of the fire area, and the smoke is removed by using downsampling and rolling guided filtering algorithms to generate smoke-free images. These images are then fused with infrared images, and an improved target detection model is used to detect the location of trapped personnel.

Benefits of technology

It enables rapid and accurate identification of trapped personnel at high-rise fire scenes, reduces the time spent by rescuers entering dangerous areas, improves rescue efficiency and reduces safety risks, and provides comprehensive and accurate information on the fire area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953575A_ABST
    Figure CN120953575A_ABST
Patent Text Reader

Abstract

The invention provides a fire accurate rescue implementation method, system and device based on an unmanned aerial vehicle and a medium, and belongs to the technical field of emergency rescue, and the method comprises the steps that a dual-light image of a fire area shot by an imaging assembly on the unmanned aerial vehicle is acquired, and the dual-light image comprises an original visible light image and an original infrared image; according to a down-sampling algorithm and a rolling guide filtering algorithm, processing the original visible light image to obtain a smoke-removed image; generating a fusion image according to the smoke-removed image and the original infrared image; and based on an improved target detection model, detecting trapped persons in the fire area and positions of the trapped persons according to the fused image. According to the invention, rescue workers can comprehensively and accurately grasp fire area information, and the rescue efficiency and safety are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to a method, system, equipment, and medium for precise fire rescue based on unmanned aerial vehicles (UAVs). Background Technology

[0002] With the acceleration of urbanization, the number of high-rise and super high-rise buildings has increased dramatically. Although these buildings improve space utilization efficiency, their prominent height, complex structure, and difficulty in evacuating people make them extremely prone to causing major casualties and property losses in the event of a fire.

[0003] Current fire rescue methods have many limitations when dealing with high-rise building fires. For example, ground rescue equipment (such as ladder trucks) has limited operating height, making it difficult to cover the area where a high-rise fire occurs. High-pressure water cannons have limited range, making them ineffective at extinguishing fires at heights. Simultaneously, fire areas are often accompanied by dense smoke and low visibility, making it difficult for rescuers to see the situation clearly and accurately locate trapped individuals, severely hindering rescue operations. Furthermore, the complex internal structure of high-rise buildings and the presence of toxic gases and other hazards further increase the difficulty of rescue and pose a significant threat to the lives of rescue personnel. While current image enhancement technologies (such as dark channel prior dehazing algorithms and multi-scale Retinex methods) have made progress in fire rescue research, their robustness and real-time performance in real fire environments remain insufficient, failing to meet the demands of complex and ever-changing fire zones.

[0004] In conclusion, it is urgent to develop a more efficient, intelligent, and mobile positioning system for high-rise building fire rescue in order to cope with the complex situation of urban high-rise fire rescue, improve fire rescue efficiency, and protect the safety of life and property. Summary of the Invention

[0005] In view of this, it is necessary to provide a method, system, equipment and medium for precise fire rescue based on drones, in order to solve the technical problem that existing technologies cannot effectively and efficiently carry out rescue in complex and ever-changing urban high-rise fire areas.

[0006] To address the aforementioned technical problems, in a first aspect, the present invention provides a method for precise fire rescue based on unmanned aerial vehicles (UAVs), comprising: Acquire dual-light images of the fire area captured by the imaging components on the drone, the dual-light images including a raw visible light image and a raw infrared image; The original visible light image is processed using a downsampling algorithm and a rolling guided filtering algorithm to obtain a smoke-free image. A fused image is generated based on the smoke-removed image and the original infrared image; Based on the improved target detection model, the trapped personnel and their locations in the fire area are detected according to the fused image.

[0007] In one possible implementation, processing the original visible light image to obtain the smoke-free image according to the downsampling algorithm and the rolling guided filtering algorithm includes: Multiple resolution target visible light images are obtained by downsampling the original visible light image based on the Gaussian pyramid. The target visible light image is processed based on the rolling guided filtering algorithm, and the smoke-removed image is obtained based on the processed image and the original target visible light image.

[0008] In one possible implementation, the step of processing the target visible light image based on the rolling guided filtering algorithm, and obtaining the smoke-removed image based on the processed image and the original target visible light image, includes: At each resolution level, the target visible light image is subjected to multiple iterative filtering processes based on the rolling guided filtering algorithm to obtain the corresponding smoke component patch; The target component image is obtained by upsampling the smoke component patch, and the resolution of the target component image is equal to the resolution of the original visible light image; Based on the weights of the smoke component images at different resolutions, a weighted average of all the target component images is calculated to obtain the final smoke image; The smoke-free image is obtained based on the original visible light image and the final smoke image.

[0009] In one possible implementation, the step of performing multiple iterative filtering processes on the target visible light image based on a rolling guided filtering algorithm at each resolution level to obtain corresponding smoke component patches includes: Gaussian kernel smoothing is applied to multiple visible light images of the target at the specified resolutions to obtain a smoothed image; After performing multiple iterations of the smoothed image based on joint bilateral filtering, the corresponding smoke component patch is obtained by comparing the final filtered image after the last iteration with the original visible light image.

[0010] In one possible implementation, after acquiring the dual-light image of the fire area captured by the imaging component on the UAV, and before processing the original visible light image according to a downsampling algorithm and a rolling guided filtering algorithm to obtain the smoke-removed image, the process includes: The original visible light image and the original infrared image are spatially aligned and then temporally aligned.

[0011] In one possible implementation, generating the fused image based on the smoke-removed image and the original infrared image includes: The first feature map of the smoke-removed image and the second feature map of the original infrared image are extracted based on the dual-branch encoder. Initial fused features are obtained based on the first feature map and the second feature map; The first feature map, the second feature map, and the initial fused feature are input into the cross-scale iterative attention decoder to generate intermediate fused features; Based on the cross-scale iterative attention decoder, the intermediate fusion features are transformed according to the activation function to generate the fused image; The probability values ​​of whether the fused image is similar to the original infrared image and the smoke-removed image are evaluated based on two discriminators respectively. The parameters of the cross-scale iterative attention decoder and the discriminator are optimized until the probability value reaches a preset threshold. The final fused image is then obtained based on the optimized cross-scale iterative attention decoder.

[0012] In one possible implementation, the step of detecting trapped personnel and their locations in the fire area based on the fused image using an improved target detection model includes: A hybrid attention mechanism residual network is added to the YOLOv11 network model to construct the improved object detection model. The hybrid attention mechanism residual network includes a first residual block and a second residual block, both of which include a first network branch. Each first network branch includes a first convolutional layer, a depthwise separable convolutional layer, an attention mechanism network, and a second convolutional layer. The spatial dimensions of the first and second convolutional layers are 1×1, and the spatial dimension of the depthwise separable convolutional layer is 3×3. The output of the first convolutional layer, the input of the depthwise separable convolutional layer, and the input of the second convolutional layer are connected sequentially. The input of the attention mechanism network is... The first residual block is connected to the output of the depth-separable convolutional layer and the input of the second convolutional layer, respectively. The first residual block also includes a first weight layer and a second network branch. The output of the second convolutional layer of the first residual block is connected to the input of the first weight layer. The second network branch includes a residual connection layer and a second weight layer. The input of the residual connection layer is connected to the input of the first convolutional layer. The output of the residual connection layer is connected to the input of the second weight layer. The output of the first weight layer is connected to the output of the second weight layer. The first weight layer and the second weight layer respectively set the weight ratio of the output of the first network branch and the second network branch. The fused image is input into the improved target detection model to output person recognition results; Wherein, the first network branch of the first residual block corresponds to the first weight value, the second network branch corresponds to the second weight value, and the sum of the first weight value and the second weight value is equal to 1.

[0013] Secondly, the present invention also provides a precision fire rescue system based on unmanned aerial vehicles (UAVs), including a UAV and a display device connected to the UAV. The drone is used to capture dual-light images of the fire area, the dual-light images including a raw visible light image and a raw infrared image; the raw visible light image is processed according to a downsampling algorithm and a rolling guided filtering algorithm to obtain a smoke-free image; a fused image is generated based on the smoke-free image and the raw infrared image; and based on an improved target detection model, trapped personnel and their locations are detected in the fire area according to the fused image. The display device is used to display the test results.

[0014] Thirdly, the present invention also provides an electronic device, including a memory and a processor, wherein, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the drone-based precision fire rescue method described in any of the above implementations.

[0015] Fourthly, the present invention also provides a computer-readable storage medium for storing a computer-readable program or instruction, which, when executed by a processor, can implement the steps in the above-described method for precise fire rescue based on unmanned aerial vehicles.

[0016] The beneficial effects of this invention are as follows: The method for precise fire rescue based on drones provided by this invention first acquires a dual-light image, including the original visible light image and the original infrared image, from the visible light camera and infrared thermal imaging module installed on the drone. Then, a downsampling algorithm and a rolling guided filtering algorithm are used to remove smoke areas from the original visible light image to generate a smoke-free image. The image features of the smoke-free image and the image features of the original infrared image are fused to generate a fused image with complementary information. Then, the fused image is input into an improved target detection model, which detects and identifies trapped personnel in the fire area and locates their positions. Based on the dual-light image captured by the drone, which can quickly fly to a high-rise fire area and capture images, the detection and location of trapped personnel can help rescuers understand the dangerous areas of the fire area in advance, avoiding the time wasted blindly entering dangerous areas such as high temperature and dense smoke for reconnaissance, and reducing the safety risks to rescuers themselves. Furthermore, this invention employs a downsampling algorithm and a rolling guided filtering algorithm to remove smoke from the original visible light image, resulting in a clearer image. This allows the fused image generated from the dual-light images (including the original visible light image and the original infrared image) to more comprehensively and accurately reflect the situation at the fire scene, providing richer information for subsequent detection of trapped personnel. Moreover, because this invention uses an improved target detection model to detect and locate trapped personnel from the fused image, it reduces the difficulty of accurately detecting and locating trapped personnel due to dense smoke in the fire area, minimizing the possibility of misjudgment. This provides rescuers with comprehensive and accurate information about the fire area, thereby assisting them in planning rescue strategies more quickly and effectively, and improving rescue efficiency. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic flowchart of an embodiment of the drone-based precision fire rescue method provided by the present invention; Figure 2 For the present invention Figure 1 A schematic diagram of an embodiment of S200; Figure 3 For the present invention Figure 2 A schematic diagram of an embodiment of S220; Figure 4 For the present invention Figure 3A schematic diagram of an embodiment of generating a smoke-removed image after a series of processing steps on the original visible light image; Figure 5 For the present invention Figure 1 A schematic diagram of an embodiment of S300; Figure 6 This is a schematic diagram of the CrossFuse model provided by the present invention; Figure 7 This is a schematic diagram of the structure of the hybrid attention mechanism residual network provided by the present invention; Figure 8 A schematic diagram of an embodiment of the electronic device provided by the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0020] In the description of the embodiments of the present invention, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0021] The terms "first," "second," etc., used in the embodiments of this invention are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a technical feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.

[0022] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0023] Before demonstrating the embodiments, the following terms will be explained.

[0024] Rolling Guidance Filter (RGF) is an advanced image multi-scale decomposition and filtering technique, primarily used to smooth images, remove noise and small-scale textures, while precisely preserving and enhancing significant edges and structures.

[0025] A Gaussian pyramid is a method for multi-resolution image representation. It generates image sequences by progressively reducing the image resolution through low-pass filtering (usually Gaussian filtering) and downsampling. The Gaussian filtering in the pyramid is used for multi-scale image representation. It constructs each layer of the image pyramid through smoothing and downsampling. It is a linear filter that uses a fixed Gaussian kernel to convolve the image, primarily to remove high-frequency information (smooth the image) and prevent aliasing during downsampling.

[0026] Gaussian kernel smoothing is a linear filtering method that uses a Gaussian kernel function as the convolution kernel to perform a weighted average on the image, thereby achieving blurring and noise reduction. Gaussian filtering in Gaussian kernel smoothing is used for adaptive smoothing of local regions and is a nonlinear filter.

[0027] Joint Bilateral Filtering is a filtering method based on a guide image. It introduces an additional guide image to guide the filtering process on the basis of traditional bilateral filtering. It uses information from the guide image (such as edges and textures) to guide the filtering process, thereby better preserving edges and details.

[0028] HSV (Hue, Saturation, Value) is a color model used to describe the three basic attributes of color: H for hue (representing the "type" or "name" of the color), S for saturation (representing the purity or intensity of the color), and V for value (representing the lightness or darkness of the color).

[0029] Histogram equalization is a non-linear transformation method used to improve image contrast. It is specifically designed to improve the contrast and brightness of images, and is especially suitable for low-contrast, low-brightness images caused by fire smoke. It enhances the overall contrast of the image by redistributing the gray-level distribution in the image, making the pixel values ​​more evenly distributed.

[0030] The CrossFuse model is a novel algorithm for fusing infrared and visible light images. It introduces a Cross Attention Mechanism (CAM) to enhance the complementary (uncorrelated) features between different modalities (infrared and visible light) while reducing redundant information. This mechanism differs from traditional attention mechanisms that focus on the correlation between features; instead, it focuses on extracting and enhancing unrelated features between modalities, thereby improving the information richness and visual quality of the fused image.

[0031] This invention provides a method, system, equipment, and medium for precise fire rescue based on unmanned aerial vehicles (UAVs), which will be described below.

[0032] Figure 1 This is a schematic flowchart of an embodiment of the drone-based precision fire rescue method provided by the present invention, as shown below. Figure 1 As shown, the method for achieving precise fire rescue based on drones includes the following steps: S100. Acquire a dual-light image of the fire area captured by the imaging component on the drone, the dual-light image including the original visible light image and the original infrared image.

[0033] It should be noted that the fire area includes both the outer perimeter of the fire scene and the area close to the fire scene. The drone is equipped with a visible light camera, an infrared thermal imaging module, an onboard processor, and a communication module. All three components are connected to the communication module, which receives control commands from the remote controller carried by rescue personnel to execute flight and imaging tasks. Rescue personnel can control the drone to quickly reach the fire area. The drone's flight path and imaging position in high-rise building fire rescue need to be dynamically adjusted according to the specific situation at the fire scene. Typically, the drone will begin reconnaissance from the perimeter of the fire scene, gradually approaching it. Once the drone reaches the fire area, it can activate the visible light camera to capture raw visible light images from multiple angles and the infrared thermal imaging module to capture raw infrared images from multiple angles, focusing on capturing images of the burning floor and the possible locations of trapped individuals.

[0034] In some embodiments of this invention, the communication module can be a Lightbridge module. The Lightbridge module supports 1080p full HD resolution, ensuring the transmission of clear image data to meet the needs of accurate identification of trapped personnel and the situation on-site in fire scenarios. Simultaneously, the Lightbridge module's transmission latency is controlled within 200ms, ensuring real-time synchronization between the acquired images and the drone's flight status, enabling rescue personnel to monitor and analyze the data promptly, providing strong support for rescue decisions. In open environments, the Lightbridge module's effective transmission distance can reach 5 kilometers, thus comprehensively covering the fire scene and meeting the continuous data acquisition needs of different locations in high-rise buildings. Furthermore, the Lightbridge module utilizes adaptive frequency hopping technology to address multipath effects and electromagnetic interference, possessing strong anti-interference performance. Even in complex fire scene environments, it maintains signal stability, greatly reducing the risk of image interruption. The Lightbridge module also adopts a lightweight and low-power design, weighing less than 50g and consuming less than 3W, effectively reducing the load on the drone and extending the single acquisition operation time by approximately 15%-20%, significantly improving rescue efficiency. Finally, the compact, integrated packaging of the Lightbridge module is compatible with mainstream industry drones and will not adversely affect the drone's flight maneuverability, ensuring stable operation of the drone during flight.

[0035] S200. The original visible light image is processed according to the downsampling algorithm and the rolling guided filtering algorithm to obtain the smoke-free image.

[0036] It should be noted that the airborne processor can connect to an infrared thermal imaging module and a visible light camera. This allows the airborne processor to acquire raw infrared and visible light images from these devices. Alternatively, the infrared thermal imaging module and visible light camera can directly send the raw infrared and visible light images to the communication module, enabling the electronic device to receive the dual-light images. Then, the airborne processor (or electronic device) can perform smoke removal processing on the raw visible light image using downsampling and rolling guided filtering algorithms. This removes smoke areas from the raw visible light image, resulting in a smoke-free image. The downsampling algorithm is selected based on different scenarios and requirements. Downsampling algorithms include bicubic interpolation, Gaussian downsampling, median downsampling, Fourier transform downsampling, wavelet transform downsampling, and deep learning-based downsampling algorithms.

[0037] S300. Generate a fused image based on the smoke-removed image and the original infrared image.

[0038] It should be noted that: Fusion images can be generated using pixel-level fusion methods, where each pixel value of the smoke-removed image and the original infrared image is weighted and averaged to obtain a weighted pixel value, and then the corresponding fusion image is generated based on this weighted pixel value. Alternatively, fusion images can be generated using multi-scale transformation methods, where the smoke-removed image and the original infrared image are separately decomposed using wavelet transform to extract detailed information at different scales and directions. The decomposed coefficients are then fused, and finally, the fusion image is reconstructed using inverse wavelet transform. Finally, fusion images can also be generated based on feature extraction methods, where the pixel values ​​of the smoke-removed image and the original infrared image are used as feature vectors input into a PCA model, the principal components are extracted using PCA, and then the fusion image is reconstructed based on these principal components.

[0039] S400. Based on the improved target detection model, detect the trapped personnel and their locations in the fire area according to the fused image.

[0040] It's important to note that a suitable base model for object detection should be chosen, such as YOLO (You Only LookOnce), SSD (Single Shot MultiBox Detector), or Faster R-CNN. Since fire scenes may contain significant interference from smoke, flames, and other factors, an attention mechanism can be introduced into the base model. Then, a large amount of fire scene image data is prepared, and this data undergoes preprocessing steps (including data augmentation, normalization, and annotation). The preprocessed image data includes trapped individuals and their annotation information (such as bounding boxes and category labels). The base model with the introduced attention mechanism is trained using this preprocessed image data to obtain an improved object detection model. Then, the trained improved object detection model is used to detect the trapped individuals in the fused image, obtaining their bounding boxes and confidence scores. Furthermore, based on the detected bounding boxes, the location of the trapped individuals in the fused image is determined and converted into actual location information (such as latitude, longitude, or distance).

[0041] In summary, the UAV-based precise fire rescue method provided by this invention utilizes dual-light images captured by UAVs quickly reaching high-rise fire areas to detect and locate trapped personnel. This helps rescuers understand the danger zones of the fire area in advance, avoiding the time wasted blindly entering high-temperature, dense smoke, and other dangerous areas for reconnaissance, and reducing the safety risks to rescuers. Furthermore, this invention employs downsampling and rolling guided filtering algorithms to remove smoke from the original visible light images, making the original visible light images clearer. Thus, the fused image generated from the dual-light images (including the original visible light image and the original infrared image) can more comprehensively and accurately reflect the situation at the fire scene, providing richer information for subsequent trapped personnel detection. Moreover, since this invention uses an improved target detection model to detect and locate trapped personnel from the fused image, it reduces the difficulty of accurately detecting and locating trapped personnel due to dense smoke in the fire area, reducing the possibility of misjudgment. This provides rescuers with comprehensive and accurate information about the fire area, thereby assisting them in planning rescue plans more quickly and effectively, and improving rescue efficiency.

[0042] In some embodiments of the present invention, such as Figure 2 As shown, step S200 includes: S210. The original visible light image is downsampled based on the Gaussian pyramid to obtain target visible light images of multiple resolutions; S220. The target visible light image is processed based on the rolling guided filtering algorithm, and the smoke-removed image is obtained based on the processed image and the original target visible light image.

[0043] Image size is typically expressed as "width × height," while resolution refers to the product of the image's width and height, which is the total number of pixels in the image. Based on a Gaussian pyramid, the original visible light image is downsampled multiple times (alternating row and column sampling) to obtain multiple target visible light images of different resolutions. The resolution of any target visible light image is smaller than that of the original visible light image. Visible light images of different resolutions retain information at different frequencies. Higher-resolution original and target visible light images retain high-frequency information (e.g., edges, small-scale details), while lower-resolution images retain low-frequency information (e.g., general outlines, overall brightness, smoke, and other large-scale structures). The original and target visible light images are arranged in a pyramid-like structure, arranged sequentially from largest to smallest image size (i.e., resolution), from bottom to top (or top to bottom). Each layer of the image is obtained by sequentially applying Gaussian filtering and downsampling to the previous layer, with the resolution halved with each layer, resulting in increasingly smaller image sizes. After multiple downsamplings of the original visible light image to obtain target visible light images of multiple resolutions, the rolling guided filtering algorithm is used to process the target visible light images of multiple lower resolutions, so as to remove smoke without blurring the edges, textures and details of the image to obtain the smoke-removed image.

[0044] It should be noted that the number of downsampling times can be set while balancing the sampling effect and the running time. If the number of downsampling times is too large or too small, no effective components will be extracted in the end. Since downsampling takes a long time, the number of downsampling times is set to 3 to 4 times in this invention.

[0045] This invention employs a Gaussian pyramid to perform Gaussian filtering and downsampling on the original visible light image, generating multiple lower-resolution target visible light images. This provides a multi-scale image representation, allowing for the separation of smoke at different scales and reducing the computational cost of subsequent Redirect Filtering (RGF). Then, RGF is used to filter the target visible light images, effectively removing low-frequency smoke while preserving high-frequency information (e.g., edges, textures, fine structures) at each resolution scale. Therefore, by combining the Gaussian pyramid and the rolling guided filtering algorithm, the smoke-free image obtained after processing the original visible light image shows significant improvements in contrast, sharpness, and color reproduction, effectively eliminating smoke interference while avoiding image quality degradation caused by excessive smoothing. This improves the image quality of the subsequent fused image generated from the smoke-free image, thereby enhancing the accuracy and robustness of detecting trapped personnel based on the fused image.

[0046] In some embodiments of the present invention, such as Figure 3As shown, step S220 includes: S221. At each resolution level, the target visible light image is iteratively filtered multiple times based on the rolling guided filtering algorithm to obtain the corresponding smoke component patch. S222. Upsample the smoke component image to obtain a target component image, wherein the resolution of the target component image is equal to the resolution of the original visible light image; S223. Based on the weights of the smoke component images at different resolutions, a weighted average of all the target component images is calculated to obtain the final smoke image; S224. Obtain the smoke-removed image based on the original visible light image and the final smoke image.

[0047] Specifically, since smoke appears as a large, hazy area in an image, belonging to the low-frequency component of the image, and the pixel values ​​of smoke and background are similar, direct subtraction or single-pass filtering is insufficient for accurate separation. Therefore, this application uses a rolling guided filtering algorithm to filter target visible light images at multiple resolutions, iteratively extracting smoke component patches through multiple iterations.

[0048] Since smoke component patches at different resolutions are obtained through Gaussian pyramid and rolling guided filtering algorithms, and the resolution of these smoke component patches is lower than that of the original visible light image, it is necessary to restore the resolution of the smoke component patches to the size of the original image in order to remove the smoke component patches and obtain the smoke-free image. This invention upsamples the smoke component patches extracted from each layer of the Gaussian pyramid to generate corresponding target component images, with the resolution of the target component images being consistent with that of the original visible light image. The upsampling process includes bicubic interpolation, nearest neighbor interpolation, bilinear interpolation, and deep learning-based super-resolution reconstruction methods (such as SRCNN, ESPCN, etc.).

[0049] In step S221 above, multiple resolution visible light images of the target were obtained through Gaussian pyramids, and corresponding smoke component patches were extracted for each. These smoke component patches reflect smoke information at different scales. To obtain a more accurate smoke distribution, the target component images corresponding to these smoke component patches at different resolutions are weighted and fused, that is, each resolution smoke component patch is assigned a weight. Thus, the target component images upsampled from smoke component patches at different resolutions each correspond to a weight, and the sum of the weights corresponding to all resolutions of smoke component patches equals 1. The weights can be determined based on the sharpness or signal-to-noise ratio of the smoke components; the weights are positively correlated with sharpness and also positively correlated with the signal-to-noise ratio. The pixel values ​​of all pixels in the current target component image are multiplied by their corresponding weights to calculate the product. The product results of all target component images are summed and averaged to obtain the final smoke image. That is, the pixel value of each pixel in the final smoke image is equal to the weighted average of the pixel values ​​of all target component images at that pixel. Then, the original visible light image and the final smoke image are subjected to image subtraction processing, that is, the pixel value of each pixel in the original visible light image is subtracted from the pixel value of the corresponding pixel in the final smoke component image, so as to obtain the smoke-free image.

[0050] In this embodiment of the invention, a rolling guided filtering algorithm is used to perform multiple iterative filtering processes on target visible light images at various resolutions, thereby obtaining smoke component patches corresponding to each resolution level. This allows for the capture of smoke features at different scales, avoiding confusion between smoke and background at a single scale, and thus enabling accurate smoke extraction while preserving important details such as edges and textures, avoiding over-smoothing, and improving the stability and accuracy of smoke component patch extraction. Furthermore, this invention upsamples the smoke component patches extracted at each resolution level to the resolution of the original visible light image using an upsampling algorithm. This upsampling process preserves the spatial details of the smoke components as much as possible, avoiding information loss. Moreover, unifying all smoke component patches to the original image resolution facilitates subsequent weighted average calculation and subtraction processing to obtain the final smoke image. Since smoke component patches at different resolutions contain smoke information at different scales, a pixel-by-pixel weighted average calculation is performed on all target component images based on the weights of the smoke component patches at different resolutions to obtain the final smoke image. This fully utilizes the advantages of each scale, improving the completeness and accuracy of the generated final smoke image. By employing direct subtraction or fusion subtraction, the final smoke image is removed from the original visible light image to generate a smoke-free image. This method can preserve the real scene details in the original image to the greatest extent, avoid over-processing, and significantly reduce the impact of smoke on image quality. It also improves the clarity and visibility of the smoke-free image, thereby providing a more reliable input for subsequent target detection, recognition, and localization, and significantly improving detection accuracy and robustness.

[0051] In some embodiments of the present invention, step S221 includes: S2211. Perform Gaussian kernel smoothing on the multiple target visible light images of the specified resolutions to obtain a smoothed image; S2212. After performing multiple iterative processing on the smoothed image based on joint bilateral filtering, the corresponding smoke component patch is obtained based on the final filtered image after the last iteration and the original visible light image.

[0052] Specifically, the target visible light image I is first subjected to Gaussian smoothing to obtain an initial estimate J. (t) The pixel value of each pixel, based on the current estimate J (t) To guide the process, joint bilateral filtering is performed on the visible light image I of the target to obtain a new estimate J. (t +1) Based on the above formula (2), the smooth image is iteratively filtered until the convergence condition is met (such as the number of iterations of the joint bilateral filtering process reaches a set value, or the image change is less than a threshold).

[0053] Gaussian smoothing is to use the following formula (1) to smooth multiple target visible light images in order to remove small-scale details (such as noise and fine textures) in the target visible light images and retain large-scale structures (such as smoke and object outlines) to obtain a smooth image.

[0054] (1) Where J0(p) is the pixel value of pixel p in the smoothed image obtained after Gaussian kernel smoothing, and K p Here, σ is the normalization coefficient, N(p) is the neighborhood of pixel p, exp() is the natural exponential function, |pq| is the spatial distance between pixel p and pixel q, and σ is the normalization coefficient. s I(q) is the standard deviation of the Gaussian kernel in the spatial domain to control the degree of smoothing, and I(q) is the gray value of pixel q in the visible light image of the target.

[0055] The joint bilateral filtering process uses the following formula (2) to filter the smoothed image, so as to remove the smoke while retaining the edge information of the image. It also considers the spatial distance and pixel value difference, and avoids the edge being blurred by calculating the joint weight of spatial and pixel intensity.

[0056] (2) Among them, J (t+1) (p) is the pixel value of pixel p after the (t+1)th iteration, J (t) (p) is the pixel value of pixel p after the t-th iteration, J (t) (q) is the pixel value of pixel q after the t-th iteration, σr It is the standard deviation of the Gaussian kernel in the pixel value range to control the impact of pixel value differences on the weights.

[0057] The embodiments of the present invention, through Gaussian kernel smoothing and joint bilateral filtering, combined with the aforementioned Gaussian pyramid, allow the target visible light image to be iteratively processed at different scales, thereby effectively separating smoke and background in the target visible light image.

[0058] Because visible light cameras and infrared thermal imaging modules differ in their imaging principles, focal lengths, field of view, and shooting times, the position of the same object in the two images may deviate. This causes feature points in the visible light image to not correspond to feature points in the infrared image, resulting in inaccurate target positions in the subsequent fused image. To address this problem, in some embodiments of the present invention, after S100 and before S200, the following steps are included: S150. Spatially align the original visible light image and the original infrared image, and then temporally align them.

[0059] Specifically, based on time and space alignment mechanisms, it is ensured that the original visible light image and the original infrared image of each set of dual-light images sent to the airborne processor or electronic equipment are time-aligned and spatially aligned.

[0060] Spatial alignment involves precisely matching the original visible light image and the original infrared image in spatial location, ensuring that the same object in both images is completely identical in pixel position. Spatial alignment can be based on feature-based image registration. This involves using feature detection algorithms (such as SIFT, SURF, and ORB) to extract feature points from both the visible light and infrared images, and using feature matching algorithms (such as FLANN and BFMatcher) to find corresponding feature point pairs in the two images. Based on the matched feature point pairs, a spatial transformation matrix (such as affine transformation or perspective transformation) is calculated to transform the original infrared image (or original visible light image) to align it with the other image. Alternatively, spatial alignment can be based on region-based image registration. This involves selecting corresponding regions (such as building edges or windows) in the original visible light and infrared images, using cross-correlation or mutual information to measure region similarity, and then using optimization algorithms (such as gradient descent) to find the optimal alignment position.

[0061] The process involves synchronizing the acquisition time of the original visible light and infrared images to ensure that both images capture the same scene at the same moment. Time alignment can be based on hardware synchronization, using a visible light camera and infrared thermal imaging module that support synchronized triggering. This can be achieved by simultaneously triggering both the camera and module via the drone's remote control or electronic devices, ensuring that the timestamps of the two images are consistent. Alternatively, time alignment can be based on software synchronization, where the acquisition timestamps of the original visible light and infrared images are recorded separately, and the two images with the closest timestamps are selected for alignment.

[0062] In this embodiment of the invention, the original visible light image and the original infrared image are first spatially aligned to ensure that the same object is in the same position in the two images, thereby improving the quality of the fused image and the accuracy of target detection. Then, they are temporally aligned to ensure that the two images were captured at the same time, thereby improving the realism of the fused image and the accuracy of target detection. In this way, spatial alignment and temporal alignment together ensure the accuracy, realism and reliability of the subsequently generated fused image, providing high-quality data support for subsequent target detection and rescue decisions.

[0063] Images caused by fire smoke are typically low in brightness and lack contrast, and even after smoke removal processing, the visual effect may still be poor. In some embodiments of the present invention, S224 is followed by: S225. Convert the smoke-removed image to the HSV color space and perform histogram equalization.

[0064] Specifically, histogram equalization is a non-linear transformation method used to improve image contrast. It enhances detail by adjusting the grayscale distribution of an image, making pixel values ​​more evenly distributed. Its calculation formula is as follows:

[0065] Among them, S k The mapped pixel values ​​are L, where L is the pixel value range (usually 256), and p r (r j ) represents the probability density of the gray level, and k represents the current gray level.

[0066] This invention converts the smoke-removed image to the HSV color space, enhances the V (luminance) channel, and uses histogram equalization to optimize the luminance distribution for enhanced brightness, making the smoke-removed image more natural. Alternatively, CLAHE (contrast-limited adaptive histogram equalization) can be provided to avoid localized overexposure problems that may occur with traditional histogram equalization.

[0067] For example, such as Figure 4As shown, H×W×C represents the image width, height, and number of channels. The original 1920×1080×3 visible light image is downsampled three times. Specifically, the first downsampling of the original 1920×1080×3 visible light image yields a 960×540×3 target visible light image, which is then subjected to multiple iterative filtering processes to obtain a 960×540×3 smoke component patch. The second downsampling of the 960×540×3 target visible light image yields a 480×280×3 target visible light image, which is then subjected to multiple iterative filtering processes to obtain a 480×280×3 smoke component patch. The third downsampling of the 480×280×3 target visible light image yields a 240×135×3 target visible light image, which is then subjected to multiple iterative filtering processes to obtain a 240×135×3 smoke component patch. Then, the smoke component images are upsampled to obtain the target component images. A weighted average of all target component images is then calculated to obtain the final smoke image. The original visible light image is subtracted from the final smoke image to obtain the smoke-free image. Histogram equalization is then applied to the smoke-free image to obtain the final smoke-free image.

[0068] In some embodiments of the present invention, such as Figure 5 As shown, S300 includes: S310. Based on the dual-branch encoder, extract the first feature map of the smoke-removed image and the second feature map of the original infrared image respectively; S320. Obtain initial fused features based on the first feature map and the second feature map; S330. Input the first feature map, the second feature map and the initial fused feature into the cross-scale iterative attention decoder to generate intermediate fused features; S340. Based on the cross-scale iterative attention decoder, the intermediate fusion features are transformed according to the activation function to generate the fused image; S350. Evaluate the probability values ​​of whether the fused image is similar to the original infrared image and the smoke-removed image based on two discriminators respectively; S360. Optimize the parameters of the cross-scale iterative attention decoder and the discriminator until the probability value reaches a preset threshold, and obtain the final fused image based on the optimized cross-scale iterative attention decoder.

[0069] Specifically, such as Figure 6As shown, the CrossFuse dual-light fusion algorithm is a highly efficient image fusion technique that aims to generate a fused image that combines thermal radiation features with visual details by fusing a smoke-removed image and the original infrared image. This algorithm is particularly suitable for complex environments, such as dense smoke and low-light scenes, and can significantly improve the accuracy and robustness of image recognition. The dual-light fusion algorithm uses the CrossFuse model as its core. A dual-branch encoder extracts multi-scale first and second feature maps, and then generates initial fusion features based on the first and second feature maps. A cross-scale iterative attention decoder combined with a cross-modal attention mechanism and activation function outputs the fused image. The discriminator optimizes and adjusts the image to obtain a high-quality fused image that combines visible light details and infrared thermal information.

[0070] like Figure 6 As shown, each branch of the dual-branch encoder consists of four multi-scale convolutional blocks (MCB1, MCB2, MCB3, and MCB4) for extracting scale features. The first branch encoder 11 extracts the second feature map of the original infrared image, denoted as... The second branch encoder 12 extracts the first feature map of the smoke-free image, denoted as... Where l = 1, 2, 3, 4. The first branch encoder 11 and the second branch encoder 12 typically consist of multiple convolutional layers and pooling layers, outputting first and second feature maps at multiple scales, respectively. Each MCB contains two convolutional layers with a kernel size of 3×3 and strides of 1 and 2, respectively. Simultaneously, their filter banks are set to 16×L and 16×L, respectively, to adapt to different scales. Then, the features extracted from the two branches are initially fused as input to the subsequent attention mechanism. This initial fusion involves concatenating or weighting the first and second feature maps at each scale to obtain the initial fused features.

[0071] like Figure 6As shown, the cross-scale iterative attention decoder 20 is used for feature reconstruction, in which four cross-modal attention ensemble modules (CAIMs) with upsampling operations are designed to connect cross-scale features. The cross-modal attention ensemble modules (CAIMs) include channel-independent paths and spatial-independent paths. The channel attention mechanism is the core component of the channel-independent path. It analyzes the importance of each feature channel (i.e., which channels contain more useful information), extracts statistical information of the channel dimension through global pooling (such as average pooling and max pooling), and then generates channel attention weights through convolutional layers and activation functions. The spatial attention mechanism is the core component of the spatial-independent path. It analyzes the importance of each spatial location in the image (i.e., which regions in the image are more worthy of attention), extracts statistical information of the spatial dimension through spatial pooling (such as average pooling and max pooling), and then generates spatial attention weights through convolutional layers and activation functions. First feature maps at different scales are shown. Second feature map and its initial fusion characteristics Input is fed into CAIM at the same level to generate intermediate fusion features. This serves as the input source for the next CAIM, thus transforming the intermediate fusion features into a fused image based on the activation function.

[0072] In the channel-independent path, firstly, max pooling and average pooling operations are used to convert the initial fused features into initial channel attention vectors. Then, these vectors pass through two convolutional layers and a PReLU activation layer, are concatenated, and fed into a convolutional layer to generate channel attention vectors. Its expression is shown in formula (3).

[0073] (3) Where Conv and ρ represent convolution and PReLU activation operations, respectively, and MP(·) and AP(·) represent global max pooling and average pooling operations, respectively.

[0074] Similarly, in the spatially independent path, max pooling and average pooling operations are used to obtain the initial spatial attention matrix, and then they are concatenated and fed into a convolutional layer to generate the spatial attention matrix l, saF∈R1×H×W, as shown in Equation (4).

[0075] (4) Subsequently, the channel attention vectors are multiplied element-wise by the spatial attention matrix to obtain the initial fused feature attention maps. Then, the sigmoid activation function is used to normalize these attention maps to generate the corresponding attention weights, the expression of which is shown in Equation (5).

[0076] (5) in, This represents the Sigmoid activation function.

[0077] Finally, the attention weights (Right now Assign it to the infrared path, and Assigned to the visible light path. Intermediate fusion features. It can be calculated using formula (6).

[0078] (6) Subsequently, the cross-scale iterative attention decoder 20, based on the intermediate fused features... Using the Sigmoid or Tanh activation function, the output value is restricted to the range of [0, 1] or [-1, 1], which is used as the fused pixel value of the fused image. The fused image is then generated based on the fused pixel value.

[0079] The two discriminators (D_ir and D_vis) share the same network framework, consisting of four convolutional layers. The kernel size and stride are 3×3 and 2, respectively, with output channels of 16, 32, 64, and 128, respectively. The first three layers use LeakyReLU activation, and the last layer uses the Tanh function. During training, the discriminators are trained using real images (the original infrared image and the smoke-removed image) and the generated fused image. During training, D_ir is used to distinguish between the original infrared image and the fused image, while D_vis is used to distinguish between the fused image and the smoke-removed image. Two discriminators evaluate the similarity between the fused image and the infrared and visible light images, respectively. The discriminator loss and the generator (i.e., the cross-scale iterative attention decoder) loss are calculated. The parameters of the discriminator and the generator are alternately optimized and the training is repeated until the probability value output by the discriminator reaches a preset threshold. In this way, the generator can generate a fused image that approximates the real image. At this point, the parameters of the cross-scale iterative attention decoder 20 are fixed, and the original infrared image and the smoke-removed image are input into the cross-scale iterative attention decoder 20 with fixed parameters to generate the final fused image.

[0080] In the proposed CrossFuse model, the total loss function of the cross-scale iterative attention decoder 20 consists of two parts: content loss and adversarial loss, and its expression is shown in Equation (7).

[0081] (7) Where LG represents the total loss function, Lcon and Ladv represent the content loss function and adversarial loss function, respectively, and the parameter λ1 is used to control the balance between them.

[0082] For the content loss function, an intensity loss was first designed to constrain the pixel intensity similarity between the fused result and the source image. L1 norm and L2 norm were used for infrared and visible light images, respectively. Therefore, the formula for the intensity loss is shown in Equation (8).

[0083] (8) Where If, ​​Iir, and Ivis represent the fused image, the original infrared image, and the smoke-removed image, respectively. and Let L1 and L2 represent the L1 norm and L2 norm respectively, and λ2 be a weighting coefficient.

[0084] Next, a texture loss is introduced to assist the intensity loss, forcing the fused image to retain more texture details. The formula for the texture loss is shown in Equation (9).

[0085] (9) Here, ∇ represents the gradient operator.

[0086] Finally, the content loss function is a weighted combination of intensity loss and texture loss, and its expression is shown in Equation (10).

[0087] (10) Here, λ3 is a weighting coefficient.

[0088] In the discriminator, two discriminators (D_ir and D_vis) are used to distinguish the data distribution of the fused image If from the visible light image Ivis and the infrared image Iir. Meanwhile, the loss function of the adversarial network is shown in Equation (11).

[0089] (11) Furthermore, the loss functions of the two discriminators are shown in Equation (12) and Equation (13), respectively.

[0090] (12) (13) Here, the first and second terms represent the Wasserstein distance estimate and gradient penalty, respectively, and λ4 is the regularization parameter.

[0091] In this embodiment, the input images are the smoke-removed image and the original infrared image. The smoke-removed image and the original infrared image are fused together by the dual-branch encoder and cross-scale iterative attention decoder 20 in the CrossFuse network to generate a new fused image. The objects in the fused image are clearer, the details are more complete, the brightness levels are richer, the infrared target and background details are highlighted, and unnecessary noise is reduced. Moreover, the color and grayscale information are more natural, avoiding the over-enhancement of single channel information.

[0092] In some embodiments of the present invention, S400 includes: S410. Add a hybrid attention mechanism residual network to the YOLOv11 network model to construct the improved object detection model; the hybrid attention mechanism residual network includes a first residual block and a second residual block, both of which include a first network branch. The first network branch includes a first convolutional layer, a depthwise separable convolutional layer, an attention mechanism network, and a second convolutional layer. The spatial dimensions of the first and second convolutional layers are 1×1, and the spatial dimension of the depthwise separable convolutional layer is 3×3. The output of the first convolutional layer, the input of the depthwise separable convolutional layer, and the input of the second convolutional layer are connected sequentially. The attention mechanism network... The input and output terminals are respectively connected to the output terminal of the depth-separable convolutional layer and the input terminal of the second convolutional layer. The first residual block also includes a first weight layer and a second network branch. The output terminal of the second convolutional layer of the first residual block is connected to the input terminal of the first weight layer. The second network branch includes a residual connection layer and a second weight layer. The input terminal of the residual connection layer is connected to the input terminal of the first convolutional layer. The output terminal of the residual connection layer is connected to the input terminal of the second weight layer. The output terminal of the first weight layer is connected to the output terminal of the second weight layer. The first weight layer and the second weight layer respectively set the weight ratio of the output of the first network branch and the second network branch. S420. Input the fused image into the improved target detection model to output the person recognition result; Wherein, the first network branch of the first residual block corresponds to the first weight value, the second network branch corresponds to the second weight value, and the sum of the first weight value and the second weight value is equal to 1.

[0093] Specifically, the structure of a hybrid attention mechanism residual network is as follows: Figure 7 As shown, while increasing the number of intermediate channels and introducing a Shuffle Attention mechanism (i.e., an attention mechanism network) on top of the YOLOv11 network model improves feature representation capabilities, it also significantly increases network complexity, leading to slower convergence and larger model size. This invention will... Figure 7 The hybrid attention mechanism residual network shown is added to the YOLOv11 network model to construct the improved object detection model. This is equivalent to introducing a learnable gating mechanism to enhance the network's flexibility and expressiveness, and accelerate its convergence. The gating parameters of the hybrid attention mechanism residual network are learnable, and these parameters can dynamically adjust the weights of each path, thereby effectively controlling the information flow between the attention-enhanced features and the residual connection layers. This allows the model to focus more on highly relevant information while reducing interference from unnecessary paths, thus improving the model's efficiency and performance.

[0094] This invention introduces a hybrid attention mechanism residual network into the YOLOv11 network model, comprising two key residual blocks (a first residual block and a second residual block). Each key residual block includes a first network branch, which consists of a first convolutional layer → a depthwise separable convolutional layer → an attention mechanism network → a second convolutional layer. The first residual block also includes a first weight layer and a second network branch connected to the output of the second convolutional layer. The second network branch consists of a residual connection layer → a second weight layer. The first and second weight layers control the output weight ratio of the two branches, respectively, and their weights sum to 1. Since the attention mechanism network is embedded after the depthwise separable convolutional layer, it enhances feature extraction capabilities, focusing on key regions (such as trapped individuals).

[0095] While increasing the number of intermediate channels and introducing an attention mechanism in the backbone structure of the YOLOv11 network model improves feature representation, it also significantly increases network complexity, leading to slower convergence and larger model size. Therefore, the residual block was optimized by introducing a learnable gating mechanism to enhance the network's flexibility and expressiveness, and accelerate convergence. Specifically, learnable gating parameters were introduced into the first residual block. This involves dynamically adjusting the weights of each path through the first and second weight layers, effectively controlling the information flow between the attention mechanism and the residual connection layers. This allows the model to focus more on highly relevant information while reducing interference from unnecessary paths, thereby improving computational efficiency, robustness, and accuracy. It also enables the model to adaptively balance the importance of different branches, enhancing its adaptability to complex environments. Through dynamic adjustment of the weight layers, the model can adaptively focus on more important features, improving feature discriminative power. Furthermore, the residual connection layers, first and second weight layers enable flexible feature fusion, avoiding the vanishing gradient problem and improving model stability and convergence speed. In this way, by inputting a fused image (composed of a visible light image after smoke removal and the original infrared image) into the improved target detection model, and outputting the detection results, the improved target detection model can not only ensure the safety of rescue personnel, but also effectively improve rescue efficiency and the survival probability of trapped personnel.

[0096] It should be noted that, in addition to the location information of the trapped personnel, the detection results can also include core data such as the number of trapped personnel, detection time, and confidence level, to help rescuers dynamically grasp the detection progress. After the mission is completed, a statistical report containing indicators such as the number of targets and the detection time will be generated, providing a quantitative basis for the evaluation of rescue effectiveness and algorithm optimization.

[0097] This invention relates to a drone-based precision fire rescue system, primarily composed of drones and electronic equipment. Through close collaboration between modules, it achieves efficient reconnaissance of high-rise building fire scenes, precise location of trapped personnel, and crucial support for rescue decision-making, significantly improving the efficiency and safety of high-rise building fire rescue. The system utilizes drones equipped with advanced image acquisition devices to obtain visible light and infrared image data of the fire scene, transmitting it in real-time to terminal software. Unique image processing algorithms are used for smoke removal enhancement, dual-light fusion, and target detection, providing rescuers with a clear view of the trapped personnel's location. Unlike traditional rescue methods, this system leverages the flexibility and maneuverability of drones, overcoming the limitations of traditional ground rescue equipment in terms of operating height and low dispatch efficiency, greatly expanding the rescue coverage and response speed. In image processing, the system employs Rolling Guided Filter (RGF) combined with multi-scale image pyramid technology and the CrossFuse dual-light fusion algorithm to effectively remove dense smoke interference, generating high-quality fused images and providing a clear and accurate data foundation for target detection. From a cost and efficiency perspective, this system replaces manual reconnaissance at high-risk fire scenes with drones, reducing the risk of injury to rescue personnel and lowering labor costs. Moreover, drones can quickly reach the fire scene and rapidly conduct reconnaissance work, significantly shortening the rescue response time and improving the efficiency of single fire response compared to traditional rescue methods. In practical applications, with the acceleration of urbanization and the increasing number of high-rise buildings, this system shows broad application prospects in the field of urban smart fire protection and public safety management, and is expected to generate good economic benefits. In addition, the core technology of the system has good scalability, not only applicable to high-rise building fire rescue, but also applicable to various high-risk emergency scenarios such as earthquake search and rescue and chemical leak monitoring. Through coordination with other emergency rescue systems, the intelligence level of the entire emergency rescue system can be further improved. For example, in earthquake search and rescue, the mobility and image recognition capabilities of drones can be used to quickly search for people trapped in rubble; in chemical leak monitoring, by equipping special sensors and combining them with the drone-based precise fire rescue method of this invention, the precise location of the leak source and rapid assessment of the danger zone can be achieved. In summary, this invention has many advantages, such as adapting to complex fire environments, accurately and efficiently locating trapped personnel, reducing rescue costs and risks, and expanding application scenarios. Through intelligent technology, it achieves computer adaptive processing and analysis, which has extremely high feasibility and is of great significance and value in improving emergency rescue capabilities and protecting people's lives and property.

[0098] To better implement the drone-based precise fire rescue method in the embodiments of the present invention, the present invention also provides a drone-based precise fire rescue system, including a drone and a display device connected to the drone. The drone is used to capture dual-light images of the fire area, the dual-light images including a raw visible light image and a raw infrared image; the raw visible light image is processed according to a downsampling algorithm and a rolling guided filtering algorithm to obtain a smoke-free image; a fused image is generated based on the smoke-free image and the raw infrared image; and based on an improved target detection model, trapped personnel and their locations are detected in the fire area according to the fused image. The display device is used to display the test results.

[0099] Specifically, the drone hardware includes an image acquisition module, an onboard processor, a communication module, and a device connection and power supply module. The image acquisition module is used to acquire visible light and infrared images of the fire scene. Equipped with a high-performance visible light camera and infrared thermal imaging module, it can clearly capture image information of the fire scene in complex environments. The onboard processor has the ability to quickly process image data, reducing the burden on the flight control system and is suitable for lightweight AI inference, computer vision, and embedded control scenarios. The communication module enables stable communication between the terminal software and the drone, supports real-time data transmission, and ensures that rescue personnel can obtain information from the fire scene in a timely manner. The device connection and power supply module ensures the integrity of the system and its long-term operating capability, providing continuous support for fire rescue missions. Through the collaborative work of these modules, this invention can provide strong technical support for fire rescue, improving rescue efficiency and safety.

[0100] In some optional implementations, visible light cameras are used for image acquisition at high-rise fire scenes, clearly capturing details such as the location and posture of trapped personnel. Infrared cameras are also included to ensure image quality even at night or in low-light conditions caused by dense smoke at the fire scene. Visible light cameras are lightweight and compact, and their integration into drone platforms does not affect flight stability. They also support the Raspberry Pi 3B+ / 4B CSI interface, allowing for direct image acquisition and processing using Python or OpenCV, demonstrating good adaptability. Simultaneously, an infrared thermal imaging module with an optional long-wave infrared sensor is selected to meet the real-time imaging needs of the drone during flight, avoiding image ghosting. Image acquisition facilitates real-time monitoring of the fire scene, enabling subsequent software data processing and providing strong support for fire rescue decision-making.

[0101] In some optional implementations, the Raspberry Pi 4B, a compact embedded computing platform, is selected as the external onboard processor. It features a quad-core Arm Cortex-A76 CPU (2.4GHz) and a VideoCore VII GPU, supporting OpenGL ES 3.1 and Vulkan 1.2, providing some graphics processing capabilities. It also comes equipped with 4GB / 8GB LPDDR4X memory, sufficient for moderate-load real-time data processing. Measuring 85.6mm × 56.5mm, it uses a standard 40-pin GPIO interface for easy integration and expansion. The Raspberry Pi 4B is suitable for lightweight AI inference, computer vision, and embedded control scenarios.

[0102] In some optional implementations, the communication module employs the Lightbridge image transmission module, which possesses several key characteristics: support for 1080p full HD resolution, meeting the needs for accurate identification of trapped personnel and on-site conditions in fire scenarios; transmission latency controlled within 200ms, ensuring real-time synchronization between the acquired images and the drone's flight status, providing timely information for rescue decisions; an effective transmission distance of up to 5 kilometers in open environments, capable of comprehensively covering the fire scene; adaptive frequency hopping technology to address multipath effects and electromagnetic interference, possessing strong anti-interference performance and reducing the risk of image interruption; a lightweight and low-power design to reduce the load on the drone, extend the single acquisition operation time, and improve rescue efficiency; and a compact, integrated package structure, compatible with mainstream industry drones, ensuring stable operation of the drone during flight. Through the collaborative work of these modules, this invention can provide strong technical support for fire rescue, improving rescue efficiency and safety.

[0103] In some optional implementations, the device connectivity and power supply modules must ensure the integrity of the drone equipment and enable long-duration flight. Through the coordinated operation of these modules, the safe and stable operation of the drone is guaranteed, providing strong technical support for fire rescue and improving rescue efficiency and safety.

[0104] The UAV-based fire precision rescue system provided in the above embodiments can realize the technical solutions described in the above UAV-based fire precision rescue method embodiments. The specific implementation principles of each module or unit can be found in the corresponding content in the above UAV-based fire precision rescue method embodiments, which will not be repeated here.

[0105] like Figure 8 As shown, the present invention also provides an electronic device 800. The electronic device 800 includes a processor 801, a memory 802, and a display 803. Figure 8Only some components of the electronic device 800 are shown, but it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented instead.

[0106] In some embodiments, processor 801 may be a central processing unit (CPU), microprocessor, or other data processing chip, used to run program code stored in memory 802 or process data, such as the drone-based precision fire rescue method of the present invention.

[0107] In some embodiments, processor 801 may be a single server or a group of servers. The server group may be centralized or distributed. In some embodiments, processor 801 may be local or remote. In some embodiments, processor 801 may be implemented on a cloud platform. In one embodiment, the cloud platform may include a private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, intranet, multi-cloud, etc., or any combination thereof.

[0108] In some embodiments, memory 802 may be an internal storage unit of electronic device 800, such as a hard disk or memory of electronic device 800. In other embodiments, memory 802 may also be an external storage device of electronic device 800, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on electronic device 800.

[0109] Furthermore, the memory 802 may include both internal storage units of the electronic device 800 and external storage devices. The memory 802 is used to store application software and various types of data installed on the electronic device 800.

[0110] In some embodiments, display 803 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. Display 803 is used to display information from electronic device 800 and to display a visual user interface. Components 801-803 of electronic device 800 communicate with each other via a system bus.

[0111] In one embodiment, when the processor executes the personnel location detection program in memory 802, a method for precise fire rescue based on unmanned aerial vehicles (UAVs) can be implemented, including: Acquire dual-light images of the fire area captured by the imaging components on the drone, the dual-light images including a raw visible light image and a raw infrared image; The original visible light image is processed using a downsampling algorithm and a rolling guided filtering algorithm to obtain a smoke-free image. A fused image is generated based on the smoke-removed image and the original infrared image; Based on the improved target detection model, the trapped personnel and their locations in the fire area are detected according to the fused image.

[0112] It should be understood that when the processor 801 executes the personnel positioning and detection program in the memory 802, in addition to the functions mentioned above, it can also perform other functions, as can be found in the description of the corresponding method embodiments above.

[0113] Furthermore, this embodiment of the invention does not specifically limit the type of electronic device 800 mentioned. Electronic device 800 can be a mobile phone, tablet computer, personal digital assistant (PDA), wearable device, laptop computer, or other portable electronic device. Exemplary embodiments of portable electronic devices include, but are not limited to, portable electronic devices running iOS, Android, Microsoft, or other operating systems. The aforementioned portable electronic device can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the invention, electronic device 800 may not be a portable electronic device, but rather a desktop computer with a touch-sensitive surface (e.g., a touch panel).

[0114] Accordingly, this application also provides a computer-readable storage medium for storing computer-readable programs or instructions. When the programs or instructions are executed by a processor, they can implement the steps or functions of the drone-based fire precision rescue method provided in the above-described method embodiments.

[0115] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.), and the computer program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.

[0116] The above provides a detailed description of the method, system, equipment, and medium for precise fire rescue based on unmanned aerial vehicles (UAVs) provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A method for precise fire rescue based on unmanned aerial vehicles (UAVs), characterized in that, include: Acquire dual-light images of the fire area captured by the imaging components on the drone, the dual-light images including a raw visible light image and a raw infrared image; The original visible light image is processed using a downsampling algorithm and a rolling guided filtering algorithm to obtain a smoke-free image. A fused image is generated based on the smoke-removed image and the original infrared image; Based on the improved target detection model, the trapped personnel and their locations in the fire area are detected according to the fused image.

2. The method for precise fire rescue based on unmanned aerial vehicles (UAVs) according to claim 1, characterized in that, The step of processing the original visible light image to obtain the smoke-free image according to the downsampling algorithm and the rolling guided filtering algorithm includes: Multiple resolution target visible light images are obtained by downsampling the original visible light image based on the Gaussian pyramid. The target visible light image is processed based on the rolling guided filtering algorithm, and the smoke-removed image is obtained based on the processed image and the original target visible light image.

3. The method for precise fire rescue based on unmanned aerial vehicles (UAVs) according to claim 2, characterized in that, The process of processing the target visible light image based on the rolling guided filtering algorithm, and obtaining the smoke-removed image based on the processed image and the original target visible light image, includes: At each resolution level, the target visible light image is subjected to multiple iterative filtering processes based on the rolling guided filtering algorithm to obtain the corresponding smoke component patch; The target component image is obtained by upsampling the smoke component patch, and the resolution of the target component image is equal to the resolution of the original visible light image; Based on the weights of the smoke component images at different resolutions, a weighted average of all the target component images is calculated to obtain the final smoke image; The smoke-free image is obtained based on the original visible light image and the final smoke image.

4. The method for precise fire rescue based on unmanned aerial vehicles (UAVs) according to claim 3, characterized in that, At each resolution level, the target visible light image is iteratively filtered multiple times based on a rolling guided filtering algorithm to obtain the corresponding smoke component patch, including: Gaussian kernel smoothing is applied to multiple visible light images of the target at the specified resolutions to obtain a smoothed image; After performing multiple iterations of the smoothed image based on joint bilateral filtering, the corresponding smoke component patch is obtained by comparing the final filtered image after the last iteration with the original visible light image.

5. The method for precise fire rescue based on unmanned aerial vehicles (UAVs) according to claim 1, characterized in that, After acquiring the dual-light image of the fire area captured by the imaging component on the drone, and before processing the original visible light image according to the downsampling algorithm and the rolling guided filtering algorithm to obtain the smoke-removed image, the process includes: The original visible light image and the original infrared image are spatially aligned and then temporally aligned.

6. The method for precise fire rescue based on unmanned aerial vehicles (UAVs) according to claim 5, characterized in that, The step of generating a fused image based on the smoke-removed image and the original infrared image includes: The first feature map of the smoke-removed image and the second feature map of the original infrared image are extracted based on the dual-branch encoder. Initial fused features are obtained based on the first feature map and the second feature map; The first feature map, the second feature map, and the initial fused feature are input into the cross-scale iterative attention decoder to generate intermediate fused features; Based on the cross-scale iterative attention decoder, the intermediate fusion features are transformed according to the activation function to generate the fused image; The probability values ​​of whether the fused image is similar to the original infrared image and the smoke-removed image are evaluated based on two discriminators respectively. The parameters of the cross-scale iterative attention decoder and the discriminator are optimized until the probability value reaches a preset threshold. The final fused image is then obtained based on the optimized cross-scale iterative attention decoder.

7. The method for precise fire rescue based on unmanned aerial vehicles (UAVs) according to claim 1, characterized in that, The method of detecting trapped personnel and their locations in the fire area based on the improved target detection model and the fused image includes: A hybrid attention mechanism residual network is added to the YOLOv11 network model to construct the improved object detection model. The hybrid attention mechanism residual network includes a first residual block and a second residual block, both of which include a first network branch. Each first network branch includes a first convolutional layer, a depthwise separable convolutional layer, an attention mechanism network, and a second convolutional layer. The spatial dimensions of the first and second convolutional layers are 1×1, and the spatial dimension of the depthwise separable convolutional layer is 3×3. The output of the first convolutional layer, the input of the depthwise separable convolutional layer, and the input of the second convolutional layer are connected sequentially. The input of the attention mechanism network is... The first residual block is connected to the output of the depth-separable convolutional layer and the input of the second convolutional layer, respectively. The first residual block also includes a first weight layer and a second network branch. The output of the second convolutional layer of the first residual block is connected to the input of the first weight layer. The second network branch includes a residual connection layer and a second weight layer. The input of the residual connection layer is connected to the input of the first convolutional layer. The output of the residual connection layer is connected to the input of the second weight layer. The output of the first weight layer is connected to the output of the second weight layer. The first weight layer and the second weight layer respectively set the weight ratio of the output of the first network branch and the second network branch. The fused image is input into the improved target detection model to output person recognition results; Wherein, the first network branch of the first residual block corresponds to the first weight value, the second network branch corresponds to the second weight value, and the sum of the first weight value and the second weight value is equal to 1.

8. A precision fire rescue system based on unmanned aerial vehicles (UAVs), comprising a UAV and a display device connected to the UAV, characterized in that, The drone is used to capture dual-light images of the fire area, the dual-light images including a raw visible light image and a raw infrared image; the raw visible light image is processed according to a downsampling algorithm and a rolling guided filtering algorithm to obtain a smoke-free image; a fused image is generated based on the smoke-free image and the raw infrared image; and based on an improved target detection model, trapped personnel and their locations are detected in the fire area according to the fused image. The display device is used to display the test results.

9. An electronic device, characterized in that, Including memory and processor, among which, The memory is used to store programs; The processor, coupled to the memory, is used to execute the program stored in the memory to implement the steps in the method for precise fire rescue based on unmanned aerial vehicles as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store computer-readable programs or instructions, which, when executed by a processor, can implement the steps in the drone-based precision fire rescue method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Image recognition method, electronic equipment, storage medium, program product and chip

    CN121330465A