Flame detection method and equipment based on flame specificity and multi-exposure image fusion

By constructing a flame detection model based on flame specificity and multi-exposure image fusion, and using the P-DenseBlock encoder and DenseNet decoder for feature extraction and fusion, the problems of high misjudgment rate and high cost of flame detection in existing technologies are solved, and high-precision flame recognition is achieved under extreme exposure conditions.

CN120656042AActive Publication Date: 2025-09-16UNIV OF SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511120766.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-16
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing flame detection technology has a high misjudgment rate under extreme exposure conditions of visible light images, and infrared image datasets are scarce, resulting in high detection accuracy and cost. Existing image fusion methods cannot effectively overcome interference sources and extreme exposure problems.

Method used

A flame recognition model is constructed based on flame specificity and multi-exposure image fusion. The image fusion module EMEIF and the flame detection network are used. The P-DenseBlock encoder and the DenseNet decoder are used for feature extraction and fusion. A three-stage training strategy is designed, including the structural similarity loss between images and the bounding box regression loss, to achieve adaptive fusion of multi-exposure images and flame recognition.

Benefits of technology

Effectively integrate the features of images with different exposures, reduce the impact of extreme exposure on detection accuracy, improve the accuracy and generalization ability of flame detection, and reduce detection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656042A_ABST
    Figure CN120656042A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image recognition, in particular to a flame detection method and device based on flame specificity and multi-exposure image fusion. A flame recognition model constructed by the method comprises an image fusion module EMEIF and a flame detection network. The image fusion module generates a fusion image according to the multiple continuous multi-exposure images, and the flame detection network generates a flame recognition result according to the fusion image. The EMEIF module is composed of an encoder, a fusion unit and a decoder; the encoder performs feature extraction on the input image; the fusion unit performs weighted fusion on the extracted feature information by using a self-adaptive fusion weight; the fusion features are processed by a decoder to obtain a fusion image; for a newly designed network model, a three-stage training strategy is introduced, so that the model can be transited from reference shot flame data feature learning to flame data feature learning after multi-exposure fusion. According to the invention, the defects of the existing scheme in the aspects of detection precision and cost are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and in particular to a flame detection method based on flame specificity and multi-exposure image fusion, and corresponding computer program products, storage media and computer equipment. Background Art

[0002] Traditional solutions mostly use temperature sensors, smoke sensors, etc. to achieve fire detection and early warning. However, such solutions not only have a high misjudgment rate but also cannot guarantee real-time performance. To solve this problem, a series of computer vision-based methods have been proposed and applied to fire detection. Among them, the YOLO series is widely used because of its lightweight and fast response speed. Its main method is to put the labeled flame dataset into the model, and the model learns the characteristics of the flame through layers of networks to achieve flame detection and recognition. However, although such solutions perform well in experiments, they perform less well in real-world detection. The reasons for this are as follows: (1) The flame image dataset is relatively scarce. Currently, most fire experiments use only a few thousand datasets, which lack sample data covering various interference sources (such as car lights, glass reflections, etc.). In real-world use, the images captured by the camera contain many interference sources, which makes it easy for the model to make misjudgments. (2) The sample images in actual detection are easily affected by environmental factors, resulting in overexposure and underexposure. The existing network model has poor recognition ability for flame targets in these two extreme low-quality images.

[0003] Given the high cost of expanding flame datasets, researchers have proposed a series of methods to improve network detection performance based on small sample sizes. Among them, Liu Jiali et al. proposed flame detection using multi-information fusion in the Ohta color space. This method analyzes the saturation of the red component in the flame color space, assigning a threshold RT to the red component to ensure its saturation is greater than RT. A threshold is also set on the flame saturation to mitigate the effects of background lighting. This method improves detection accuracy to a certain extent. Tang Linfeng et al. introduced illumination factors into the model, allowing it to learn the difference between day and night, and controlled the soft selection of visible light and infrared by adjusting weights. However, this approach still did not significantly improve when using flame images with extreme exposure conditions. While numerous improved methods have indeed improved flame detection accuracy, few have effectively addressed extreme exposure conditions. Furthermore, most improved methods still fail to effectively overcome the problem of interference sources in visible light flame image detection.

[0004] For complex object detection tasks, researchers have proposed image fusion detection. Fusion detection can be broadly categorized into three types: early-stage fusion, mid-stage fusion, and late-stage fusion. Mid-stage and late-stage fusion achieves the most significant results. However, mid-stage and late-stage fusion often requires a dual-branch or even multi-branch backbone network, resulting in models sacrificing time for accuracy. Early-stage fusion, on the other hand, fuses visible light and infrared images before the image is input into the backbone network. This type of fusion hinders the interaction between visible and infrared image features, resulting in inferior results compared to mid-stage and late-stage fusion. Inspired by the use of fusion detection for pedestrian images, flame detection has also seen a rise in the number of methods that fuse visible and infrared images. These methods have attracted widespread attention due to the high saliency of flames in infrared images and their ability to eliminate certain interference sources. However, flame datasets are scarce, and corresponding infrared datasets are even more scarce. Even with datasets that consume significant resources and capture both visible and infrared images simultaneously, data annotation, processing, and alignment present challenges. Furthermore, the model is expensive to use and requires cameras capable of capturing both visible and infrared images. Even if the resulting detection results are significant, this hinders widespread adoption. Summary of the Invention

[0005] In order to solve the problems of high accuracy and cost in the flame recognition technology based on infrared and visible light image fusion in the existing technology, the present invention provides a flame detection method based on flame specificity and multi-exposure image fusion, and its corresponding computer program product, storage medium and computer device.

[0006] The technical solutions provided by the present invention are as follows: A flame detection method based on flame specificity and multi-exposure image fusion includes the following steps: Construct a flame recognition model including image fusion module EMEIF and flame detection network. In the flame recognition model, the image fusion module is used to n A continuous multi-exposure image is generated into a fusion image, and the flame detection network is used to generate the flame recognition result based on the fusion image. Among them, the image fusion module is composed of an encoder, a fusion unit and a decoder; the encoder extracts features from the input multi-exposure image and obtains feature maps F1~F n The fusion unit is based on the preset fusion weights W1~W based on relative exposure. n F1~ F n After weighting, the channels are concatenated to obtain fused features, which are then processed by the decoder to produce a fused image.

[0007] The initial model is composed of EMEIF and a pre-trained flame detection network; it is trained as follows: (1) Set the loss including the structural similarity between images LSSIM and weight loss L wight The loss function L 1. Perform one round of training on the initial model and iteratively update the weight coefficients of the fusion unit in EMEIF. (2) Set the bounding box regression loss L box , target confidence loss L obj , category classification loss L cls The combined loss L 2. Perform a second round of training on the model after the previous round of training; iteratively update the model parameters of the flame detection network. (3) Set the L 1. L box 、 L obj and L cls The combined loss L 3. Perform three rounds of training on the model after the previous round of training to iteratively update the model parameters of the EMEIF and flame detection network; The collected multiple continuous multi-exposure images are input into the flame recognition model after three rounds of training to achieve flame recognition.

[0008] As a further improvement of the present invention, the flame detection network adopts a pre-trained network based on YOLOv5.

[0009] As a further improvement of the present invention, the encoder in the image fusion module uses a P-DenseBlock encoder. The P-DenseBlock encoder sequentially includes a first convolutional layer PConv, which uses a pinwheel convolution module with a convolution kernel of 3×3, and a second convolutional layer DPC1, a third convolutional layer DPC2, and a fourth convolutional layer DPC3, which use depthwise separable convolution with a convolution kernel of 3×3. The output of each convolutional layer in the P-DenseBlock encoder is residually connected with the output of each subsequent convolutional layer.

[0010] As a further improvement of the present invention, the decoder in the image fusion module adopts a DenseNet decoder, and the DenseNet decoder sequentially includes four convolution layers with a convolution kernel of 3×3.

[0011] As a further improvement of the present invention, in the fusion unit of the image fusion module, the fusion weights W1~W n The generation method is as follows: S01: Calculate the feature map extracted from the original image at each exposure using the following formula F i Exposure value Ei : ; In the above formula, I ( p ) represents the first p The normalized intensity value of pixels; N Represents the total number of pixels in the image.

[0012] S02: The exposure value of each image is converted into E i Convert to relative exposure R i : .

[0013] S03: The relative exposure is converted to R i Convert to dynamic weights W i : ; In the above formula, and are two learnable weight coefficients in the fusion unit; among them, increasing The weight of low-exposure images will increase significantly, and the weight of high-exposure images will approach 0; The relative exposure threshold can be adjusted to affect the average point of weight distribution.

[0014] As a further improvement of the present invention, a round of training uses a data set consisting of real sample images.

[0015] The second round of training uses a dataset consisting of real sample images and fused images output by EMEIF in the first round of training.

[0016] The three rounds of training use a dataset consisting of real sample images and their enhanced images.

[0017] As a further improvement of the present invention, the loss function L The expression for 1 is as follows: ; In the above formula, I i Indicates the input i The original image at different exposure levels; represents the fused image; is an evaluation function used to measure the structural similarity between the fused image and the input image; Indicates the i The original image structure adjustment factor under different exposure levels; Used to calculate the square of the L2 norm of each weight; represents a parameter factor, =0.0001.

[0018] As a further improvement of the present invention, the loss function L The expression for 2 is as follows: L 2 =L box + L obj + L cls .

[0019] As a further improvement of the present invention, the loss function L The expression for 3 is as follows: L 3 =ρL 1 +L box + L obj + L cls ; In the above formula, ρ represents the balance factor, ρ =0.1.

[0020] As a further improvement of the present invention, the original input of the flame recognition model is three consecutive multi-exposure images; in the inference stage, the three consecutive multi-exposure images are continuously acquired by the camera at a sampling interval of 0.2s, and the exposure of each image is adjusted by changing the exposure time.

[0021] The present invention also includes a computer program product, which includes a computer program. When the computer program is executed by a processor, it creates a trained flame recognition model as in the aforementioned flame detection method based on flame specificity and multi-exposure image fusion, thereby realizing the recognition of the flames contained in multiple continuous multi-exposure images based on input.

[0022] The present invention also includes a storage medium storing a computer program. When the computer program is executed by a processor, a trained flame recognition model is created as described above in the flame detection method based on flame specificity and multi-exposure image fusion, thereby realizing the recognition of the flames contained in the input multiple continuous multi-exposure images.

[0023] The present invention also includes a computer device, which includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the computer program is executed by the processor, a trained flame recognition model is created as in the aforementioned flame detection method based on flame specificity and multi-exposure image fusion, thereby realizing the recognition of the flames contained in multiple continuous multi-exposure images based on input.

[0024] The present invention has the following beneficial effects: Based on the specificity of flame detection, this paper constructs a novel flame recognition model consisting of an image fusion module and a flame detection network. This model uses continuously sampled multi-exposure images to achieve flame recognition. The image fusion module in this method adaptively fuses multiple images of varying exposures taken with a short time difference. The resulting fused image effectively integrates feature information from images of varying exposures, improving the balance between dark and bright areas in the original image. This provides the flame detection network with data more suitable for flame detection, effectively overcoming the impact of under- and over-exposure on detection accuracy, and preventing the network from misjudging images with extreme exposures.

[0025] To improve the training of flame recognition models using the EMFIF module, the present invention employs a unique three-stage training strategy, enabling the model to transition from learning the characteristics of flame data from baseline shots to learning the characteristics of flame data fused from multiple exposures. For flame detection, the early fusion of multiple exposures is a novel approach to expanding the flame dataset, and its phased mixed data training method further improves the generalization capabilities of the flame detection network.

[0026] The image fusion strategy provided by the present invention belongs to an early fusion strategy, which can fuse multiple exposure images continuously taken by a single camera to obtain a fused image with richer information and suitable for flame target detection, overcoming the problem of limited information contained in the sample image in the traditional algorithm that adjusts the image exposure in the later stage. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a network architecture diagram of the flame recognition model provided in Example 1 of the present invention.

[0028] Figure 2 This is a schematic diagram of the image fusion module in the flame recognition model of Example 1 of the present invention.

[0029] Figure 3 This is the module structure of the encoder part of the image fusion module in Example 1 of the present invention.

[0030] Figure 4 This is the module structure of the decoder part of the image fusion module in Example 1 of the present invention.

[0031] Figure 5 These are two typical sets of flame data taken at different times and distances in the test experiment and the images obtained using the pyramid fusion method.

[0032] Figure 6 For testing experiments Figure 5 Flame detection results of two fused images.

[0033] Figure 7 Some cases of missed detection and false detection in the test experiment.

[0034] Figure 8 This is a partial image fusion effect diagram of the model after training with the solution of the present invention in the test experiment. DETAILED DESCRIPTION

[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "or / and" as used herein includes any and all combinations of one or more of the associated listed items.

[0037] Example 1 Image fusion techniques are commonly used in image-based object recognition, classification, and segmentation tasks. Image fusion involves integrating information from multiple source images (possibly from different sensors, different viewpoints, and different time points) into a single output image. This process aims to effectively combine complementary, redundant, or unique information from each source image to produce a comprehensive image that captures all key details while also enhancing its features. Generally speaking, image fusion can be categorized into pixel-level fusion, feature-level fusion, and decision-level fusion. Pixel-level fusion, however, is unsuitable for object detection due to limitations such as high computational complexity, high hardware requirements, and susceptibility to pixel information contamination. With increasing interest in machine learning, feature-level fusion after feature extraction using convolutional neural networks has become increasingly mainstream, giving rise to multimodal object detection. These fusion methods are primarily categorized into early-stage fusion, mid-stage fusion, and late-stage fusion. Mid-stage and late-stage fusion primarily involves extracting features from dual backbone branches and then fusing them after the backbone network or at the neck of the network, primarily to improve the network's learning capabilities. Although most experiments have shown that the effect of the early fusion method is not as good as the mid- and late-stage fusion method, the early fusion method is mainly aimed at providing data with richer features. Because it contains more interference information, it is not effective in some target detection tasks.

[0038] In their research in the field of flame detection, technicians in this embodiment have discovered that when shooting flame images, when the brightness is too high or overexposed, some areas in the image will be highlighted, details and color features will be lost, and a large number of interference sources will appear (such as glass reflections and white walls in some areas). And because the flame itself has a strong brightness, it is more difficult to obtain relatively high-quality flame data in different environments. For example, when shooting flame data in a fire laboratory, it is easy to overexpose the image because there is a lot of refracted light indoors. In addition, when shooting flame data outdoors under strong sunlight, the high-brightness environment causes the color features of the captured flames to be weakened, and the texture details of the flame edges are also easily blurred, which can easily lead to poor image quality and model misjudgment.

[0039] In response to the aforementioned quality issues with visible light images of flames, fusing visible light and infrared images of flames can solve the problem to a certain extent. However, when the visible light image itself is overexposed and the color and texture features of the flame are relatively weak, the fusion model tends to focus on infrared, which mainly focuses on the shape features of the flame. However, in real-world shooting, it was also found that when visible light is overexposed, large areas of near-infrared images will also be overexposed. In addition, the cost of deploying edge devices (such as cameras) that align visible light and infrared is high. At the same time, because mid- and late-stage fusion mainly involves splicing and fusing features after passing through the backbone network for post-detection, it is impossible to effectively handle interference such as overexposure during the fusion stage.

[0040] In addition, the technicians of this embodiment also found that flame data has a certain specificity in target detection. This specificity makes the early fusion very suitable for the flame field. Specifically, compared with conventional dynamic target detection tasks (such as pedestrian and vehicle detection), the specificity of flame recognition is manifested as follows: Pedestrian and vehicle target detection is achieved using video stream data captured by a camera. The processing process requires frame extraction from the video stream data and detection using each frame. Since the detection target of this task is a dynamic target, when using some existing multi-exposure fusion methods to generate new images, it is necessary to ensure that the captured images are aligned and the shooting time is consistent to ensure that the pedestrian and vehicle maintain consistent coverage alignment after fusion. For example, if the time is inconsistent, the fused pedestrian images may partially overlap, resulting in significant differences between the pedestrian and the pedestrian learned by the detection network, causing the network to be unable to make predictions or reducing its confidence. Therefore, when performing multi-exposure fusion on moving targets, the only way is to adjust the original images through a post-exposure adjustment algorithm to obtain images of different exposures in the same scene. However, the image quality obtained by adjusting the original images to obtain images of different exposures is generally poor, and there is information loss.

[0041] Based on this, the researchers noted that flame shapes can change rapidly for a variety of reasons, and fire datasets are inherently scarce. However, fire detection models using smaller datasets can still identify irregularly shaped flames under certain conditions with appropriate exposure. This demonstrates that existing flame target detection networks have a strong generalization ability for flame shapes. The researchers speculate that this may be because flame target detection networks perceive flames as fluids like water, so even when flames with different shapes in different images are mixed, they still appear as flames.

[0042] Based on the above-discovered patterns of the effects of exposure and morphological fusion on flame target detection accuracy, this embodiment proposes a flame detection method based on flame specificity and multi-exposure image fusion. During the image acquisition phase, this method directly captures multiple continuous visible light images of different exposures using hardware devices such as cameras. The method then fuses each continuous image in the early stages, and finally uses the flame detection network to achieve flame recognition based on the fused image. This new solution is expected to reduce detection costs while effectively reducing the impact of overexposure on flame detection accuracy in various environments. Specifically, the flame detection method provided in this embodiment includes the following steps: 1. Construction of flame recognition model First, this embodiment uses the existing flame detection network with the newly designed image fusion module EMEIF to build a Figure 1The flame recognition model of the created architecture is shown in the figure. The flame detection network can be based on the YOLOv5 network model, or other network models with different architectures that have efficient flame recognition performance, such as DERT, Faster R-CNN, RetinaNet, etc. Figure 1 In the flame recognition model, the image fusion module receives multiple (assuming n, n ≥ 2) consecutive multi-exposure images and fuses them into a single fused image. The fused image generated by the image fusion module is input into the flame detection network, which then identifies the flame based on the fused image.

[0043] Specifically, in this embodiment, the continuous image input to the image fusion module should be multiple flame images captured at different exposures using the same camera at a fixed location. Considering that during the capture phase, image brightness is primarily dependent on exposure and ISO (light sensitivity). Exposure refers to the amount of light received by the photosensitive element during the exposure time, and controlling exposure is one of the primary ways a camera controls light. Exposure is equal to the product of the speed at which the photosensitive element receives light and the exposure time (shutter time). When exposure is constant, image brightness can be increased by increasing the ISO value of the photosensitive element. However, increasing ISO can easily generate noise, so exposure is generally adjusted to increase image brightness. Overexposure is also known as images captured with excessive exposure. To address the specific challenges of flame detection, this embodiment uses the camera's automatic exposure control (AE) to select exposure intervals and adjusts the brightness of each captured image by adjusting the exposure time. Furthermore, for multiple images in the same detection task, this embodiment sets the time interval between captured images at different exposures to less than one second to prevent significant differences in the captured flame morphology, which would be detrimental to fusion and detection.

[0044] Furthermore, for the flame recognition model, the number of multi-exposure input images should not be too large. This is because the most effective flame-related image information is typically present within a narrow exposure range. Introducing too many images at different exposures may cause the fused image to contain excessive interference information, hindering detection. Therefore, in the typical solution of this embodiment, the original input of the flame recognition model is three consecutive multi-exposure images. During the inference phase, the three consecutive multi-exposure input images are continuously acquired by the camera at a sampling interval of 0.2 seconds, and the exposure of each image is adjusted by varying the exposure duration.

[0045] In order to achieve efficient fusion of multiple images with different exposures, the image fusion module designed in this embodiment is as follows: Figure 2 As shown in Figure 1, it consists of three parts: encoder, fusion unit and decoder. The encoder extracts features from the input multi-exposure images and obtains feature maps F1~F nThe fusion unit is used to calculate the relative exposure-based fusion weights W1~W n F1~ F n After weighting, each feature image is then channel-joined to obtain the fusion feature. Finally, the fusion feature output by the fusion unit is processed by the decoder to obtain the fusion image. .

[0046] In practical applications, there are many encoder options. Considering that the EMEIF module uses multi-branch fusion during training, it is necessary to preserve as much contextual information as possible while also ensuring inference speed. Traditional convolution has a limited receptive field and cannot capture all information. The encoder used in this embodiment replaces the standard convolution (Conv) with the pinwheel convolution (PConv) based on the DenceBlock encoder, expanding the receptive field while minimally increasing the number of parameters. To distinguish it from the original encoder, this embodiment is named P-DenseBlock.

[0047] like Figure 3 As shown, the P-DenseBlock encoder includes, in sequence, a first convolutional layer PConv using a windmill-shaped convolution module with a convolution kernel of 3×3, and a second convolutional layer DPC1, a third convolutional layer DPC2, and a fourth convolutional layer DPC3 using a depthwise separable convolution with a convolution kernel of 3×3. In the P-DenseBlock encoder of this embodiment, the number of input channels and the number of output channels of the first convolutional layer are 1 and 16 respectively; the number of input channels and the number of output channels of the second convolutional layer are 16 and 16 respectively; the number of input channels and the number of output channels of the third convolutional layer are 32 and 16 respectively; the number of input channels and the number of output channels of the fourth convolutional layer are 48 and 16 respectively. In addition, the output of each convolutional layer in the P-DenseBlock encoder is residually connected with the output of each subsequent convolutional layer. Therefore, each feature map F1~F processed by the encoder is n The number of channels is 64.

[0048] On this basis, the fusion unit is based on the preset fusion weights W1~W n For F1~F n After weighting and channel concatenation, the number of channels of the fused features is also 64. During the fusion stage, this embodiment designs an adaptive weighting mechanism related to image exposure to adaptively fuse the three features so that the fused features, when restored by the decoder, produce a fused image that contains the most suitable feature information for flame detection.

[0049] Among them, considering that after the image passes through the encoder, brightness information and the like will be distributed in the image features, if the image features are simply spliced ​​or adaptively fused, the brightness information finally fused may not be suitable for flame detection. For targets such as flames that have relatively high brightness, lower exposure is more suitable for their detection. In low-exposure images, because the flame itself has a high brightness, when other things lose texture and shape details, the flame may gain more texture features due to the reduction in exposure. Therefore, the image fusion module of this embodiment is more inclined to lower the exposure of the fused image, but it cannot be too low to retain only the flame features, otherwise the use of the obtained data for training will reduce its robustness, so it is necessary to ensure that the scene details are not lost as much as possible. To address this problem, this embodiment introduces fusion weights W1~W based on relative exposure in the image fusion module. n .

[0050] Specifically, in the fusion unit of the image fusion module provided in this embodiment, the designed fusion weights W1~W n The generation method is as follows: S01: Calculate the feature map extracted from the original image at each exposure using the following formula F i Exposure value E i : ; In the above formula, I ( p ) represents the first p The normalized intensity value of pixels, I ( p )∈[0,1]; N Represents the total number of pixels in the image.

[0051] S02: The exposure value of each image is converted into E i Convert to relative exposure R i : .

[0052] Analyzing the above formula, we can find that for any image, when its corresponding exposure value E i The higher the value, the relative exposure obtained after the above processing R i The lower the value of , the higher the relative exposure value of the image with the highest exposure is. The relative exposure value of the image with the highest exposure is 0, which can ensure that the low-exposure image obtains a higher score in the fusion process of this embodiment.

[0053] S03: The relative exposure is converted to R i Convert to dynamic weights W i : ; In the above formula, and are two learnable weight coefficients in the fusion unit. The weight of low-exposure images will increase significantly, and the weight of high-exposure images will approach 0; The relative exposure threshold can be adjusted to affect the average point of weight distribution.

[0054] It can be seen that in the generation mechanism designed in this embodiment, the fusion weights W1~W n The value of will be adaptively adjusted with the relative exposure of each original image input to the network model, and will also be affected by the learnable weight coefficient and In the subsequent process of this embodiment, the image fusion model in the network model can be trained in a targeted manner to obtain the optimal value of the weight coefficient; to ensure the recognition accuracy of the network model for multi-exposure images in different scenes, thereby improving the generalization of the solution.

[0055] In the decoder part, considering that the encoder already has multiple branches and some parameters are added, based on the need for lightweight design, the decoder of this embodiment adopts the DenseNet decoder. Figure 4 As shown in the figure, the DenseNet decoder sequentially includes four convolutional layers with a 3×3 convolution kernel. The first convolutional layer has 64 input channels and 64 output channels, respectively; the second convolutional layer has 64 input channels and 32 output channels, respectively; the third convolutional layer has 32 input channels and 16 output channels, respectively; and the fourth convolutional layer has 16 input channels and 1 output channel, respectively. This shows that the image fusion module of this embodiment can output a fused image with one channel that contains rich information from the original images at different exposure levels.

[0056] 2. Training of flame recognition model In this embodiment, in order to ensure that the newly designed flame recognition model can achieve the best target detection performance, this embodiment also designs a corresponding multi-stage training strategy for it. Specifically, this embodiment selects the EMEIF module and the pre-trained flame detection network to form the initial model. Among them, the flame detection network can adopt various existing pre-trained network models and retain their model parameters. In the EMEIF module, this embodiment uses the weight coefficient The initial setting is 0.6, Initial setting is 0.5.

[0057] In practical applications, this embodiment designs a three-stage training strategy. The first stage is mainly used to train the EMEIF module, the second stage is used to update the model parameters of the flame detection network, and the third stage is used to train the flame recognition model that includes the EMEIF module and the flame detection network. The training process is as follows: (1) Setting the loss of structural similarity between images L SSIM and weight loss L wight The loss function L 1. Perform one round of training on the initial model and iteratively update the weight coefficients of the fusion unit in EMEIF. One round of training uses a dataset consisting of real sample images; the loss function used in this stage is L The expression for 1 is as follows: ; In the above formula, I i Indicates the input i The original image at different exposure levels; represents the fused image; is an evaluation function used to measure the structural similarity between the fused image and the input image; Indicates the i The original image structure adjustment factor under different exposure levels; Used to calculate the square of the L2 norm of each weight; represents a parameter factor. In this embodiment, =0.0001.

[0058] (2) Setting the bounding box regression loss L box , target confidence loss L obj , category classification loss L cls The combined loss L 2. Perform a second round of training on the model trained in the previous round; iteratively update the model parameters of the flame detection network. The second round of training uses a dataset consisting of real sample images and fused images output by EMEIF in the first round of training. The loss function used in this stage is L The expression for 2 is as follows: L 2 =L box + L obj +L cls .

[0059] (3) Setting includes L 1. L box 、 L obj and L cls The combined loss L 3. Perform three rounds of training on the model after the previous round of training, and iteratively update the model parameters of the EMEIF and flame detection network. The three rounds of training use a data set consisting of real sample images and their enhanced images. The sample images obtained through data enhancement include new sample images obtained by cutting the original image, as well as new sample images constructed by reconstructing the fused image obtained in the first stage with two of the original sample images. The loss function for this stage is L The expression for 3 is as follows: L 3 =ρL 1 +L box + L obj + L cls ; In the above formula, ρ represents the balance factor, ρ =0.1.

[0060] 3. Application of flame recognition model After the three rounds of training described above, a flame recognition model that meets the requirements for detection accuracy and robustness is obtained. This example saves the parameters of the network model that has been verified to have the best performance. In actual applications, multiple consecutive multi-exposure images are then input into the flame recognition model after three rounds of training. The model then performs inference and recognizes the flames contained therein.

[0061] Example 2 The flame detection method based on flame specificity and multi-exposure image fusion provided in Example 1 is essentially a data processing method. In order to better apply the solution in Example 1, this embodiment further provides a computer program product, storage medium and computer device that can implement this method.

[0062] Specifically, the computer program product provided in this embodiment includes a computer program. When the computer program is executed by a processor, it creates a trained flame recognition model as in the flame detection method based on flame specificity and multi-exposure image fusion in Example 1, thereby realizing the recognition of the flames contained in the input multiple continuous multi-exposure images.

[0063] The storage medium provided in this embodiment stores a computer program. When the computer program is executed by a processor, a trained flame recognition model is created as in the flame detection method based on flame specificity and multi-exposure image fusion in Example 1, thereby realizing the recognition of the flames contained in multiple continuous multi-exposure images based on the input.

[0064] The computer device provided in this embodiment includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the computer program is executed by the processor, a trained flame recognition model is created as in the flame detection method based on flame specificity and multi-exposure image fusion in Example 1, thereby realizing the recognition of the flames contained in multiple continuous multi-exposure images based on the input.

[0065] In actual applications, this computer device can be an embedded device and deployed in certain terminal devices with fire monitoring capabilities to support the aforementioned data processing. It can also be used as a standalone computer device to support data processing requirements for flame detection on multi-exposure images collected by the front end in certain scenarios. This non-embedded computer device can be a laptop, tablet, desktop computer, or a medium-to-large computer device such as a rack-mounted server, blade server, tower server, or cabinet server (including standalone servers or server clusters consisting of multiple servers) capable of executing computer programs.

[0066] Specifically, the computer device of this embodiment includes at least, but is not limited to, a memory and a processor that are communicatively connected to each other via a system bus. In this embodiment, the memory (i.e., readable storage medium) includes flash memory, a hard disk, a multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, magnetic disks, optical disks, and the like. In some embodiments, the memory may be an internal storage unit of the computer device, such as the computer device's hard disk or internal memory. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, and the like. Of course, the memory may also include both the internal storage unit and external storage devices of the computer device. In this embodiment, the memory is typically used to store the operating system and various application software installed on the computer device. Furthermore, the memory may also be used to temporarily store various types of data that has been output or is about to be output.

[0067] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of a computer device.

[0068] Test Experiment 1. Theoretical Verification of the Strategy in the Solution of the Present Invention 1. Verify the impact of flame shape on detection accuracy This experiment first pre-trained YOLOv5 using a dataset of real flames, resulting in a flame detection network based on YOLOv5. Subsequently, a 20-minute video of a flame burning in a laboratory was captured, with exposure controlled to ensure clear flames. During the capture process, the program was set to take an image every 0.3 seconds, resulting in 1,000 consecutive images. Two flame images separated by a 0.3-second interval were then combined using a pyramid fusion method to produce 500 fused flame images. Due to the time interval between these images, the shape of the flames in the fused images changed slightly from the original images. These 500 images were fed into the initially trained flame detection network, and the network's recognition results showed that it still accurately identified flames with minimal difference in confidence. Subsequently, the experiment re-shot the data using three images separated by a total of 0.6 seconds. The fused images were then fed into the flame detection network, and the model continued to successfully identify flames.

[0069] Figure 5 Part (a) and Figure 5 Part (b) shows two sets of flame data with the same exposure taken at different times and distances, and the images obtained by pyramid fusion. The detection results obtained by using the pre-trained flame detection weights are as follows: Figure 6 Part (a) and Figure 6 This is shown in part (b) of the image. The figure shows that the detection results under different conditions still have a high degree of confidence. Based on the above experimental content, it can be confirmed that for visible light flame target detection, the detection network has strong generalization ability for flame shapes.

[0070] 2. Impact of image sampling method on detection accuracy This experiment varied the time interval between image acquisitions and modified the number of fused images. The resulting 3,000 flame data images were fed into YOLO v5 for training. The trained weights were then used to detect the original flame data. The network successfully identified the flames, and the confidence level was similar to that during training and verification.

[0071] Combined with experimental data, it was finally found that when the wind direction did not change significantly, the image obtained by fusing flame images within a 1s interval was the best. The maximum number of flame images was fused, and increasing the number would result in a large difference between the fused image and the original image, while also increasing the amount of calculation.

[0072] 3. Impact of sample image overexposure on detection accuracy This experiment used overexposed images to detect flame targets and found that the network could not recognize them regardless of whether they were fused or not. At the same time, at certain exposure levels, some lights or reflections were easily misdetected after fusion. When the image exposure is too high, the flame features will be weakened and missed or misdetected. At too low exposure, although the flame features are more obvious, the surrounding environment features are weakened, resulting in misdetection of objects with similar colors, such as lights, at this exposure level. Figure 7 A series of missed or false detections are shown.

[0073] The above three experiments also confirm that flame target detection is specific. Unlike pedestrian or vehicle inspections, which have high shape requirements and weak shape generalization, the flame target detection network naturally has a strong generalization ability for the flame shape in the detection target. Therefore, performing non-real-time image fusion on visible light images has minimal impact on detection accuracy. This shows that the strategy proposed in this invention of using short-term continuous acquisition of multiple exposure images to achieve flame detection is feasible and has theoretical support. At the same time, the above experiments also confirm that for self-luminous detection targets such as flames, it is necessary to overcome or adjust the exposure when designing the fusion module.

[0074] II. Performance comparison test between the present invention and existing solutions 2.1 Model Training This experiment is based on the previous Figures 1-4 The scheme shown here builds and trains a flame recognition model. To obtain an effective dataset, this experiment captured approximately 6,000 flame images in simple indoor scenes. During the capture process, the surrounding scene remained static or exhibited only minor changes, except for the changes in the flame itself. Over 20,000 flame images were also captured in complex outdoor scenes, including images of flames that changed shape under varying wind speeds.

[0075] In the first round of training, this experiment divided 27,456 data sets into 9,152 groups of three scenes, corresponding to 8,219 annotations. Each data set used one set of annotations, and a quarter of the data, or 2,288 groups, was used as the validation set. Approximately 2,000 groups of simple indoor scenes were selected and fed into the EMEIF module for pre-training, with 50 epochs.

[0076] In the second round of training, one data point was randomly sampled from each of the 9,152 original data sets as the baseline data, resulting in 9,152 data points. Another 2,288 data points were randomly sampled from these fused data points in the second phase. The two sampling results were then mixed and shuffled, and a quarter was used as the validation set for training in YOLO V5s. The data augmentation method used was the simplest cropping, with an epoch of 80 and a learning rate of 0.01. The training focused on improving flame detection capabilities in complex scenarios and enhancing its robustness.

[0077] In the third phase of training, each scene data set consists of at least three images. The 9,152 images fused from the second phase of training are used as one image in each scene data set. Two images are then extracted from each of the original 9,152 sets, combined to create a new set of 9,152 scene data sets, which are then fed into the entire network for training. The weights used in the fusion and detection modules are the same as those obtained in the first two training phases. After the fusion phase, the 9,152 images obtained are used directly in the detection module training for 150 epochs, with the weights obtained in the third training phase being used as the primary weight.

[0078] Figure 8 The image fusion effect of the trained model of the present invention is demonstrated. The first three columns are three input flame data with short time differences and different exposures, with the exposure levels of the three being high to low. The last column, output, is the output image of the exposure fusion module of the present invention. By observing the images, it can be found that the solution of the present invention can fully solve the problem of environmental feature disorder caused by overexposure. It can retain some of the flame features when overexposed while avoiding the weakening of environmental features when underexposed.

[0079] 2.2, comparative experiment: In order to verify the advantages of the scheme provided by the present invention in flame detection performance, technicians conducted a comparative experiment. Because flame datasets are relatively scarce and do not meet the experimental requirements of the scheme of the present invention for multiple exposures, this experiment used a dataset taken by the technicians themselves for performance testing, and selected multiple existing multi-exposure fusion algorithms as a control group to compare the prediction performance with this scheme.

[0080] (2.2.1) Dataset.

[0081] Research in the field of video fire detection has been ongoing for over two decades, but there is currently no fire video database of suitable size, rich content, and uniform format available for researchers to use. The video and image data used by researchers comes from four main sources: First, several relatively small public fire video databases. These databases generally have relatively small data volumes and are not uniform in size, format, or scene. Second, fire video data is collected from relevant fire prevention agencies, such as local fire departments and fire research institutes. This video data is real fire records and is most suitable for video fire detection research. However, the current collaboration between relevant research institutions is not close enough, and this data needs to be mined. Third, relevant fire video image data is captured from the Internet. Images captured online using keywords such as "smoke" and "fire" are of varying content, mostly news illustrations or artistic photos, which differ significantly from real fire images. Fourth, researchers design their own experiments, simulate fire scenes, and capture fire video data.

[0082] Here are some publicly available fire video image datasets: (1) Fire video dataset provided by Bilkent University in Turkey. This dataset is the most commonly used dataset by early fire video detection researchers. It contains multiple flame and smoke videos, as well as interference videos. However, most of these videos were recorded by early monitoring systems and have a low resolution of only 240 pixels × 320 pixels.

[0083] (2) The fire video dataset released by the Computer Vision and Pattern Recognition Laboratory of Keimyung University in South Korea. This dataset contains flame and smoke videos, and the scenes also include both indoor and outdoor scenes.

[0084] (3) A fire database established by the Environmental Science Laboratory of the University of Corsica, France. This database was created for a vegetation fire modeling project and primarily contains images of vegetation flames, including visible light and near-infrared images. You need to register an account to download the data; the website also provides an upload function, allowing users to upload their own fire video image data.

[0085] (4) Forest fire database compiled by the Faculty of Electrical Engineering, University of Split, Croatia, which provides forest smoke video with a very wide field of view.

[0086] (5) The forest fire database released by the Institute of Microelectronics in Seville, Spain. This database was established for the V-MOTE project, which aims to use wireless sensor networks to build a video detection system and apply it to areas such as forest fire detection, building monitoring, and border control.

[0087] (6) The fire database published by Professor Yuan Feiniu's research group at Jiangxi University of Finance and Economics. The website provides some smoke videos and interference videos, as well as four smoke image libraries and a video smoke detection demonstration program.

[0088] (7) Video fire data provided by the State Key Laboratory of Fire Science, University of Science and Technology of China. This dataset contains fire data used in related articles by the fire detection research group of the laboratory, including videos and images.

[0089] There are also fire video image data used in research papers that have not yet been publicly released online. For example, Professor Zhou Zhiqiang's research team at Beijing Institute of Technology used high-quality forest smoke videos in their paper. Due to the diversity of fire scenarios and the wide range of shooting conditions, it is not possible to simply compile all available fire video image data to create a single database for all researchers in this field. Currently, researchers often have only a small amount of fire video image data available for their own research. While they achieve good recognition results with their own video image data, the lack of a unified database makes it difficult to rationally evaluate the performance of algorithms and hinders the training of robust and generalizable video fire detection algorithms. Furthermore, because our invention focuses on the fusion detection of distinct flames with multiple time differences, these datasets do not meet the experimental requirements. Therefore, this experiment opted to conduct its own data collection.

[0090] The acquisition device used in this experiment was a Hikvision MV CE200 10GC industrial camera. The experimenters divided the dataset into FIRE_IN and FIRE_OUT, respectively, for indoor and outdoor use. FIRE_IN consisted of approximately 12,000 video frames, while FIRE_OUT consisted of approximately 18,000 frames. Approximately 2,000 frames were subsequently filtered to account for significant flame image shifts caused by camera movement. During capture, the exposure gate was set to 15, and the exposure time was adjusted to adjust the exposure. The camera supports a maximum exposure time of 700,000 μs. The initial exposure time was automatically selected by the camera based on the scene. Subsequently, the exposure time was adjusted by 150,000 μs between each two exposures to capture consecutive images at different exposure levels (controlled by a script).

[0091] This experiment ultimately yielded two datasets, which were used to compare the performance of different solutions. FIRE_IN was shot indoors, with differences in the ignition platform, shooting angle, and burning material, resulting in a relatively simple overall environment. FIRE_OUT was shot outdoors, capturing multiple scenes and creating a more complex scene.

[0092] (2.2.2) Experimental results and analysis.

[0093] Because the present invention is a post-fusion detection, that is, it is first fused into a piece of data before being put into the detection network, the final detection result will also be presented in the fused image. In order to be able to compare with the ordinary non-fusion direct training detection, this experiment sets the final detection to be mapped to the original three images, that is, the detection results on the fused image will be mapped to the original three images. The following is a control experiment. The final test results are as follows: Among them, Table 1 is the performance test of the ablation and replacement detection network, and Table 2 is the results of using the same data, the same detection network (YOLOV5n) but different fusion methods, including traditional pixel-level methods and deep learning-based methods.

[0094] Table 1: Performance comparison between the present invention and the replacement detection network solution

[0095] Table 2: Performance comparison of the fusion module in this invention and other fusion methods

[0096] Analysis of the above experimental data shows that the present invention achieves the best results in detecting indoor flame data, i.e., medium and close distance flames, compared with traditional methods and other image fusion methods, and also has more outstanding performance in complex scenes such as outdoor.

[0097] The above-described embodiment merely represents one embodiment of the present invention. While the description is relatively specific and detailed, it should not be construed as limiting the scope of the invention. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, and these modifications and improvements fall within the scope of the present invention. Therefore, the scope of the present invention shall be determined by the appended claims.

Claims

1. A flame detection method based on flame specificity and multi-exposure image fusion, characterized in that: It includes: Construct a flame recognition model including the image fusion module EMEIF and the flame detection network; EM EIF is used according to n A continuous multi-exposure image is generated into a fused image; the flame detection network is used to generate the flame recognition result based on the fused image; wherein, EMEIF is composed of an encoder, a fusion unit and a decoder; the encoder extracts features from the input multi-exposure image to obtain multiple feature maps F1~F n The fusion unit is based on the preset fusion weights W1~W based on relative exposure. n F1~ F n After weighting, the channels are spliced ​​to obtain fused features; the fused features are processed by the decoder to obtain a fused image; The initial model is composed of EMEIF and a pre-trained flame detection network; it is trained as follows: the loss of structural similarity between images is set L SSIM and weight loss L wight The loss function L 1. Perform one round of training on the initial model and iteratively update the weight coefficients of the fusion unit in EMEIF; set the bounding box regression loss L box , target confidence loss L obj , category classification loss L cls The combined loss L 2. Perform a second round of training on the model after the previous round of training; iteratively update the model parameters of the flame detection network; set L 1. L box 、 L obj and L cls The combined loss L 3. Perform three rounds of training on the model after the previous round of training to iteratively update the model parameters of the EMEIF and flame detection network; The collected multiple continuous multi-exposure images are input into the flame recognition model after three rounds of training to achieve flame recognition.

2. The flame detection method based on flame specificity and multi-exposure image fusion according to claim 1, characterized in that: The flame detection network uses a pre-trained network based on YOLOv5; and / or The encoder in the image fusion module adopts a P-DenseBlock encoder; the P-DenseBlock encoder sequentially includes a first convolution layer PConv using a pinwheel convolution module with a convolution kernel of 3×3, and a second convolution layer DPC1, a third convolution layer DPC2, and a fourth convolution layer DPC3 using depthwise separable convolution with a convolution kernel of 3×3; the output of each convolution layer in the P-DenseBlock encoder is residually connected with the output of each subsequent convolution layer; and / or The decoder in the image fusion module adopts a DenseNet decoder, which sequentially includes four convolution layers with a convolution kernel of 3×3.

3. The flame detection method based on flame specificity and multi-exposure image fusion according to claim 2, characterized in that: In the fusion unit of the image fusion module, the fusion weights W1~W n The generation method is as follows: S01: Calculate the feature map extracted from the original image at each exposure using the following formula F i Exposure value E i : ; In the above formula, I ( p ) represents the first p The normalized intensity value of pixels; N Represents the total number of pixels in the image; S02: The exposure value of each image is converted into E i Convert to relative exposure R i : ; S03: The relative exposure is converted to R i Convert to dynamic weights W i : ; In the above formula, and are two learnable weight coefficients in the fusion unit, increasing The weight of low-exposure images will increase significantly, and the weight of high-exposure images will approach 0; The relative exposure threshold can be adjusted to affect the average point of weight distribution.

4. The flame detection method based on flame specificity and multi-exposure image fusion according to claim 1, characterized in that: One round of training uses a dataset consisting of real sample images; The second round of training uses a dataset composed of real sample images and fused images output by EMEIF in the first round of training; The three rounds of training use a dataset consisting of real sample images and their enhanced images.

5. The flame detection method based on flame specificity and multi-exposure image fusion according to claim 4, characterized in that: Loss Function L The expression for 1 is as follows: ; In the above formula, I i Indicates the input i The original image at different exposure levels; represents the fused image; is an evaluation function used to measure the structural similarity between the fused image and the input image; Indicates the i The original image structure adjustment factor under different exposure levels; Used to calculate the square of the L2 norm of each weight; represents a parameter factor, =0.0001.

6. The flame detection method based on flame specificity and multi-exposure image fusion according to claim 5, characterized in that: Loss Function L The expression for 2 is as follows: L 2 =L box + L obj + L cls ; Loss Function L The expression for 3 is as follows: L 3 =ρL 1 +L box + L obj + L cls ; In the above formula, ρ represents the balance factor, ρ =0.

1.

7. The flame detection method based on flame specificity and multi-exposure image fusion according to claim 1, characterized in that: The original input of the flame recognition model is three consecutive multi-exposure images. During the inference stage, the three consecutive multi-exposure images are continuously acquired by the camera at a sampling interval of 0.2 seconds, and the exposure of each image is adjusted by changing the exposure time.

8. A computer program product comprising a computer program, characterized in that: When the computer program is executed by a processor, it creates a trained flame recognition model in the flame detection method based on flame specificity and multi-exposure image fusion as described in any one of claims 1 to 7, thereby realizing the recognition of flames contained in multiple continuous multi-exposure images based on input.

9. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it creates a trained flame recognition model in the flame detection method based on flame specificity and multi-exposure image fusion as described in any one of claims 1 to 7, thereby realizing the recognition of flames contained in multiple continuous multi-exposure images based on input.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the computer program is executed by a processor, it creates a trained flame recognition model in the flame detection method based on flame specificity and multi-exposure image fusion as described in any one of claims 1 to 7, thereby realizing the recognition of flames contained in multiple continuous multi-exposure images based on input.

Citation Information

Patent Citations

  • Flame detection method based on visible light and infrared feature fusion network

    CN118799718A

  • Grayscale image flame identification method and system based on DeOldify and YOLOv8, and storage medium

    CN118865110A

  • Lightweight fire-DET flame detection method and system

    WO2022105143A1