Smoke and fire detection method, electronic equipment and computer readable storage medium

By combining the feature fusion processing of visible light images and thermal imaging images, the problems of low sensitivity and high false alarm rate of traditional smoke and fire detection methods are solved, and high-precision smoke and fire detection is achieved in all weather conditions.

CN120976844APending Publication Date: 2025-11-18ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510920632.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Traditional smoke detection methods have low sensitivity, slow response, and high false alarm rate, which cannot meet the monitoring needs of complex scenarios. Furthermore, visible light images perform poorly at night or in low visibility environments.

Method used

By combining visible light images and thermal imaging images, features are extracted from each through a feature extraction network, and feature fusion processing is performed. Then, a multimodal attention mechanism and a feature pyramid network are used for smoke detection.

Benefits of technology

It improves the accuracy and positioning precision of fireworks detection, reduces the false alarm rate, and enables all-weather fireworks detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976844A_ABST
    Figure CN120976844A_ABST
Patent Text Reader

Abstract

The invention discloses a smoke and fire detection method, electronic equipment and a computer readable storage medium, and the method comprises the steps: obtaining a visible light image and a thermal imaging image of the same scene, and enabling pixel points in the visible light image and the thermal imaging image to be in one-to-one correspondence; respectively inputting the visible light image and the thermal imaging image into a feature extraction network to obtain visible light image features and thermal imaging image features; performing fusion processing on the visible light image features and the thermal imaging image features to obtain target fusion features; and performing smoke and fire detection processing based on the target fusion feature to obtain a smoke and fire detection result. According to the scheme, the smoke and fire detection accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method for detecting smoke and fire, an electronic device, and a computer-readable storage medium. Background Technology

[0002] Fire detection uses intelligent methods to accurately identify and warn of flames and smoke at an early stage, which can reduce the possibility of fires.

[0003] Traditional smoke and fire detection methods primarily rely on smoke and heat detectors. However, this approach depends on physical signals and suffers from low sensitivity, response delays (requiring the fire to spread before triggering), and a high false alarm rate. Furthermore, it only provides point-based alarms and cannot directly pinpoint the fire source. Therefore, it is difficult to meet the monitoring needs of complex scenarios (such as outdoor equipment or smokeless fires).

[0004] With the rapid development of computer vision technology, smoke detection methods based on visible light images have emerged. However, visible light images mainly rely on visual information, and their performance is poor or even completely ineffective in low-visibility environments such as nighttime or dense fog. Summary of the Invention

[0005] The main technical problem addressed by this application is to provide a method, electronic device, and computer-readable storage medium for detecting fireworks, which can improve the accuracy of fireworks detection.

[0006] To address the aforementioned technical problems, this application provides a smoke and fire detection method, comprising: acquiring a visible light image and a thermal imaging image of the same scene, wherein the pixels in the visible light image and the thermal imaging image correspond one-to-one; inputting the visible light image and the thermal imaging image into a feature extraction network respectively to obtain visible light image features and thermal imaging image features; performing fusion processing on the visible light image features and the thermal imaging image features to obtain target fusion features; and performing smoke and fire detection processing based on the target fusion features to obtain a smoke and fire detection result.

[0007] To solve the above-mentioned technical problems, another technical solution adopted in this application is: to provide an electronic device, including a memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to execute the above-mentioned smoke detection method.

[0008] To solve the above-mentioned technical problems, another technical solution adopted in this application is to provide a computer-readable storage medium including program data, which, when executed by a processor, is used to implement the above-mentioned smoke detection method.

[0009] The smoke detection method of this application acquires visible light and thermal imaging images of the same scene, with each pixel in the visible light and thermal imaging images corresponding one-to-one. The visible light and thermal imaging images are then input into a feature extraction network to obtain visible light image features and thermal imaging image features, respectively. These features are then fused to obtain target fusion features. Smoke detection is performed based on these target fusion features to obtain the smoke detection result. This scheme, by fusing visible light and thermal imaging features at the feature level, can learn rich multimodal features, thereby accurately detecting and locating smoke, and reducing the false alarm rate. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:

[0011] Figure 1 This is a schematic flowchart of an exemplary embodiment of the fireworks detection method shown in this application;

[0012] Figure 2 yes Figure 1 A flowchart illustrating an exemplary embodiment of step S120 in the smoke detection method is shown.

[0013] Figure 3 This is a schematic diagram of the framework of an exemplary embodiment of the weight prediction module shown in this application;

[0014] Figure 4 This is a schematic diagram of the framework of an exemplary embodiment of the feature fusion process shown in this application;

[0015] Figure 5 This is a schematic diagram of a framework of an exemplary embodiment of the fireworks detection method shown in this application;

[0016] Figure 6 This is a schematic diagram of the framework of an exemplary embodiment of the verification module shown in this application;

[0017] Figure 7 This is a schematic diagram of an exemplary embodiment of a smoke detection device;

[0018] Figure 8 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application;

[0019] Figure 9 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all structures. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0021] Once a fire spreads, it can cause immeasurable and serious consequences. Therefore, it is crucial to detect and stop a fire in its early stages. Common methods often involve using smoke detectors or temperature detectors to monitor the environment for smoke or high temperatures in real time. These methods are suitable for fire detection in small environments, but they suffer from problems such as delays and a high false alarm rate.

[0022] Visible light image detection is more accurate than smoke and heat detectors. However, visible light image detection depends on visibility. In environments with low visibility, it is difficult to determine the presence of smoke and the specific location of the smoke based solely on visible light images.

[0023] Based on this, this application proposes a smoke detection method, electronic device, and computer-readable storage medium that combines visible light images and thermal imaging images for smoke detection. For details, please refer to [link / reference needed]. Figure 1 , Figure 1 This is a schematic flowchart of an exemplary embodiment of the fireworks detection method shown in this application.

[0024] The execution entity of the smoke and fire detection method can be a terminal device, a server, or other processing equipment. The terminal device can be a user equipment (UE), computer, mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, vehicle-mounted device, wearable device, etc. The execution entity of the smoke and fire detection method can also be a smoke and fire detection device. In some possible implementations, the smoke and fire detection method can be implemented by a processor calling computer-readable instructions stored in memory.

[0025] Specifically, the smoke detection method in this embodiment includes the following steps:

[0026] S110: Acquire visible light image and thermal imaging image of the same scene, with each pixel in the visible light image and thermal imaging image corresponding to the previous one.

[0027] The same scene can be a scene captured by the acquisition device from the same observation angle. Acquiring different images based on the same scene can make the pixels in the different images correspond one-to-one as much as possible, that is, pixels at the same location indicate the same scene. For example, a dual-spectrum device integrating visible light and thermal imaging can be used to align with the same area to obtain visible light images and thermal images of the same scene.

[0028] Visible light images are images formed by capturing visible light reflected or emitted by an object using a photosensor. For example, a visible light camera can be used to photograph a target area to obtain a visible light image of that area. The visible light camera may include a photosensor. Visible light images can realistically reproduce the environment as observed by the eye, clearly showing the details of the target area. The target area can be an indoor environment, a forested area, a road, a house, or other areas where fire detection is needed.

[0029] Thermal imaging refers to the process of capturing infrared radiation using an infrared sensor and converting it into a visible image. For example, a thermal imaging camera can be used to capture images of a target area, obtaining a thermal image of that area. This thermal imaging camera may contain an infrared sensor. Since the imaging principle of thermal imaging is based on the relationship between the infrared radiation emitted by different objects and their surface temperatures, and uses different colors or grayscale levels to represent different temperature regions to display the temperature distribution of the target object, a visual representation of heat distribution is achieved. Therefore, thermal imaging can be used to detect smoke and fire conditions in a real-time scene.

[0030] In some embodiments, visible light acquisition devices and thermal imaging acquisition devices can be used simultaneously to acquire visible light images and thermal images of the same scene, respectively. In other embodiments, a dual-spectrum integrated device for visible light and thermal imaging can be used to simultaneously acquire visible light images and thermal images of the same scene. To ensure a one-to-one correspondence between pixels in the visible light and thermal images, hardware-level triggering is used to synchronously acquire visible light and thermal images before acquiring images using the dual-spectrum integrated device, eliminating data misalignment caused by differences in acquisition timing. The focal length and distortion parameters of the visible light and thermal imaging acquisition devices are calibrated using a checkerboard calibration board or active infrared marking method, and lens distortion is corrected using a nonlinear optimization algorithm to obtain the corrected image acquired by the dual-spectrum integrated device. The visible light and thermal images acquired using the corrected image acquired by the dual-spectrum integrated device have higher image quality and better matching.

[0031] A pixel is the basic building block of an image; multiple pixels combine to form an image. In this embodiment, the visible light image and the thermal imaging image are obtained by a dual-spectrum acquisition device capturing the same scene from the same angle. Therefore, the pixels in the visible light image and the thermal imaging image correspond one-to-one, corresponding to the same location in the scene. It should be noted that a one-to-one pixel correspondence can mean that pixels at the same location in the two images indicate the same scene. Alternatively, it can mean that the corresponding pixels have a certain range of positional offset in distance, as long as a large positional offset does not affect the fusion of the two images.

[0032] S120: Input the visible light image and the thermal imaging image into the feature extraction network respectively to obtain the visible light image features and the thermal imaging image features.

[0033] Feature extraction networks are used to extract image features from visible light images and thermal imaging images. In some embodiments, the feature extraction network may include, but is not limited to, VGG (Visual Geometry Group) networks and ResNet (Residual Network). In other embodiments, the feature extraction network may be a multimodal feature network used to perform deep fusion of visible light images and thermal imaging images to obtain visible light image features and thermal imaging image features.

[0034] Visible light image features can be mathematical expressions extracted from visible light images that describe the essential information of the visible light image content. Thermal imaging image features can be mathematical expressions extracted from thermal imaging image features that describe the essential information of the thermal imaging image content. To improve the collaboration between visible light images and thermal imaging images, multi-layer feature interaction fusion can be used during feature extraction, so that the output visible light image features contain image information from the thermal imaging image, and vice versa.

[0035] S130: The visible light image features and thermal imaging image features are fused to obtain the target fused features.

[0036] After obtaining the visible light image features and thermal imaging image features output by the feature extraction network, the smoke detection device performs a fusion process on the visible light image features and thermal imaging image features to obtain fused features. In some embodiments, the feature fusion process can employ a multimodal attention mechanism to adjust the weights of the visible light image features and thermal imaging image features to obtain the target fused features. This achieves the integration of image information from the visible light image and the thermal imaging image.

[0037] S140: Fire detection is performed based on target fusion features to obtain fire detection results.

[0038] Smoke detection primarily involves detecting the smoke produced by burning flames. The smoke detection results include the area and location information of the smoke-covered region. For example, a pre-trained detection module can be used to process the target fusion features to obtain the smoke detection results. This detection module can include, but is not limited to, commonly used YOLO detection modules, such as YOLO-V8, YOLO-V5, YOLO-V4, YOLO-V7, and PP-YOLOv2.

[0039] The smoke and fire detection device detects smoke and fire based on target fusion features. It can combine more abundant target area information obtained from visible light images and thermal imaging images to achieve accurate smoke and fire detection.

[0040] As can be seen, the smoke detection method in this embodiment acquires visible light and thermal imaging images of the same scene, with each pixel in the visible light and thermal imaging images corresponding one-to-one. The visible light and thermal imaging images are then input into a feature extraction network to obtain visible light image features and thermal imaging image features, respectively. These features are then fused to obtain target fused features. Smoke detection is performed based on these target fused features to obtain the smoke detection result. By fusing visible light and thermal imaging features at the feature level, rich multimodal features can be learned, thereby accurately detecting and locating smoke, reducing the false alarm rate.

[0041] Based on the above embodiments, the embodiments of this application adopt... Figure 2 The flowchart details how to extract features from visible light and thermal imaging images to obtain the actual visible light and thermal images. Please refer to [link / reference needed]. Figure 2 , Figure 2 yes Figure 1 The illustrated flowchart shows an exemplary embodiment of step S120 in the smoke detection method. Specifically, step S120, which inputs the visible light image and the thermal imaging image into the feature extraction network to obtain the visible light image features and the thermal imaging image features, includes the following steps:

[0042] First, it should be noted that the feature extraction network includes a weight prediction module and a feature fusion module. The weight prediction module is used to predict the fusion weights of the visible light image and the thermal imaging image, and the feature fusion module is used to perform feature fusion processing on the visible light image and the thermal imaging image to obtain the visible light image features and the thermal imaging image features.

[0043] S210: Input the visible light image into the weight prediction module to obtain the first fusion weight corresponding to the visible light image and the second fusion weight corresponding to the thermal imaging image.

[0044] The weight prediction module can generate a first fusion weight corresponding to the visible light image and a second fusion weight corresponding to the thermal imaging image based on the visible light image. In some embodiments, the weight prediction module can also generate the first fusion weight corresponding to the visible light image and the second fusion weight corresponding to the thermal imaging image based on the thermal imaging image. In other embodiments, the weight prediction module can also determine the first fusion weight of the visible light image and the thermal imaging image by combining the visible light image and the thermal imaging image. In still other embodiments, the first fusion weight of the visible light image and the second fusion weight of the thermal imaging image can be preset based on experience.

[0045] The weight prediction module can evaluate image quality when determining fusion weights based on visible light images and / or thermal imaging images. In some embodiments, a first fusion weight for the visible light image and a second fusion weight for the thermal imaging image are determined based on the image quality of the visible light image. Specifically, if the image quality of the visible light image is greater than a preset image quality, a first preset fusion weight is determined as the first fusion weight for the visible light image, and a second fusion weight for the thermal imaging image is determined based on a preset value and the first fusion weight; if the image quality of the visible light image is less than or equal to the preset image quality, a second preset fusion weight is determined as the first fusion weight for the visible light image, and a second fusion weight for the thermal imaging image is determined based on a preset value and the first fusion weight, wherein the first preset fusion weight is greater than the second preset fusion weight. Alternatively, a mapping relationship can be established between the image quality of the visible light image and the first fusion weight, for example, a proportional relationship can be established, where a higher image quality corresponds to a higher first fusion weight. In other embodiments, the determination of the first fusion weight of the visible light image and the second fusion weight of the thermal imaging image based on the image quality of the thermal imaging image can be performed in the manner described above. In still other embodiments, the determination of the first fusion weight of the visible light image and the second fusion weight of the thermal imaging image is based on the image quality of both the visible light image and the thermal imaging image. For example, a first fusion weight and a second fusion weight are determined based on the ratio of the image quality of the visible light image to the image quality of the thermal imaging image.

[0046] The weight prediction module includes a first feature extraction layer and a regression processing layer. As an example, a visible light image is input to the first feature extraction layer for feature extraction, resulting in first visible light image features. These first visible light image features are then input to the regression processing layer for processing, yielding a first fusion weight corresponding to the visible light image. A second fusion weight corresponding to the thermal imaging image is determined based on a preset value and the first fusion weight. By determining the fusion weight using the visible light image, the fusion process can focus more on image features with better image quality, enhancing the reliability of the fused features and thus improving the accuracy of smoke and fire detection.

[0047] The first feature extraction layer is used to extract image features from the visible light image, obtaining the first visible light image features. For example, the first feature extraction layer can be a VGG network or a ResNet network, and the ResNet network can be a ResNet18 network. The first feature extraction layer may include, but is not limited to, multiple convolutional layers, which perform convolution processing on the visible light image to obtain the image features of the visible light image.

[0048] The regression processing layer is used to process the features of the first visible light image to obtain the first fusion weights corresponding to the visible light image. For an example, please refer to [link to example]. Figure 3 The regression processing layer includes a pooling layer, a batch normalization layer, a dropout layer, a fully connected layer (FC), and an activation function (Sigmoid). The visible light image is first processed by a ResNet18 network to obtain the first visible light image features. The first visible light image features are then processed sequentially through the pooling layer, batch normalization layer, regularization layer, fully connected layer, and activation function to obtain the first fusion weights of the visible light image.

[0049] After obtaining the first fusion weight of the visible light image, the second fusion weight corresponding to the thermal imaging can be determined based on a preset value and the first fusion weight. The difference between the preset value and the first fusion weight can be defined as the second fusion weight. The preset value can be set to 1.

[0050] In other embodiments, the thermal imaging image can be input into a weight prediction module for processing to obtain a second fusion weight for the thermal imaging image. Then, a first fusion weight for the visible light image is determined based on a preset value and the second fusion weight.

[0051] In other embodiments, random forest models, principal component analysis models, etc., can also be used as weight prediction modules to analyze visible light images and thermal imaging images to obtain their respective fusion weights.

[0052] S220: Input the visible light image, the first fusion weight, the thermal imaging image, and the second fusion weight into the feature fusion module to obtain the visible light image features and the thermal imaging image features.

[0053] The feature fusion module is used to perform feature fusion processing on visible light images and thermal imaging images using a first fusion weight and a second fusion weight to obtain visible light image features and thermal imaging image features.

[0054] In some embodiments, the structure of the feature fusion module can refer to a feature pyramid network constructed following the Path Aggregation Network (PAN) concept. The feature pyramid includes two stages: top-down and bottom-up. First, preliminary feature extraction is performed on the visible light image and the thermal imaging image to obtain initial visible light image features and initial thermal imaging image features. The initial visible light image features and initial thermal imaging image features are then used as inputs to the feature pyramid network. The feature pyramid network fuses the initial visible light image features and initial thermal imaging image features to obtain the visible light image features and thermal imaging image features.

[0055] In other embodiments, the feature fusion module can also be a multi-layer feature interaction fusion structure. Specifically, the feature fusion module may further include a second feature extraction layer, a third feature extraction layer, and a feature fusion layer. The visible light image and the thermal imaging image are input to the second feature extraction layer to obtain second visible light image features and second thermal imaging image features, respectively. The second visible light image features, a first fusion weight, the second thermal imaging image features, and the second fusion weight are input to the feature fusion layer for fusion processing to obtain a first fused feature. The first fused feature and the second visible light image features are input to the third feature extraction layer to obtain visible light image features. The first fused feature and the second thermal imaging image features are input to the third feature extraction layer to obtain thermal imaging image features. Thus, by adding fusion features to each feature extraction layer, deep fusion at the feature level is achieved, fully absorbing the feature information of the visible light image and the thermal imaging image, providing richer features.

[0056] As an example, the visible light image and the thermal imaging image are respectively input into the second feature extraction layer for feature extraction, resulting in a second visible light image feature A and a second thermal imaging image feature B. The second visible light image feature A, the first fusion weight, the second thermal imaging image feature B, and the second fusion weight are input into the feature fusion layer, which fuses the second visible light image feature A and the second thermal imaging image feature B based on the first fusion weight and the second fusion weight, resulting in a first fused feature C. The second visible light image feature A and the first fused feature C are then concatenated or added together to obtain A+C, which is then input into the third feature extraction layer for feature extraction, resulting in the visible light image feature. The second thermal imaging image feature B and the first fused feature C are then concatenated or added together to obtain B+C, which is then input into the third feature extraction layer for feature extraction, resulting in the thermal imaging image feature.

[0057] If the feature fusion module includes a second feature extraction layer, a third feature extraction layer, a fourth feature extraction layer, a first feature fusion layer, and a second feature fusion layer, then the visible light image and the thermal imaging image are respectively input into the second feature extraction layer for feature extraction to obtain the second visible light image feature A and the second thermal imaging image feature B; the second visible light image feature A, the first fusion weight, the second thermal imaging image feature B, and the second fusion weight are input into the first feature fusion layer, which fuses the second visible light image feature A and the second thermal imaging image feature B based on the first fusion weight and the second fusion weight to obtain the first fused feature C; the second visible light image feature A and the first fused feature C are spliced ​​or added to obtain A+C, and A+C is input into the third feature extraction layer for feature extraction to obtain the third visible light image feature A'; the second thermal imaging image feature B and the first fused feature C are then combined. The three features are: First, a stitching or addition process is performed to obtain B+C. Then, B+C is input to the third feature extraction layer for feature extraction, resulting in the third thermal imaging image feature B'. Next, the third visible light image feature A' and the third thermal imaging image feature B' are input to the second fusion layer. The second feature fusion layer fuses the third visible light image feature A' and the third thermal imaging image feature B' based on the first and second fusion weights, resulting in the second fused feature C'. Then, the third visible light image feature A' and the second fused feature C' are stitched or added to obtain A'+C'. Then, A'+C' is input to the fourth feature extraction layer for feature extraction, resulting in the visible light image feature. Finally, the third thermal imaging image feature B' and the second fused feature C' are stitched or added to obtain B'+C'. Then, B'+C' is input to the fourth feature extraction layer for feature extraction, resulting in the thermal imaging image feature. Similarly, a feature fusion layer is added after each feature extraction layer to fuse the visible light image features and thermal imaging image features output by the feature extraction layer. The output of the feature fusion layer is then fused with the output of the previous feature extraction layer and input into the next feature extraction layer for feature extraction, until the final visible light image features and thermal imaging image features are obtained.

[0058] In this feature fusion module, each feature extraction layer can have the same structure or different structures. To improve feature extraction efficiency, each feature extraction layer can include a first sub-feature extraction layer and a second sub-feature extraction layer. The first sub-feature extraction layer is used to extract features from the visible light image, and the second sub-feature extraction layer is used to extract features from the thermal imaging image. The first and second sub-feature extraction layers can have the same structure, or different structures can be designed according to the different characteristics of the two images. To reduce structural complexity, each feature extraction layer can also include only one feature extraction layer, performing feature extraction on the visible light image and the thermal imaging image sequentially.

[0059] The feature fusion layer fuses the corresponding visible light image features and the corresponding thermal imaging image features output from the previous feature extraction layer based on a first fusion weight and a second fusion weight to obtain the corresponding fused features. As an example, the steps of the feature fusion layer in fusing the second visible light image features and the second thermal imaging image features may include: weighting the second visible light image features and the second thermal imaging image features according to the first fusion weight corresponding to the second visible light image features and the second fusion weight corresponding to the second thermal imaging image features to obtain weighted features; performing deep fusion processing on the weighted features to obtain deep fused features; and summing the weighted features and the deep fused features to obtain the first fused feature. Feature fusion based on fusion weights allows the fused features to focus more on image features with high fusion weights, reducing the false alarm rate of subsequent smoke detection.

[0060] The weighted feature is a feature obtained by weighting the second visible light image feature and the second thermal imaging image feature. For example, the first fusion weight can be used as the weight of the second visible light image feature, and the second fusion weight can be used as the weight of the second thermal imaging image feature to perform weighted processing on the second visible light image feature and the second thermal imaging image feature, thus obtaining the fused feature. The weighting processing method may include, but is not limited to, weighted summation, weighted averaging, etc.

[0061] Deep fusion features are features obtained by deeply fusing weighted features. For example, weighted features can be processed through multiple fully connected layers to obtain deep fusion features. As an example, the deep fusion process sequentially includes a first fully connected layer, a batch normalization layer, a regularization layer, a second fully connected layer, and an activation function. The weighted features are processed sequentially through the first fully connected layer, batch normalization layer, regularization layer, second fully connected layer, and activation function to obtain the final deep fusion feature. Of course, in other embodiments, the number of fully connected layers in the deep fusion process can be set according to actual needs; for example, a third fully connected layer may also be included.

[0062] After obtaining the deep fusion features, the weighted features are summed with the deep fusion features to obtain the first fusion feature. This method of multiple fusions of features from the two data sources allows for the full absorption of the strengths of both during subsequent fireworks detection, resulting in more robust fireworks detection results.

[0063] Please refer to Figure 4 , Figure 4This is a schematic diagram of an exemplary embodiment of the feature fusion process shown in this application. The second visible light image features and the second thermal imaging image features are weighted and averaged using a first fusion weight and a second fusion weight to obtain a weighted feature. The weighted feature is then processed sequentially through a first fully connected layer, a batch normalization layer, a regularization layer, a second fully connected layer, and an activation function to obtain a deep fusion feature. The weighted feature and the deep fusion feature are then added together to obtain the first fusion feature. It should be noted that this example only uses the second visible light image features, the second thermal imaging image features, and the first fusion feature as an example of feature interaction fusion; feature interaction fusion methods for other layers can refer to this example.

[0064] As can be seen, the role of visible light images differs under varying visibility conditions. Therefore, in this embodiment, the visible light image is input into a weight prediction module to obtain a first fusion weight corresponding to the visible light image and a second fusion weight corresponding to the thermal imaging image. Then, the visible light image, the first fusion weight, the thermal imaging image, and the second fusion weight are input into a feature fusion module to obtain visible light image features and thermal imaging image features. The weight prediction module can process the visible light image, outputting different fusion weights depending on whether the lighting conditions are good or the visibility is low, thus improving the flexibility and adaptability of the fusion weights. For example, the first fusion weight output can be greater than the second fusion weight when the lighting conditions are good, making feature fusion more biased towards the visible light image; conversely, the first fusion weight output can be less than the second fusion weight when the visibility is low, making feature fusion more biased towards the thermal imaging image. This enables all-weather detection of fireworks.

[0065] After obtaining the visible light image features and thermal imaging image features, the visible light image features and thermal imaging image features are fused to obtain the target fused features. In some embodiments, the visible light image features and thermal imaging image features can be added together to obtain the target fused features. In other embodiments, the visible light image features and thermal imaging image features can also be fused using the first fusion weight corresponding to the visible light image and the second fusion weight corresponding to the thermal imaging image output by the weight prediction module to obtain the target fused features. Specifically, the feature extraction network includes a weight prediction module and a feature fusion module. The visible light image is input into the weight prediction module to obtain the first fusion weight corresponding to the visible light image and the second fusion weight corresponding to the thermal imaging image; the visible light image, the first fusion weight, the thermal imaging image, and the second fusion weight are input into the feature fusion module to obtain the visible light image features and the thermal imaging image features; the visible light image features and the thermal imaging image features are weighted and summed according to the first fusion weight and the second fusion weight to obtain the target fused features.

[0066] To improve the fusion effect of visible light and thermal images, the process of acquiring visible light and thermal images of the same scene can involve image registration to ensure data spatial alignment between the two images. Specifically, initial visible light and initial thermal images of the same scene are acquired; matching feature points are obtained from the initial visible light and initial thermal images; and registration is performed on the initial visible light and initial thermal images based on the matching feature points to ensure a one-to-one correspondence between the pixels of the initial visible light and initial thermal images, thus obtaining the visible light and thermal images.

[0067] Initial visible light and initial thermal imaging images of the same scene can be acquired using pre-calibrated dual-spectral acquisition devices. Pre-calibration can be a hardware registration process. Hardware registration methods include, but are not limited to: eliminating data misalignment caused by acquisition timing through hardware-level triggering; calibrating the focal length and distortion parameters of the visible light and thermal imaging acquisition devices using a checkerboard calibration board or active infrared marking method; correcting lens distortion using nonlinear optimization algorithms; and establishing a unified coordinate system by using 3D calibration objects or natural scene feature points (e.g., edges, corners) and calculating the rotation and translation matrix between the visible light and thermal imaging acquisition devices through feature matching. Preliminary viewpoint matching is then performed through internal and external parameter alignment.

[0068] After acquiring the initial visible light image and the initial thermal imaging image, the smoke detection device can also perform pixel-level matching between the two images through software registration. Software registration can be cross-modal image registration, used to solve the data spatial alignment between heterogeneous sensors (e.g., visible light images and thermal imaging images), overcoming geometric translations, rotations, and scaling between the initial visible light image and the initial thermal imaging image. For example, affine transformation can be used for feature matching to achieve cross-modal registration. Thus, by using a two-step approach of hardware and software registration, the simplification and accuracy of registration between the visible light image and the thermal imaging image are greatly improved.

[0069] Matching feature points indicate the location of the same scene in two images. The selection criteria for matching feature points include that they are easily detectable in both images, are significant, and will not be confused with other points. For example, a dual-channel model can be constructed to extract edge contour features from the initial visible light image and the initial thermal image respectively; a feature point detection algorithm can be used to focus on capturing common geometric structural features in both images to obtain feature points, such as device edges or corners; and the feature points in the two images can be matched to obtain matching feature points.

[0070] The feature point detection algorithms include the improved FAST algorithm, the Harris Corner detection algorithm, and nonmaximum suppression. Furthermore, to address cross-modal differences, the principal direction of the matched feature points can be corrected by adjusting the edge tangent direction, thereby enhancing feature rotation invariance.

[0071] The feature point matching process between the two images can further include: using BRIEF (Binary Robust Independent Elementary Features) descriptors to binary encode each feature point, obtaining feature point descriptors for each feature point; acquiring the feature similarity between the feature point descriptors in the initial visible light image and the feature point descriptors in the initial thermal imaging image, where the feature similarity is calculated using methods such as Hamming distance; and determining the corresponding two feature points as matching feature points for the initial visible light image and the initial thermal imaging image if the feature similarity between the two feature point descriptors is greater than a preset similarity (which can take any value between 0.7 and 0.8). In some embodiments, the RANSAC (Random Sample Consensus) algorithm can be further used to detect anomalies between two feature points with a feature similarity greater than the preset similarity, and the two feature points that do not exhibit anomalies as matching feature points are used as the matching feature points.

[0072] After obtaining the matching feature points, the initial visible light image and the initial thermal imaging image are registered based on these feature points. In some embodiments, an affine transformation model can be established based on the matching feature points. The formula for the affine transformation model is as follows:

[0073] X′=s·R·X+T

[0074] Where X′ represents the transformed image, s represents the scaling factor, R represents the rotation matrix, X represents the image to be transformed (which can be either the initial visible light image or the initial thermal image), and T represents the translation vector. The scaling factor, rotation matrix, and translation vector can be solved using the least squares method. During the solution process, a composite loss function can be used to simultaneously optimize geometric alignment accuracy and contour similarity to improve the accuracy of the affine transformation.

[0075] Furthermore, the reprojection error can be calculated to assess the registration quality. If the confidence of the matched feature points is lower than the threshold (which can be arbitrarily selected between 0.7 and 0.9), the affine transformation matrix can be iteratively optimized through local feature fine-tuning or parameter fine-tuning so that the image transformed based on the affine transformation model can achieve image space alignment between the two modes, resulting in aligned visible light and thermal images.

[0076] To elaborate on the smoke detection method of this application, Figure 5 The framework diagram shown below provides further explanation, as detailed below:

[0077] The system acquires visible light and thermal images of the same scene. Specifically, the fireworks detection device uses a pre-calibrated dual-spectrum acquisition device to acquire initial visible light and initial thermal images of the same scene; it detects feature points in the initial visible light and initial thermal images; it calculates the feature similarity between feature points in the initial visible light and initial thermal images, and determines matching feature points based on the relationship between the feature similarity and a preset similarity; it establishes an affine transformation model based on the matching feature points, and performs transformation processing on either the initial visible light or initial thermal image based on the affine transformation model, for example, transforming the initial thermal image to the plane of the initial visible light image; it optimizes the affine transformation matrix based on the reprojection error to complete the spatial alignment of the initial visible light and initial thermal images, thus obtaining the visible light and thermal images.

[0078] The visible light image is input into the weight prediction module for weight prediction processing, resulting in a first fusion weight corresponding to the visible light image and a second fusion weight corresponding to the thermal imaging image. The weight prediction module, also known as an adaptive weight prediction network, takes the visible light image as input and passes through multiple convolutional and fully connected layers to obtain a two-dimensional one-hot encoding. This value represents the first fusion weight corresponding to the visible light image and the second fusion weight corresponding to the thermal imaging image. These fusion weights can be used in subsequent feature interaction fusion and weighted summation modules to fuse the data from the visible light image and the thermal imaging image.

[0079] Visible light images and thermal imaging images are input into a feature fusion module. The feature fusion module includes a multi-layer feature extraction module and a multi-layer feature interaction fusion module. The multi-layer feature extraction module includes at least one feature extraction layer, and the multi-layer feature interaction fusion module includes at least one feature fusion layer. The feature fusion layer is used to fuse the visible light image features and thermal imaging image features extracted by each feature extraction layer using a first fusion weight and a second fusion weight. The fused features are then fused with the visible light image features and thermal imaging image features output by the previous feature extraction layer as input to the next feature extraction layer, until the visible light image features and thermal imaging features output by the last feature extraction layer are obtained.

[0080] The visible light image features and thermal imaging image features are weighted and summed using a first fusion weight and a second fusion weight to obtain the target fusion feature. This target fusion feature is then input into the detection module for smoke and fire detection processing to obtain the smoke and fire detection result. The weight prediction module, feature fusion module, and weighted summation can be collectively used as a multimodal feature extraction and fusion network.

[0081] In some embodiments, the fireworks detection device can input target fusion features into the detection module to obtain the pixel region, category, and confidence level corresponding to the fireworks; the pixel region and category with a confidence level greater than a preset confidence level can be used as the final fireworks detection result.

[0082] In other embodiments, the fireworks detection device can also perform initial detection processing on the target fusion features to obtain the initial pixel region where the fireworks are located; and perform verification processing on the initial pixel region to obtain the target pixel region of the fireworks. This verification process can reduce the false alarm rate for fireworks.

[0083] The initial pixel region is the area of ​​pixels obtained from the initial detection. This initial pixel region can be represented by a rectangle; the detected fireworks are selected using a rectangle to obtain the initial pixel region. There can be one or more initial pixel regions. The initial pixel regions are then validated to determine the accuracy of the initial fireworks detection.

[0084] The verification process involves extracting the initial pixel region and re-detecting it to confirm whether it contains smoke or fire. Specifically, the initial pixel regions in the visible light image and the corresponding initial pixel regions in the thermal imaging image are subjected to feature extraction and fusion processing to obtain region fusion features. These region fusion features are then classified to obtain the classification results for the initial pixel regions. If the classification results indicate the presence of smoke or fire in the initial pixel region, then that initial pixel region is identified as the target pixel region. Since the initial pixel region is a local area in the image, re-extracting and detecting this local region allows for the observation of more fine-grained information, thereby improving the accuracy of smoke and fire detection.

[0085] Visible light images and thermal imaging images have a pixel-level correspondence, meaning each pixel corresponds one-to-one. The initial detection identifies a corresponding initial pixel region in both images. If multiple initial pixel regions exist, the corresponding initial pixel regions in the visible light image and the thermal imaging image are extracted. Regional features are then extracted from these initial pixel regions in both images to obtain visible light region features and thermal imaging region features. These features are then fused to obtain region fusion features. Finally, the region fusion features are classified to determine whether the corresponding initial pixel region represents smoke or fire. If the classification result indicates the presence of smoke or fire in the initial pixel region, the initial pixel region is considered successfully verified and designated as the target pixel region. Otherwise, the initial pixel region is considered to have failed verification and is removed from the smoke and fire detection results, indicating a false alarm.

[0086] For example, see Figure 6 , Figure 6 This is a schematic diagram of a framework of an exemplary embodiment of the verification module shown in this application. The verification module includes, as follows: Figure 5 The multimodal feature extraction and fusion network and dimensionality reduction network are shown. Initial pixel regions from the visible light image and the thermal imaging image are input into the multimodal feature extraction and fusion network for feature extraction and fusion processing to obtain region fusion features. The structure of the multimodal feature extraction and fusion network will not be elaborated further here. Since the features of the initial pixel regions are relatively simple, a dimensionality reduction network can be added after the multimodal feature extraction and fusion network to reduce the dimensionality of the region fusion features, resulting in dimensionality-reduced region fusion features. The dimensionality reduction network includes two fully connected layers: a third fully connected layer, a batch normalization layer, a regularization layer, and a fourth fully connected layer connected sequentially. Then, an activation function is used to classify the dimensionality-reduced region fusion features to obtain the classification results of the initial pixel regions. If the initial pixel regions contain smoke, it is a positive sample; if they do not, it is a false alarm. This greatly reduces false alarms and plays a role in false alarm suppression.

[0087] Please see Figure 7 , Figure 7 This is a schematic diagram of an exemplary embodiment of the fireworks detection device shown in this application. The fireworks detection device 700 includes an acquisition module 710, a feature extraction module 720, a fusion module 730, and a detection module 740. The acquisition module 710 is used to acquire a visible light image and a thermal imaging image of the same scene, wherein the pixels in the visible light image and the thermal imaging image correspond one-to-one. The feature extraction module 720 is used to input the visible light image and the thermal imaging image into a feature extraction network respectively to obtain visible light image features and thermal imaging image features. The fusion module 730 is used to fuse the visible light image features and the thermal imaging image features to obtain target fused features. The detection module 740 is used to perform fireworks detection processing based on the target fused features to obtain fireworks detection results.

[0088] The above scheme involves a smoke detection device acquiring visible light and thermal images of the same scene, with each image containing a corresponding pixel. The visible light and thermal images are then input into a feature extraction network to obtain visible light and thermal image features, respectively. These features are then fused to obtain target fusion features. Finally, smoke detection is performed based on these target fusion features to obtain the smoke detection result. This scheme, by fusing visible light and thermal image features at the feature level, learns rich multimodal features, thereby accurately detecting and locating smoke and reducing the false alarm rate.

[0089] The functions of each module can be found in the implementation examples of the fireworks detection method, and will not be repeated here.

[0090] To implement the smoke detection method of the above embodiments, this application proposes another electronic device, please refer to [link / reference needed]. Figure 8 , Figure 8 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application.

[0091] Electronic device 800 includes memory 810 and processor 820, wherein memory 810 and processor 820 are coupled together.

[0092] The memory 810 is used to store program data, and the processor 820 is used to execute the program data to implement the smoke detection method of the above embodiment.

[0093] In this embodiment, processor 820 can also be referred to as CPU (Central Processing Unit). Processor 820 may be an integrated circuit chip with signal processing capabilities. Processor 820 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor can be a microprocessor, or processor 820 can be any conventional processor.

[0094] This application also provides a computer-readable storage medium, such as Figure 9 As shown, the computer-readable storage medium 900 is used to store program data 910, which, when executed by a processor, is used to implement the smoke detection method as described in the method embodiments of this application.

[0095] The methods involved in the embodiments of the fireworks detection method of this application, when implemented as software functional units and sold or used as independent products, can be stored in a device, such as a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0096] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for detecting smoke and fire, characterized in that, The smoke detection method includes: Acquire visible light and thermal imaging images of the same scene, wherein the pixels in the visible light image and the thermal imaging image correspond one-to-one; The visible light image and the thermal imaging image are respectively input into a feature extraction network to obtain visible light image features and thermal imaging image features; The visible light image features and the thermal imaging image features are fused to obtain the target fused features; Smoke and fire detection is performed based on the target fusion features to obtain smoke and fire detection results.

2. The smoke detection method according to claim 1, characterized in that, The feature extraction network includes a weight prediction module and a feature fusion module. The step of inputting the visible light image and the thermal imaging image into the feature extraction network to obtain visible light image features and thermal imaging image features respectively includes: The visible light image is input into the weight prediction module to obtain the first fusion weight corresponding to the visible light image and the second fusion weight corresponding to the thermal imaging image; The visible light image, the first fusion weight, the thermal imaging image, and the second fusion weight are input into the feature fusion module to obtain the visible light image features and the thermal imaging image features.

3. The smoke detection method according to claim 2, characterized in that, The weight prediction module includes a first feature extraction layer and a regression processing layer. The step of inputting the visible light image into the weight prediction module to obtain a first fusion weight corresponding to the visible light image and a second fusion weight corresponding to the thermal imaging image includes: The visible light image is input into the first feature extraction layer for feature extraction to obtain the first visible light image features; The first visible light image features are input into the regression processing layer for processing to obtain the first fusion weights corresponding to the visible light image; The second fusion weight corresponding to the thermal imaging image is determined based on the preset value and the first fusion weight.

4. The smoke detection method according to claim 2, characterized in that, The feature fusion module further includes a second feature extraction layer, a third feature extraction layer, and a feature fusion layer. The step of inputting the visible light image and its corresponding first fusion weight, and the thermal imaging image and its corresponding second fusion weight into the feature fusion module to obtain the visible light image features and the thermal imaging image features includes: The visible light image and the thermal imaging image are respectively input into the second feature extraction layer to obtain the second visible light image features and the second thermal imaging image features; The second visible light image feature, the first fusion weight, the second thermal imaging image feature, and the second fusion weight are input into the feature fusion layer for fusion processing to obtain the first fused feature; The first fused feature and the second visible light image feature are input into the third feature extraction layer to obtain the visible light image feature; The first fusion feature and the second thermal imaging image feature are input into the third feature extraction layer to obtain the thermal imaging image feature.

5. The smoke detection method according to claim 4, characterized in that, The step of inputting the second visible light image features and the corresponding first fusion weight, and the second thermal imaging image features and the corresponding second fusion weight into the feature fusion layer for fusion processing to obtain the first fused feature includes: The second visible light image features and the second thermal imaging image features are weighted according to the first fusion weight corresponding to the second visible light image features and the second fusion weight corresponding to the second thermal imaging image features to obtain weighted features; The weighted features are subjected to deep fusion processing to obtain deep fused features; The weighted features and the deep fusion features are summed to obtain the first fusion feature.

6. The smoke detection method according to claim 1, characterized in that, The smoke detection result includes the target pixel region where the smoke is located. The step of performing smoke detection processing based on the target fusion features to obtain the smoke detection result includes: The initial detection process is performed on the target fusion features to obtain the initial pixel region where the fireworks are located; The initial pixel region is verified to obtain the target pixel region of the fireworks.

7. The smoke detection method according to claim 6, characterized in that, The step of verifying the initial pixel region to obtain the target pixel region of the fireworks includes: The initial pixel regions corresponding to the visible light image and the initial pixel regions corresponding to the thermal imaging image are subjected to feature extraction and fusion processing to obtain region fusion features; The region fusion features are classified to obtain the classification result of the initial pixel region; In response to the classification result indicating the presence of smoke in the initial pixel region, the corresponding initial pixel region is determined as the target pixel region.

8. The smoke detection method according to claim 1, characterized in that, The step of acquiring visible light images and thermal imaging images of the same scene, wherein the pixels in the visible light images and the thermal imaging images correspond one-to-one, includes: Acquire initial visible light images and initial thermal images of the same scene; Matching feature points are obtained from the initial visible light image and the initial thermal imaging image; Based on the matching feature points, the initial visible light image and the initial thermal imaging image are registered to ensure that the pixels of the initial visible light image and the initial thermal imaging image correspond one-to-one, thus obtaining the visible light image and the thermal imaging image.

9. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to perform the method as claimed in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, include: The system stores program data, which, when executed by a processor, is used to implement the method as described in any one of claims 1-8.