Fire detection method, main control equipment, fire monitoring system and storage medium

By combining multi-model analysis and layered verification of visible light and infrared images, the false alarm and underreporting problems of traditional fire detection systems are solved, and efficient fire intelligent early warning is achieved.

CN120544331APending Publication Date: 2025-08-26YANTAI RAYTRON TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510792231.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Traditional fire alarm systems cannot quickly collect smoke temperature changes information, making it difficult to meet the requirements of early detection and forecasting of fires, and a single detection mode is prone to false alarms and missed alarms.

Method used

A dual-light camera is used to collect visible light images and infrared images simultaneously, detect firework targets in the visible light images through the first recognition method, and detect high-temperature abnormal areas in combination with infrared images, and verify using a pre-trained object detection model to achieve triple verification to improve detection accuracy.

Benefits of technology

While ensuring real-time, it significantly reduces the false alarm rate and improves the accuracy and reliability of fire detection. It is suitable for high-demand scenarios such as forest fire prevention and chemical plant monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544331A_ABST
    Figure CN120544331A_ABST
Patent Text Reader

Abstract

The invention provides a fire detection method, main control equipment, a fire monitoring system and a storage medium. The method comprises the following steps: acquiring a visible light image and an infrared image of the same scene; under the condition that a smoke and fire target is detected in the visible light image in a first identification mode and a high-temperature abnormal area is detected in the infrared image of the same scene, verifying whether the smoke and fire target exists in the visible light image or a fusion image of the visible light image and the infrared image by adopting a pre-trained target detection large model; and if the target detection large model verifies that the visible light image or the fused image has the smoke and fire target, determining that a fire occurs in the current scene. According to the method, through the layered detection mechanism, the real-time performance can be guaranteed, meanwhile, the false alarm rate is greatly reduced, the defects existing in a single detection mode are overcome, and the intelligent fire early warning effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of image processing technology and fire protection technology, and in particular to a fire detection method, a main control device, a fire monitoring system and a storage medium. Background Art

[0002] Fires are sudden, random, and have a short destructive time, so it is very important to detect fires in a timely manner and issue fire warnings.

[0003] Traditional fire alarm systems typically rely on infrared and smoke sensors to detect parameters such as smoke, temperature, and light generated during a fire. After signal processing, comparison, and judgment, they issue a fire alarm signal. However, their drawback is their inability to rapidly capture information about temperature fluctuations in smoke emitted by a fire, making them inadequate for early detection and forecasting of such fires. In recent years, infrared thermal imaging and visible light image detection have seen some application in fire and smoke detection. However, due to their inherent imaging and detection principles, these single detection modes are prone to false alarms and missed detections. Summary of the Invention

[0004] In order to solve the existing technical problems, the present application provides a fire detection method, a main control device, a fire monitoring system and a storage medium that can improve the accuracy of fire detection.

[0005] In a first aspect, a fire detection method is provided, the method comprising:

[0006] Collect visible light and infrared images of the same scene;

[0007] When a fireworks target is detected in the visible light image by the first recognition method and an abnormally high temperature area is detected in the infrared image of the same scene, a pre-trained large target detection model is used to verify whether the fireworks target is present in the visible light image or in a fusion image of the visible light image and the infrared image;

[0008] If the visible light image is verified by the target detection model, or the fireworks target exists in the fused image, it is determined that a fire occurs in the current scene.

[0009] In a second aspect, a main control device is provided, comprising a processor and a memory connected to the processor, wherein the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, the steps of the fire detection method described in the above embodiment are implemented.

[0010] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the fire detection method described in the above embodiment are implemented.

[0011] In a fourth aspect, a fire monitoring system is provided, comprising a dual-light camera and the main control device described in the above embodiment.

[0012] The fire detection method provided in the above embodiment utilizes a multi-model analysis and hierarchical verification method, and solves the false detection problem of the traditional single detection mode by combining visible light image analysis, infrared thermal imaging auxiliary detection and auxiliary verification of a large model. By performing smoke and fire target detection based on visible light images, non-fire scenes can be quickly filtered, and high temperature detection based on infrared images can eliminate visual similarity interference (such as reflections). By finally calling the large model for confirmation, the false alarm rate of complex scenes is reduced. Through the above-mentioned hierarchical detection mechanism, the false alarm rate can be greatly reduced while ensuring real-time performance, making up for the shortcomings of a single detection mode, so as to achieve the effect of intelligent fire warning. This fire detection method can be widely used in high-demand scenarios such as forest fire prevention and chemical plant monitoring.

[0013] The main control device, fire monitoring system and storage medium provided in the above embodiments are of the same concept as the corresponding fire detection method embodiments, and thus have the same technical effects as the corresponding fire detection method embodiments, and will not be described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 FIG. 1 is a schematic diagram of the system architecture of a fire monitoring system in one embodiment.

[0015] Figure 2 FIG. 4 is a flow chart of steps of a fire detection method in one embodiment.

[0016] Figure 3 FIG. 4 is a flow chart of steps of a fire detection method in another embodiment.

[0017] Figure 4 FIG. 4 is a flow chart of steps of a fire detection method in another embodiment.

[0018] Figure 5 FIG. 4 is a flow chart of steps of a fire detection method in another embodiment.

[0019] Figure 6 Schematic diagram of the structure of the first target detection model in one embodiment.

[0020] Figure 7 1 is a flowchart of a method for detecting whether an infrared image contains an abnormally high temperature area in one embodiment.

[0021] Figure 8 Schematic diagram of the structure of a large target detection model in one embodiment. DETAILED DESCRIPTION

[0022] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0024] In the following description, the expression "some embodiments" is involved, which describes a subset of all possible embodiments. It should be noted that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.

[0025] In the following description, the terms "first, second, and third" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first, second, and third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0026] A fire monitoring system, such as Figure 1 As shown, it includes: a dual-light camera 10, a main control device 20, an alarm device 30 and a fire extinguishing device 40.

[0027] The dual-light camera 10 is equipped with both a visible light imaging system and an infrared imaging system, capable of capturing visible light images and infrared images of the same scene simultaneously or on demand. The dual-light camera 10 is connected to the main control device 20 and sends the visible light image and infrared image to the main control device 20.

[0028] The main control device 20 includes a processor and a memory connected to the processor. The memory stores a computer program executable by the processor. When executed by the processor, the computer program implements the fire detection method of the present application and determines whether a fire has occurred in the current scene. The main control device 20 is also connected to an alarm device 30 and a fire extinguishing device 40. After the main control device 20 determines that a fire has occurred in the current scene, it controls the alarm device 30 to sound an alarm and simultaneously controls the fire extinguishing device 40 to activate.

[0029] In large venues, bifocal cameras and fire-fighting equipment 40 are deployed in each fire zone within the scene. After the master control device 20 determines a fire has occurred in a zone based on the visible and infrared images from a bifocal camera 10, it activates the corresponding fire-fighting equipment 40. This enables targeted and efficient firefighting within the fire scene, minimizing fire losses and ensuring site safety and stability. The fire-fighting equipment 40 may include high-pressure water mist nozzles and / or perfluorohexanone gas fire-fighting modules.

[0030] In one embodiment, if Figure 2 As shown, the present application provides a fire detection method, comprising:

[0031] Step 202: Capture a visible light image and an infrared image of the same scene.

[0032] Specifically, a dual-light camera is used to capture the monitored scene, obtaining both visible light and infrared images of the same scene. The dual-light camera can capture both visible light and infrared video during the monitoring period, obtaining both visible light and infrared images of the same scene. The dual-light camera can also capture both visible light and infrared images at a preset frequency during the monitoring period, for example, every 5 seconds.

[0033] In step 204 , fireworks targets are detected in the visible light image using a first recognition method, and an abnormally high temperature area is detected in the infrared image of the same scene.

[0034] By using the first recognition method to detect fireworks and fire targets in visible light images, scene fire analysis based on visual modalities can be achieved. There are two implementations of this first recognition method: one is a traditional image recognition method that uses manual feature extraction (such as color threshold segmentation and texture feature analysis) combined with a classifier to achieve fireworks and fire detection; the other is a lightweight target detection model based on deep learning.

[0035] The second method uses a convolutional neural network structure. By pre-training on images labeled with fire and smoke targets, the lightweight object detection model is equipped with the ability to identify fire and smoke targets. In one embodiment, fire and smoke targets can include smoke targets and flame targets. The first recognition method can be used to detect the presence of smoke or flame targets in visible light images. If smoke or flame targets are detected in the visible light image, a preliminary determination can be made that a fire is suspected in the current scene.

[0036] However, fire analysis based solely on visual modalities is subject to errors. For example, visual modal analysis can easily misdetect interference targets that resemble fireworks as fireworks, leading to false detections.

[0037] Infrared imaging is primarily based on infrared radiation emitted by objects. Infrared radiation is a type of electromagnetic wave with a wavelength longer than visible light and is invisible to the human eye. According to Planck's radiation law, all objects with a temperature above absolute zero (-273.15°C) emit infrared radiation. The intensity of infrared radiation emitted by an object is closely related to its temperature and surface characteristics.

[0038] An infrared detector is the core component of infrared imaging, receiving infrared radiation emitted by an object. The detector typically contains sensitive components such as thermopiles, thermistors, or quantum well infrared detectors, which convert infrared radiation into electrical signals. This electrical signal is amplified, filtered, and processed to produce an infrared image. Therefore, each pixel in an infrared image represents the temperature of a specific area on the object's surface. Thus, infrared images can directly reflect temperature information.

[0039] When a fire breaks out, it's often accompanied by a violent oxidation reaction, releasing large amounts of energy and causing a significant increase in the local ambient temperature. Based on this physical property, infrared image detection technology can be used to precisely locate high-temperature areas by capturing differences in infrared radiation emitted by surfaces, enabling preliminary fire detection.

[0040] This method has the advantage of non-contact measurement, which allows the surface temperature to be measured without touching the object being measured. However, fire detection based solely on high temperature is prone to false detection.

[0041] In this embodiment, a preliminary detection is performed by combining visible light images and infrared images. Only when the visible light detects a firework target and the infrared image detects an abnormally high temperature area is further large-scale model verification triggered. This can eliminate the interference of a single firework target or a single abnormally high temperature area and improve the accuracy of the preliminary detection.

[0042] In step 206 , a pre-trained target detection model is used to verify whether the visible light image or the fusion image of the visible light image and the infrared image contains fireworks. If so, step 208 is executed.

[0043] In this embodiment, the target detection large model is called for verification only when both the visible light image detects a fireworks target and the infrared image of the same scene detects an abnormal high temperature area. This allows the large model to only process the visible light image that meets the requirements, or the fusion image of the visible light image and the infrared image, reducing the amount of data processing required for verification by the large model and improving the real-time performance of the large model.

[0044] Large-scale object detection models are those with a deep neural network structure. These models are trained using a large amount of diverse data, resulting in large scale (large number of parameters and data requirements) and strong capabilities (high detection accuracy). These models are capable of detecting fireworks. For example, a large-scale model for detecting fireworks is trained using the Grounding DINO framework.

[0045] The target detection model can verify the presence of fireworks targets based on visible light, or based on a fusion image of visible light and infrared images.

[0046] Step 208: Determine whether a fire occurs in the current scene.

[0047] In this embodiment, if the determination result of step 206 is yes, that is, if the visible light image is verified by the target detection large model, or if there is a fire target in the fused image, it is determined that a fire has occurred in the current scene.

[0048] In this embodiment, the conclusion that a fire has occurred in the current scene is based on three verifications: detecting fireworks targets based on visible light images, detecting abnormally high temperature areas based on infrared images, and detecting fireworks targets using a large target inspection model on visible light images, or on a fusion image of visible light images and infrared images.

[0049] It should be noted that in this embodiment, the order of verifying fireworks targets based on visible light images and detecting abnormally high temperature areas based on infrared images is not limited. These two verification methods can be performed simultaneously. Alternatively, fireworks targets can be detected based on visible light images first, followed by abnormally high temperature areas based on infrared images. Alternatively, abnormally high temperature areas can be detected based on infrared images first, followed by fireworks targets based on visible light images. Only when fireworks targets are detected in the visible light image using the first recognition method, and abnormally high temperature areas are detected in the infrared image of the same scene, is the target detection model invoked for the third level of verification. Only when this third level of verification passes is a fire determined to have occurred in the current scene. In other words, a fire is determined to have occurred in the current scene only when all three levels of verification pass.

[0050] The fire detection method mentioned above utilizes multi-model analysis and hierarchical verification methods, and solves the false detection problem of the traditional single detection mode by combining visible light image analysis, infrared thermal imaging auxiliary detection and large model auxiliary verification. By performing smoke and fire target detection based on visible light images, non-fire scenes can be quickly filtered out, and high temperature detection based on infrared images can eliminate visual similarity interference (such as reflections). By finally calling the large model for confirmation, the false alarm rate in complex scenes is reduced. Through the above-mentioned hierarchical detection mechanism, while ensuring real-time performance, the false alarm rate can be greatly reduced, which makes up for the shortcomings of a single detection mode and achieves the effect of intelligent fire warning. This fire detection method can be widely used in high-demand scenarios such as forest fire prevention and chemical plant monitoring.

[0051] If no fireworks are detected in the visible light image using the first recognition method, or no abnormally high temperature areas are detected in the infrared image, or if the large target detection model verifies that no fireworks are present in the visible light image or fused image, then the scene is deemed fire-free. In other words, if any of these three verification steps fails, the scene is deemed fire-free. This reduces false detections caused by single-dimensional detection methods and, by leveraging the large target detection model, lowers the false detection rate.

[0052] In one embodiment, a method for determining whether a fireworks target is detected in a visible light image by a first recognition method and a high temperature abnormality area is detected in an infrared image of the same scene can be: using the first recognition method to detect whether there is a fireworks target in the visible light image; if a fireworks target is detected in the visible light image by the first recognition method, then detecting whether there is a high temperature abnormality area in the infrared image of the same scene.

[0053] In this embodiment, Figure 3 As shown in the figure, the conclusion that a fire has occurred in the current scene needs to undergo three verifications, namely:

[0054] First level of verification: detecting fireworks targets based on visible light images.

[0055] Second level of verification: Detection of abnormally high temperature areas based on infrared images.

[0056] The third level of verification: The target inspection large model detects fireworks targets using visible light images, or fusion images of visible light images and infrared images.

[0057] A first verification step is performed on the visible light image and infrared image of the current scene captured, using a first recognition method to detect whether there are fireworks in the visible light image. If no fireworks are detected in the visible light image using the first recognition method, it is determined that no fire has occurred.

[0058] If the first recognition method is used to detect the presence of fireworks in the visible light image, that is, if the first verification is passed, the second verification is started to detect whether there is an abnormally high temperature area in the infrared image.

[0059] If the second verification fails, it means that the first recognition method detected fireworks in the visible light image, but the infrared image of the same scene did not detect abnormally high temperature areas. This double verification eliminates visual interference, such as interference from objects that resemble fire sources. The first verification method uses the first recognition method, which may mistakenly identify objects that resemble fire sources (such as infrared clothing, reflective surfaces, lights, or water vapor) as fireworks. However, the second verification eliminates this interference and determines that there is no fire in the current scene.

[0060] If the second verification passes, it means that the first recognition method has detected fireworks in the visible light image, and the infrared image of the same scene has detected an abnormally high temperature area. If the second verification passes, the third verification is initiated. In other words, the third verification is initiated only after both the first and second verifications pass.

[0061] If the third level of verification fails, this means that the first recognition method detected fireworks in the visible light image, and the infrared image of the same scene detected an abnormally high temperature area, but the large target detection model verified that there were no fireworks in the visible light image or the fused image. In this case, the large target detection model can further eliminate complex interfering objects and determine that there is no fire in the current scene.

[0062] Large object detection models are trained using large amounts of diverse data. They are characterized by large scale (large number of parameters and data requirements) and high capabilities (high detection accuracy). Using large object detection models for third-level verification can reduce false alarm rates in complex scenarios and improve detection accuracy. For example, industrial areas may contain various equipment and materials whose appearance and thermal characteristics may resemble those of fires, resulting in false detections of fires despite passing first and second-level verification. Large object detection models can eliminate interference from similar situations and improve detection accuracy.

[0063] If the third level of verification passes, it means that the first recognition method detected the presence of fireworks in the visible light image, the infrared image of the same scene detected an abnormally high temperature area, and the large target detection model verified the presence of fireworks in the visible light image or fused image. In other words, only when all three levels of verification pass is a fire determined to have occurred in the current scene. This layered detection mechanism significantly reduces false alarms while ensuring real-time performance, compensating for the shortcomings of a single detection mode and achieving intelligent fire early warning.

[0064] In this embodiment, the detection of fireworks targets based on visible light images serves as the first level of verification, while the detection of abnormally high temperature areas based on infrared images serves as the second level of verification. Relatively speaking, the accuracy of fire determination based on visual images is higher than that of fire determination based on temperature. Therefore, using visible light image detection as the first level of verification allows for the initial screening of suspected fire scenes with relatively high accuracy. Then, using infrared image detection of abnormally high temperature areas as the second level of verification eliminates interference from fire-like objects from the high temperature dimension, thereby improving detection accuracy. Finally, after both the first and second level of verification have passed, a large-scale (large number of parameters, large data requirements) and powerful (high detection accuracy) target detection model is invoked for verification. This allows the large model to process only visible light images that have passed both the first and second level of verification, or a fusion of visible light and infrared images. This reduces the amount of data required for verification and improves detection accuracy and real-time performance.

[0065] Through the above-mentioned layered detection mechanism, the system greatly reduces the false alarm rate while ensuring real-time performance, making up for the shortcomings of a single detection mode to achieve the effect of intelligent fire warning.

[0066] In one embodiment, a fireworks target is detected in a visible light image by a first recognition method, and a high temperature abnormal area is detected in an infrared image of the same scene. The determination method can be: detecting whether there is a high temperature abnormal area in the infrared image; if the high temperature abnormal area is detected in the infrared image, using the first recognition method to detect whether there is a fireworks target in the visible light image of the same scene.

[0067] In this embodiment, Figure 4 As shown in the figure, the conclusion that a fire has occurred in the current scene is based on three verifications:

[0068] The first level of verification: Detecting abnormally high temperature areas based on infrared images.

[0069] Second level of verification: Detecting fireworks targets based on visible light images.

[0070] The third level of verification: The target inspection model detects fireworks targets using visible light images, or fusion images of visible light images and infrared images.

[0071] Figure 3 The approach shown here first uses visible light detection to perform a broad preliminary screening of the scene, eliminating most non-firework targets and narrowing the area of ​​interest. For potential firework areas detected by visible light, infrared detection is then used for high-temperature verification. This reduces false alarms caused by interference from other heat sources when using infrared detection directly, improving detection accuracy and reliability.

[0072] Infrared detection is mainly based on thermal radiation. In addition to flames, many other heat sources may also be detected, such as heat dissipation from industrial equipment, automobile exhaust emissions, solar water heaters, etc. These heat sources may appear as high-temperature areas in infrared images, but they are not caused by fire. Therefore, Figure 4 In the approach shown, when infrared high-temperature detection is performed first, a large number of non-fire high-temperature targets may be detected, resulting in an increase in the number of targets requiring further verification. When these infrared-detected high-temperature areas are subsequently verified as fire targets using visible light detection, the large number of false high-temperature targets significantly increases the amount of data required for processing and analysis. This undoubtedly increases the system burden, prolongs the entire detection and verification process, and thus affects the real-time performance of fire detection.

[0073] therefore, Figure 3 Compared with the multiple verification method shown in the embodiment Figure 4 The verification method shown in the embodiment can more effectively reduce the interference of false targets and reduce the workload of the second verification, thereby better ensuring the real-time and accuracy of fire detection. Figure 5 As shown, in one embodiment, a fireworks target is detected in a visible light image by a first recognition method, and a determination method for detecting an abnormally high temperature area in an infrared image of the same scene can be: using the first recognition method to detect whether there is a fireworks target in the visible light image, and at the same time detecting whether there is an abnormally high temperature area in the infrared image of the same scene.

[0074] This approach simultaneously detects fireworks and fire targets based on visible light images and high-temperature areas based on infrared images. This allows the first recognition method to detect fireworks and fire targets in visible light images, and also detect abnormally high-temperature areas in infrared images of the same scene. However, this approach requires detecting both visible light and infrared images of the scene, significantly increasing the amount of data to be processed and analyzed. This undoubtedly increases the burden on the system, prolongs the entire detection and verification process, and thus affects the real-time performance of fire detection.

[0075] It can be seen that priority is given to Figure 3 The verification method shown here uses the first recognition method to quickly screen visible light images to reduce computational load, followed by progressively more refined verification using infrared and large-scale models. This balances real-time performance with accuracy, reduces false alarms and missed alerts, and provides strong technical support for fire prevention and control.

[0076] By improving the accuracy of fireworks detection based on visible light images, the problem of missed detection can be reduced. This can be achieved by using a target detection model with higher recognition accuracy.

[0077] In one embodiment, the first identification method is used to identify whether a fireworks target exists in the visible light image, including using a first target detection model to identify whether a fireworks target exists in the visible light image.

[0078] In one embodiment, the structure of the first target detection model is as follows Figure 6 As shown, it includes an input layer 60, a backbone network 61, a neck network 62 and a prediction module 63.

[0079] Among them, the backbone network 61 includes a feature extraction module 610, a feature map generation module 611, a feature fusion module 612, a pooling module 612 and an attention module 614 which are connected in sequence.

[0080] The neck network 62 includes a primary fusion module 620 , a semantic extraction module, and a fusion module 622 , which are connected in sequence.

[0081] Specifically, the first identification method is used to identify whether a fireworks target exists in the visible light image, including:

[0082] Input the visible light image into a pre-trained lightweight first object detection model; the first object detection model includes an input layer 60, a backbone network 61, a neck network 62, and a prediction module 63 connected in sequence;

[0083] Preprocessing the visible light image through the input layer 60;

[0084] The preprocessed visible light image is input into the backbone network 61. The backbone network includes a feature extraction module 610, a feature map generation module 611, a feature fusion module 612, a pooling module 613 and an attention module 614 connected in sequence. The feature extraction module 611 performs slicing and convolution operations on the visible light image to extract a first preliminary feature map of the visible light image. The feature map generation module 611 performs convolution operations on the first preliminary feature map to generate a second preliminary feature map. The feature fusion module 612 performs convolution and residual processing to fuse the features of the second preliminary features at different levels to obtain an enhanced feature map. The pooling module 613 performs multi-scale pooling processing on the enhanced feature map to generate feature maps of different scales. The attention module 614 uses an attention enhancement mechanism to process the feature maps of different scales to obtain multi-scale image features of the visible light image.

[0085] The multi-scale image features are input into the neck network 62, which includes an initial fusion module 620, a semantic extraction module 621 and a fusion module 622 connected in sequence. The multi-scale image features are preliminarily fused through the initial fusion module 620 to obtain a first fusion feature; the semantic extraction module 621 generates a multi-scale feature map with semantic information based on the first fusion feature through a top-down path and lateral connections; the fusion module 622 fuses the multi-scale feature map with semantic information through a bottom-up path to obtain a fusion feature with positioning information and semantic information.

[0086] The prediction module 63 predicts the position and category of the fireworks target in the visible light image based on the fused features.

[0087] In this embodiment, a first target detection model is utilized. The first target detection model includes, in sequence, an input layer 60, a backbone network 61, a neck network 62, and a prediction module 63, which are respectively responsible for input image preprocessing, feature extraction, feature fusion, and output detection information. The input of each module is derived from the output of the previous module. The input layer 60 primarily preprocesses the image. Preprocessing may include scaling the input image to the network input size, performing normalization operations, and using data augmentation operations such as Mosaic. The backbone network 61 processes the input image through a feature extraction module, a feature map generation module, a feature fusion module, a pooling module, and an attention module, gradually reducing the size of the feature map while increasing the number of channels to retain and extract important features in the image. In the neck network 62, an initial fusion module designed based on CSPnet is adopted to enhance the network's feature fusion capabilities. The semantic extraction module can capture strong semantic features from the top down, while the fusion module conveys strong positioning features from the bottom up. By combining these two modules, the fusion of shallow graphic features and deep semantic features is achieved, effectively completing the functions of fireworks target detection and positioning.

[0088] It is understood that in order for the first object detection model to have the capabilities of fireworks detection and location, a neural network model must be pre-trained using the fireworks target detection dataset. This neural network model can adopt a convolutional neural network architecture. The images in the fireworks target detection dataset are visible light images. The annotation file includes the category of the object in the image and the bounding box coordinates of the fireworks target.

[0089] Through training, the first target detection model is able to detect fireworks targets and their locations from visible light images.

[0090] In one embodiment, the method of detecting whether there is an abnormally high temperature area in the infrared image is as follows: Figure 7 Shown, including:

[0091] Step 702 : According to a preset mapping relationship between grayscale value and temperature, the grayscale value of each pixel in the infrared image is converted into temperature to obtain the temperature value of each pixel in the infrared image.

[0092] The grayscale value of each pixel in an infrared image is related to the intensity of infrared radiation emitted by the object's surface, which in turn is directly related to the object's temperature. The intensity of an object's radiation increases with increasing temperature. Therefore, by calibrating the infrared camera, we can pre-establish a mapping between grayscale value and temperature.

[0093] By using the preset mapping relationship between grayscale value and temperature, the grayscale value of each pixel point in the infrared image can be converted into temperature to obtain the temperature value of each pixel point in the infrared image.

[0094] For example, the relationship expression of the mapping relationship between grayscale value and temperature is as follows:

[0095] T(x,y)=kI(x,y)+b+ΔT env

[0096] Where T(x,y) is the calibration temperature, I(x,y) is the pixel value, k,b are the calibration parameters, ΔT env Compensation values ​​for ambient temperature, humidity, distance, etc.

[0097] Step 704 : extracting high-temperature suspected areas based on the temperature value of each pixel in the infrared image.

[0098] Specifically, image segmentation (e.g., thresholding and edge detection) is used to extract suspected high-temperature areas. This extraction has the following benefits: It can eliminate isolated noise and single-pixel high-temperature points (e.g., sensor noise); verify target continuity, as real fire sources typically have spatial extension (area > 10px); and reduce computational complexity, allowing subsequent temperature analysis to be performed only on valid areas.

[0099] Step 706: Obtain the maximum temperature of the suspected high-temperature area.

[0100] Step 708: Determine whether the maximum temperature is greater than the temperature threshold of the current scene. If so, proceed to step 710: Determine whether there is an abnormally high temperature area in the infrared image.

[0101] The temperature threshold can be adjusted based on the actual application scenario. The maximum temperature of the suspected high-temperature area is compared with the safety temperature threshold set for the current scenario. If the maximum temperature is greater than the temperature threshold for the current scenario, the suspected high-temperature area is determined to be a true high-temperature area (such as a fire source). This can help reduce false alarms of low-temperature heat sources (such as human body temperature of 37°C).

[0102] If the maximum temperature is less than the temperature threshold of the current scene, the suspected high-temperature area is excluded as a real high-temperature area.

[0103] In this embodiment, image segmentation is performed to extract suspected high-temperature areas. The maximum pixel temperature within the suspected area is then calculated and compared with a temperature threshold to determine whether it represents a fire source. This utilizes infrared thermal imaging technology to detect temperature anomalies in a scene, complementing visual image-based smoke and fire detection.

[0104] Specifically, single sensors often have limitations. For example, relying solely on visual imagery for fire and smoke detection can severely impact detection at night or in dimly lit environments due to insufficient visible light. While relying solely on infrared thermal imaging can detect high-temperature areas from a temperature perspective, interference can be generated by non-fire sources such as sunlight reflection and high-temperature equipment. The two complement each other: infrared thermal imaging compensates for the lack of visible light at night, while visual image detection effectively identifies interference from non-fire sources. Together, they address the limitations of single-sensor fire detection.

[0105] In one embodiment, a method for training a large object detection model includes:

[0106] Obtain a pre-trained initial visual model, a fireworks target detection dataset, and an annotation file; the images in the fireworks target detection dataset are visible light images or fused images of visible light images and infrared images; the annotation file includes the category of the target in the image and the bounding box coordinates of the fireworks target;

[0107] The pre-trained initial visual large model is fine-tuned according to the fireworks target detection dataset to obtain the target detection large model; the target detection large model is used to predict the category of the target in the visible light image or the fused image, and the bounding box coordinates of the fireworks target.

[0108] The target detection large model of this application is obtained by fine-tuning the pre-trained initial visual large model using the fireworks target detection dataset, which can improve development efficiency. By fine-tuning the pre-trained initial visual large model, the target detection large model is obtained, which can quickly adapt to specific scenarios. Specifically, by fine-tuning on a large number of annotated fireworks datasets, the model can learn the characteristics of fireworks in different scenes (such as night, day), lighting (such as strong light, weak light), and angles (such as front, side). Through dynamic anchor frame allocation, the model can locate the position of fireworks more accurately and reduce false detections (such as misjudging red lights as fireworks).

[0109] The pre-trained initial visual large model can be any one of the Grounding DINO model, DINO model, Qianwen large model, MiniCPM and Florence-2.

[0110] In one embodiment, the structure of the pre-trained initial visual model is as follows: Figure 8 As shown in the figure, the pre-trained initial visual large model consists of an image backbone network, a text backbone network, a feature enhancer, a language-guided query selection, and a cross-modal decoder, which are responsible for extracting rich visual features, extracting text features, feature-enhancing visual information and text information, and dynamically selecting language information or generating visual queries to focus on visual areas related to language descriptions, fuse information from different modalities (such as vision and language), and decode it into specific outputs (such as text, detection boxes, or classification results).

[0111] The image backbone network uses a powerful convolutional neural network (CNN) or a Vision Transformer (ViT) variant, such as the Swin Transformer. It outputs a series of feature maps of varying resolutions, encompassing information ranging from low-level textures to high-level semantics. The text backbone network typically uses a pre-trained Transformer model, such as BERT or its variants. This model deeply understands the semantics of text and the relationships between words, outputting a vector representation of each word (token) and a comprehensive representation of the entire sentence.

[0112] The feature enhancer is used for cross-modal fusion. It has the following functions:

[0113] Image feature enhancement uses text information to guide image feature extraction, allowing the model to focus more on image regions relevant to the text description. Text feature enhancement uses image information to fine-tune the understanding of text features, particularly when processing referential or spatial relationships. Multi-scale feature fusion fuses feature maps from different layers of the image backbone network and allows these fused features to interact with text information. This is typically achieved through a cross-modal attention mechanism.

[0114] Language-guided query selection, which selects features that are more relevant to the input text as decoder queries based on the cosine similarity of image and text features;

[0115] The cross-modal decoder consists of multiple decoder layers, each of which uses self-attention on queries and cross-attention between queries and fused features. Each object query ultimately outputs a bounding box and a corresponding confidence score through a feed-forward neural network (FFN).

[0116] The pre-trained Initial Vision Large Model utilizes a dual encoder-single decoder architecture and over 20 million images for training, significantly improving its ability to recognize objects of unknown categories. The image and text backbone networks of the pre-trained Initial Vision Large Model typically utilize model weights pre-trained on ultra-large-scale datasets. These pre-trained models have already acquired a wealth of general vision and language knowledge. Fine-tuning the pre-trained Initial Vision Large Model on this basis allows it to more quickly and effectively learn the correspondence between images and text, resulting in stronger generalization capabilities.

[0117] Specifically, by fine-tuning the model on a large dataset of annotated fireworks detection, the model learns the characteristics and variations of fireworks in different scenes, lighting conditions, and angles. This allows for more accurate identification of fireworks and reduces false detections in real-world detections. For example, in complex backgrounds, the fine-tuned model can more accurately distinguish fireworks from non-firework objects of similar shape or color.

[0118] Fine-tuning based on the pre-trained model eliminates the need to train the entire model from scratch. Instead, it only requires optimized training for the fireworks target detection task based on the original model. This allows for rapid acquisition of a fireworks detection model with improved performance, accelerating the development and deployment of fireworks detection projects.

[0119] The images in the fireworks target detection dataset can be visible light images or a fusion of visible light and infrared images. To improve the generalization of the model, data augmentation operations such as random cropping, rotation, and flipping can be performed on the original training data. This enhances the training data and improves the generalization of the model.

[0120] The pre-trained initial visual large model itself has a certain zero-shot detection capability, but fine-tuning can enable it to quickly learn key features on a limited fireworks dataset. Compared with training the model from scratch, it can achieve better detection results even with less data, greatly reducing the dependence on large-scale labeled data and shortening the model development cycle and cost.

[0121] In one embodiment, the samples in the fireworks target detection dataset are fused images of visible light and infrared images. The samples can also be labeled with fire severity. For example, fire severity can be determined and labeled based on quantitative indicators such as fire area, flame height, and temperature range. The fireworks target detection dataset is used to fine-tune the pre-trained initial visual model, enabling it to identify fire severity.

[0122] After a fire is detected, the visible light image and infrared image are fused in real time to produce a fused image. In complex environmental conditions, such as dense smoke and insufficient or excessive lighting, the quality and information integrity of single-modality images can be affected. After fusion, even if the information from one modality is limited, the other can still provide effective information, ensuring the stable operation of the detection system. The fused image contains more comprehensive information, facilitating a deeper understanding and analysis of the scene. For example, in a fire scene assessment, visible light images can be used to observe building structures and the path of fire spread, while infrared images can identify high-temperature areas and the distribution of heat sources. The combination of the two can more accurately determine the severity of a fire.

[0123] Among them, the visible light image and the infrared image can be registered first to align the two in space. Then, the grayscale values ​​or color values ​​of the corresponding pixels in the visible light image and the infrared image can be integrated according to certain rules or algorithms (such as weighted averaging method, principal component analysis method, etc.) to generate a fused image.

[0124] The large object detection model is used to identify the severity of the fire based on the fused image. Changes in fire severity are then used to determine development trends. For example, if fire severity increases within a short period of time, the development trend can be determined to be rapid. If the fire severity decreases, the fire can be determined to be under control.

[0125] By reporting fire severity and trends to emergency command and other relevant management platforms, relevant personnel can quickly formulate appropriate firefighting strategies, rationally allocate rescue resources, improve the efficiency and effectiveness of fire emergency response, and minimize fire losses. Reports can include real-time fire severity levels, specific numerical indicators, and detailed descriptions of development trends, providing firefighters and emergency decision-makers with comprehensive, accurate, and real-time fire information.

[0126] The fire detection method of the present application collects visible light images and infrared images of the same scene and uses the visible light images and infrared images to perform fire detection.

[0127] Specifically, visible light is used to capture the visual characteristics of flames / smoke, and infrared images are used to detect temperature anomalies, forming a complementary verification mechanism.

[0128] Specifically, the approach uses a tiered detection process, including triple validation.

[0129] First verification: The first object detection model is used to detect flames / smoke in visible light images in real time and quickly filter out non-fire scenes.

[0130] Second level of verification: For infrared images of the same scene that passed the first level of verification of visible light images, detect whether there are any areas of abnormally high temperature, and eliminate interference from areas that look similar but do not have high temperatures (such as light reflections and clouds).

[0131] Third-level verification: If fireworks are detected in the visible light image after first and second-level verification, and an abnormally high temperature area is detected in the infrared image of the same scene, a pre-trained large-scale object detection model is used to verify the presence of fireworks in the visible light image, or in a fusion of visible light and infrared images. This further reduces the false alarm rate.

[0132] This method has the following effects:

[0133] 1. Combining visible light imaging (for detecting flames / smoke) and infrared thermal imaging (for detecting temperature anomalies) to build dual-modal data input and achieve multi-level verification through a layered detection process (three-tier verification mechanism).

[0134] 2. Use visible light to capture visual features and infrared to detect temperature anomalies. The two work together to solve the limitations of a single sensor (such as insufficient visible light at night or interference from high-temperature non-fire sources).

[0135] 3. Layered filtering. A lightweight model is used to quickly and initially screen visible light images, reducing computational load. Infrared and large-scale models are then used for progressively more refined verification, balancing real-time performance with accuracy, reducing false alarms and missed alerts, and providing strong technical support for fire prevention and control.

[0136] 4. Through the three core innovations of multimodal data fusion, layered detection process and lightweight-large model collaboration, an efficient and robust fire detection system has been built, which significantly improves the detection accuracy and real-time performance, while solving the problem of false detection in the traditional single detection mode. It is suitable for high-demand scenarios such as forest fire prevention, industrial safety, and urban security.

[0137] In another aspect, the present application also provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the various processes of the fire detection method embodiment and achieves the same technical effects. To avoid repetition, the details are not described here. The computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0138] On the other hand, an embodiment of the present application further provides a main control device, which includes a processor and a memory connected to the processor, wherein the memory stores a computer program that can be executed by the processor. When the computer program is executed by the processor, the various processes of the fire detection method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0139] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0140] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0141] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A fire detection method, characterized in that: The method comprises: Collect visible light and infrared images of the same scene; When a fireworks target is detected in the visible light image by the first recognition method and an abnormally high temperature area is detected in the infrared image of the same scene, a pre-trained large target detection model is used to verify whether the fireworks target is present in the visible light image or in a fusion image of the visible light image and the infrared image; If the visible light image is verified by the target detection model, or the fireworks target exists in the fused image, it is determined that a fire occurs in the current scene.

2. The fire detection method according to claim 1, characterized in that: The method further comprises: If the fireworks target is not detected in the visible light image through the first recognition method, or the abnormal high temperature area is not detected in the infrared image, or the fireworks target is not present in the visible light image or the fused image as verified by the target detection large model, it is determined that no fire has occurred in the current scene.

3. The fire detection method according to claim 1, characterized in that: The method for determining that a firework target is detected in the visible light image by the first recognition method and an abnormally high temperature area is detected in the infrared image of the same scene includes any one of the following methods: the first method: Using a first recognition method to detect whether there is a fireworks target in the visible light image; If a fireworks target is detected in the visible light image by the first recognition method, detecting whether there is an abnormally high temperature area in the infrared image of the same scene; The second type: Detect whether there are abnormally high temperature areas in infrared images; If an abnormally high temperature area is detected in the infrared image, a first recognition method is used to detect whether there is a fireworks target in the visible light image of the same scene; The third type: The first recognition method is used to detect whether there are fireworks targets in the visible light image, and at the same time, the infrared image of the same scene is detected whether there are high temperature abnormal areas.

4. The fire detection method according to claim 1 or 3, characterized in that: Identifying whether a fireworks target exists in the visible light image using a first identification method includes: Inputting the visible light image into a pre-trained lightweight first object detection model; the first object detection model includes an input layer, a backbone network, a neck network, and a prediction module connected in sequence; Preprocessing the visible light image through the input layer; Inputting the preprocessed visible light image into the backbone network, and extracting multi-scale image features of the visible light image through the backbone network; Inputting the multi-scale image features into the neck network, and processing the neck network to obtain fused features that fuse positioning information and semantic information; The prediction module predicts the position and category of the fireworks target in the visible light image based on the fused features.

5. The fire detection method according to claim 1 or 3, characterized in that: The method of detecting whether the infrared image contains an abnormally high temperature area includes: According to a preset mapping relationship between grayscale value and temperature, the grayscale value of each pixel of the infrared image is converted into temperature to obtain the temperature value of each pixel in the infrared image; Extracting high-temperature suspected areas according to the temperature value of each pixel in the infrared image; Obtaining the maximum temperature of the suspected high-temperature area; If the maximum temperature is greater than the temperature threshold of the current scene, it is determined that there is a high temperature abnormality area in the infrared image.

6. The fire detection method according to claim 1, characterized in that: Methods for training the target detection model include: Obtain a pre-trained initial visual large model, a fireworks target detection dataset, and an annotation file; the images in the fireworks target detection dataset are visible light images or fused images of visible light images and infrared images; the annotation file includes the category of the target in the image and the bounding box coordinates of the fireworks target; The initial visual large model is fine-tuned according to the fireworks target detection dataset to obtain the target detection large model; the target detection large model is used to predict the category of the target in the visible light image or the fused image, and the bounding box coordinates of the fireworks target.

7. The fire detection method according to claim 1, 2, 3 or 6, characterized in that: The method further comprises: After a fire is detected, the target detection model is called to determine the severity of the fire based on a fusion image of the real-time visible light image and the infrared image of the current scene; Determining trends based on changes in the severity of said fires; Report the severity of the fire and / or the development trend of the fire.

8. A main control device, comprising a processor and a memory connected to the processor, characterized in that: The memory stores a computer program that can be executed by the processor, and when the computer program is executed by the processor, the steps of the fire detection method according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the fire detection method according to any one of claims 1 to 7 are implemented.

10. A fire monitoring system comprising a dual-light camera and the main control device according to claim 8; The dual-light camera is connected to the main control device.

11. The fire monitoring system according to claim 10, characterized in that: The system further comprises an alarm device and / or a fire extinguishing device connected to the main control device; The main control device controls the alarm device to sound an alarm and / or controls the fire extinguishing device to start in response to the fire target detected by the real-time image.

Citation Information

Cited By

  • Forest fire monitoring method, device and equipment and readable storage medium

    CN121505752A