A method for monitoring fire in urban underground comprehensive pipe gallery
By constructing a fire infrared monitoring image dataset and improving the YOLOv8n model, combined with infrared and visible light cameras, the problems of complex lighting and large model computation in underground utility tunnel fire detection were solved, achieving high-precision detection and edge deployment of early-stage small-target fires.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JILIN JIANZHU UNIVERSITY
- Filing Date
- 2026-04-01
- Publication Date
- 2026-07-24
AI Technical Summary
Fire monitoring in urban underground utility tunnels faces challenges such as poor lighting and complex environments, making it difficult to detect early-stage small-target fires. Existing models also suffer from large parameter counts, high computational demands, and difficulty in edge deployment, resulting in poor detection accuracy.
We constructed a fire infrared monitoring image dataset, improved the YOLOv8n model, introduced spatial-to-depth convolution and attention mechanisms, adopted a lightweight convolution method, and combined it with the TensorRT framework to accelerate model inference. We then used infrared and visible light cameras for fire detection.
It enables accurate detection of early-stage small-target fires, improves fire detection accuracy, and can be deployed in real time on edge devices to ensure accurate assessment of the fire scene.
Smart Images

Figure CN122454497A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of safety monitoring technology for urban underground utility tunnels, and specifically relates to a fire monitoring method for urban underground utility tunnels. Background Technology
[0002] Urban underground utility tunnels contain a large amount of fire load. Furthermore, due to the narrow spatial structure and semi-enclosed conditions of the tunnels, a fire could lead to serious casualties and property damage.
[0003] Visible light imaging is a commonly used method for fire monitoring in urban underground utility tunnels. However, due to poor lighting and complex environments, including unavoidable interference from water mist and dust, detecting the color, shape, and deformation characteristics of fires in underground utility tunnels is challenging. Furthermore, while visible images are effective for identifying open flames, they are less effective at detecting early-stage smoldering fires and small-target flames.
[0004] In addition, existing fire detection models generally have a large number of parameters and a large amount of computation, which makes it difficult to deploy them at the edge when applied to fire detection in urban underground utility tunnels, and also results in poor detection accuracy. Summary of the Invention
[0005] The purpose of this invention is to provide a method for monitoring fires in urban underground utility tunnels, which can detect early-stage small-target fires and improve the accuracy of fire detection in urban underground utility tunnels.
[0006] The technical solution provided by this invention is as follows: A method for fire monitoring in urban underground utility tunnels includes the following steps: Step 1: Construct a fire infrared monitoring image dataset; Step 2: Improve the YOLOv8n model and train the improved YOLOv8n model using the fire infrared image dataset to obtain a fire monitoring model for urban underground utility tunnels. The methods for improving the YOLOv8n model are as follows: pruning the P5 detection layer in the detection layer structure of the YOLOv8n model, introducing spatial-to-depth convolution in the feature extraction network, replacing the Conv in the Neck network with GSConv, replacing the C2f module in the Neck network with the VoV-GSCSP module, and introducing an attention mechanism in the feature extraction network. Step 3: Acquire infrared images of the monitoring environment in the underground utility tunnel using a thermal imaging camera, and input the infrared images of the monitoring environment into the urban underground utility tunnel fire monitoring model. The urban underground utility tunnel fire monitoring model outputs the fire detection results and marks the flame location. If the underground utility tunnel fire monitoring model detects a fire, it will capture a visible light image of the fire scene using a visible light camera and transmit it to the monitoring host computer.
[0007] Preferably, the spatial-to-depth convolution consists of a spatial-to-depth layer SDP and a non-strut convolutional layer Conv.
[0008] Preferably, the spatial-to-depth convolution is positioned between the first two C2f modules in the feature extraction network.
[0009] Preferably, an attention mechanism is introduced in the layer preceding the SPPF spatial pyramid pooling module of the feature extraction network.
[0010] Preferably, the attention mechanism employs the SimAM attention mechanism.
[0011] Preferably, in step three, the thermal imaging camera has a response range of 8μm to 14μm in the infrared band.
[0012] Preferably, in step one, the fire infrared image dataset consists of infrared flame images taken on-site, as well as non-fire infrared images selected from the M3FD dataset and LLVIP dataset retrieved from the internet.
[0013] Preferably, it also includes using the TensorRT framework to accelerate the inference process of the urban underground utility tunnel fire monitoring model.
[0014] The beneficial effects of this invention are: The fire monitoring method for urban underground utility tunnels provided by this invention can detect early-stage small-target fires, improving the accuracy of fire detection in urban underground utility tunnels; and by obtaining fire location images through model detection, it can accurately determine the situation at the fire scene. Attached Figure Description
[0015] Figure 1 This is a framework diagram of spatial-to-depth convolution as described in this invention.
[0016] Figure 2 This is a schematic diagram of the GSConv structure described in this invention.
[0017] Figure 3 This is a schematic diagram of the SimAM parameterless attention mechanism described in this invention.
[0018] Figure 4 This is a performance comparison chart of the different attention mechanisms described in this invention.
[0019] Figure 5 This is a framework diagram of the improved YOLOv8n model described in this invention. Detailed Implementation
[0020] The present invention will now be described in further detail so that those skilled in the art can implement it based on the description.
[0021] This invention provides a method for fire monitoring in urban underground utility tunnels, aiming to solve the problem of difficult detection of small-target fires in the early stages of fires or when pipeline conditions are poor. Utilizing the advantages of infrared visual imaging technology in low-light conditions, an infrared fire detection dataset is established for model training, constructing a lightweight fire monitoring model for urban underground utility tunnels that balances high detection accuracy and fast detection speed for deployment at edge devices. Simultaneously, when the fire monitoring model detects a fire, a visible light camera captures images of the fire location and transmits them to the host computer at the monitoring center. This allows monitoring personnel to further determine the actual situation at the fire scene, avoid model misjudgments, and organize rescue efforts based on the fire situation. The following is a detailed description of the fire monitoring method for urban underground utility tunnels provided by this invention.
[0022] I. Constructing a Fire Infrared Monitoring Image Dataset To ensure the superior performance of the final trained model, this invention independently builds an infrared vision flame image acquisition environment to obtain a large number of infrared small target flame images for training.
[0023] The infrared vision equipment mainly includes a dual-spectrum integrated camera, combustion discs of different sizes, and other auxiliary equipment. The dual-spectrum integrated camera has two lenses: one for visible light and one for thermal imaging. The infrared thermal imaging lens has a resolution of 256×192 pixels and a response range of 8μm-14μm in the infrared band, which meets the requirements for acquiring infrared small target flame images in this invention's dataset. In use, the visible light lens and the infrared thermal imaging lens can be activated independently.
[0024] To better simulate the internal environment of underground utility tunnels, this invention constructed an infrared small-target flame image acquisition environment. A 6m long, 3.5m wide, and 3m high enclosed indoor environment was selected as the experimental site. All fire experiments were conducted with fire extinguishers provided and under the on-site supervision of staff to ensure the safety of the experimental process. Alcohol was chosen as the ignition source for the small targets, along with smoldering wet wood and other combustible materials. Environmental interference factors such as light sources and heat sources were incorporated during the combustion process, and the effects of dust, water mist, and smoke were considered. For experimental safety, fine dust was selected as the dust interference factor.
[0025] Each experiment maintained a combustion time of approximately 3 minutes. The interfering heat source was an electric kettle that was gradually heated and kept at a constant temperature, simulating both fixed and mobile heat sources in the environment. The dual-spectrum integrated camera was installed at a stable height of approximately 2.1 meters. During the experiments, 8cm×8cm and 16cm×16cm combustion plates were used for ignition tests to simulate the early stages of a fire, including small, spreading fire sources. Simultaneously, the distance of the combustion plate was controlled to vary in all directions, increasing the diversity of fire scenes captured by the camera's monitoring perspective.
[0026] In the early stages of a fire, due to the presence of complex factors such as heat sources, ambient heat sources, dust, or water mist, small point-like fire sources or spreading fire sources exhibit diverse fire characteristics in infrared images captured by surveillance cameras. By capturing infrared videos of small target flames in different fire scenarios, to avoid the problem of high similarity in fire features between adjacent frames, the videos were extracted into flame images using a frame-by-frame capture method. One image was captured at intervals of three frames. After similarity deduplication, a total of 9000 images were obtained from the above experimental conditions as positive samples for the dataset. LabelImg software was used to label the images.
[0027] LabelImg is a rectangular bounding box annotation tool commonly used for image annotation in target recognition and detection. Labels can be generated in YOLO, PASCAL VOC, and CreateML formats. Labelme is a polygonal bounding box annotation tool that accurately marks target outlines, generating labels in JSON, VOC, and COCO formats. To improve the efficiency of manual image annotation in the dataset and reduce the time spent converting data label formats for subsequent comparative experiments, LabelImg software was chosen to annotate the images in the dataset. A target category was set, labeled "fire," and the annotations were saved as a .txt file used by the YOLO model, including information such as the target category, quantity, and the location of fire target features in the image.
[0028] Since the background information during the acquisition of infrared small target flame images is relatively simple, in order to enhance the generalization ability of the model in complex backgrounds, we searched the M3FD (Multi-Modal Multi-Scale Fusion Dataset) and LLVIP (Low-Light Visible-Infrared Pair Dataset) infrared datasets through the network. We selected 9,000 images from a large number of non-flaming thermal infrared images as negative samples of the dataset and labeled them as empty txt files to help the model identify which features are effective and which are irrelevant during training.
[0029] The final fire infrared image dataset contains 18,000 images, divided in an 8:1:1 ratio. The dataset was augmented using Mosaic9 data augmentation and will be used for model training, validation, and testing experiments.
[0030] According to the definition of relative scale, a small target is generally considered to have a width-to-height ratio of less than 0.1 to the entire image. In the fire infrared image dataset constructed in this invention, over 70% of the image annotations conform to the definition of a small target, indicating that small targets (flames) are predominant, thus meeting the dataset training requirements for the infrared visual detection model of this invention.
[0031] II. Constructing a Fire Monitoring Model for Urban Underground Utility Tunnels When the model needs to be applied in practice, the large number of parameters and computational load still need to be considered. In addition, the detection performance in infrared fire scenarios needs to be improved. This invention improves the YOLOv8n model to construct a lightweight urban underground integrated pipe gallery fire monitoring model, which can be deployed and applied to various hardware devices and edge devices while ensuring detection accuracy and speed.
[0032] (1) Improved detection layer structure Since the flames of small targets in infrared images are relatively small, the P5 large target detection layer used to detect 20×20 feature maps becomes a redundant layer. Pruning this layer is an effective way to reduce the number of parameters.
[0033] To demonstrate the effectiveness of the P5 layer detection, a comparative experiment was conducted. In the first group, the detection layer structure remained unchanged; in the second group, the P5 detection layer was trimmed; in the third group, a P2 detection layer was added; and in the fourth group, a P2 detection layer was added while the P5 detection layer was trimmed.
[0034] The results show that after pruning the P5 large target detection layer in the second group of experiments, the average accuracy (mAP) improved by 0.2%, the number of parameters decreased by 33%, the computational cost decreased by 0.8 GFLOPs, and the detection speed reached 80 FPS. The method of adding a P2 small target detection layer in the third group of experiments also achieved an mAP of 97.3%, with a similar number of parameters to the basic model, but the computational cost was the highest among the four groups. However, the fourth group of experiments, which combined shallow and deep feature maps, actually saw a 0.6% decrease in accuracy. This indicates that increasing network depth while pruning redundant detection layers can easily introduce irrelevant interference information during the feature fusion and output detection of global context information. This not only makes the method unsuitable for the fire infrared monitoring image dataset of this invention but also affects the improvement of the model's detection capability. Clearly, the second group of experiments, which pruned the redundant P5 large target detection layer, had significantly fewer parameters and less computational cost than the other groups, and the model had the fastest detection speed, making it more suitable for the fire infrared monitoring image dataset of this invention.
[0035] (2) Improvement of feature extraction network Urban underground utility tunnels often suffer from poor lighting, and during a fire, the environment may contain interfering factors such as light sources, heat sources, dust, or water mist. In the event of a fire, infrared thermal radiation is first absorbed or refracted by interfering media in the tunnel atmosphere. This attenuated thermal radiation results in poor image quality for early infrared small target flames. Furthermore, infrared thermal imaging itself has low resolution, and the target area for early flame features is small, limiting the information available for fire detection models to learn. When early infrared small target flame features coexist with numerous interfering targets such as light sources and heat sources in the same environmental background, large interfering targets form fire-like spots in the infrared thermal imager, generating redundant feature information similar to the small target flame representation. This makes it easy for the model to lose early point-like or spreading small target flame features during continuous downsampling of the feature extraction network, resulting in the early small target flames not being detected correctly.
[0036] When performing traditional convolution operations on input feature maps, the stride value is usually set to be greater than 1 to perform downsampling. However, this operation may lead to partial overlap or loss of downsampled information when processing low-resolution images and small target features. To adapt to the correct learning process of infrared small target flame features and solve the problem of interference from redundant feature information, the Space-to-Depth Convolution (SPD-Conv) is introduced into the YOLOv8n model. SPD-Conv consists of a Space-to-Depth (SPD) layer and a non-strut convolutional layer (Conv).
[0037] Unlike the sliding sampling method of traditional convolutional kernels, SPD-Conv cleverly combines SPD layers and Conv layers, such as... Figure 1 As shown, the input feature map is W×H×C1. First, an SPD layer is used to transform the spatial dimension to the channel dimension, resulting in four W / 2×H / 2×C1 feature maps. The compressed feature map has a smaller spatial dimension size while the number of channels remains the same, thus preserving more dimensional feature information. Second, the feature maps are rearranged and concatenated according to the channel dimension to obtain a W / 2×H / 2×4C1 feature map. With the number of channels increased to 4C1, it is easier to obtain rich feature information through convolution operations. Further, a non-strut Conv layer is used for convolution operations. Non-strut convolution, by performing convolution operations on each pixel or feature map without sliding on the feature map, helps to reduce the oversampling problem that may occur during feature extraction in the SPD layer and retains more fine-grained information. Finally, the input feature map is output as a new feature map with a width of W / 2, a height of H / 2, and the number of channels becomes C2.
[0038] To retain more spatial information and add more channel information, the model feature extraction network reduces the loss of some useful information and filters redundant information during downsampling. This invention introduces a spatial-to-depth convolutional SPD-Conv operation between the first two C2f modules in the YOLOv8n model's feature extraction network. This can better capture detailed features and contextual information in infrared flame images, helping to distinguish infrared small target flame features and suppress interference from fire-like spots. It is suitable for the detection of infrared small target flame images.
[0039] (3) Improvements to Neck networks Due to the high requirements for real-time performance and detection accuracy in fire detection within the complex environment of urban underground utility tunnels, target detection models composed of numerous standard convolutions may suffer from excessive parameter counts and slow inference speeds in practical applications. Therefore, this invention introduces a lightweight GSConv method to replace standard convolutions in the Neck network of the YOLOv8n model, achieving both model lightweighting and high-accuracy inference, making the fire detection model suitable for edge computing devices. GSConv is a novel lightweight convolution method that combines standard convolution (SC), depthwise separable convolution (DSC), concat operations, and shuffle operations. This reduces the number of parameters and computational load without compromising detection accuracy due to the use of depthwise separable convolutions, aiming to achieve the same feature extraction capabilities as ordinary convolutions. Figure 2 As shown.
[0040] Specifically, the feature map with C1 channels is first grouped using Conv to output a feature map with C2 / 2 channels. Then, a DSC operation is performed to output the results channel by channel. This operation is to reduce the number of parameters and computational cost. After that, a Concat operation is performed, followed by a Shuffle operation to obtain the final output feature map. The channel shuffle operation evenly scrambles and merges the channel information, improving the fusion and interaction process between channels, preserving multi-channel feature information, enhancing the model's ability to extract image features, and improving model performance. Using Depthwise Separable Convolution (DSC) to replace the traditional Standard Convolution (SC) significantly reduces the computational cost of convolution operations. The aim is to improve detection speed while maintaining detection accuracy by reducing the number of parameters and computational cost of convolution operations, and to allow the model to be applied to more devices with different computing power. The principle of reducing the number of parameters and computational cost of Depthwise Separable Convolution is explained below.
[0041] The standard convolution calculation process is as follows: Assuming the input feature map has 3 channels, the size of the convolution kernel is 3×3. Four convolution kernels are used to calculate the feature map. After convolution, the output is a 4-channel feature map. The number of parameters is 3×4×3×3=108, and the number of calculations is 108×3×3=972.
[0042] The computational process of depthwise separable convolution is mainly divided into two parts: depthwise convolution and pointwise convolution. Depthwise convolution is used to extract the spatial features of each channel. It uses three 3×3 convolution kernels to generate three feature maps from the channel dimension. The number of parameters is 3×3×3=27, and the computation is 27×3×3=243. Pointwise convolution is used to extract channel features. After the feature maps in the previous step are concatenated and superimposed, four 1×1 convolution kernels are used to perform convolution to obtain a new 4-channel feature map. The number of parameters is 3×4×1×1=12, and the computation is 12×3×3=108. The total number of parameters in the whole process is 27+12=39, and the computation is 243+108=351.
[0043] It can be seen that while completing the mapping process from feature information to a new feature space, the DSC operation makes it easier to adjust the number of channels in the feature map. Compared with the traditional standard convolution operation, it reduces the number of parameters and computational cost, while achieving the same feature extraction effect as standard convolution as much as possible. However, this also means that the overall performance of depthwise separable convolution is not very stable. The channel-wise and point-wise convolution method reduces the feature extraction capability and may not fully capture the relevant information between the channels of the feature map. Therefore, in many practical applications, it is necessary to comprehensively consider the scenario requirements and task difficulty to determine the specific usage of depthwise separable convolution.
[0044] Due to the downsampling process of the feature extraction network, the feature map is gradually compressed and the channel dimension changes layer by layer to obtain deeper semantic information. At this time, the redundant features learned by the model in the feature fusion stage of the Neck network are less. Therefore, this invention uses GSConv and VoV-GSCSP modules to construct a lightweight Neck network in the Neck network (feature fusion network).
[0045] GSbottleNeck is an innovative architecture that uses two GSConv modules concatenated together, with GSConv replacing the ordinary convolutions in the bottleneck residual module. The VoV-GSCSP (Cross-Level Partial Network) module is a further extension of GSbottleNeck, employing a one-time aggregation method to concatenate feature maps from previous layers with those from subsequent layers. This design allows for more diverse information exchange between feature maps from different layers, thus enriching the diversity of feature representations. After concatenation, a new feature map is output through convolution operations to achieve more powerful feature extraction capabilities.
[0046] This invention, based on a model that prunes the redundant P5 detection layer and adds SPD-Conv, replaces the ordinary Conv operation in the original Neck network with GSConv and replaces the C2f module with the VoV-GSCSP module, resulting in a reduction in both the number of parameters and computational cost. Replacing either GSConv or the VoV-GSCSP module individually reduces the number of parameters. When both modules are used simultaneously to form a lightweight Neck structure, the number of parameters and computational cost are significantly reduced, with a 2.96% reduction in parameters and a 0.5 GFLOPs reduction in computation. This lightweight approach maximizes the learning capability of the feature sampling process while achieving the best effect in global feature information fusion, further improving the overall performance of the model.
[0047] (4) Introducing attention mechanism Attention mechanisms essentially mimic the human visual system's focus on important features or regions, enhancing the network model's ability to perceive image information. In their development, attention mechanisms can be categorized into: channel domain attention mechanisms, spatial domain attention mechanisms, and hybrid attention mechanisms.
[0048] SENet (Squeeze-and-Excitation Networks) attention mechanism, a classic example of channel-domain attention, adaptively adjusts the weights for different parts of the input feature map to better learn the importance of channel features. After the input image is transformed into a feature map through convolution, global average pooling compresses the feature vectors from the channel dimension. Next, fully connected layers and activation functions are used to learn channel dependencies, generating weights for each feature channel to achieve accurate extraction and enhancement of image information. Finally, the feature map and weight vector for each channel are weighted. Its advantage is that it helps the model focus on important channel features and reduces the model's reliance on redundant features. However, the SE attention mechanism also has limitations; it ignores spatial dimension feature information, which may affect the model's final inference and recognition performance.
[0049] The CBAM (Convolutional Block Attention Module) attention mechanism is a further improvement on SENet's attention mechanism, integrating feature information from both channel and spatial dimensions. First, the input feature map passes through the channel attention module, enhancing the model's sensitivity to information across different channels and generating a channel-weighted feature map. Then, this feature map further passes through the spatial attention module to highlight information from important regions, improving the model's focus on key spatial locations. By combining channel and spatial attention mechanisms, its inference and recognition performance is improved.
[0050] The channel attention module performs max pooling and average pooling on the input feature maps. The two feature maps are then passed through a shared fully connected layer. The result of the fully connected layer is summed and activated, outputting a channel attention feature map containing weight values. This is then multiplied with the original input feature map to generate the feature map needed for the next step, spatial attention. The spatial attention module performs max pooling and average pooling on the initial feature map, passes it through a convolutional layer and activates it to obtain a feature map containing spatial dimension weights. This spatial attention feature map is then multiplied with the feature map from the channel attention module to obtain the final output feature map.
[0051] ECA (Efficient Channel Attention) performs global average pooling on the input feature map without reducing the channel dimensionality. The model adaptively selects the kernel size K and performs one-dimensional convolution operations accordingly, obtaining the channel weights through the sigmoid activation function. Finally, it uses the different channel weights to perform corresponding multiplication operations with the input feature map to generate the final output feature map. ECA attention achieves local cross-channel information interaction. Although it avoids the impact of dimensionality reduction on the model from a channel perspective, it does not consider the importance of the spatial dimensional relationships of the feature maps to the model's learning ability.
[0052] The Coordinate Attention (CA) mechanism consists of two steps: coordinate information embedding and coordinate attention generation. The input feature map is average-pooled along both the horizontal (X) and vertical (Y) directions. The two feature maps are then concatenated and convolved using a 1×1 kernel, changing the number of channels to C / r. After batch normalization and non-linear regression, complex non-linear features are learned. The feature map is then split back into two directional feature maps and convolved again using a 1×1 kernel, adjusted to the initial input channel number. An activation function is then applied to obtain attention weights in the height and width directions. These weights are multiplied by the corresponding channels of the initial input feature map to emphasize key regions, resulting in the final feature map. This process simultaneously captures both spatial and positional information.
[0053] SimAM (Simple Parameter-Free Attention Module) is a parameter-free attention mechanism that can significantly improve performance when added to appropriate locations in a model's network structure. SimAM determines the importance of feature regions by calculating 3D weights, without adding extra parameters or altering the model's complexity, thus reducing the inference burden. These characteristics make it well-suited for lightweight YOLO model network structures. Unlike channel attention mechanisms such as SE and ECA, which focus only on channel features, and unlike attention mechanisms like CBAM, which prioritize channel features before spatial features, SimAM calculates a solution to a defined energy function, considering both channel and spatial weights of the feature map. It treats each neuron equally, inferring 3D weights and assigning them to the feature map, effectively focusing on key regions and suppressing irrelevant features. Figure 3 As shown.
[0054] The appropriate placement of the attention mechanism depends on the specific fire detection task; otherwise, it will hinder the model's accuracy in locating fire features and balance overall performance. In the YOLOv8n model, attention mechanisms can be added in three locations: the feature extraction network, the feature fusion network, and the detection layer. This invention incorporates an attention mechanism in the feature extraction network. To demonstrate the effectiveness of this attention mechanism, four comparative experiments were conducted. The control group used the improved YOLOv8n model without any attention mechanism. The other three experimental groups added the SimAM attention mechanism to the feature extraction network, the feature fusion network, and the detection layer, respectively. Specifically, the first experimental group added it before the SPPF spatial pyramid pooling module in the feature extraction network; the second group added it after the VoV-GSCSP module in the feature fusion network; and the third group added it before the output of the detection layer. All experiments were trained for 100 epochs under the same experimental environment and hyperparameter conditions.
[0055] Experimental results show that the first experimental group exhibited the best overall performance, achieving an average precision (mAP) of 97.7%, with no change in parameter count or computational cost. The second experimental group, due to its introduction during the feature fusion stage, saw a 0.2% decrease in mAP, a 64% increase in parameters, and a 10.3 GFLOPs increase in computational cost, indicating that inappropriate addition can degrade model performance. The third experimental group's addition method also improved model performance, but its performance in terms of mAP and parameter count was slightly inferior to the first experimental group. These results demonstrate that assigning three-dimensional weights to neurons using SimAM attention can enhance the focus on the feature regions of small infrared target flames and suppress irrelevant background information in the image. Therefore, adding the SimAM attention mechanism to the feature extraction network can more accurately extract small target fire features from infrared images.
[0056] This invention compares the impact of commonly used attention mechanisms such as SE, CA, CBAM, ECA, GCNet, GENet, and SimAM on model performance. Comparative experiments were conducted by adding each mechanism to the feature extraction network and training for 100 epochs under the same experimental environment. The results are as follows: Figure 4 As shown.
[0057] from Figure 4 As can be seen, the attention mechanism is not always effective in improving model performance. After adding it to the feature extraction network, various indicators of the model fluctuate. For example, the mean accuracy (mAP) of adding EMA and GENet attention decreased by 0.4% and 0.1% respectively, while the mAP values of CBAM, GCNet and SimAM attention were 97.5%, 97.5% and 97.7% respectively. However, in terms of the number of parameters, adding CBAM attention increased the number of parameters by 3.4%, and adding GCNet attention increased the number of parameters by 0.8%.
[0058] Experimental results show that SimAM attention has better overall performance compared with the other seven attention mechanisms, and has the greatest improvement on the overall performance of the model. Precision is improved by 0.3%, recall by 0.6%, and mean precision (mAP) is improved to 97.7%, without introducing additional parameters or computational cost, thus verifying the superiority of the SimAM attention mechanism.
[0059] The improved YOLOv8n model of this invention is as follows: Figure 5As shown, the detection layer structure prunes the redundant large target detection layer P5, outputting only the feature maps of P3 and P4 for detection. Secondly, in the feature extraction network, SPD-Conv generates an 80×80 feature map, which is then further downsampled to extract deeper features and concatenated in the feature fusion stage. This prevents the loss of fine-grained information about some flame features during the numerous convolutional operations in downsampling, ensuring that the flame features of small targets are fully extracted. Based on the results of experiments comparing the attention addition location and different attention mechanisms, the SimAM attention mechanism is chosen and embedded in the layer before the SPPF layer. This increases the attention to the flame features of small targets before the feature map enters the SPPF layer, without introducing any additional parameters or computational overhead.
[0060] Finally, in the feature fusion stage, the C2f module is replaced with the VoV-GSCSP module, and the ordinary Conv is replaced with GSConv, forming a slim-neck lightweight structure for the feature fusion network. Since there are fewer redundant interference features in the feature fusion stage, the effective features retained after VoV-GSCSP processing of the upsampled and stitched feature maps are better. At the same time, using GSConv to process the feature maps in stages can effectively improve the model's detection accuracy while reducing the number of parameters and computation.
[0061] By training the improved YOLOv8n model constructed using the method of this invention, a fire monitoring model for urban underground utility tunnels can be obtained.
[0062] To further compare the superiority of the fire monitoring model constructed in this invention, comparative experiments were conducted using mainstream single-stage target detection algorithms such as the YOLO series. The experimental hardware and software environments were kept consistent, and the IRFDD infrared fire detection dataset, divided into the same proportions, was used for training, validation, and testing. To ensure fairness in the experimental process, each comparative experiment was trained without loading pre-trained weights. The experimental results are shown in Table 1.
[0063] Table 1 Comparative experimental results of different models YOLOv5n 95.6 96.8 62 1760518 4.1 3.75 YOLOv5s 96.2 97.1 6 7012822 15.8 13.7 YOLOv6n 94.5 96.2 51 4630000 11.3 9.8 YOLOv6s 94.0 96.1 53 18500000 45.2 38.8 YOLOv7-tiny 93.4 94.9 36 607387 13.0 12.5 YOLOv8n 94.8 97.0 72 3005974 8.1 6.01 YOLOv8s 94.9 96.5 68 11128961 28.3 21.3 This invention 95.5 97.5 61 1954367 7.5 3.93 As shown in Table 1, the average accuracy of YOLOv5 is slightly better than the YOLOv8 baseline model, but its detection speed is slightly slower than the YOLOv8 series; each has its advantages. YOLOv6 and YOLOv7-tiny perform relatively poorly, with both detection accuracy and speed lagging behind similar algorithms, indicating that these two model structures are not suitable for the fire infrared monitoring image dataset of this invention.
[0064] The fire monitoring model constructed in this invention achieves an average accuracy (mAP) of 97.5%, the highest among many single-stage algorithms, with a detection speed of 61 FPS. The lightweight improvement method of the model results in minimal speed loss. The fire monitoring model constructed in this invention exhibits the best overall performance, possessing high accuracy and real-time characteristics, and its smaller weight file is more suitable for deployment on various computing devices.
[0065] The fire monitoring model constructed in this invention was deployed on the Jetson Nano platform. The TensorRT framework was used to accelerate the model's structure and quantize its accuracy. The inference speed of infrared fire videos and the impact of accuracy quantization on detection performance were tested to verify the real-time detection capability of the external fire detection model on edge devices.
[0066] This invention selects the TensorRT framework to accelerate the inference process of an urban underground utility tunnel fire monitoring model. When the accuracy is quantized to FP16, the model's inference speed reaches 15.2 FPS, essentially achieving real-time detection. In practical applications, a trade-off between detection accuracy and speed needs to be maintained based on specific scenario requirements to better leverage the model's detection performance. This invention also verifies the accuracy of the fire monitoring model on a test set of a fire infrared monitoring image dataset. The test results show that the recall rate of the FP32 model decreased by 0.3%, and the mAP value decreased by only 0.1%; the mAP value of the FP16 model decreased by only 0.2%. This indicates that the fire monitoring model constructed in this invention, using the TensorRT framework to accelerate accuracy quantization, can significantly improve the model's detection speed without significant loss of detection accuracy. Based on the experimental results of speed testing and accuracy verification, the lightweight urban underground utility tunnel fire monitoring model and its inference acceleration method are verified to be feasible for edge computing applications.
[0067] III. Applying fire detection models to fire monitoring in urban underground utility tunnels In one embodiment, a dual-spectrum integrated camera is used for actual monitoring, with only the infrared thermal imaging lens activated under normal monitoring conditions. Infrared images of the monitoring environment within the underground utility tunnel are acquired using the infrared thermal imaging lens and input into the urban underground utility tunnel fire monitoring model. The model outputs fire detection results and marks the flame location. If the underground utility tunnel fire monitoring model detects a fire, the visible light camera of the dual-spectrum integrated camera is activated to capture a visible light image of the fire scene and transmit it to the monitoring host computer.
[0068] In practical applications, separate thermal imaging cameras and visible light cameras can be set up, or existing visible light cameras in the utility tunnel can be used to capture on-site visible light images.
[0069] The fire monitoring method for urban underground utility tunnels provided by this invention constructs a fire monitoring model for urban underground utility tunnels. The lightweight structure can be deployed at the edge and is used to detect early small target fires in urban underground utility tunnels, thereby improving the accuracy of fire detection in urban underground utility tunnels. Furthermore, the fire location image obtained after detection by the model enables accurate judgment of the fire scene situation, which is conducive to subsequent rescue organization.
[0070] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.
Claims
1. A method for monitoring fires in urban underground utility tunnels, characterized in that, Includes the following steps: Step 1: Construct a fire infrared monitoring image dataset; Step 2: Improve the YOLOv8n model and train the improved YOLOv8n model using the fire infrared image dataset to obtain a fire monitoring model for urban underground utility tunnels. The methods for improving the YOLOv8n model are as follows: pruning the P5 detection layer in the detection layer structure of the YOLOv8n model, introducing spatial-to-depth convolution in the feature extraction network, replacing the Conv in the Neck network with GSConv, replacing the C2f module in the Neck network with the VoV-GSCSP module, and introducing an attention mechanism in the feature extraction network. Step 3: Acquire infrared images of the monitoring environment in the underground utility tunnel using a thermal imaging camera, and input the infrared images of the monitoring environment into the urban underground utility tunnel fire monitoring model. The urban underground utility tunnel fire monitoring model outputs the fire detection results and marks the flame location. If the underground utility tunnel fire monitoring model detects a fire, it will capture a visible light image of the fire scene using a visible light camera and transmit it to the monitoring host computer.
2. The method for fire monitoring of urban underground utility tunnels according to claim 1, characterized in that, The spatial-to-depth convolution consists of a spatial-to-depth layer SDP and a non-strut convolutional layer Conv.
3. The method for monitoring fires in urban underground utility tunnels according to claim 2, characterized in that, The spatial-to-depth convolution is set between the first two C2f modules in the feature extraction network.
4. The method for monitoring fires in urban underground utility tunnels according to claim 3, characterized in that, An attention mechanism is introduced in the layer before the SPPF spatial pyramid pooling module of the feature extraction network.
5. The method for monitoring fires in urban underground utility tunnels according to claim 4, characterized in that, The attention mechanism described uses the SimAM attention mechanism.
6. The method for fire monitoring of urban underground utility tunnels according to any one of claims 1-5, characterized in that, In step three, the thermal imaging camera has a response range of 8μm to 14μm in the infrared band.
7. The method for fire monitoring of urban underground utility tunnels according to any one of claims 6, characterized in that, In step one, the fire infrared image dataset consists of infrared flame images taken on-site, as well as non-fire infrared images selected from the M3FD dataset and LLVIP dataset retrieved from the internet.
8. The method for fire monitoring of urban underground utility tunnels according to any one of claims 7, characterized in that, It also includes using the TensorRT framework to accelerate the inference process of urban underground utility tunnel fire monitoring models.