Infrared thermal imaging gas leakage detection method based on spatial-temporal feature fusion
By adopting a spatiotemporal feature fusion method in gas leakage detection, combining multi-scale spatial feature aggregation and global attention mechanism, GasConv and RepViT-Gas modules are built, and integrated with the lightweight object detection framework Nanodet, the problem of insufficient accuracy and robustness of gas leakage detection in the prior art is solved, and more efficient and accurate gas leakage detection is achieved.
Patent Information
- Application Number
- CN202510155144.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively utilize the feature information of infrared images in gas leakage detection, especially when the gas leakage characteristics change complexly, resulting in insufficient detection accuracy and robustness.
The infrared thermal imaging gas leakage detection method based on spatiotemporal feature fusion is adopted, and the GasConv and RepViT-Gas modules are constructed by integrating multi-scale spatial feature aggregation (MSFA) and global attention mechanism (GAM), and the GasConv and RepViT-Gas modules are constructed, combining frame difference method and deformable convolution (DCN), and time feature extraction module (STFM) is constructed, and it is integrated with the lightweight object detection framework Nanodet for gas leakage detection.
It improves the accuracy and efficiency of gas leakage detection, enhances the ability to identify gas leakage characteristics, especially in real-time detection, and reduces safety hazards and environmental pollution risks caused by gas leakage.
Smart Images

Figure CN120070383A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to an infrared thermal imaging gas leakage detection method based on spatio-temporal feature fusion. Background Art
[0002] Gas leakage poses a serious threat to the environment and safety. In industrial production, gas leakage may lead to explosions and fires, especially in high-risk places such as chemical plants and oil refineries. Leaked combustible gases, such as methane or hydrogen, may cause violent explosions when encountering sparks or high temperatures, resulting in equipment damage, casualties, and property losses. In addition, the leakage of certain toxic gases (such as chlorine, ammonia) may lead to poisoning incidents, causing serious harm to the health of employees and even affecting the residents of the surrounding communities. Therefore, timely and accurate detection of gas leakage is a key measure to prevent accidents and ensure personal safety. In addition, the progress of gas leakage detection technology also helps to improve the safety of the production process, reduce environmental pollution, and enhance the overall safety management level.
[0003] In recent years, the use of deep learning methods for infrared thermal imaging gas leakage detection methods can be roughly divided into two directions, namely single feature extraction and multi-feature fusion. Single feature extraction mainly relies on extracting a single type of feature from gas leakage images or data for detection. Currently, the research on gas leakage detection using a single feature is specifically divided into improvements based on object detection frameworks (YOLO, Faster R-CNN, etc.), variants of convolutional neural network CNN, etc. However, due to the relatively unobvious features of gas infrared images in terms of texture, contrast, etc., extracting a single feature may not be able to fully utilize the feature information of infrared images and is difficult to adapt to the feature changes of gas leakage at different scales.
[0004] The multi-feature fusion method combines multiple feature information from different sources or different types, thereby improving the robustness and accuracy of detection. Currently, the research on gas leakage detection using multi-feature fusion can be specifically divided into methods based on local-global feature fusion, multi-modal data feature fusion, spatio-temporal feature fusion, etc. The Transformer technology is commonly used in local-global feature fusion methods, which can effectively capture long-range dependencies in infrared video sequences through the global self-attention mechanism. However, this technology will introduce additional computational complexity, and the introduction of global features may mask local features, resulting in a decrease in the model's ability to identify subtle gas leakage features. In addition, some studies have used the texture information of RGB images and the gas region information in thermal images to enhance gas features. However, in the detection scenario of invisible gases, the model will not be able to rely on the texture information of RGB images, affecting the detection accuracy.
[0005] In summary, given that the morphology of gas leakage and the movement characteristics of gas are extremely important, extracting the spatio-temporal characteristics of gas can help better understand the diffusion behavior and dynamic changes of gas, which is crucial for improving the performance of gas leakage detection. Therefore, it is very meaningful to provide an infrared thermal imaging gas leakage detection method based on spatio-temporal feature fusion. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide an infrared thermal imaging gas leakage detection method based on spatio-temporal feature fusion, so as to improve the accuracy and efficiency of gas leakage detection in industrial environments and reduce the safety hazards and environmental pollution risks caused by gas leakage.
[0007] The technical solution of the present invention is: an infrared thermal imaging gas leakage detection method based on spatio-temporal feature fusion, including:
[0008] S1: Construct a GasConv module by integrating multi-scale spatial feature aggregation MSFA;
[0009] S2: Introduce a global attention mechanism GAM on the RepViT architecture to construct a RepViT-Gas module for extracting the local-global features of gas;
[0010] S3: Combine the GasConv module with the RepViT-Gas module to construct an image feature extraction network GasNet to obtain three different scales of predicted feature maps;
[0011] S4: Use the frame difference method and deformable convolution DCN to construct a time feature extraction module STFM; use the module STFM to extract the time dimension features of the three different scales of predicted feature maps and fuse them with the spatial features of the key detection frames;
[0012] S5: Adopt a lightweight object detection framework Nanodet to detect gas leakage, where the image feature extraction network GasNet is used as the Backbone of Nanodet, the bidirectional feature pyramid network Ghost PAN is used as the Neck of Nanodet, and the STFM module is inserted between the Neck and the Backbone;
[0013] S6: Use the three different scales of predicted feature maps output by the image feature extraction network GasNet as the feature maps finally used for gas leakage prediction, and use Nanodet for image classification and target detection box regression;
[0014] Preferably, the GasConv module consists of 2 1×1 convolutions, 1 3×3 depthwise separable convolution DWConv, and 1 multi-scale spatial feature aggregation module MSFA; the multi-scale spatial feature aggregation module MSFA consists of 3 branches, and each branch contains a 3×3 depthwise separable convolution DWConv for aggregating the multi-scale spatial features of the image.
[0015] Preferably, in S2, based on MobileNetv3 and combined with ViT, RepViT moves the DWConv in the MobileNetv3 Block upward and uses 1 3×3 DWConv and 1 1×1 DWConv in parallel to extract the spatial and channel features of the image respectively, realizing the decoupling of spatial information and channel information;
[0016] In the inference stage, RepViT removes the 1×1 DWConv and merges the multi-branches into a single branch to achieve structural re-parameterization;
[0017] The constructed RepViT-Gas module processes the RepViT-Gas module and the original SE module in parallel.
[0018] Preferably, in S3, the GasConv module and the RepViT-Gas module are combined to construct the image feature extraction network GasNet, obtaining three different scales of predicted feature maps, including:
[0019] S31: Use the GasConv module to downsample the image twice first for image feature extraction.
[0020] S32: Use a GasConv module and a combination of three multiple RepViT-Gas modules in sequence to extract features of different scales from the image.
[0021] Preferably, in S4, the STFM module consists of frame difference operation, upsampling convolution, and deformable convolution DCN, and finally the output of DCN is spliced and fused through global pooling in the horizontal and vertical directions;
[0022] Use the module STFM to extract the time dimension features from the three different scales of predicted feature maps and fuse them with the spatial features of the key detection frames, including:
[0023] Input K consecutive feature maps of the same size into the STFM module, perform differential operations pairwise on these K feature maps in sequence, and merge all the differential results, and then perform an upsampling of 3×3;
[0024] Use the deformable convolution DCN to extract the motion features of the upsampled result;
[0025] The output of the DCN is split using global average pooling operations in the horizontal and vertical directions and fused with the spatial features of the K / 2-th frame.
[0026] The present invention provides an infrared thermal imaging gas leakage detection method based on spatio-temporal feature fusion. This method can efficiently and accurately identify and locate gas leakage sources, providing a powerful technical tool for fields such as industrial safety, environmental monitoring, and emergency response, and having broad application prospects and important social value. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 It is a flowchart of the overall framework of an infrared thermal imaging gas leakage detection method based on spatio-temporal feature fusion provided for the disclosed embodiments of the present invention;
[0030] Figure 2 It is a structural diagram of the GasConv module and the MSFA module provided for the disclosed embodiments of the present invention;
[0031] Figure 3 It is a structural diagram of the improved RepViT-Gas module provided for the disclosed embodiments of the present invention;
[0032] Figure 4 It is a structural diagram of the GasNet network provided for the disclosed embodiments of the present invention;
[0033] Figure 5 It is a structural diagram of the STFM module provided for the disclosed embodiments of the present invention;
[0034] Figure 6 It is a diagram showing the gas leakage detection results provided for the disclosed embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of systems consistent with some aspects of the present invention as detailed in the appended claims.
[0036] In view of the fact that the detection methods in the prior art do not take into account features such as the form of gas leakage and the movement of gas, the present embodiment provides an infrared thermal imaging gas leakage detection method based on spatio-temporal feature fusion, including:
[0037] S1: Design a multi-scale spatial feature aggregation method MSFA, and use this module to build a downsampling convolution module GasConv for the image;
[0038] Build a downsampling convolution module GasConv for the image using the multi-scale spatial feature aggregation method MSFA. The MSFA module consists of an operation for splitting the channels of the image feature map, as well as a combination of depthwise separable convolution and a residual structure. The MSFA module does not change the shape and size of the image feature map. GasConv is composed of depthwise separable convolution and the MSFA module to implement the downsampling feature extraction operation for the image.
[0039] S2: Introduce a global attention mechanism GAM on the basis of the original RepViT to build a RepViT-Gas module for extracting local and global features of the gas; the RepViT-Gas module is improved by introducing a global attention mechanism GAM on the basis of the original RepViT for supplementing the extraction of the global features of the gas by the model.
[0040] S3: Combine GasConv in step S1 with the RepViT-Gas module in S2 to build an image feature extraction network GasNet, and obtain three different-scale predicted feature maps;
[0041] S31: Use the GasConv module to first perform 2 times of downsampling on the image for image feature extraction.
[0042] S32: Use a combination of one GasConv module and multiple RepViT-Gas modules of three types to perform feature extraction on the image at different scales.
[0043] S4: Use the frame difference method and deformable convolution DCN to build a time feature extraction module STFM;
[0044] The STFM module consists of frame difference operation, upsampling convolution, and deformable convolution DCN. Finally, the output of DCN is spliced and fused through global pooling in the horizontal and vertical directions. The STFM module is mainly used to extract temporal dimension features from the feature maps obtained by the RepViT-Gas module and GasConv module in S3 for multi-frame images.
[0045] S5: Use the lightweight object detection framework Nanodet for gas leakage detection. Use GasNet in step S3 as the Backbone of Nanodet, use Ghost PAN as the Neck part of Nanodet, and insert the STFM module in step S4 between the Neck and the Backbone.
[0046] Use the lightweight object detection framework Nanodet as the gas leakage detection baseline. Use the GasNet feature extraction network in S3 as the Backbone structure of Nanodet, use Ghost PAN as the Neck structure of Nanodet, and connect the STFM module in S4 between the Backbone and the Neck.
[0047] S6: Take the outputs of three different scales in GasNet as the feature maps finally used for gas leakage prediction, and use Nanodet for image classification and target detection box regression.
[0048] The outputs of three different scales in the GasNet network, after passing through the STFM module, are input into the Neck, and three different scales of predicted feature maps are output. And use the detection head of Nanode to classify and predict the target box for the predicted feature maps respectively.
[0049] The following combines specific embodiments to further explain the present invention, but does not limit the protection scope of the present invention.
[0050] Embodiment 1
[0051] An infrared thermal imaging gas leakage detection method based on spatio-temporal feature fusion, the overall framework flow chart is as Figure 1 shown. Specifically, it includes the following steps:
[0052] Step 1: Convert the publicly available dataset IOD-Video dataset into the COCO data format, use continuous K frames as an input sample, and use the K / 2-th frame as the key detection frame. Among them, converting to the COCO data format is convenient for directly calling the COCO API to calculate evaluation metrics in the final verification stage.
[0053] Specifically,
[0054] 1) The original annotation data format of the publicly available dataset IOD-Video is shown in Table 1. The data combinations at the corresponding positions of train_videos and test_videos are used for model training and evaluation.
[0055] Table 1 IOD-Video Dataset Format
[0056]
[0057] 2) The COCO data format is shown in Table 2. It is necessary to convert the object detection box format in the IOD-Video dataset by changing x_max and y_max to the width and height of the image.
[0058] Table 2 COCO Data Format
[0059]
[0060] 3) Finally, we read K consecutive frames as an input sample, and use the K / 2-th frame as the key detection frame, and input the ground truth detection box bbox of this frame for the calculation of the final detection box regression loss. For the remaining K - 1 frames, only their frame data is retained, and their detection box annotation data is not used for input. Finally, the total number of samples is 141017 / K.
[0061] Step 2: By integrating the multi-scale spatial feature aggregation technique MSFA, the GasConv module is constructed. Additionally, based on the existing RepViT architecture, we introduce the global attention mechanism GAM to construct the RepViT-Gas module.
[0062] Specifically,
[0063] 1) The GasConv module is constructed as Figure 2 shown. The GasConv module consists of 2 1×1 convolutions, 1 3×3 depthwise separable convolution DWConv, and 1 multi-scale spatial feature aggregation module MSFA. Among them, for the input image, it first passes through a 1×1 convolution to increase its channel dimension, then uses a 3×3 depthwise separable convolution DWConv to perform downsampling to extract features, and passes the downsampled result through the MSFA module. The MSFA module does not change the shape and size of the input features, and finally passes through a 1×1 convolution to reduce its channel dimension.
[0064] 2) The MSFA module is constructed as Figure 2As shown in the figure. The MSFA module consists of 3 branches, and each branch contains a 3×3 depthwise separable convolution DWConv for aggregating multi-scale spatial features of the image. Among them, for the input features, MSFA first divides them into 3 parts (C1, C2, C3) along the channel dimension, and the 3 parts pass through the 3 branches in sequence according to the channel order. The output of C1 passing through the first branch DWConv is used as the residual of C2; the output of C2 combined with the residual of C1 and then passing through DWConv is used as the residual of C3; the output of C3 combined with the residual of C2 and then passing through DWConv is used as the output. Finally, the results of the three branches are concatenated along the channel dimension as the output of the MSFA module.
[0065] 3) The RepViT-Gas module is constructed as shown in Figure 3. In the original RepViT structure, its ability to capture global features is poor. We introduce the global attention mechanism GAM into the original RepViT structure to supplement its ability to extract global features. Based on MobileNetv3, RepViT combines the idea of ViT, moves the DWConv in the original MobileNetv3Block upward, and uses 1 3×3 DWConv and 1 1×1 DWConv in parallel to extract the spatial and channel features of the image respectively, realizing the decoupling of spatial information and channel information. In addition, during the inference stage, RepViT removes the above-mentioned 1×1 DWConv, merges multiple branches into a single branch, realizes structural reparameterization, and reduces the computational overhead. Our RepViT-Gas module processes the RepViT-Gas module and the original SE module in parallel to achieve the extraction of local and global features.
[0066] Step 3: Combine the GasConv and RepViT-Gas modules in Step 2 to construct the image feature extraction network GasNet, and obtain three different-scale predicted feature maps. The GasNet network is as Figure 4 shown. We first use GasConv twice to perform downsampling feature extraction on the image, and then use a GasConv module and three multiple RepViT-Gas modules in sequence to perform feature extraction on the image. As shown in Table 3, the first predicted feature map is obtained by using the RepViT-Gas module twice and 1 GasConv module; the second predicted feature map is obtained by using the RepViT-Gas module four times and 1 GasConv module; the third predicted feature map is obtained by using the RepViT-Gas module three times and 1 GasConv module.
[0067] Table 3 GasNet Structure Table
[0068]
[0069] Step 4: Use the frame difference method and deformable convolutional network (DCN) to construct the spatio-temporal feature extraction module STFM. The STFM module is as follows. Figure 5 As shown, input K consecutive feature maps of the same size into the STFM module. Perform pairwise differential operations on these K feature maps in sequence, and merge all the differential results. Then, perform an upsampling of 3×3. Use the deformable convolutional network (DCN) to extract the motion features from the upsampled result. Finally, split the output of the DCN using global average pooling operations in the horizontal and vertical directions, and fuse it with the spatial features of the K / 2-th frame. The STFM module is used to extract the spatio-temporal dimensional features from multiple frames in GasNet and fuse them with the spatial features of the key detection frame.
[0070] Step 5: Adopt the lightweight object detection framework Nanodet for gas leakage detection. Use the GasNet network in Step 3 as the backbone of Nanodet. At the same time, select Ghost PAN as the neck structure (Neck part) of Nanodet. Among them, the STFM module defined in Step 4 is embedded between the Neck and the Backbone. The overall detection framework structure is as follows. Figure 1 As shown.
[0071] Step 6: Take the outputs of three different scales in GasNet as the feature maps finally used for gas leakage prediction, and use Nanodet for image classification and target detection box regression.
[0072] Experiments prove that the infrared thermal imaging gas leakage detection method based on spatio-temporal feature fusion proposed in the present invention can effectively detect the gas leakage location. Especially in real-time detection, it has great advantages compared with other algorithms. At the same time, under the test of the IOD-Video dataset (as shown in Figure 6 ), the average precision (AP) of this method is 15.42% when the IoU threshold is 0.75, and the average detection time per frame is about 30 ms, which proves the effectiveness of the method of the present invention in infrared thermal imaging gas leakage detection.
[0073] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can still be made, and these changes and deformations should also be regarded as the protection scope of the present invention.
Claims
1. Infrared thermal imaging gas leak detection method based on spatiotemporal feature fusion, characterized in that: include: S1: Construct the GasConv module by integrating multi-scale spatial feature aggregation MSFA; S2: Introducing the global attention mechanism GAM on the RepViT architecture to construct the RepViT-Gas module, which is used to extract local global features of the gas; S3: Combining the GasConv module with the RepViT-Gas module to build the image feature extraction network GasNet, and obtain three prediction feature maps of different scales; S4: Use the frame difference method and deformable convolution DCN to build a temporal feature extraction module STFM; use the module STFM to extract the temporal dimension features of the prediction feature maps of three different scales and fuse them with the spatial features of the key detection frame; S5: A lightweight target detection framework Nanodet is used to detect gas leaks, wherein the image feature extraction network GasNet is used as the Backbone of Nanodet, Ghost PAN is used as the Neck of Nanodet, and the STFM module is inserted between the Neck and the Backbone; S6: Use the three different scales of predicted feature maps output by the image feature extraction network GasNet as the final feature maps for gas leakage prediction, and use Nanodet for image classification and target detection box regression.
2. The infrared thermal imaging gas leakage detection method based on spatiotemporal feature fusion according to claim 1 is characterized in that: The GasConv module consists of two 1×1 convolutions, one 3×3 depth-wise separable convolution DWConv, and one multi-scale spatial feature aggregation module MSFA; the multi-scale spatial feature aggregation module MSFA consists of three branches, each of which contains a 3×3 depth-wise separable convolution DWConv for aggregating multi-scale spatial features of the image.
3. The infrared thermal imaging gas leakage detection method based on spatiotemporal feature fusion according to claim 1 is characterized in that: In S2, RepViT combines ViT with MobileNetv3 to move the DWConv in the MobileNetv3 Block upwards, and uses a 3×3DWConv and a 1×1DWConv in parallel to extract the spatial and channel features of the image, respectively, to achieve the decoupling of spatial information and channel information. In the inference phase, RepViT removes the 1×1DWConv, merges multiple branches into a single branch, and achieves structural reparameterization; The constructed RepViT-Gas module processes the RepViT-Gas module in parallel with the original SE module.
4. The infrared thermal imaging gas leakage detection method based on spatiotemporal feature fusion according to claim 1 is characterized in that: In S3, the GasConv module and the RepViT-Gas module are combined to build an image feature extraction network GasNet, and three prediction feature maps of different scales are obtained, including: S31: Use the GasConv module to downsample the image twice and extract image features. S32: Use a GasConv module and three or more RepViT-Gas modules in combination to extract features of different scales from the image.
5. The infrared thermal imaging gas leakage detection method based on spatiotemporal feature fusion according to claim 1 is characterized in that: The STFM module in S4 consists of frame difference operation, upsampling convolution, and deformable convolution DCN. Finally, the output of DCN is spliced and fused through global pooling in horizontal and vertical directions; The module STFM is used to extract the temporal dimension features of the prediction feature maps of three different scales and fuse them with the spatial features of the key detection frames, including: Input K consecutive feature maps of the same size into the STFM module, perform differential operations on these K feature maps in pairs, merge all differential results, and then perform a 3×3 upsampling; The up-sampled result is used to extract its motion features using deformable convolution DCN; The output of DCN is split using global average pooling operations in horizontal and vertical directions and fused with the spatial features of the K / 2th frame.