Infrared thermal imaging gas leakage image recognition and positioning method based on deep learning
Through a deep learning-based infrared thermal imaging image recognition method, combined with preprocessing and optical flow tracking technology, the problems of difficult identification of small leaks, high false alarm rate and insufficient positioning accuracy in infrared thermal imaging gas leak detection are solved, and efficient gas leak identification and positioning are achieved.
Patent Information
- Application Number
- CN202510880079.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies for infrared thermal imaging gas leak detection have problems such as difficulty in identifying tiny leaks in low-contrast images, high false alarm rates under complex background interference, and insufficient real-time positioning accuracy.
An infrared thermal imaging image recognition method based on deep learning is adopted. The infrared thermal imaging module is used to capture the heat signal of gas leakage. Combined with the bilateral filtering and edge enhancement of the preprocessing module, the YOLOv8-SAM2 model is used for anchor frame optimization and temporal feature fusion. The optical flow tracking and positioning module is combined for three-dimensional spatial positioning. The Lucas-Kanade algorithm is used to calculate the optical flow field and GPS/IMU data for precise positioning.
It improves the accuracy of identifying tiny leaks, reduces the false alarm rate under complex backgrounds, and achieves real-time and high-precision gas leak positioning.
Smart Images

Figure CN120808006A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of gas leakage monitoring, and particularly relates to an infrared thermal imaging gas leakage image recognition and positioning method based on deep learning. BACKGROUND
[0002] For gas leakage monitoring, traditional methods (such as frame difference method and optical flow method) rely on artificial feature extraction, and are prone to missed detection or false positives for small leaks or complex background interference (such as pipeline obstruction and light changes); and traditional machine learning algorithms are complex to calculate, and are difficult to meet the real-time detection requirements of industrial scenes (such as requiring a response within 3 seconds); and deep learning requires a large amount of labeled data, but infrared gas leakage images are scarce and the labeling cost is high, and existing technologies have poor robustness to interference factors such as temperature fluctuations and wind speed changes.
[0003] Therefore, the above problems are further improved. SUMMARY
[0004] The main purpose of the present application is to provide an infrared thermal imaging gas leakage image recognition and positioning method based on deep learning, which solves the problems of difficult recognition of small leaks in low-contrast images, high false positive rate in complex background interference, and insufficient real-time positioning accuracy in infrared thermal imaging gas leakage detection.
[0005] To achieve the above purpose, the present application provides an infrared thermal imaging gas leakage image recognition and positioning method based on deep learning, comprising the following steps: Step S1: capturing the thermal signal characteristics of gas leakage through an infrared thermal imaging module, and finally generating an infrared thermal imaging image; Step S2: performing bilateral filtering and edge enhancement processing on the infrared thermal imaging image through a preprocessing module, so as to optimize the image quality and highlight the leakage features; Step S3: the deep learning detection module improves the YOLOv8-SAM2 model to optimize the anchor frame, fuse the time sequence features, and perform SAM2 segmentation optimization processing on the image processed by the preprocessing module; Step S4: the optical flow tracking and positioning module calculates the optical flow field of consecutive frames based on the Lucas-Kanade algorithm, reversely infers the coordinates of the leakage source, and realizes three-dimensional space positioning in combination with GPS / IMU data.
[0006] As a further preferred technical solution of the above technical solution, step S2 is specifically implemented as: Step S2.1: for bilateral filtering, the spatial distance and pixel value difference of the pixel points are considered at the same time, the image edge details are preserved while the noise is removed, and for each pixel point, only the points with close spatial distance and similar pixel value among the adjacent pixels are weighted and averaged to avoid edge blurring; Step S2.2: For edge enhancement, the Canny algorithm is used to enhance edge features to make the outline of the leakage area clear.
[0007] As a further preferred technical solution of the above technical solution, the bilateral filtering in step S2.1 is specifically implemented as follows: Perform bilateral filtering on the input original infrared thermal imaging image and visible light image to remove sensor noise and retain edges. The formula is: ; in, The original infrared image at pixel point The gray value at , is the standard deviation of the spatial Gaussian kernel, which controls the spatial range of the filter, , It is the standard deviation of the grayscale Gaussian kernel and determines the sensitivity of retaining edges.
[0008] As a further preferred technical solution of the above technical solution, the edge enhancement of step S2.2 is specifically implemented as follows: Perform temperature difference enhancement and strengthen the temperature difference in the leakage area through adaptive color mapping. The formula is: , is the temperature value of the leakage area, The temperature value of the background area is automatically adjusted according to the temperature distribution of the scene to highlight Areas larger than the preset value.
[0009] As a further preferred technical solution of the above technical solution, step S3 is specifically implemented as follows: Step S3.1: For anchor box optimization, YOLOv8 is used as the backbone network to optimize the anchor box design to adapt to the irregular shape of the gas plume. K-means clustering is used to redesign the anchor box size. The input of the clustering objective function is the bounding box size of the annotated gas leakage area, and the output is the optimized anchor box size that covers the irregular plume shape. Step S3.2: For time series feature fusion, the formula is: , is the hidden state at the current moment, is the optical flow field feature of the t-th frame. The input of the formula is the optical flow field of 5 consecutive frames, and the output is a time series feature vector, which is used to correct the detection result of the current frame. Step S3.3: For SAM2 segmentation optimization, where: Input: candidate regions detected by YOLOv8; Split formula: , is a binary mask, marking the leak pixels, is a candidate region image for YOLOv8 detection, is the bounding box coordinates output by YOLO.
[0010] As a further preferred technical solution of the above technical solution, step S4 is specifically implemented as: Step S4.1: Perform optical flow field calculation, assuming that the gas motion between adjacent frames satisfies the constant brightness and small displacement, then the optical flow equation is: , and is the spatial gradient of the image in the x and y directions, is the time gradient, and is the optical flow vector, and is optimized by least squares in the local window, and the solution is: ; Step S4.2: Leak source positioning is performed by backtracking, according to the optical flow field vector, from the current frame leak area to the initial diffusion point frame by frame, the formula is: ; wherein, is the center coordinate of the detected leak area in the current frame, is the optical flow vector of the kth frame, is the backtracking frame number, is the frame interval time; Three-dimensional space mapping is performed, 2D image coordinates are converted into 3D world coordinates combined with GPS / IMU data, and the conversion formula is: , is the camera rotation matrix, is the camera translation vector, is the image pixel coordinate, is the three-dimensional world coordinate of the leak source. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 is the flowchart of the present application. DETAILED DESCRIPTION
[0012] The following description is used to disclose the present application so that those skilled in the art can implement the present application. The preferred embodiments in the following description are only as examples, and other obvious modifications can be thought of by those skilled in the art. The basic principles of the present application defined in the following description can be applied to other embodiments, modifications, improvements, equivalents and other technical solutions without departing from the spirit and scope of the present application.
[0013] In preferred embodiments of the present application, the skilled person will note that the detectors and the like to which the present application relates can be considered as prior art.
[0014] Preferred embodiments.
[0015] As Figure 1 shown, the present application discloses an infrared thermal imaging gas leakage image recognition and positioning method based on deep learning, comprising the following steps: Step S1: capturing the thermal signal characteristics (thermal radiation data stream) of gas leakage through an infrared thermal imaging module (including a non-cooled infrared detector), and finally generating an infrared thermal imaging image (the temperature correction algorithm integrated with the detector can eliminate the interference of environmental thermal noise (such as direct sunlight, self-heating of equipment) in real time, ensuring the accuracy of the thermal imaging image); Step S2: through a pre-processing module, the infrared thermal imaging image is subjected to bilateral filtering and edge enhancement processing, so as to optimize the image quality and highlight the leakage features; Step S3: the deep learning detection module improves the YOLOv8-SAM2 model, and respectively performs anchor box optimization, time sequence feature fusion (LSTM module) and SAM2 segmentation optimization processing on the image subjected to the pre-processing module; Step S4: the optical flow tracking and positioning module calculates the continuous frame optical flow field based on the Lucas-Kanade algorithm, reversely infers the leakage source coordinates, and realizes three-dimensional space positioning in combination with GPS / IMU data.
[0016] Specifically, for step S2, it is embodied as: Step S2.1: for bilateral filtering, the spatial distance and pixel value difference of the pixel points are considered at the same time, the image edge details (such as the outline edge of the leaked gas) are retained while the noise is removed, and for each pixel point, only the points with close spatial distance and similar pixel value in the adjacent pixels are weighted and averaged to avoid edge blurring; Step S2.2: for edge enhancement, the edge features are enhanced through the Canny algorithm to make the outline of the leakage area (such as the edge of the gas cloud) clear.
[0017] More specifically, for step S2.1, the bilateral filtering is embodied as: The input original infrared thermal imaging image (temperature matrix) and visible light image (RGB three channels) are subjected to bilateral filtering, so as to remove sensor noise and retain edges, and the formula is: ; Wherein, is the gray value (corresponding to temperature information) of the original infrared image at the pixel point , , σs is the standard deviation of spatial Gaussian kernel, controlling the spatial range of filter (usually 3-5 pixels), , σg is the standard deviation of gray value Gaussian kernel, determining the sensitivity of edge preservation (usually 1-2 times of temperature difference), is the original input value to be filtered.
[0018] Further, the edge enhancement of step S2.2 is specifically implemented as: Temperature difference enhancement is performed to strengthen the temperature difference of the leakage area through adaptive color mapping, and the formula is: , is the temperature value of the leakage area (extracted by dynamic threshold segmentation), is the temperature value of the background area (taking the average temperature of non-leakage pixels around the leakage area), and the mapping function is automatically adjusted according to the scene temperature distribution to highlight areas greater than a preset value (2℃) (and multi-modal fusion is performed to fuse infrared and visible light images through CNN at the feature level).
[0019] Further, step S3 is specifically implemented as: Step S3.1: For anchor box optimization, YOLOv8 is used as the backbone network, the anchor box design is optimized to adapt to the irregular shape of the gas plume, and K-means clustering is used to redesign the anchor box size. The input of the clustering objective function is the labeled gas leakage area boundary box size (width, height), and the output is the optimized anchor box size covering the irregular plume shape; Step S3.2: For temporal feature fusion, the formula is: , is the hidden state of the current time, is the optical flow field feature of the t-th frame (calculated by the Lucas-Kanade algorithm), and the input of the formula is the optical flow field of the last 5 frames (calculated by the Lucas-Kanade algorithm). The output is a temporal feature vector used to correct the detection result of the current frame (LSTM module is introduced to analyze the optical flow field of consecutive frames, track the gas diffusion path, and reduce the single-frame false detection rate); Step S3.3: For SAM2 segmentation optimization, wherein: Input: YOLOv8 detected candidate region (ROI); Segmentation formula: , is a binary mask marking the leakage pixels (1 for leakage and 0 for background), is the candidate region (ROI) image detected by YOLOv8, The boundary box coordinates output by YOLO (the SAM2 algorithm is integrated to realize pixel-level leakage area segmentation, and a morphological operation (such as an opening operation) is combined to optimize the mask precision for leakage area calculation).
[0020] Preferably, step S4 is specifically implemented as: Step S4.1: Perform optical flow field calculation. If the gas movement between adjacent frames satisfies the constant brightness and small displacement, the optical flow equation is: , and is the spatial gradient of the image in the x and y directions (calculated by the Sobel operator), is the time gradient (gray scale change between adjacent frames), and is the optical flow vector (the movement speed of the pixel in the x and y directions), and is optimized by the least square method in the local window, and the solution is: ; Step S4.2: Perform leakage source positioning by backtracking. According to the optical flow field vector, the leakage area of the current frame is backtracked frame by frame to the initial diffusion point, and the formula is: ; wherein, is the center coordinate of the detected leakage area of the current frame, is the optical flow vector of the kth frame, is the backtracking frame number (usually 5-10 frames), is the frame interval time (determined by the camera frame rate, such as 30 FPS corresponding to Δt=1 / 30 seconds); Perform three-dimensional space mapping. The 2D image coordinates are converted into 3D world coordinates in combination with GPS / IMU data, and the conversion formula is: , is the camera rotation matrix (calibrated by IMU data), is the camera translation vector (calibrated by GPS / IMU data), is the image pixel coordinate, is the three-dimensional world coordinate of the leakage source.
[0021] Preferably, after the three-dimensional world coordinate positioning, a risk score is performed through a multi-level alarm system, and then sound and light alarm and hierarchical alarm pushing are performed for processing.
[0022] It is worth mentioning that the detector and other technical features involved in the present patent application should be regarded as prior art. The specific structure, working principle and possible control method and spatial arrangement method of these technical features can be selected conventionally in the art, and should not be regarded as the invention point of the present patent. The present patent will not be further expanded and detailed.
[0023] It can be understood by those skilled in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced equivalently, as long as the modifications, equivalent replacements, improvements, etc. are within the spirit and principles of the present application.
Claims
1. A method for infrared thermal imaging gas leak image recognition and positioning based on deep learning, characterized in that: The following steps are involved: Step S1: Capturing the thermal signal characteristics of the gas leak through the infrared thermal imaging module and ultimately generating an infrared thermal imaging image; Step S2: performing bilateral filtering and edge enhancement processing on the infrared thermal imaging image through a pre-processing module to optimize image quality and highlight leakage features; Step S3: The deep learning detection module is improved by the YOLOv8-SAM2 model, and the images after the preprocessing module are subjected to anchor box optimization, temporal feature fusion, and SAM2 segmentation optimization processing respectively; Step S4: The optical flow tracking and positioning module calculates the continuous frame optical flow field based on the Lucas-Kanade algorithm, reversely infers the coordinates of the leakage source, and combines GPS / IMU data to achieve three-dimensional spatial positioning.
2. The method for infrared thermal imaging gas leak image recognition and positioning based on deep learning according to claim 1 is characterized in that: The specific implementation of step S2 is as follows: Step S2.1: For bilateral filtering, both the spatial distance and pixel value differences of pixels are considered simultaneously to remove noise while preserving image edge details. For each pixel, only the pixels with similar spatial distance and pixel values to neighboring pixels are retained for weighted averaging to avoid edge blurring. Step S2.2: For edge enhancement, the Canny algorithm is used to enhance edge features to make the outline of the leakage area clear.
3. The method for infrared thermal imaging gas leak image recognition and positioning based on deep learning according to claim 2, characterized in that: The specific implementation of the bilateral filtering in step S2.1 is as follows: Perform bilateral filtering on the input original infrared thermal imaging image and visible light image to remove sensor noise and retain edges. The formula is: ; in, The original infrared image at pixel point The gray value at , is the standard deviation of the spatial Gaussian kernel, which controls the spatial range of the filter, , It is the standard deviation of the grayscale Gaussian kernel and determines the sensitivity of retaining edges.
4. The method for identifying and locating infrared thermal imaging gas leaks based on deep learning according to claim 3, characterized in that: The specific implementation of edge enhancement in step S2.2 is as follows: Perform temperature difference enhancement and strengthen the temperature difference in the leakage area through adaptive color mapping. The formula is: , is the temperature value of the leakage area, The temperature value of the background area is automatically adjusted according to the temperature distribution of the scene to highlight Areas larger than the preset value.
5. The method for identifying and locating infrared thermal imaging gas leaks based on deep learning according to claim 4, characterized in that: Step S3 is specifically implemented as follows: Step S3.1: For anchor box optimization, YOLOv8 is used as the backbone network to optimize the anchor box design to adapt to the irregular shape of the gas plume. K-means clustering is used to redesign the anchor box size. The input of the clustering objective function is the bounding box size of the annotated gas leakage area, and the output is the optimized anchor box size that covers the irregular plume shape. Step S3.2: For time series feature fusion, the formula is: , is the hidden state at the current moment, is the optical flow field feature of the t-th frame. The input of the formula is the optical flow field of 5 consecutive frames, and the output is a time series feature vector, which is used to correct the detection result of the current frame. Step S3.3: For SAM2 segmentation optimization, where: Input: candidate regions detected by YOLOv8; Split formula: , is a binary mask marking leaked pixels, It is the candidate region image detected by YOLOv8. The bounding box coordinates output by YOLO.
6. The method for identifying and locating infrared thermal imaging gas leaks based on deep learning according to claim 5, characterized in that: Step S4 is specifically implemented as follows: Step S4.1: Calculate the optical flow field. Assuming that the gas motion between adjacent frames satisfies constant brightness and small displacement, the optical flow equation is: , and is the spatial gradient of the image in the x and y directions, is the time gradient, and is the optical flow vector, and is optimized by the least squares method in the local window to obtain: ; Step S4.2: Locate the leakage source by reverse tracing. Based on the optical flow field vector, reversely trace from the leakage area of the current frame to the initial diffusion point frame by frame. The formula is: ; in, is the center coordinate of the leakage area detected in the current frame, is the optical flow vector of the kth frame, is the number of backtracking frames, is the frame interval time; Perform three-dimensional space mapping and combine GPS / IMU data to convert 2D image coordinates into 3D world coordinates. The conversion formula is: , is the camera rotation matrix, is the camera translation vector, is the image pixel coordinate, is the 3D world coordinate of the leak source.
Citation Information
Cited By
Tongue image enhancement system and method based on deep learning
CN120953107A
A tongue image enhancement system and method based on deep learning
CN120953107B
Infrared image noise reduction method and system for gas leakage detection
CN121504759A