Transformer substation leakage detection method based on YOLOv8 and background difference

By using multimodal data fusion and an improved YOLOv8 model with background difference method, the problems of low efficiency and insufficient anti-interference ability in leakage detection in substations are solved. This enables accurate detection and timely alarm of both large and small leakage targets, thereby improving the safety and reliability of substations.

CN120894533APending Publication Date: 2025-11-04HUANGHE WATER CONSERVANCY & HYDROPOWER DEV GENERAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510746817.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing technologies in substations suffer from low leakage detection efficiency, difficulty in detecting small targets, insufficient anti-interference capabilities, and data silos leading to delayed early warnings, making it difficult to meet the needs of intelligent detection.

Method used

Employing multimodal data fusion technology, images are acquired using a visible light camera and an infrared thermal imager. Illumination compensation and feature fusion are performed using the Retinex algorithm, combined with an improved YOLOv8 model and background subtraction method, to detect large and small leakage targets. Graded alarms are then generated through spatiotemporal trajectory analysis.

Benefits of technology

It improves the accuracy and real-time performance of leak detection, reduces missed and false detections, and can promptly and accurately identify potential leak risks, ensuring the safe and stable operation of substations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894533A_ABST
    Figure CN120894533A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer substation leakage detection method based on YOLOv8 and background differencing, and belongs to the field of transformer substation equipment intelligent detection.The transformer substation leakage detection method comprises the steps that firstly, multi-modal data collection and preprocessing are conducted, a transformer substation site image is collected through a visible light camera and an infrared thermal imager, the image is preprocessed, and a detection result is obtained; performing feature fusion on the pre-processed visible light image and the pre-processed infrared thermal image, and outputting a fusion feature map; then detecting a large-size leakage target by using an improved YOLOv8 model, and outputting a detection result containing a target category, a position and confidence; then dynamic background modeling and tiny target detection are carried out, a dynamic background model is constructed based on 3D-CNN, and a tiny leakage target is verified in combination with a background difference method and the temperature difference of an infrared image; and finally, carrying out spatial-temporal trajectory association and hierarchical alarm, optimizing cross-frame trajectory matching by adopting a graph neural network, and triggering hierarchical alarm according to different conditions. According to the method, the accuracy, the real-time performance and the robustness of the leakage and leakage detection of the transformer substation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent detection of substation equipment, and particularly relates to a substation running, leaking and dripping detection method based on YOLOv8 and background difference. BACKGROUND

[0002] As the core hub of the power system, substations are densely equipped and have complex operating environments, and are faced with significant running, leaking and dripping risks. Substations have various types of equipment, including primary equipment such as transformers and circuit breakers and secondary equipment such as relay protection devices, and pipeline systems including cooling water pipes and lubricating oil pipes. In the long-term operation process, factors such as vibration and corrosion easily lead to the failure of sealing systems, causing the leakage of oil, gas and water. At the same time, the high-temperature and high-humidity environment of substations accelerates metal corrosion, strong electromagnetic interference affects the stability of detection signals, and equipment vibration can cause the loosening of sealing elements, further exacerbating the risk of leakage. These leaks can not only cause equipment failure and threaten the safety of operating personnel, but also cause environmental pollution.

[0003] Currently, manual inspection is still one of the main ways of substation leakage detection, but it is inefficient, relies on the experience of inspection personnel, and is difficult to detect leaks in hidden locations, and there are also personal safety hazards. Traditional automated detection technologies also have obvious defects. Single sensors such as pressure and flow sensors cannot accurately locate or sense trace leaks, and visual detection technology is greatly affected by light and is prone to misjudgment in dynamic backgrounds, and has insufficient detection capability for multi-scale targets such as small droplets. In addition, existing detection systems generally have data island problems, early warning lags and high deployment and operation costs, and are difficult to meet the intelligent detection needs of substations.

[0004] In view of the special scene of substations, the detection technology needs to have adaptability to complex scenes, be compatible with different indoor and outdoor environments, adapt to equipment vibration, dust and other dynamic interference, and realize the detection of multiple types of leakage such as liquid, gas and solid / liquid mixtures. However, existing technologies have not effectively broken through key problems such as multi-modal data fusion, small target detection in dynamic backgrounds and anti-interference algorithm optimization, and there is an urgent need for more advanced intelligent detection solutions to improve the leakage detection effect. SUMMARY

[0005] The purpose of the present application is to provide a substation running, leaking and dripping detection method based on YOLOv8 and background difference, which improves the accuracy, real-time performance and reliability of leakage detection through multi-modal data fusion, improved deep learning models and spatio-temporal trajectory analysis.

[0006] To achieve the above purpose, the technical scheme adopted by the present application is: a substation running, leaking and dripping detection method based on YOLOv8 and background difference, comprising the following steps: S1. Simultaneously acquire substation site images using a visible light camera and an infrared thermal imager. First, use the Retinex algorithm to perform illumination compensation on the visible light image, then perform adaptive histogram equalization to adjust the image contrast. Then, perform feature fusion on the preprocessed visible light image and infrared thermal image to output a fused feature map. S2. Input the fused feature map into the improved YOLOv8 model to detect large-sized leaking targets and output the detection results including target category, location and confidence level; S3. Construct a dynamic background model based on 3D-CNN, combine background subtraction method with temperature anomaly areas of infrared thermal images to verify micro-leakage targets, and output the location and confidence level of micro-leakage targets; S4. The apparent features output by YOLOv8 are integrated with the motion features calculated by optical flow to construct a trajectory graph model. Then, cross-frame trajectory matching is optimized through graph neural networks, and the spatiotemporal trajectory correlation score between nodes is calculated. Finally, a graded alarm is triggered based on the detection confidence, trajectory duration, and temperature anomaly.

[0007] Further, in step S1, the specific implementation process of feature fusion is as follows: the preprocessed visible light image and infrared thermal image are input into the DenseNet network for feature fusion, including: the input layer receives the RGB three channels of the visible light image and the temperature single channel of the infrared thermal image, for a total of 4 channels of data; feature extraction and cross-layer interaction are performed through a densely connected convolutional network, and the input of each layer includes the output of all previous layers to achieve deep fusion of multimodal data; the output layer generates a 128-channel fused feature map, which serves as the input data for subsequent YOLOv8 detection and background subtraction.

[0008] Further, in step S2, the improved YOLOv8 model's backbone network embeds the CBAM attention mechanism, the neck network uses a BiFPN bidirectional feature pyramid, and the head network uses an elliptical anchor mechanism; the improved YOLOv8 model's total loss function is: In the formula, Indicates the loss of complete intersection and union. This represents the elliptic intersection-union ratio loss. This represents classification loss.

[0009] Furthermore, the CBAM attention mechanism includes a channel attention module and a spatial attention module. The channel attention module extracts global information of features through global average pooling and global max pooling, while the spatial attention module extracts spatial features from the feature map after channel attention through a 7×7 convolution operation.

[0010] Furthermore, the feature fusion process of the BiFPN bidirectional feature pyramid network includes: first, adding the high-level semantic features to the low-level features through learnable weights, and then adding the fused features back to the high-level features through learnable weights to achieve cross-layer feature fusion.

[0011] Furthermore, the parameters of the elliptical anchor point mechanism include: center coordinates. : Represents the pixel position of the ellipse center point in the image; Major axis a and minor axis b: Represent the lengths of the major and minor axes of the ellipse, respectively, in pixels; Rotation angle θ: Represents the rotation angle of the ellipse's major axis relative to the horizontal direction, in radians; Preset anchor point size is: [(32, 16, 0°), (64, 32, 0°), (128, 64, 0°)], used to match the elliptical shape features of the leak target.

[0012] Furthermore, in step S3, the construction of the dynamic background model involves using a 3D convolution kernel to analyze the spatiotemporal features of five consecutive video frames. The update rate of the background model is set to 0.01 to 0.1, and the foreground detection threshold is set to three times the standard deviation.

[0013] Furthermore, in step S4, the spatiotemporal trajectory association process is as follows: S41. The 128-dimensional apparent feature vector output by YOLOv8 is concatenated with the displacement vector calculated by the optical flow method to form a 130-dimensional fused feature vector. S42. Construct a trajectory graph model using the detection results of 5 consecutive frames as the node set. The node features include appearance features, spatiotemporal coordinates and confidence, while the edge features include spatiotemporal distance, feature differences and motion consistency. S43. Calculate the association score between nodes using a graph attention network, with the following formula: In the formula, Represents a node and nodes The attention weights between them are exp, which represents the exponential function, and LeakyReLU is the activation function. , , It is a learnable weight matrix. , Representing nodes respectively and nodes eigenvectors, Represents a node and nodes The feature vectors of the edges between them, where ⊕ represents the concatenation operation. This indicates a summation operation performed on the neighboring nodes of a node.

[0014] Further, in step S4, the tiered alarm mechanism includes: Level 1 alarm triggering conditions: YOLOv8 detection confidence ≥ 0.9; target area ≥ 500 pixels²; the same target is detected in 3 consecutive frames; Level 2 alarm triggering conditions: background differential detection confidence ≥ 0.6; trajectory duration ≥ 180 seconds; appearing in the same device area in 5 consecutive frames; the difference between the abnormal temperature value and the background temperature. .

[0015] The beneficial effects of the above scheme are as follows: (1) This invention utilizes industrial cameras and infrared thermal imagers to collect multimodal data, and enhances feature representation through illumination compensation and DenseNet fusion, providing a high-quality data foundation for subsequent detection. The improved YOLOv8 model, combined with its innovative structure, can more accurately detect leaking targets and reduce missed detections; then, spatiotemporal trajectory association is performed based on the cascade matching of appearance features and motion features, which can accurately distinguish different targets even in complex backgrounds, making cross-frame tracking more stable and accurate, solving the trajectory disorder problem that is prone to occur in multi-target and occluded scenes in traditional methods, and improving the reliability of trajectory association.

[0016] (2) This invention combines YOLOv8 and background subtraction method for complementary technology and collaborative verification. YOLOv8 accurately identifies large-sized liquid targets based on deep learning; background subtraction method sensitively captures small targets through pixel-by-pixel difference analysis, making up for YOLOv8's missed detection of small targets. False alarms caused by illumination interference from background subtraction method can be filtered out by the morphological features of YOLOv8, while the blind spots of YOLOv8 in low-contrast scenes are filled by the pixel-level analysis of background subtraction method, significantly improving the detection efficiency of liquid leaks in substations.

[0017] (3) This invention is based on trajectory duration, recurrence frequency, and infrared temperature anomaly triggering graded alarms. This multi-dimensional judgment method avoids false alarms that may be caused by single-factor judgment. By setting reasonable thresholds, potential leakage risks can be detected in a timely and accurate manner, effectively reducing the hazards such as equipment damage, safety accidents and environmental pollution that may be caused by leakage, and ensuring the safe and stable operation of the substation. Attached Figure Description

[0018] Figure 1 This is the overall flowchart of the present invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0020] It should be noted that, unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0021] A method for detecting leaks in substations based on YOLOv8 and background subtraction includes the following steps: S1. Simultaneously acquire substation site images using a visible light camera and an infrared thermal imager. First, use the Retinex algorithm to perform illumination compensation on the visible light image, then perform adaptive histogram equalization to adjust the image contrast. Then, perform feature fusion on the preprocessed visible light image and infrared thermal image to output a fused feature map. S2. Input the fused feature map into the improved YOLOv8 model to detect large-sized leaking targets and output the detection results including target category, location and confidence level; S3. Construct a dynamic background model based on 3D-CNN, combine background subtraction method with temperature anomaly areas of infrared thermal images to verify micro-leakage targets, and output the location and confidence level of micro-leakage targets; S4. The apparent features output by YOLOv8 are fused with the motion features calculated by optical flow to construct a trajectory graph model. Then, cross-frame trajectory matching is optimized by graph neural network (GNN) and the spatiotemporal trajectory correlation score between nodes is calculated. Finally, a graded alarm is triggered based on the detection confidence, trajectory duration and temperature anomaly.

[0022] The implementation process of each step of the present invention will be described in detail below: S1. Data Acquisition and Preprocessing.

[0023] Data collection specifically includes: Data acquisition devices, including high-definition visible light cameras and infrared thermal imagers, were deployed in the substation. By fusing visible light and infrared images, a multimodal dataset was constructed. The visible light cameras acquired high-resolution (≥1920×1080), high-frame-rate (≥25FPS) video, ensuring that the images clearly displayed the shape, size, and location of water droplets. The visible light cameras covered both indoor and outdoor areas of the substation. To improve the model's generalization ability, images were acquired under different lighting conditions and background environments. Specifically, images were acquired at different times (morning, noon, evening) and in different weather conditions (sunny, cloudy, rainy), using different types of light sources (incandescent lamps, fluorescent lamps, LED lamps) and color temperatures (warm white light, cool white light).

[0024] The infrared thermal imager simultaneously collects temperature data to detect temperature gradient changes in SF6 gas and oil leak areas. It can be used to distinguish between normal heat dissipation and abnormal leak heat sources. The infrared thermal imager has a resolution of ≥640×480.

[0025] The data preprocessing specifically includes: performing Retinex illumination compensation on visible light images, followed by adaptive histogram equalization; and performing only normalization on infrared images.

[0026] The Retinex illumination compensation process is as follows: It employs the Retinex algorithm based on multi-scale Gaussian kernels, setting multiple Gaussian kernel scales to consider both local and global illumination estimation. The calculation formula is as follows: In the formula, For the original image, To estimate the light component, As a smoothing factor, The weights for each scale, is the Gaussian kernel standard deviation.

[0027] The adaptive histogram equalization process is as follows: the image is divided into 16×16 pixel sub-blocks, the histogram cropping threshold is 0.01, and the boundary effect between sub-blocks is eliminated by bilinear interpolation to ensure smooth image transition.

[0028] In complex lighting scenarios in substations, after using Retinex for lighting compensation to eliminate uneven lighting, adaptive histogram equalization can more accurately enhance target contrast.

[0029] Next, the preprocessed visible light image and infrared thermal image are fused. The specific implementation process of feature fusion is as follows: the preprocessed visible light image and infrared thermal image are input into the DenseNet network for feature fusion, including: the input layer receives the RGB three channels of the visible light image and the temperature single channel of the infrared thermal image, for a total of 4 channels of data; then, feature extraction and cross-layer interaction are performed through a densely connected convolutional network, and the input of each layer includes the output of all previous layers to achieve deep fusion of multimodal data; the output layer generates a 128-channel fused feature map, which serves as the input data for subsequent YOLOv8 detection and background subtraction.

[0030] S2. Improved YOLOv8 for detecting large-sized leak targets.

[0031] The improved YOLOv8 model embeds the CBAM attention mechanism in the C2f module of the backbone network, adopts the BiFPN bidirectional feature pyramid in the neck network, and uses the elliptical anchor point mechanism in the head network.

[0032] The CBAM attention mechanism includes a channel attention module and a spatial attention module. The channel attention module extracts global information of features through global average pooling and global max pooling, while the spatial attention module extracts spatial features from the feature map after channel attention through a 7×7 convolution operation.

[0033] The feature fusion process of the BiFPN bidirectional feature pyramid network includes a top-down path and a bottom-up path. The top-down path adds high-level semantic features to low-level features using learnable weights, while the bottom-up path adds the fused features back to high-level features using learnable weights. Multi-scale feature fusion is achieved through cross-layer connections and weight normalization, improving the efficiency of multi-scale feature fusion, especially enhancing the detection accuracy of tiny droplets.

[0034] The parameters of the elliptical anchor point mechanism include: center coordinates. : Represents the pixel position of the ellipse center point in the image; Major axis a and minor axis b: Represent the lengths of the major and minor axes of the ellipse, respectively, in pixels; Rotation angle θ: Represents the rotation angle of the ellipse's major axis relative to the horizontal direction, in radians; Preset anchor point size is: [(32, 16, 0°), (64, 32, 0°), (128, 64, 0°)], used to match the elliptical shape features of the leak target and reduce the false negative rate.

[0035] The improved YOLOv8 model's total loss function is: In the formula, Indicates the loss of complete intersection and union. This represents the elliptic intersection-union ratio loss. This represents classification loss.

[0036] The perfect intersection-union ratio (MIU) loss function is: In the formula, IoU is the intersection-union ratio of the predicted bounding box and the ground truth bounding box. The square of the distance between the center point of the predicted bounding box and the center point of the ground truth bounding box. This represents the bounding box predicted by the model. Represents the true bounding box. Let α be the diagonal length of the minimum bounding rectangle containing the two boxes, and let α be the shape matching weight. This is the shape difference parameter.

[0037] The elliptic intersection-union ratio loss function is: In the formula, and These are the predicted ellipse and the actual ellipse, respectively. and These are the predicted and actual values ​​of the rotation angle, respectively. This is the angle penalty coefficient.

[0038] The classification loss function is: In the formula, N represents the number of samples in the batch; C represents the number of target categories (such as "oil leak", "water leak", "gas leak" etc.). This indicates the true label (0 or 1) that sample i belongs to category c; This represents the probability that the model predicts sample i belongs to class c, which is output through the Softmax activation function.

[0039] By jointly optimizing the loss function, the improved YOLOv8 model can simultaneously enhance the positioning accuracy, shape matching accuracy, and category recognition capability of leak targets, meeting the detection needs in complex substation scenarios.

[0040] S3, Dynamic Background Modeling and Detection of Targets with Small Leaks.

[0041] A dynamic background model is constructed using a 3D convolutional neural network (3D-CNN). By analyzing the spatiotemporal characteristics of consecutive video frames, the dynamic change patterns of the background are automatically learned. The grayscale difference between corresponding pixels in the current frame and the background frame is calculated. Pixels with a grayscale difference exceeding a threshold are identified as foreground, and candidate regions for minor leaks are initially marked.

[0042] Specifically, a dynamic background model is constructed by learning spatiotemporal features from a 5-frame continuous video sequence (256×256 pixels) using a 3D convolutional kernel (kernel size 3×3×3, stride 1×1×1). Foreground targets are extracted using background subtraction. The background model update rate α is set to 0.01–0.1, dynamically adjusted according to the equipment vibration state: α = 0.1 when the equipment vibrates violently, and α = 0.01 when the vibration is stable. The foreground detection threshold is set to 3 standard deviations.

[0043] Then, the micro-leakage target is verified by combining the temperature anomaly region of the infrared thermal image, and the location and confidence level of the micro-target are output. Specifically, infrared thermal images of the same scene are acquired, and temperature data is obtained using an infrared thermal imager. The temperature difference between each pixel and the temperature of the area during normal operation is calculated, and a temperature difference threshold of ≥2℃ is set to filter out temperature anomaly regions. The foreground candidate region obtained by the background subtraction method is spatially matched with the temperature anomaly region of the infrared thermal image. If the spatial overlap and temperature anomaly features match, it is determined to be a real micro-leakage target, and the location information and confidence level are output; otherwise, it is considered interference.

[0044] S4. Spatiotemporal trajectory correlation and hierarchical alarm.

[0045] The spatiotemporal trajectory association process is as follows: S41. First, feature fusion is performed. The 128-dimensional apparent feature vector output by YOLOv8 is concatenated with the displacement vector (u,v) calculated by the optical flow method to form a 130-dimensional fused feature vector.

[0046] The appearance features include the target's visual semantic information, position, and morphological features; motion features reflect the target's displacement trend in consecutive frames, distinguishing background noise from real leaked targets with continuous motion trajectories; fusing the two can filter out false detections caused by device vibration, enhance the traceability of small targets' trajectories through motion features, and improve the robustness of cross-frame tracking. Optical flow is used to detect the continuous motion trajectory of small targets (e.g., area < 50 pixels²), and can be calculated using sparse optical flow or dense optical flow methods; the specific calculation process will not be elaborated further.

[0047] S42. Construct a trajectory graph model using the detection results of 5 consecutive frames as the node set. The node features include appearance features, spatiotemporal coordinates and confidence, while the edge features include spatiotemporal distance, feature differences and motion consistency. S43. Calculate the association score between nodes using a graph attention network, with the following formula: In the formula, Represents a node and nodes The attention weights between them are exp, which represents the exponential function, and LeakyReLU is the activation function. , , It is a learnable weight matrix. , Representing nodes respectively and nodes eigenvectors, Represents a node and nodes The feature vectors of the edges between them, where ⊕ represents the concatenation operation. This indicates a summation operation performed on the neighboring nodes of a node.

[0048] The tiered alarm mechanism includes: Level 1 alarm trigger conditions: YOLOv8 detection confidence ≥ 0.9; target area ≥ 500 pixels²; the same target is detected for 3 consecutive frames; Level 2 alarm trigger conditions: background differential detection confidence level ≥ 0.6; trajectory duration ≥ 180 seconds; five consecutive frames appearing in the same device area verified by trajectory association algorithm; difference between temperature anomaly and background temperature. .

[0049] Specifically, the triggering process for the Level 2 alarm is as follows: A region with a pixel difference of 40px² is detected in 5 consecutive frames using 3D-CNN and identified as a candidate for a small target; the displacement vector of the region (u=1.2, v=0.8) is calculated using optical flow; the region is input into an improved YOLOv8 model, which outputs a 128-dimensional feature vector f with a confidence level of 0.7; the feature vector f and the displacement vector are concatenated, and then cross-frame trajectory matching is optimized using a graph neural network. If the matching is successful in 5 consecutive frames and the displacement trend is consistent, the Level 2 alarm is triggered.

[0050] Furthermore, graph neural network trajectory optimization also includes occlusion handling and trajectory repair: when a target is detected to disappear in a frame, virtual nodes are generated to represent the potential trajectory continuation. The features of the virtual nodes are predicted by Kalman filtering, and the formula is: In the formula, This represents the predicted state of the target at time k. Represents the state transition matrix. express The actual state of the target at any given moment. Represents the control input matrix. This represents the control input at time k.

[0051] For trajectory segments with GNN scores below a threshold (set to 0.5 in this embodiment), the labels of high-scoring trajectories are transferred to low-scoring segments using a graph diffusion algorithm, as shown in the formula: + In the formula, Let S represent the label matrix for the (t+1)th iteration, and S be the similarity matrix. Let be the label matrix of the t-th iteration. For smoothing parameters, This represents the labeled matrix of high-confidence trajectories.

[0052] In substation equipment obstruction scenarios, by using graph neural networks to generate virtual nodes, the trajectory can be reasonably continued to avoid trajectory breakage due to obstruction, thereby reducing target loss and misjudgment caused by obstruction. For trajectory breakage or errors caused by detection mistakes, noise, etc., the trajectory can be repaired and the true movement path of the target can be restored by propagating reliable tags through graph neural networks, providing a basis for accurately judging the development trend of leakage.

[0053] In summary, this invention, through multimodal data fusion and combining the advantages of visible light and infrared thermal imaging, can more accurately detect different types of leaks, especially small targets and concealed leaks. The improved YOLOv8 model utilizes the CBAM attention mechanism, BiFPN feature pyramid, and elliptical anchor point mechanism to enhance the feature extraction and localization capabilities of targets, reducing the false negative and false positive rates.

[0054] The spatiotemporal trajectory association process of this invention, optimized by GNN, effectively solves the trajectory disorder problem in multi-target occlusion scenarios, improving trajectory stability and accuracy, and providing a reliable basis for subsequent alarms and analysis. A reasonable hierarchical alarm mechanism can promptly and accurately issue alarms of different levels based on the severity and characteristics of the leaking target, facilitating rapid response and handling by maintenance personnel, and improving the operational safety and reliability of substation equipment.

[0055] The method of this invention fully considers the complex operating environment of substations, such as changes in lighting and equipment vibration, in terms of data preprocessing, model improvement and algorithm design, and has strong robustness and adaptability.

[0056] Finally, it should be noted that any parts of this invention not described in detail are prior art. Those skilled in the art will understand that the above descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples or make equivalent substitutions for some of the technical features. All modifications and equivalent substitutions made within the spirit and principles of the invention should be included within the scope of protection of the invention.

Claims

1. A method for detecting leaks in substations based on YOLOv8 and background difference, characterized in that, Includes the following steps: S1. Simultaneously acquire substation site images using a visible light camera and an infrared thermal imager. First, use the Retinex algorithm to perform illumination compensation on the visible light image, then perform adaptive histogram equalization to adjust the image contrast. Then, perform feature fusion on the preprocessed visible light image and infrared thermal image to output a fused feature map. S2. Input the fused feature map into the improved YOLOv8 model to detect large-sized leaking targets and output the detection results including target category, location and confidence level; S3. Construct a dynamic background model based on 3D-CNN, combine background subtraction method with temperature anomaly areas of infrared thermal images to verify micro-leakage targets, and output the location and confidence level of micro-leakage targets; S4. The apparent features output by YOLOv8 are integrated with the motion features calculated by optical flow to construct a trajectory graph model. Then, cross-frame trajectory matching is optimized through graph neural networks, and the spatiotemporal trajectory correlation score between nodes is calculated. Finally, a graded alarm is triggered based on the detection confidence, trajectory duration, and temperature anomaly.

2. The method for detecting leaks in substations based on YOLOv8 and background difference as described in claim 1, characterized in that, In step S1, the specific implementation process of feature fusion is as follows: the preprocessed visible light image and the infrared thermal image are input into the DenseNet network for feature fusion, including: The input layer receives RGB three channels of visible light image and temperature single channel of infrared thermal image, for a total of 4 channels of data; Feature extraction and cross-layer interaction are achieved through densely connected convolutional networks, with the input of each layer containing the outputs of all previous layers, thus realizing deep fusion of multimodal data; The output layer generates a 128-channel fused feature map, which serves as input data for subsequent YOLOv8 detection and background subtraction.

3. The method for detecting leaks in substations based on YOLOv8 and background difference as described in claim 1, characterized in that, In step S2, the improved YOLOv8 model's backbone network embeds the CBAM attention mechanism, the neck network uses a BiFPN bidirectional feature pyramid, and the head network uses an elliptical anchoring mechanism; the improved YOLOv8 model's total loss function is: In the formula, Indicates the loss of complete intersection and union. This represents the elliptic intersection-union ratio loss. This represents classification loss.

4. The method for detecting leaks in substations based on YOLOv8 and background difference as described in claim 3, characterized in that, The CBAM attention mechanism includes a channel attention module and a spatial attention module. The channel attention module extracts global information of features through global average pooling and global max pooling, while the spatial attention module extracts spatial features from the feature map after channel attention through a 7×7 convolution operation.

5. The method for detecting leaks in substations based on YOLOv8 and background difference as described in claim 3, characterized in that, The feature fusion process of the BiFPN bidirectional feature pyramid network includes: adding high-level semantic features to low-level features through learnable weights, and adding the fused features back to high-level features through learnable weights.

6. The method for detecting leaks in substations based on YOLOv8 and background difference as described in claim 3, characterized in that, The parameters of the elliptical anchor point mechanism include: center coordinates : Indicates the pixel position of the center point of the ellipse in the image; Major axis a and minor axis b: These represent the lengths of the semi-major and semi-minor axes of the ellipse, respectively, in pixels; Rotation angle θ: represents the rotation angle of the major axis of the ellipse relative to the horizontal direction, in radians; The preset anchor point size is: [(32, 16, 0°), (64, 32, 0°), (128, 64, 0°)], which is used to match the elliptical shape characteristics of the leak target.

7. The method for detecting leaks in substations based on YOLOv8 and background difference as described in claim 1, characterized in that, In step S3, the dynamic background model is constructed by using a 3D convolution kernel to analyze the spatiotemporal features of 5 consecutive video frames. The update rate of the background model is set to 0.01 to 0.1, and the foreground detection threshold is set to 3 times the standard deviation.

8. The method for detecting leaks in substations based on YOLOv8 and background difference as described in claim 1, characterized in that, In step S4, the spatiotemporal trajectory association process is as follows: S41. Concatenate the 128-dimensional apparent feature vector output by YOLOv8 with the displacement vector calculated by the optical flow method to form a 130-dimensional fused feature vector. S42. Construct a trajectory graph model using the detection results of 5 consecutive frames as the node set. The node features include appearance features, spatiotemporal coordinates and confidence, while the edge features include spatiotemporal distance, feature differences and motion consistency. S43. Calculate the association score between nodes using a graph attention network, with the following formula: In the formula, Represents a node and nodes The attention weights between them are exp, which represents the exponential function, and LeakyReLU is the activation function. , , It is a learnable weight matrix. , Representing nodes respectively and nodes eigenvectors, Represents a node and nodes The feature vectors of the edges between them, where ⊕ represents the concatenation operation. This indicates a summation operation performed on the neighboring nodes of a node.

9. The method for detecting leaks in substations based on YOLOv8 and background difference as described in claim 1, characterized in that, In step S4, the tiered alarm mechanism includes: Level 1 alarm trigger conditions: YOLOv8 detection confidence ≥ 0.9; target area ≥ 500 pixels²; the same target is detected for 3 consecutive frames; Level 2 alarm trigger conditions: background differential detection confidence level ≥ 0.6; trajectory duration ≥ 180 seconds; 5 consecutive frames appearing in the same device area; temperature anomaly value and the difference between the background temperature. .