A defect identification method and device for substation equipment and electronic equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUBEI CENT CHINA TECH DEV OF ELECTRIC POWER
- Filing Date
- 2025-09-02
- Publication Date
- 2026-06-02
Smart Images

Figure CN121121478B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image recognition, specifically to a method, apparatus, and electronic device for defect identification of substation equipment. Background Technology
[0002] In the complex operating scenarios of substations, equipment is subjected to a superposition of multiple factors such as strong electromagnetic interference, high voltage, temperature difference and strong light change for a long time. Its operating status is often accompanied by multimodal interference and dynamic uncertainty.
[0003] Existing defect detection methods based on single-modal images typically rely on visible light or infrared images for target identification. However, under conditions of strong electromagnetic radiation and complex lighting, image signals are prone to false edges, reflection noise, and texture loss, leading to unstable target features. Furthermore, during substation equipment operation, structural obstructions such as busbar crossings, switch contact blockages, and insulator stacking are common. Conventional two-dimensional detection networks struggle to recover complete structural contours within obstructed areas, thus affecting the accuracy of defect identification. Therefore, the aforementioned methods are unsuitable for defect identification in substation equipment.
[0004] Therefore, there is an urgent need for a method, device, and electronic equipment for defect identification in substation equipment. Summary of the Invention
[0005] This application provides a method, apparatus, and electronic device for defect identification of substation equipment, which facilitates defect identification of substation equipment.
[0006] The first aspect of this application provides a defect identification method for substation equipment. The method includes: acquiring infrared images, electric field leakage maps, and visible light images collected by a multi-channel imaging system deployed at the substation site, and generating a multi-channel image tensor; inputting the multi-channel image tensor into a multi-path recognition model to obtain a fused feature map; inputting the fused feature map into a YOLOv8 backbone detection network, and combining it with the prior distribution structure map of the target substation equipment, constructing a joint region of interest in the infrared high-temperature focusing area, the electric field leakage strong response area, and the optical high-frequency edge area, forming a candidate set containing high-confidence target candidate boxes, and generating high-confidence detection results through non-maximum suppression; constructing an inter-frame residual tensor for the continuous image frame sequence of the target substation equipment corresponding to the high-confidence detection results, and obtaining enhanced detection results through temporal modeling; for target substation equipment with complex occlusion in the enhanced detection results, extracting edge response maps and inputting them into a Transformer-based edge prediction path for structural contour completion, and outputting the target boundary and defect location information of the repaired target substation equipment.
[0007] A second aspect of this application provides a defect identification device for substation equipment. The device includes an acquisition module and a processing module. The acquisition module acquires infrared images, electric field leakage maps, and visible light images collected by a multi-channel imaging system deployed at the substation site, generating a multi-channel image tensor. The processing module inputs the multi-channel image tensor into a multi-path recognition model to obtain a fused feature map. The processing module further inputs the fused feature map into a YOLOv8 backbone detection network, combining it with the prior distribution structure map of the target substation equipment, and identifies defects in the infrared high-temperature focusing region, the electric field leakage strong response region, and the optical high-frequency region. A joint region of interest is constructed in the edge region to form a candidate set containing high-confidence target candidate boxes, and high-confidence detection results are generated through non-maximum suppression. The processing module is also used to construct inter-frame residual tensors for the continuous image frame sequence of the target substation equipment corresponding to the high-confidence detection results, and obtain enhanced detection results through temporal modeling. The processing module is also used to extract edge response maps for target substation equipment with complex occlusion in the enhanced detection results, input the Transformer-based edge prediction path for structural contour completion, and output the target boundary and defect location information of the repaired target substation equipment.
[0008] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, and both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method described above.
[0009] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described above.
[0010] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages:
[0011] First, by acquiring infrared images, electric field leakage maps, and visible light images and generating multi-channel image tensors in a unified spatial domain, the synergistic utilization of multimodal information is achieved. This avoids the problem of insufficient feature representation caused by relying on a single modality in existing technologies, significantly improving robustness under electromagnetic interference and complex lighting conditions. Second, by extracting infrared, electromagnetic, and visible light features through a multi-path recognition model and introducing channel attention and spatial attention mechanisms in the fusion module, the multimodal features can achieve complementarity and alignment during the fusion process, avoiding interference from pseudo-textures and modal shifts, thereby enhancing the saliency of equipment defect areas. Third, by combining the YOLOv8 backbone detection network with the prior distribution structure map of the target substation equipment, a joint region of interest is constructed in the infrared high-temperature focusing area, the electric field leakage strong response area, and the optical high-frequency edge area. This ensures that the detection results not only rely on depth features but also incorporate prior knowledge of electrical equipment operation, strengthening key areas during candidate box generation and improving the accuracy and reliability of high-confidence candidate boxes.
[0012] Next, by constructing an inter-frame residual tensor from consecutive frame sequences of high-confidence detection results and inputting it into the temporal modeling module, minute deformations and dynamic anomalies can be captured. This solves the problem that single-frame detection struggles to identify subtle defects such as slight bulges and slow expansions, thereby improving detection sensitivity and stability. Finally, in the presence of complex occlusion, the missing structural contours are completed using Transformer-based edge prediction paths, and the complete target boundary is restored under shape consistency constraints. This enables complete identification of partially occluded equipment, ensuring the accuracy of defect location information. Therefore, it facilitates defect identification of substation equipment. Attached Figure Description
[0013] Figure 1 A flowchart illustrating a defect identification method for substation equipment provided in this application embodiment;
[0014] Figure 2 A schematic diagram of a defect identification device for substation equipment provided in an embodiment of this application;
[0015] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0016] Explanation of reference numerals in the attached figures: 21. Acquisition module; 22. Processing module; 31. Processor; 32. Communication bus; 33. User interface; 34. Network interface; 35. Memory. Detailed Implementation
[0017] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0018] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0019] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0020] To address the aforementioned technical problems, this application provides a method for defect identification in substation equipment, referring to... Figure 1 , Figure 1 This is a flowchart illustrating a defect identification method for substation equipment provided in an embodiment of this application. The method is applied to a server and includes steps S110 to S150, as follows:
[0021] S110: Acquire infrared images, electric field leakage maps, and visible light images collected by a multi-channel imaging system deployed at the substation site, and generate a multi-channel image tensor.
[0022] Specifically, a server refers to an industrial computing node operating at or near the substation, possessing image data access, time synchronization, and batch matrix computation capabilities. It receives raw data from the multi-channel imaging system and performs preprocessing and data structuring. For example, deploying an edge server with a gigabit Ethernet interface, hardware timing module, and graphics processing unit in the control building of a 500 kV substation can meet the requirements for long-term uninterrupted access and real-time computing. Deployment at the substation site means that the sensors of the multi-channel imaging system are installed above or to the side of target areas such as main transformer bays, gas-insulated metal-enclosed switchgear bays, and near busbar bridges, and their fixed calibration posture and unified timing are completed to ensure observation coverage of critical components and reduce obstruction. For example, a gimbal can be erected 30 meters above a gas-insulated metal-enclosed switchgear bay, allowing visible light imagers, infrared imagers, and electric field detection arrays to jointly view the same busbar connection point.
[0023] A multi-channel imaging system refers to an imaging and sensing combination that works collaboratively under the same clock and spatial reference, comprising a visible light imager, an infrared imager, and an electric field detection array. The visible light imager provides texture and geometric edge information, the infrared imager provides radiation temperature distribution in the 8-14 micrometer band, and the electric field detection array provides the amplitude and phase distribution of electric field leakage through electric field intensity sensing and coherent demodulation. For example, near the junction of the insulator string and the lead wire, the electric field detection array forms a two-dimensional sampling grid to facilitate subsequent interpolation and reconstruction of the electric field spatial distribution. Acquisition refers to the synchronous acquisition of raw frames of each mode under unified timing and hardware triggering, along with timestamps, trigger numbers, and extrinsic parameter identifiers to ensure subsequent pixel-level alignment. For example, triggering occurs every 33 milliseconds, simultaneously recording the timestamps of three frames, with the time difference between the three frames not exceeding one millisecond.
[0024] Visible light images refer to two-dimensional color images output by visible light imagers or brightness maps calculated from them, carrying high-frequency texture information such as equipment appearance, nameplates, bolts, and edges. For example, at switch contacts, visible light images can distinguish contact contours and wear scratches. Infrared images refer to radiation brightness maps output by infrared imagers or temperature field images after radiometric calibration, used to reflect hot spots and thermal gradients. For example, in the early stages of thermal runaway at cable joints, infrared images show localized temperature anomalies of three to five degrees Celsius. Electric field leakage maps, also known as electromagnetic maps, are two-dimensional distribution maps obtained by reconstructing discrete samples from an electric field detection array onto the imaging plane through spatial interpolation and phase unwrapping. Each pixel contains at least one or both of electric field amplitude and electric field phase, used to characterize localized electric field anomalies in the housing gaps or lead ends of gas-insulated metal-enclosed switchgear. For example, at the tiny air gaps of busbar flanges, the electric field leakage map shows stable high-amplitude patches, accompanied by an increase in the phase consistency of the 50 Hz fundamental wave.
[0025] The acquisition process involves the server accessing the multi-channel imaging system via network and time synchronization interfaces, writing the three raw frames and metadata to a cache, and then performing radiometric calibration, distortion correction, noise suppression, and extrinsic parameter projection to form physically comparable data that can be aligned. For example, blackbody two-point calibration is performed on infrared images, distortion correction on visible light images, and probe gain equalization on electric field leakage maps. A multi-channel image tensor is a three-dimensional data volume (height × width × number of channels) constructed according to a uniform spatial domain size, used to express the pixel-level correspondences of different modalities in parallel, enabling downstream networks to perform convolution and attention operations.
[0026] In one possible implementation, infrared images, electric field leakage maps, and visible light images acquired by a multi-channel imaging system deployed at a substation are obtained, and a multi-channel image tensor is generated. Specifically, this includes: triggering the synchronous acquisition of infrared images, electric field leakage maps, and visible light images by establishing a clock synchronization mechanism for the multi-channel imaging system; eliminating acquisition time differences using a unified timing source and hard trigger pulses and recording timestamps and trigger numbers as synchronization references; performing multi-pose acquisition based on a calibration target with high-contrast corner points; solving for the intrinsic and extrinsic parameter matrices and distortion coefficients of the infrared imager, visible light imager, and electric field imager; establishing cross-modal projection mapping relationships to obtain geometric calibration results; and performing blackbody-based calibration on the infrared images. The two-point radiometric calibration performs exposure response linearization based on a grayscale standard plate on the visible light image and amplitude and phase correction based on a standard field source on the electric field leakage spectrum to eliminate intensity scale differences between modes. Based on the geometric calibration results, the infrared image and the electric field leakage spectrum are reprojected onto the visible light image plane, and pixel-level alignment is performed by solving the homography matrix based on a robust matching method. A unified spatial domain raster is established according to the resolution of the visible light image, and the infrared image and the electric field leakage spectrum are mapped to the unified spatial domain raster through an interpolation method that preserves edge features to form a multimodal aligned image in a unified spatial domain. The multimodal aligned image is then normalized and stacked to construct a multi-channel image tensor.
[0027] Specifically, after establishing a clock synchronization mechanism for the multi-channel imaging system, a unified timing source and hard trigger pulse are used as the trigger starting point to synchronously acquire infrared images, electric field leakage spectra, and visible light images. The timestamp and trigger sequence number of each frame are written into the synchronization metadata as the time reference for subsequent processing. The synchronization metadata is used as input to perform time difference verification on the three frames, using a threshold constraint criterion.
[0028]
[0029] in Indicates the first The timestamp of the triggered infrared image. Indicates the first The timestamp of the visible light image triggered the next time. Indicates the first Timestamp of the triggered electric field leakage pattern. This represents the synchronization threshold. When the criterion is met, the three original sequences that have passed the time consistency check are output as the input for geometric calibration. When the criterion is not met, the trigger sequence number is used for alignment and out-of-bounds frames are removed to ensure the temporal consistency of subsequent registration.
[0030] Using three time-consistent original sequences and a calibration target image with high-contrast corner points as input, the intrinsic and extrinsic parameter matrices and distortion coefficients of the infrared imager, visible light imager, and electric field imager are solved respectively. A combined perspective projection and radial-tangential distortion model is used to estimate the minimum reprojection error of the calibration target corner points. The projection model is as follows:
[0031]
[0032] in The homogeneous coordinates of the three-dimensional calibration points. and For external parameter rotation and translation, This is the intrinsic parameter matrix. It is a projection operator that includes distortion correction and normalization to the pixel plane.
[0033] The optimization objective is:
[0034]
[0035] in For the first The pixel coordinates of each observed corner point Predict pixel coordinates for the projection model; output geometric calibration results and cross-modal projection mapping relationship as input for radiation and response consistency calibration.
[0036] Using the geometric calibration results and the three original sequences as input, the modal response is radiometrically calibrated, linearized, and amplitude-phase corrected to eliminate intensity scale differences. The infrared image is mapped from its original grayscale to temperature using the blackbody two-point method. The parameters are solved from the reference temperature pair and the corresponding response pair, and the mapping is as follows:
[0037]
[0038] in For the original infrared response, For calibration temperature, and The results were obtained from two sets of blackbody temperature and response pairs.
[0039] Visible light images are derived from a grayscale standard plate to determine the linear domain brightness, and power-law correction is applied.
[0040]
[0041] in These are the original pixel values. For linear brightness, The exposure response index is used; the electric field leakage spectrum is obtained by calculating the amplitude gain and phase zero point using a standard field source, and corrected as follows:
[0042] ,
[0043] in and For the original amplitude and phase, and To correct the results, and For amplitude gain and bias, The phase zero point is used; the output calibrated three frames are used as inputs for reprojection and pixel-level alignment.
[0044] Using calibrated infrared images, electric field leakage maps, visible light images, and geometric calibration results as input, the infrared images and electric field leakage maps are reprojected onto the coordinate plane of the visible light image. Robust feature matching is performed on the rigid texture of the device to refine alignment, and homography constraints are used for local plane correction. The mapping relationship is as follows:
[0045]
[0046] in and These are the homogeneous pixel coordinates of the source plane and the target plane, respectively. The homography matrix is obtained by minimizing the correspondence error estimate with the Huber kernel. The error term is:
[0047]
[0048] in For Huber's losses, The robust threshold is used; the output pixel-level aligned infrared image and electric field leakage spectrum to the visible light coordinate system are used as input to the unified spatial domain raster mapping.
[0049] S120. Input the multi-channel image tensor into the multi-path recognition model to obtain the fused feature map.
[0050] Specifically, the multi-path recognition model refers to a deep network structure that sets up multiple feature extraction branches and multi-level fusion paths in parallel within the same computational graph, and achieves modal complementarity through mechanisms such as cross-attention and prior guidance. Preferably, it includes visible light feature extraction branches, infrared feature extraction branches, and electromagnetic spectrum feature extraction branches. Each branch uses residual convolution stacking, dilated convolution, and feature pyramids to simultaneously characterize fine-grained edges and large-scale context, and suppresses irrelevant responses through channel attention and spatial attention in the middle section. Subsequently, cross-attention interactions are set across branches to align infrared temperature anomalies and electric field leakage patterns to the visible light texture framework, and a spatially guided mask is generated by introducing the prior distribution structure map of the target substation equipment to enhance the discriminative power of the infrared high-temperature focusing area, the electric field leakage strong response area, and the optical high-frequency edge area. For example, in the scenario of a small air gap in a bus flange, the electromagnetic spectrum branch provides amplitude and phase anomalies, and the infrared branch provides temperature gradients. The two are superimposed on the edge direction of the visible light branch through cross-attention, thus highlighting the true defect contour in a strongly reflective background. Fusion feature maps refer to the unified representation output by a multi-path recognition model after completing cross-modal alignment and attention guidance. They are used for downstream tasks such as detection, segmentation, or ranging, and are decoupled and matched with the multi-scale head of the detection backbone in terms of spatial resolution and channel dimension.
[0051] In one possible implementation, a multi-channel image tensor is input into a multi-path recognition model to obtain a fused feature map. Specifically, this includes: inputting the multi-channel image tensor into the infrared feature extraction branch, the electromagnetic spectrum feature extraction branch, and the visible light feature extraction branch of the multi-path recognition model to extract infrared features, electromagnetic features, and visible light features; performing channel-by-channel and position-by-position weighting on the infrared features, electromagnetic features, and visible light features to obtain weighted features; performing band correction on the weighted features for high-frequency details and low-frequency background, and balancing the energy distribution of different modal features in the spatial domain through residual gating to output spectrum-corrected fused features; based on the spectrum-corrected fused features, using visible light features as the query and infrared and electromagnetic features as keys and values to perform multi-head interaction to capture the complementary relationship between thermal anomalies and electric field leakage on optical textures, and maintaining the original geometric layout through residual connections to obtain complementary enhancement features; and outputting the fused feature map based on the complementary enhancement features and the prior distribution structure map of the target substation equipment.
[0052] Specifically, a multi-channel image tensor is used as input, fed into an infrared feature extraction branch, an electromagnetic spectrum feature extraction branch, and a visible light feature extraction branch, respectively. All three branches employ a multi-scale structure combining residual convolution stacking and dilated convolution to extract local edges and global context, and each outputs a size of [size missing]. , , Infrared, electromagnetic and visible light characteristics.
[0053] Within each branch, a lightweight channel selector is first constructed using channel compression and activation. The channel selector employs the following methods:
[0054]
[0055] in Indicates the first The feature tensor of each modality Indicates global average pooling. and For learnable linear mappings, It is a linear rectifier unit. For the Sigmoid function, This is the channel attention weight vector.
[0056] Simultaneously, a spatial attention map is generated using shared 2D convolutions:
[0057]
[0058] in Spatial attention weights, For size The convolution kernel.
[0059] The output is a bit-scaled feature:
[0060]
[0061] in For pixel coordinates, For channel indexing, As input for subsequent weighted fusion. For example, in the early anomaly scenario at the end of the surge arrester lead, visible light features emphasize fine textures, infrared features have significantly increased weight in the hotspot neighborhood channel, and electromagnetic features have significantly increased spatial attention at the amplitude and phase anomaly location.
[0062] Infrared, electromagnetic, and visible light features are weighted channel-by-channel and position-by-position. First, band-based characterization is performed in the frequency domain to separate high-frequency details from low-frequency background. The formulas for constructing Butterworth high-pass and low-pass filters, representing the two-dimensional discrete Fourier transform, are as follows:
[0063]
[0064] in For frequency domain coordinates, The frequency radius to the origin, The cutoff frequency, The filter order is given; perform the following operations for each mode:
[0065]
[0066] in This represents Hadamard multiplication. Indicates the inverse transform; in the spatial domain, residual-gated equilibrium is used to balance the characteristic energy distributions of different modes.
[0067]
[0068] in This indicates a channel-level cascade connection. The residual gate weights are generated by 1×1 convolution and normalized by Sigmoid, taking values ∈ [0,1], and aligned with the channel dimension. These are the modal characteristics after spectral correction and energy balance.
[0069] Then, weighted coefficients are applied channel-by-channel and position-by-position. and Perform final recalibration and obtain weighted features:
[0070]
[0071] Among them, the output The normalized visible light feature tensor. The normalized infrared feature tensor. This is the normalized electromagnetic characteristic tensor.
[0072] For example, at the highly reflective busbar flange, after the high-frequency band suppresses the mirror pseudo-high frequency, the infrared high-frequency edge and electromagnetic high-frequency stripe are preserved with high fidelity.
[0073] Based on the spectrally corrected fusion characteristics, a multi-head interaction is performed using visible light features as queries and infrared and electromagnetic features as keys and values, flattening the spatial dimensions. Each token is set to have multiple heads. .
[0074] For the first Based on the size, the following calculations were made:
[0075]
[0076]
[0077] in , , Let be the learnable projection matrix of the m-th head. and Q represents the subspace dimension of keys and values, used to avoid the softmax function becoming too large in high-dimensional space, leading to gradient vanishing. It is typically taken as an integer value related to the number of feature channels. m For querying the matrix, Km V is the key matrix. m For the value matrix, Attn m Let be the attention weight matrix for the m-th head.
[0078] Concatenate the outputs from each header and perform a linear mapping:
[0079]
[0080] in For a learnable mapping, O represents the joint output of all attention heads, and h represents the number of attention heads.
[0081] Finally, by maintaining the original geometric layout through residual connections and normalization, we obtain:
[0082]
[0083] in For layer normalization, This is a complementary enhancement feature.
[0084] This process uses visible light texture as a reference coordinate to align infrared thermal anomalies with electric field leakage patterns and inject them into semantically consistent spatial locations. For example, in the scenario of slow expansion of the transformer oil bladder, the visible light contour is enhanced by the infrared temperature gradient and the amplitude and phase anomalies of the electric field, thus highlighting minute deformations.
[0085] To enhance complementary features, a priori distribution structure map of the target substation equipment is introduced as a spatial guide to suppress irrelevant regions and highlight key structural regions. Let the priori distribution structure map be located in the reference coordinate system as follows: ,make It is flattened out and gating fusion is used to generate a fused feature map.
[0086] First, an additive prior bias is used to enhance the focus of attention on the prior region, as shown in the following formula:
[0087]
[0088] in For the prior strength coefficient, Let be the dimension of the key vector. It is a vector of all ones.
[0089] Subsequently, prior gating was used to fuse complementary enhanced features in the spatial domain and modal aggregation after spectral correction.
[0090]
[0091] in To concatenate the three-modal features in the channel dimension and then... Alignment mapping obtained from convolution and nonlinear activation This refers to the fused feature map; in implementation, it can be... Multiscale pyramids This serves as input to the subsequent detection backbone, ensuring that high-resolution branches preserve edge details while low-resolution branches carry global semantics. For example, in scenarios where busbars intersect and contact occlusion coexist... By focusing joint attention on the interface region between conductive connections and insulation, the signal-to-noise ratio of the fused feature map is significantly improved in key parts of the structure.
[0092] S130. Input the fused feature map into the YOLOv8 backbone detection network, and combine it with the prior distribution structure map of the target substation equipment to construct a joint region of interest in the infrared high-temperature focusing area, the electric field leakage strong response area and the optical high-frequency edge area, forming a candidate set containing high-confidence target candidate boxes, and generating high-confidence detection results through non-maximum suppression.
[0093] Specifically, the YOLOv8 backbone detection network refers to an end-to-end object detection structure composed of a backbone feature extraction network, a bidirectional feature pyramid, and a decoupled detection head. The backbone network is responsible for hierarchical semantic extraction, the feature pyramid is responsible for multi-scale aggregation, and the decoupled detection head outputs class confidence scores through a classification branch and bounding box center, width, height, and quality scores through a regression branch. The fused feature maps are fed into the detection head at each scale to generate multi-scale candidate responses. For example, under the occlusion view of an insulator string, the low-resolution layer of the backbone network captures the overall outline, while the high-resolution layer retains small gaps. After the two are fused in the feature pyramid, the detection head provides candidate boxes and confidence scores.
[0094] The prior distribution structure map of the target substation equipment refers to a spatial probability map generated based on the equipment layout diagram, historical inspection statistics, and structural topology. It is used to constrain the detection focus area and suppress regional responses unrelated to the equipment. The infrared high-temperature focusing area refers to the high-response area selected by the infrared temperature channel threshold or gradient criterion on the corresponding coordinates of the fused feature map. The electric field leakage strong response area refers to the abnormally high-response area selected by the electric field amplitude or phase consistency on the corresponding coordinates of the fused feature map. The optical high-frequency edge area refers to the significant edge area selected by the visible light high-frequency energy or edge operator response on the corresponding coordinates of the fused feature map. Constructing a joint domain of interest means normalizing and weighting the prior distribution structure map, the infrared high-temperature focusing area, the electric field leakage strong response area, and the optical high-frequency edge area in the same coordinate system to obtain a spatial weight map for the intermediate features of the gated detection head. For example, in the neighborhood of the moving and stationary contacts of the disconnector switch, the joint domain of interest concentrates the weight on the contact joint and insulation interface, thereby improving the score of the candidate box at that location and suppressing the interference of the background steel structure.
[0095] High-confidence target candidate boxes refer to candidate bounding boxes generated by the decoupled detection head whose class confidence and bounding box quality scores both exceed a threshold; the candidate set refers to the set of high-confidence target candidate boxes summarized across all scales, and has undergone minimum size filtering and prior consistency verification; non-maximum suppression refers to the stepwise removal of redundant candidate boxes within the same class based on the overlap metric between candidate bounding boxes to retain the most representative detection results.
[0096] In one possible implementation, the fused feature map is input into the YOLOv8 backbone detection network. Combined with the prior distribution structure map of the target substation equipment, a joint region of interest is constructed in the infrared high-temperature focusing area, the electric field leakage strong response area, and the optical high-frequency edge area. This forms a candidate set containing high-confidence target candidate boxes. High-confidence detection results are generated through non-maximum suppression. Specifically, this includes: extracting multi-scale features from the fused feature map using the feature pyramid and path aggregation network in the YOLOv8 backbone detection network and inputting them into a decoupled detection head. The decoupled detection head simultaneously outputs bounding box position parameters and target category confidence. The prior distribution structure map of the target substation equipment is used as input. A spatial guiding mask is generated and combined with the attention weights of the fused feature map to map prior constraints onto the multi-scale feature map to generate prior weighted features. In the prior weighted features, the spatial overlap ratio and average response intensity between the candidate box and the response map of the infrared high-temperature focusing area, the response map of the electric field leakage strong response area, and the response map of the optical high-frequency edge area are calculated to obtain the multimodal weighted confidence of the candidate box. A weighted candidate box set is formed based on the multimodal weighted confidence. Non-maximum suppression is performed on the weighted candidate box set. A suppression metric based on distance intersection-union ratio is used to remove candidate boxes with overlap exceeding the threshold and retain the candidate box with the highest fused confidence to output a high-confidence detection result.
[0097] Specifically, after the fused feature map is fed into the YOLOv8 backbone detection network as input, it first undergoes multi-scale transformation through a feature pyramid and path aggregation network, and then enters the decoupled detection head at each scale to simultaneously output bounding box position parameters and target class confidence. Let the fused feature map be denoted as... Multiscale mapping is Decoupling detection head in scale The output class probability tensor and regression parameter tensor are in the following form:
[0098]
[0099] in Indicates the first Feature maps at various scales Represents a classification branch mapping. This represents the regression branch mapping. This represents the probability distribution of the target class at each grid location. Indicates the number of target categories. This represents the bounding box parameters for each grid location. Center-point anchorless parameterization is used, with the grid center... With characteristic step size For reference, the bounding box is recovered using the following formula:
[0100]
[0101] in They are respectively The four components, The coordinates of the bounding box center and its width and height. For scale Space dimensions, For the scale number, For the number of channels, Enter the dimensions.
[0102] Using the prior distribution structure map of the target substation equipment as input, a spatial guiding mask is generated and combined with the attention weights of the fused feature map. This prior constraint is then mapped onto the multi-scale feature map to obtain prior weighted features. Let the prior distribution structure map be denoted as... The scale is obtained after bilinear scaling and alignment. Prior mask Let the fused feature map be scaled. The attention weights are , with coefficient The spatial weights are obtained by combining the two:
[0103]
[0104] in This indicates linear normalization to an interval. , Spatial guidance weights are used. Prior weighted features are obtained using a position-by-position gating method:
[0105]
[0106] in Represents pixel coordinates. In the above symbols... It originates from the probability of equipment layout and structural topology. This originates from channel attention or spatial attention within the fused feature map. Control the ratio of prior knowledge weights to data-driven weights.
[0107] When calculating the multimodal weighted confidence of candidate boxes in prior weighted features, the spatial overlap ratio and average response intensity of the candidate boxes are first obtained based on the response maps of the infrared high-temperature focusing region, the electric field leakage strong response region, and the optical high-frequency edge region, respectively. The candidate boxes are then denoted as... The set of pixels is The infrared response diagram is The electric field response diagram is as follows The optical edge response map is The overlap ratio and average response are defined as follows:
[0108]
[0109]
[0110]
[0111] in The threshold for each response graph, Indicates the number of elements in the set. This represents the pixel location. The multimodal weighted confidence score of the candidate box is obtained by fusing the class probability output by the detector head with the bounding box quality score and the multimodal response:
[0112]
[0113] in Candidate boxes Category The probability, Scoring the quality of the bounding box. For the Sigmoid function, These are the modal weighting coefficients. This is a bias term. Using a threshold... The filtered result is a weighted set of candidate boxes.
[0114] When performing non-maximum suppression on the weighted candidate box set, a suppression metric based on distance intersection-union ratio is used to remove candidate boxes with overlap exceeding a threshold and retain the candidate box with the highest fusion confidence. Let any two boxes... and The distance intersection-union ratio is:
[0115]
[0116] in This represents the intersection-union ratio of the two frames. Indicates the coordinates of the bounding box center. Represents Euclidean distance. This represents the diagonal length of the smallest bounding rectangle that simultaneously encloses both boxes. The preservation rule for nonmaximum suppression is:
[0117]
[0118] in To retain the final set, This is the non-maximum suppression threshold. Candidate boxes in Outputting from high to low confidence levels yields the high-confidence detection results. (The symbols mentioned above...) Control the intensity of inhibition. The ratio of the intersection area to the union area of the two frames is used to determine the intersection area. and A normalized term used to measure the distance between geometric centers relative to the circumscribed scale is used to more stably remove redundancy in elongated and adjacent but low-overlapping scenarios.
[0119] S140. For the continuous image frame sequence of the target substation equipment corresponding to the high confidence detection results, construct the inter-frame residual tensor, and obtain the enhanced detection results through temporal modeling.
[0120] Specifically, a continuous image frame sequence refers to an ordered set of images formed by continuously acquiring images of the same target substation equipment over time. Each frame corresponds to the state of the same scene at different points in time, maintaining a unified spatial perspective and modal information input. For example, when a drone photographs the moving contact of a circuit breaker along a fixed trajectory, a series of continuous image frames spaced tens of milliseconds apart can be obtained. Each frame records the state of the contact during slight vibrations or temperature changes. The inter-frame residual tensor refers to a multi-dimensional data structure obtained through pixel-level or feature-level difference calculations between adjacent image frames. Essentially, it is a tensor form representing dynamic information that changes over time. For example, when observing a slight expansion of the casing of an oil-immersed transformer in continuously captured image frames, the inter-frame residual tensor will highlight the boundary displacement of the casing area while suppressing the background, thereby enhancing sensitivity to minor anomalies.
[0121] Temporal modeling refers to the process of constructing time-dependent relationships using continuous image frame sequences and inter-frame residual tensors. It employs recurrent neural networks, temporal convolutional networks, or temporal attention mechanisms to capture dynamic patterns evolving over time. This allows detection to rely not only on static spatial features but also on temporal variations across frames. For example, in detecting minor deformations in the oil conservator of a breathing transformer, a single frame image struggles to capture minute changes. However, through temporal modeling, residual information from multiple frames can be accumulated, resulting in deformation trajectories with higher confidence. Enhanced detection results refer to the dynamic optimization and supplementation of candidate bounding boxes and class confidence in the output after combining temporal modeling. This overcomes the shortcomings of single-frame detection, which is easily affected by changes in illumination, occlusion, or noise. The results are more robust under temporal consistency constraints. For instance, during nighttime inspections of surge arrester leaks, a single frame infrared image may contain random thermal noise. Enhanced detection results can utilize multi-frame temporal consistency to remove noise, ultimately retaining only the persistent, true defect areas.
[0122] In one possible implementation, for the continuous image frame sequence of the target substation equipment corresponding to the high-confidence detection results, an inter-frame residual tensor is constructed, and enhanced detection results are obtained through temporal modeling. Specifically, this includes: using the target bounding box in the high-confidence detection results as input, establishing a cross-frame trajectory sequence of the target substation equipment on the time axis, and performing geometric alignment and scale normalization on the target regions of each frame based on the cross-frame trajectory sequence to generate a registered image sequence; calculating the optical flow field of adjacent frames based on the registered image sequence, and generating a motion-compensated image sequence under the constraints of the optical flow field, and calculating the difference between the registered image sequence and the motion-compensated image sequence in the intensity domain, gradient domain, and edge domain respectively, and stacking them to form an inter-frame residual tensor; performing temporal modeling on the inter-frame residual tensor through the residual temporal path and the deformation spatiotemporal path, and fusing them to generate a temporal semantic tensor; using the temporal semantic tensor as input, performing spatiotemporal consistency constraints in combination with the topological prior of the target substation equipment, and performing temporal smoothing and dynamic enhancement on the target bounding box confidence in the output stage to output the enhanced detection results.
[0123] Specifically, using the target bounding box from the high-confidence detection results as input, cross-frame association is performed on the same target substation equipment on the time axis to establish a cross-frame trajectory sequence. Based on this, geometric alignment and scale normalization are performed on the target regions of each frame to generate a registered image sequence; let the original target cropping of frame t be... And the reference frame is Using similarity transformation The registered image is obtained by mapping to a unified reference coordinate system. Its pixel coordinate transformation is as follows:
[0124]
[0125] in Represents the original pixel coordinates. Indicates the pixel coordinates after registration. Indicates the scale factor. Represents a two-dimensional rotation matrix. This represents the translation vector.
[0126] The similarity transformation parameters are obtained by minimizing the bounding box alignment error, and the objective function is:
[0127]
[0128] in This represents the set of pixels within the target bounding box in frame t. Represent the L2 norm; output the registered image sequence. The similarity transformation parameter set serves as the input for subsequent optical flow estimation and motion compensation.
[0129] The optical flow fields of adjacent frames are calculated based on the registered image sequence, and a motion-compensated image sequence is generated under the constraint of the optical flow field. Simultaneously, the registration differences are calculated in the intensity, gradient, and edge domains to form an inter-frame residual tensor; the registration reference coordinates are defined from... arrive Optical flow field The TV-L1 energy is used for estimation, and the formula is as follows:
[0130]
[0131] in Indicates pixel position, Represents the regularization coefficient. Represents the spatial gradient operator, It represents the first norm.
[0132] Obtain the motion-compensated image:
[0133] Based on this, three types of residuals are defined: strength residuals. Gradient residual Edge residual ,in This represents edge operators such as Sobel magnitudes; finally, the three types of residuals are stacked in the channel dimension to obtain the inter-frame residual tensor. and along with the optical flow field As input for time series modeling.
[0134] Inter-frame residual tensors are temporally modeled and fused using residual temporal paths and deformation spatiotemporal paths to generate temporal semantic tensors. The residual temporal path takes the global description vector of the inter-frame residual tensor as input and uses a gated recurrent unit to extract slowly evolving anomalous trends. Given the input after global average pooling, the gated recurrent unit is updated as follows:
[0135]
[0136]
[0137] in This represents the Sigmoid function. Represents the hyperbolic tangent function. This represents element-wise multiplication. Represents the time-domain hidden state. For learnable parameters; the deformable spatiotemporal path is input with a pixel-level description of the inter-frame residual tensor and cross-frame local deformation is aligned using deformable spatiotemporal attention, defining the position... Output:
[0138]
[0139] in Indicates the radius of the time window. This indicates the number of sampling points at each location. Represent attention weights and satisfy all The sum of is one. This represents the spatial offset predicted by the small convolutional network; the outputs of the two paths are fused in the pixel domain to obtain the temporal semantic tensor.
[0140]
[0141] in The gate weights are generated by a 1x1 convolution and then normalized using a sigmoid function. Indicates the vector Expanding to spatial dimensions with Same size; output As input for spatiotemporal consistency constraints and dynamic enhancements.
[0142] Using a temporal semantic tensor as input, spatiotemporal consistency constraints are applied in conjunction with the topological priors of the target substation equipment. In the output phase, temporal smoothing and dynamic enhancement are performed on the target bounding box confidence to output enhanced detection results. The equipment structure is abstracted as a graph. And based on node features The graph Laplacian matrix represents the pooling result of the temporal semantic tensor in each component region. in Weighted adjacency matrix Given a degree matrix, the closed-form smoothing of the consistency regularization is:
[0143]
[0144] in This represents the node features after topological prior smoothing. Represents the prior strength coefficient. Represents the identity matrix; the class confidence of the target bounding box. With bounding box quality score Temporal smoothing and dynamic enhancement are performed using an exponential moving average:
[0145]
[0146] And the dynamic gain driven by residual energy:
[0147]
[0148] Increase confidence
[0149]
[0150] in Represents the smoothing coefficient. Indicates the gain coefficient. This represents the set of pixels covered by the target bounding box. Describes the norm 1. This means cropping the value to the range of zero to one; ultimately... The results, along with the smoothed bounding box coordinates, are output as enhanced detection results, which are then used as input for subsequent structure restoration or alarm assessment.
[0151] S150. For target substation equipment with complex occlusion in the enhanced detection results, extract the edge response map and input the edge prediction path based on Transformer to complete the structural contour, and output the target boundary and defect location information of the repaired target substation equipment.
[0152] Specifically, complex occlusion refers to situations where the critical structure of the target substation equipment is partially covered by other components or environmental factors, resulting in an incomplete visible area. This includes types such as lateral occlusion, overlapping occlusion, and self-occlusion. An edge response map is a two-dimensional function graph used to characterize the geometric contours and material boundary intensity distribution of the target substation equipment. It is a fusion expression derived from multiple sources, including visible light texture gradient, infrared temperature gradient, and electric field amplitude and phase gradient. The construction method involves using the Sobel or Scharr operator to obtain the texture gradient amplitude for visible light brightness, calculating the temperature gradient amplitude for the infrared temperature field, and calculating the gradient amplitudes for the electric field amplitude and phase field, then weighted and superimposed in the same coordinate system. For example, in a scenario with a small air gap in a busbar flange, the edge response map will show significant brightening at the outer edge of the flange and the bolt edge, while forming a high-response band at locations where temperature rise and amplitude / phase abnormally coincide.
[0153] The Transformer-based edge prediction path encodes discrete edge fragments in the edge response map into token sequences and uses an attention mechanism to retrieve information about missing contours from the context of fused features, thereby predicting continuous edges in occluded areas. Its core is a multi-head attention interaction of query, key, and value, using visible edge segments as queries and infrared and electric field contextual representations as keys and values to achieve cross-modal completion. For example, in a circuit breaker scenario where moving and stationary contacts are partially obscured, the edge prediction path can align the contact boundary direction using infrared temperature gradients and electric field amplitude-phase anomalies to infer the contour of a contact segment obscured by leads. Structural contour completion, based on the edge prediction path output, constrains the curvature continuity and normal smoothness of the predicted missing curves and seen curves, and combines this with the equipment shape prototype and assembly tolerances for geometric consistency, ensuring the final contour matches the actual structure in shape. For example, when completing the contour of a through-wall bushing's skirt, the spacing and curvature of adjacent skirts are constrained to continuously vary within tolerance ranges.
[0154] The target boundary refers to the geometric boundary description of the target substation equipment after structural contour completion and sub-pixel boundary refinement. It is expressed using closed multi-segment splines or polygon vertex sequences and can be used for subsequent dimensional measurement and pose estimation. For example, the target boundary output for the surge arrester body can be directly used to calculate the angle between the body and the ground and the gap with adjacent connecting parts. Defect location information refers to the result of coordinate-based and semantic description of the defect area relative to the target boundary. It includes at least the defect category label, the outer bounding box or pixel mask of the defect connected domain, the coordinates of the defect center point, the confidence level, and the relative topological relationship with the target boundary. It can also include quantitative indicators such as the thermal anomaly amplitude and the electric field anomaly amplitude. For example, in the early heating scenario of the surge arrester joint, the defect location information will provide the pixel mask and temperature rise amplitude of the joint and mark its location at the junction of the joint and the lead.
[0155] In one possible implementation, for target substation equipment with complex occlusions in the enhanced detection results, an edge response map is extracted and input into a Transformer-based edge prediction path for structural contour completion. The output is the target boundary and defect location information of the repaired target substation equipment. Specifically, this includes: generating an infrared gradient response map, an electric field amplitude-phase change response map, and a visible light texture gradient response map within the corresponding target area based on the enhanced detection results; obtaining an edge response map in a unified spatial domain through weighted fusion; performing threshold adaptive processing and edge refinement processing on the edge response map to form an edge fragment set; calculating the curvature continuity and endpoint pairing compatibility of the edge fragment set to generate an edge fragment map; and then, based on the Transformer... In the edge prediction path of r, the edge fragment map is tokenized and its position is encoded. A cross-attention mechanism is used to perform connection inference on the edge fragment map in the context of fused features, and the predicted contour of the missing edge is output. Shape consistency constraint optimization is performed on the predicted contour, and the predicted contour is non-rigidly aligned with the prior distribution structure map of the target substation equipment. Abnormal contours are corrected under edge normal continuity and electrical connection topology constraints to generate structural contour completion results. Boundary refinement and defect response fusion are performed on the structural contour completion results to generate the target boundary of the repaired target substation equipment. Texture residuals, thermal anomaly residuals and electric field anomaly residuals are calculated in the narrow band region of the target boundary to form a defect response map, so as to obtain the repaired target boundary and defect location information.
[0156] Specifically, based on the enhanced detection results, an infrared gradient response map, an electric field amplitude-phase abrupt change response map, and a visible light texture gradient response map are generated within the corresponding target area, and an edge response map in a unified spatial domain is obtained through weighted fusion; let the pixel coordinates in the unified spatial domain be... The infrared temperature field is The electric field amplitude and phase are respectively and Visible light brightness is The infrared gradient response is defined as follows: The visible light texture gradient response is defined as The combined response to the abrupt change in electric field amplitude and phase is:
[0157]
[0158] in For spatial gradient operators, It is a 2-norm. For the Heaviside function, For phase through " "Local differential amplitude after packaging" This is the phase change threshold. These are the modal weighting coefficients; further based on the confidence plot... , , Forming normalized fusion weights:
[0159]
[0160] The final edge response map is obtained as follows:
[0161]
[0162] in To unify the edge response intensity in the spatial domain; in the above symbols, weight It reflects the reliability of each modality at that pixel location. It reflects the proportions of the contributions of electric field amplitude, phase gradient, and phase abrupt change to the electric field response.
[0163] Threshold adaptive processing and edge refinement are performed on the edge response map to form an edge fragment set. The curvature continuity and endpoint pairing compatibility of the edge fragment set are then calculated to generate an edge fragment map. A local statistical threshold is used.
[0164]
[0165] Will The pixel is set to one to obtain a binary edge map. ,in and respectively with The mean and standard deviation of the center window. For threshold gain; Perform morphological skeletonization operator Get refined edges And decompose the edge fragment set by connected components. For each fragment After parameterizing the arc length, the discrete curvature is calculated as follows:
[0166]
[0167] And provide a curvature continuity score:
[0168]
[0169] in To prevent division by zero, For normalization term, For the curvature smoothing scale; for the endpoints of any two fragments , Define endpoint pair compatibility:
[0170]
[0171] in For the endpoint coordinates, The endpoint tangential unit vector, A distance- and direction-compatible scale is used; based on this, an edge fragment graph is constructed with fragments as vertices and endpoint pairings as edge weights. .
[0172] In the Transformer-based edge prediction path, the edge fragment map is tokenized and its position encoded. A cross-attention mechanism is used to perform connection reasoning on the edge fragment map within the context of fused features, outputting the predicted contours of missing edges. Each fragment... Uniform sampling Tokens and form a query matrix ( ), its first The row token feature is obtained by concatenating local edge direction, response intensity, modality source embedding, and position encoding; the position encoding uses:
[0173]
[0174] in For frequency groups. Through the image decoder... above As a node feature, maximize the probability of path coherence:
[0175]
[0176] Finally, a set of predicted contours across occlusion connections is obtained. ,in For a set of connection paths, Embedded for nodes, To connect the rating matrix, For endpoint-compatible weights, This is the Sigmoid function.
[0177] Shape consistency constraint optimization is performed on the predicted contour, and the predicted contour is non-rigidly aligned with the prior distribution structure diagram of the target substation equipment. Abnormal contours are corrected under edge normal continuity and electrical connection topology constraints to generate a structural contour completion result. Let the shape prototype of the corresponding component in the prior distribution structure diagram be... ( (For the arc length parameter), using a thin plate spline field Achieve non-rigid alignment and optimize energy:
[0178]
[0179] in
[0180]
[0181]
[0182] For curvature, It is the normal unit vector. For a topological Laplace constructed based on electrical connections, Stack vectors for key point coordinates. For topological constraint expectation terms, The weights are used to calculate the structural contour completion result. .
[0183] The structural contour completion results are refined by boundary refinement and defect response fusion to generate the target boundary of the repaired target substation equipment. Texture residuals, thermal anomaly residuals, and electric field anomaly residuals are calculated within the narrow band region of the target boundary to form a defect response map, thus obtaining the repaired target boundary and defect location information. As the initial boundary, construct a width of [value] in its normal direction. narrow band area Subpixel-level refinement is achieved using active contour energy, specifically:
[0184]
[0185] in For first-order and second-order smoothing weights, For edge attraction weights, and For the first and second derivatives; in Internally, define texture residuals Thermal anomaly residual Abnormal residual of electric field ,in and The mean and median are within the narrow band. The phase residuals are weighted; after normalization, they are fused into a defect response map.
[0186]
[0187] in Indicated to Normalization, To integrate weights, For bias, For the Sigmoid function; Thresholding and The geometric relationships are combined to output the repaired target boundary and defect location information, including defect pixel mask or bounding box, defect category and confidence level, and relative position with the target boundary, so as to complete the closed loop of structural contour completion and fine positioning.
[0188] In one possible implementation, a multimodal feature fusion network is used to classify the target boundary and defect location information to obtain a classification result. Based on the classification result, the defect category label and defect severity level are output by combining the infrared temperature anomaly distribution feature, electric field leakage phase anomaly feature and visible light texture anomaly feature, and a defect identification result with category information and severity level is output.
[0189] Specifically, the server takes the target boundary and defect location information as input. First, it performs region clipping and alignment on the defect location area surrounded by the target boundary within a unified spatial domain. Visible light images, infrared images, and electric field leakage maps are sampled and extracted using region alignment under the same reference coordinates to extract multimodal region features. Simultaneously, geometric attributes such as the shape description, size ratio, and positional relationship of the target boundary are extracted. Then, in the multimodal feature fusion network, fine-grained texture features, temperature anomaly distribution features, and phase anomaly response features are extracted using visible light texture feature branches, infrared temperature feature branches, and electric field amplitude and phase feature branches, respectively. Then, channel attention and spatial attention are used to perform weighted fusion channel by channel and position by position. At the same time, the geometric attributes of the target boundary and the regional attributes of the defect location information are introduced as prior embeddings to achieve structure-guided discrimination. Finally, in the discrimination head, the fused multimodal region features and prior embeddings are hierarchically aggregated and semantically summarized to output the discrimination result and complete the determination of the defect category label.
[0190] In the severity assessment head, the peak amplitude, spatial expansion, and boundary fit of the infrared temperature anomaly distribution features are quantified based on the discrimination results. The phase consistency and abrupt change stability of the electric field leakage phase anomaly features are quantified. The edge damage and texture disorder of the visible light texture anomaly features are quantified. The above quantified components are weighted and integrated with the acquisition confidence and temporal stability to generate the defect severity level. Finally, the server packages the defect category label, defect severity level, target boundary, and defect location information together into the defect identification result, and provides the category information, severity level, spatial location, and multimodal evidence summary to the outside world. This ensures that the output is traceable and interpretable under the constraints of multimodal consistency, structural prior consistency, and temporal stability.
[0191] This application also provides a defect identification device for substation equipment, referring to... Figure 2 , Figure 2This is a schematic diagram of a defect identification device for substation equipment provided in an embodiment of this application. The device is a server, comprising an acquisition module 21 and a processing module 22. The acquisition module 21 acquires infrared images, electric field leakage maps, and visible light images collected by a multi-channel imaging system deployed at the substation site, generating a multi-channel image tensor. The processing module 22 inputs the multi-channel image tensor into a multi-path recognition model to obtain a fused feature map. The processing module 22 further inputs the fused feature map into a YOLOv8 backbone detection network, combining it with the prior distribution structure map of the target substation equipment, to construct a joint focus area in the infrared high-temperature focusing region, the electric field leakage strong response region, and the optical high-frequency edge region. The processing module 22 is used to form a candidate set containing high-confidence target candidate boxes and generate high-confidence detection results through non-maximum suppression. The processing module 22 is also used to construct inter-frame residual tensors for the continuous image frame sequence of the target substation equipment corresponding to the high-confidence detection results and obtain enhanced detection results through temporal modeling. The processing module 22 is also used to extract edge response maps for target substation equipment with complex occlusion in the enhanced detection results and input the Transformer-based edge prediction path for structural contour completion, and output the target boundary and defect location information of the repaired target substation equipment.
[0192] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0193] This application also provides an electronic device, with reference to... Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: at least one processor 31, at least one network interface 34, a user interface 33, a memory 35, and at least one communication bus 32.
[0194] The communication bus 32 is used to enable communication between these components.
[0195] The user interface 33 may include a display screen and a camera. Optionally, the user interface 33 may also include a standard wired interface and a wireless interface.
[0196] The network interface 34 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0197] The processor 31 may include one or more processing cores. The processor 31 connects to various parts of the server via various interfaces and lines, executing instructions, programs, code sets, or instruction sets stored in the memory 35, and calling data stored in the memory 35 to perform various server functions and process data. Optionally, the processor 31 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 31 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 31 and may be implemented as a separate chip.
[0198] The memory 35 may include random access memory (RAM) or read-only memory. Optionally, the memory 35 may include a non-transitory computer-readable storage medium. The memory 35 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 35 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 35 may also be at least one storage device located remotely from the aforementioned processor 31. Figure 3 As shown, the memory 35, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a defect identification method for substation equipment.
[0199] exist Figure 3In the electronic device shown, the user interface 33 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 31 can be used to call an application program stored in the memory 35 for a defect identification method for substation equipment. When executed by one or more processors, the electronic device executes one or more methods as described in the above embodiments.
[0200] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0201] This application also provides a computer-readable storage medium storing instructions. When executed by one or more processors, these instructions cause an electronic device to perform one or more of the methods described in the above embodiments.
[0202] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truth. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A method for defect identification in substation equipment, characterized in that, The method includes: Acquire infrared images, electric field leakage maps, and visible light images collected by a multi-channel imaging system deployed at the substation site, and generate multi-channel image tensors; The multi-channel image tensor is input into the multi-path recognition model to obtain the fused feature map; The fused feature map is input into the YOLOv8 backbone detection network. Combined with the prior distribution structure map of the target substation equipment, a joint region of interest is constructed in the infrared high-temperature focusing area, the electric field leakage strong response area, and the optical high-frequency edge area to form a candidate set containing high-confidence target candidate boxes. High-confidence detection results are generated through non-maximum suppression. For the continuous image frame sequence of the target substation equipment corresponding to the high-confidence detection results, an inter-frame residual tensor is constructed, and enhanced detection results are obtained through temporal modeling. For target substation equipment with complex occlusion in the enhanced detection results, the edge response map is extracted and the edge prediction path based on Transformer is input to complete the structural contour, and the target boundary of the repaired target substation equipment is output. Texture residuals, thermal anomaly residuals, and electric field anomaly residuals are calculated within a narrow band region of the target boundary to form a defect response map, thereby obtaining defect location information.
2. The defect identification method for substation equipment according to claim 1, characterized in that, The process of acquiring infrared images, electric field leakage maps, and visible light images collected by a multi-channel imaging system deployed at the substation site, and generating a multi-channel image tensor, specifically includes: By establishing a clock synchronization mechanism for the multi-channel imaging system, the synchronous acquisition of infrared images, electric field leakage spectra, and visible light images is triggered. A unified timing source and hard trigger pulse are used to eliminate the acquisition time difference and record the timestamp and trigger sequence number as a synchronization reference. Multi-pose acquisition is performed based on a calibration target with high contrast corners. The intrinsic and extrinsic parameter matrices and distortion coefficients of the infrared imager, visible light imager and electric field imager are solved to establish a cross-modal projection mapping relationship in order to obtain geometric calibration results. Radiometric calibration based on the blackbody two-point method is performed on the infrared image, exposure response linearization based on the grayscale standard plate is performed on the visible light image, and amplitude and phase correction based on the standard field source is performed on the electric field leakage spectrum to eliminate intensity scale differences between modes. Based on the geometric calibration results, the infrared image and the electric field leakage spectrum are reprojected onto the visible light image plane, and pixel-level alignment is performed by solving the homography matrix based on the robust matching method. A unified spatial domain grid is established based on the resolution of the visible light image. The infrared image and the electric field leakage spectrum are mapped to the unified spatial domain grid using an interpolation method that preserves edge features to form a multimodal aligned image in the unified spatial domain. The multimodal aligned image is then normalized and stacked to construct the multichannel image tensor.
3. The defect identification method for substation equipment according to claim 1, characterized in that, The step of inputting the multi-channel image tensor into the multi-path recognition model to obtain the fused feature map specifically includes: The multi-channel image tensor is input into the infrared feature extraction branch, electromagnetic spectrum feature extraction branch and visible light feature extraction branch of the multi-path recognition model respectively to extract infrared features, electromagnetic features and visible light features; The infrared feature, the electromagnetic feature, and the visible light feature are weighted channel-by-channel and position-by-position to obtain the weighted feature. The weighted features are subjected to band-separated correction for high-frequency details and low-frequency background, and the energy distribution of different modal features is balanced by residual gating in the spatial domain to output spectrally corrected fused features. Based on the spectrally corrected fusion features, the visible light features are used as queries, and the infrared features and electromagnetic features are used as keys and values to perform multi-head interactions, capturing the complementary relationship between thermal anomalies and electric field leakage on optical textures, and maintaining the original geometric layout through residual connections to obtain complementary enhancement features; Based on the complementary enhancement features and the prior distribution structure map of the target substation equipment, the fused feature map is output.
4. The defect identification method for substation equipment according to claim 1, characterized in that, The process involves inputting the fused feature map into the YOLOv8 backbone detection network, combining it with the prior distribution structure map of the target substation equipment, constructing a joint region of interest in the infrared high-temperature focusing region, the electric field leakage strong response region, and the optical high-frequency edge region, forming a candidate set containing high-confidence target candidate boxes, and generating high-confidence detection results through non-maximum suppression. Specifically, this includes: The multi-scale features of the fused feature map are extracted through the feature pyramid and path aggregation network in the YOLOv8 backbone detection network and input into the decoupled detection head. The decoupled detection head simultaneously outputs the bounding box position parameters and the target category confidence. Using the prior distribution structure map of the target substation equipment as input, a spatial guidance mask is generated and combined with the attention weight of the fused feature map to map the prior constraints onto the multi-scale feature map to generate prior weighted features. In the prior weighted features, the spatial overlap ratio and average response intensity of the candidate box with the infrared high-temperature focusing area response map, the electric field leakage strong response area response map and the optical high-frequency edge area response map are calculated to obtain the multimodal weighted confidence of the candidate box, and a weighted candidate box set is formed based on the multimodal weighted confidence. Non-maximum suppression is performed on the weighted candidate box set. A suppression metric based on distance intersection-union ratio is used to remove candidate boxes with overlap exceeding a threshold and retain the candidate box with the highest fusion confidence to output the high-confidence detection result.
5. The defect identification method for substation equipment according to claim 1, characterized in that, The process involves constructing an inter-frame residual tensor for the continuous image frame sequence of the target substation equipment corresponding to the high-confidence detection results, and obtaining enhanced detection results through temporal modeling. Specifically, this includes: Using the target bounding box in the high-confidence detection result as input, a cross-frame trajectory sequence of the target substation equipment is established on the time axis, and geometric alignment and scale normalization are performed on the target region of each frame based on the cross-frame trajectory sequence to generate a registered image sequence. Based on the registered image sequence, the optical flow field of adjacent frames is calculated, and a motion-compensated image sequence is generated under the constraint of the optical flow field. The difference between the registered image sequence and the motion-compensated image sequence is calculated in the intensity domain, gradient domain and edge domain respectively, and stacked to form the inter-frame residual tensor. The inter-frame residual tensor is used to perform temporal modeling through the residual temporal path and the deformation spatiotemporal path, and then fused to generate a temporal semantic tensor. Using the temporal semantic tensor as input, spatiotemporal consistency constraints are performed in conjunction with the topological priors of the target substation equipment. In the output stage, temporal smoothing and dynamic enhancement are performed on the confidence of the target bounding box, and the enhanced detection result is output.
6. The defect identification method for substation equipment according to claim 1, characterized in that, For the target substation equipment with complex occlusions in the enhanced detection results, the edge response map is extracted and input into the Transformer-based edge prediction path for structural contour completion. The target boundary and defect location information of the repaired target substation equipment are output, specifically including: Based on the enhanced detection results, an infrared gradient response map, an electric field amplitude-phase abrupt change response map, and a visible light texture gradient response map are generated in the corresponding target area, and an edge response map in a unified spatial domain is obtained by weighted fusion. Threshold adaptive processing and edge refinement processing are performed on the edge response map to form an edge fragment set, and the curvature continuity and endpoint pairing compatibility of the edge fragment set are calculated to generate an edge fragment map; In the Transformer-based edge prediction path, the edge fragment map is tokenized and its position is encoded. The edge fragment map is then connected and inferred in the context of fused features using a cross-attention mechanism, and the predicted contour of the missing edge is output. The predicted contour is optimized by shape consistency constraints. The predicted contour is non-rigidly aligned with the prior distribution structure diagram of the target substation equipment. Abnormal contours are corrected under edge normal continuity and electrical connection topology constraints to generate structural contour completion results. The boundary refinement and defect response fusion of the structural contour completion result are performed to generate the target boundary of the repaired target substation equipment.
7. The defect identification method for substation equipment according to claim 1, characterized in that, The method further includes: Using target boundary and defect location information as input, the defect location area surrounded by the target boundary is first cropped and aligned in a unified spatial domain. Visible light image, infrared image and electric field leakage spectrum are sampled and extracted in the same reference coordinates to extract multimodal region features. At the same time, the shape description, size ratio and positional relationship of the target boundary are extracted. Subsequently, in the multimodal feature fusion network, fine-grained texture features, temperature anomaly distribution features, and phase anomaly response features are extracted by visible light texture feature branch, infrared temperature feature branch, and electric field amplitude and phase feature branch, respectively. Then, channel attention and spatial attention are used to perform weighted fusion channel by channel and position by position. At the same time, the geometric attributes of the target boundary and the regional attributes of defect localization information are introduced as prior embedding to achieve structure-guided discrimination. Next, the fused multimodal region features and prior embeddings are hierarchically aggregated and semantically summarized in the discriminant head, and the discriminant results are output and the defect category label is determined. In the severity assessment head, the peak amplitude, spatial expansion and boundary fit of the infrared temperature anomaly distribution feature are quantified based on the discrimination results, the phase consistency and abrupt change stability of the electric field leakage phase anomaly feature are quantified, and the edge damage and texture disorder of the visible light texture anomaly feature are quantified. The quantified components are then weighted and integrated with the acquisition confidence and temporal stability to generate the defect severity level. Finally, the defect category label, defect severity level, target boundary, and defect location information are packaged together into the defect identification result.
8. A defect identification device for substation equipment, characterized in that, The device is used to perform the defect identification method for substation equipment as described in any one of claims 1 to 7, the device comprising an acquisition module (21) and a processing module (22), wherein, The acquisition module (21) is used to acquire infrared images, electric field leakage spectra and visible light images collected by the multi-channel imaging system deployed at the substation site, and generate multi-channel image tensors; The processing module (22) is used to input the multi-channel image tensor into the multi-path recognition model to obtain a fused feature map; The processing module (22) is also used to input the fused feature map into the YOLOv8 backbone detection network, combine it with the prior distribution structure map of the target substation equipment, construct a joint domain of interest in the infrared high temperature focusing area, the electric field leakage strong response area and the optical high frequency edge area, form a candidate set containing high confidence target candidate boxes, and generate high confidence detection results through non-maximum suppression; The processing module (22) is also used to construct an inter-frame residual tensor for the continuous image frame sequence of the target substation equipment corresponding to the high confidence detection result, and obtain the enhanced detection result through time series modeling; The processing module (22) is also used to extract the edge response map of the target substation equipment with complex occlusion in the enhanced detection result, input the edge prediction path based on Transformer to complete the structural outline, and output the target boundary of the repaired target substation equipment. The processing module (22) is also used to calculate the texture residual, thermal anomaly residual and electric field anomaly residual in the narrow band area of the target boundary to form a defect response map in order to obtain defect location information.
9. An electronic device, characterized in that, The electronic device includes a processor (31), a memory (35), a user interface (33), and a network interface (34). The memory (35) is used to store instructions. The user interface (33) and the network interface (34) are both used to communicate with other devices. The processor (31) is used to execute the instructions stored in the memory (35) to cause the electronic device to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Defect detection method for high-voltage equipment based on deep learning and multispectral image fusion
CN120355722A
Substation equipment defect real-time identification and early warning system based on image fusion technology
CN120525879A