SAM-based machine tour video segmentation method and system

CN122799331APending Publication Date: 2026-09-22GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610936080.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0003]然而,在电力设备中的避雷器伞裙检测中,表面附着物的存在会直接干扰对伞裙边缘和背景的准确区分,这种干扰会因为附着物覆盖面积的不同而呈现出不同的影响程度,例如覆盖面积较小时可能仅影响局部特征,而覆盖面积稍大时则可能导致整体特征的混淆,这种由覆盖面积变化引发的特征区分困难,成为识别过程中一个独特的挑战,如何从复杂的视频画面中精准识别伞裙并分析其状态,始终是该领域亟待突破的重点,当前,尽管许多方法能够对巡检视频中的设备进行初步识别,但这些方法往往在面对复杂环境干扰时表现不佳

Benefits of technology

从无人机采集的巡检视频中分割出避雷器伞裙区域的像素级标注图;对所述像素级标注图中的伞裙区域进行附着物识别,进而确定附着物的覆盖比例,当所述覆盖比例超过预设阈值时,增强边缘轮廓,得到边缘特征图;通过所述边缘特征图对所述巡检视频中的视频帧进行附着物背景分离,得到附着物的干扰分布区域;基于预设的SAM模型对所述干扰分布区域进行特征优化重构,当重构特征值低于伞裙匹配阈值时,进行局部放大,得到优化后的特征重构图;根据所述特征重构图分析附着物面积变化引起的动态干扰模式,得到附着物面积变化的量化指标图;对所述量化指标图和所述像素级标注图进行匹配校准,并根据匹配校准结果确定避雷器伞裙的评估报告;根据所述评估报告更新所述巡检视频的元数据标签。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799331A_ABST
    Figure CN122799331A_ABST
Patent Text Reader

Abstract

The application provides a SAM-based machine tour video segmentation method and system, relates to the technical field of video analysis, and achieves pixel-level annotation of a lightning arrester umbrella skirt area by segmenting the lightning arrester umbrella skirt area from a tour video; determines the coverage ratio of the attached object, enhances the edge contour when the coverage ratio exceeds a preset threshold, and obtains an edge feature map; separates the attached object background from the video frames in the tour video, and obtains the interference distribution area of the attached object; optimizes and reconstructs the features of the interference distribution area, and obtains an optimized feature reconstruction map; analyzes the dynamic interference mode caused by the area change of the attached object, and obtains a quantitative index map of the area change of the attached object; matches and calibrates the quantitative index map and the pixel-level annotation map, determines the evaluation report of the lightning arrester umbrella skirt, and updates the metadata label of the tour video according to the evaluation report. The application can accurately distinguish the lightning arrester umbrella skirt from the background in the tour video, and effectively deal with the feature interference caused by the area change of the attached object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video analysis technology, and more specifically, to a machine-based video segmentation method and system based on SAM. Background Technology

[0002] In the field of power equipment inspection, the application of video analytics technology is of paramount importance in ensuring the safe operation of the power grid. With the popularization of drone inspection, real-time monitoring of equipment status through video data has become a key direction for industry development. This technology can not only improve inspection efficiency but also effectively reduce the risks of manual inspection.

[0003] However, in the inspection of surge arrester skirts in power equipment, the presence of surface deposits directly interferes with the accurate differentiation between the skirt edges and the background. This interference varies in degree depending on the area covered by the deposits; for example, a small coverage area may only affect local features, while a larger coverage area may lead to confusion of overall features. This difficulty in feature differentiation caused by changes in coverage area poses a unique challenge in the identification process. How to accurately identify the skirts and analyze their condition from complex video footage remains a key area requiring breakthroughs in this field. Currently, although many methods can perform preliminary identification of equipment in inspection videos, these methods often perform poorly when faced with complex environmental interference. Especially under the influence of natural factors, various deposits may appear on the equipment surface, leading to a decrease in identification accuracy. Existing technologies lack adaptability to dynamic environmental changes when dealing with this type of interference, especially when the characteristics of the interfering object are highly similar to those of the target object, easily leading to misjudgments and affecting the reliability of subsequent analysis. Therefore, how to accurately distinguish surge arrester skirts from the background in inspection videos while effectively addressing feature interference caused by changes in the coverage area of ​​deposits has become a key problem that this research urgently needs to solve. Summary of the Invention

[0004] This application provides a method and system for video segmentation based on SAM (Surf Arrester Aperture Modulation), which can accurately distinguish between the lightning arrester skirt and the background in the inspection video, while effectively dealing with feature interference caused by changes in the coverage area of ​​attached objects.

[0005] Firstly, this application provides a machine-guided video segmentation method based on SAM, comprising the following steps: Pixel-level labeled images of the lightning arrester skirt area were segmented from inspection videos collected by drones; The umbrella-shaped region in the pixel-level annotation map is subjected to attachment identification, and the coverage ratio of the attachment is determined. When the coverage ratio exceeds a preset threshold, the edge contour is enhanced to obtain an edge feature map. The edge feature map is used to separate the background of the attachments in the video frames of the inspection video to obtain the interference distribution area of ​​the attachments. Based on the preset SAM model, the interference distribution area is reconstructed by feature optimization. When the reconstructed feature value is lower than the umbrella skirt matching threshold, local magnification is performed to obtain the optimized feature reconstruction map. Based on the feature reconstruction map, the dynamic interference pattern caused by the change in the area of ​​the attached material is analyzed, and a quantitative index map of the change in the area of ​​the attached material is obtained. The quantitative index map and the pixel-level annotation map are matched and calibrated, and the evaluation report of the surge arrester skirt is determined based on the matching and calibration results; Update the metadata tags of the inspection video according to the assessment report.

[0006] In this embodiment, the pixel-level annotation map segmented from the inspection video collected by the drone specifically includes: Extracting time-series continuous image sequences from inspection video data collected by drones; Based on the preset SAM model, the lightning arrester skirt region is segmented for each video frame in the time-series continuous image sequence to obtain the binary mask of the lightning arrester skirt region. The lightning arrester skirt area is pixel-level labeled according to the binary mask to obtain a pixel-level labeled map of the lightning arrester skirt area.

[0007] In this embodiment, identifying attachments in the umbrella-shaped area of ​​the pixel-level annotation map and then determining the coverage ratio of the attachments specifically includes: The attachments in the umbrella skirt region of the pixel-level annotation map are identified pixel by pixel, and the umbrella skirt body pixels and attachment pixels are distinguished to obtain the attachment pixel set and the umbrella skirt body pixel set. The coverage ratio of the attachment is determined based on the set of pixels of the attachment and the set of pixels of the umbrella skirt body.

[0008] In this embodiment, when the coverage ratio exceeds a preset threshold, enhancing the edge contour to obtain an edge feature map specifically includes: When the coverage ratio exceeds the preset threshold, interference compensation is activated; The optimal response scale is determined based on the density distribution of edge points in each scale layer of the multi-scale edge point set in the umbrella skirt region. By connecting the edges of the skirt region using the optimal response scale, a continuous skirt outline is obtained; The umbrella skirt region is expanded based on the umbrella skirt outline to obtain an edge feature map.

[0009] In this embodiment, the process of separating the background of the attachments in the video frames of the inspection video using the edge feature map to obtain the interference distribution area of ​​the attachments specifically includes: The edge feature map is used to perform brightness equalization fusion on the video frames in the inspection video to obtain a fusion result map. Based on the fusion result image, background separation is performed using color clustering to obtain the foreground umbrella skirt image; The interference pixels of the soiled umbrella skirt are extracted from the foreground umbrella skirt image to obtain the interference pixel distribution map of the attached material; The interference distribution area of ​​the attachment is determined based on the interference pixel distribution map.

[0010] In this embodiment, the interference distribution area is reconstructed and its features are optimized based on a preset SAM model. When the reconstructed feature value is lower than the skirt matching threshold, local magnification is performed to obtain the optimized feature reconstruction map, specifically including: Based on the preset SAM model, the interference distribution area is reconstructed by feature optimization to obtain a reconstructed feature map; The reconstructed feature map is matched with a preset umbrella skirt standard feature template to obtain the global matching degree; Based on the global matching degree, the reconstructed feature map is marked with low-matching regions to obtain a low-matching region mask; The low-matching region mask is locally magnified, and the magnified region is edge-enhanced to obtain an optimized feature reconstruction map.

[0011] In this embodiment, the dynamic interference pattern caused by the change in the area of ​​the attached material is analyzed based on the feature reconstruction map, and the resulting quantitative index map of the change in the area of ​​the attached material specifically includes: Extract the outline of the attachment region based on the reconstructed feature map; A temporal variation sequence of the area covered by the attachment region is constructed by the outline of the attachment region; The boundaries of the attachments in each frame of the time-series change sequence are tracked to obtain a boundary change trend map; A quantitative index map of the change in the area of ​​the attachment is determined based on the boundary change trend map.

[0012] In this embodiment, the matching calibration of the quantitative index map and the pixel-level annotation map, and the determination of the evaluation report of the surge arrester skirt based on the matching calibration results, specifically include: Spatial registration is performed between the quantized index map and the pixel-level annotation map to obtain the matching calibration result; Based on the matching calibration results, a decision tree-based classification of the contamination status is performed to obtain the contamination category identifier of the surge arrester skirt; An alarm is determined based on the soiling category identifier. When the soiling category identifier is a severe soiling category, a skirt soiling alarm signal is generated, and the trigger time of the skirt soiling alarm signal is recorded. Based on the pollution category identifier and the matching calibration results, the location number, pollution level, attachment coverage area ratio and attachment expansion rate of the surge arrester skirt are summarized to obtain the skirt pollution summary information. The arrester skirt contamination alarm signal, the trigger time, and the skirt contamination summary information are filled in according to a preset report template to obtain an assessment report of the arrester skirt.

[0013] In this embodiment, updating the metadata tags of the inspection video according to the evaluation report specifically includes: Based on the contamination level, coverage ratio, expansion speed and umbrella skirt location number in the assessment report, metadata tags are generated for the inspection video, resulting in an inspection video file with metadata tags. The sampling interval of subsequent inspection video processing tasks is adjusted based on the metadata tags to obtain inspection video processing tags that match the state of dirtiness.

[0014] Secondly, this application provides a machine-guided video segmentation system based on SAM, used to execute a machine-guided video segmentation method based on SAM, the machine-guided video segmentation system comprising: The pixel annotation module is used to segment out pixel-level annotation maps of the lightning arrester skirt area from inspection videos collected by drones; The attachment recognition and edge enhancement module is used to identify attachments in the umbrella skirt area of ​​the pixel-level annotation map, and then determine the coverage ratio of the attachments. When the coverage ratio exceeds a preset threshold, the edge contour is enhanced to obtain an edge feature map. The attachment background separation module is used to separate the attachment background from the video frames in the inspection video using the edge feature map to obtain the interference distribution area of ​​the attachment. The feature reconstruction optimization module is used to perform feature optimization and reconstruction on the interference distribution area based on the preset SAM model. When the reconstructed feature value is lower than the umbrella skirt matching threshold, local magnification is performed to obtain the optimized feature reconstruction map. The attachment area dynamic analysis module analyzes the dynamic interference pattern caused by the change in attachment area based on the feature reconstruction map, and obtains a quantitative index map of the change in attachment area. The evaluation report generation module is used to match and calibrate the quantitative index map and the pixel-level annotation map, and determine the evaluation report of the surge arrester skirt based on the matching and calibration results; The inspection video update module is used to update the metadata tags of the inspection video according to the evaluation report.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: Pixel-level annotations of the surge arrester skirt area are segmented from inspection videos collected by drones. Attachments are identified within the skirt area of ​​the pixel-level annotations to determine their coverage ratio. When the coverage ratio exceeds a preset threshold, the edge contour is enhanced to obtain an edge feature map. The edge feature map is used to separate the attachment background from the video frames in the inspection video, revealing the interference distribution area of ​​the attachments. Based on a preset SAM model, the interference distribution area is reconstructed using feature optimization. When the reconstructed feature value is lower than the skirt matching threshold, local magnification is performed to obtain an optimized feature reconstruction map. The dynamic interference pattern caused by changes in attachment area is analyzed based on the feature reconstruction map to obtain a quantitative index map of attachment area changes. The quantitative index map and the pixel-level annotations are matched and calibrated, and an evaluation report for the surge arrester skirt is determined based on the matching calibration results. The metadata tags of the inspection video are updated based on the evaluation report.

[0016] Therefore, in this application, the metadata tags of the inspection video can be updated according to the evaluation report. Firstly, by segmenting the pixel-level annotation map of the lightning arrester skirt area from the inspection video collected by the drone, the target area of ​​the skirt can be determined in complex video footage, providing an accurate spatial basis for attachment identification and background interference elimination, reducing the impact of irrelevant background on skirt edge recognition. Secondly, by identifying attachments in the skirt area of ​​the pixel-level annotation map and determining the attachment coverage ratio, the edge contour is enhanced when the coverage ratio exceeds a preset threshold. This adaptively strengthens the skirt boundary features according to the degree of change in attachment coverage area, avoiding local missegmentation caused by small-area attachments, while mitigating the problem of skirt contour confusion with background features caused by larger-area attachments. Furthermore, the attachment background is separated from the video frame using the edge feature map to obtain the interference distribution area of ​​the attachments, and the interference distribution area is reconstructed based on a preset SAM model. When the reconstructed feature value is low... By performing local magnification at the umbrella skirt matching threshold, targeted recovery of umbrella skirt features obscured or weakened by attached objects can be achieved, improving the continuous segmentation capability and feature expression stability of the umbrella skirt region under complex interference. Furthermore, by analyzing the dynamic interference patterns caused by changes in the area of ​​attached objects based on the feature reconstruction map, a quantitative index map of the changes in the area of ​​attached objects is obtained. This transforms the impact of changes in the coverage area of ​​attached objects on umbrella skirt recognition into a calibrable quantitative basis, thereby enhancing the adaptability of the segmentation process to changes in the dynamic environment. Finally, by matching and calibrating the quantitative index map and the pixel-level annotation map, and determining the evaluation report of the surge arrester umbrella skirt based on the matching and calibration results, and then updating the metadata tags of the inspection video, a closed-loop correction can be performed on the umbrella skirt segmentation results and the attached object interference analysis results. This reduces the probability of misjudgment when the attached object and umbrella skirt features are highly similar, thereby improving the overall accuracy of distinguishing the surge arrester umbrella skirt from the background in the inspection video, the reliability of the status assessment, and the traceability of subsequent video data management.

[0017] In summary, the technical solution adopted in this application can accurately distinguish between the lightning arrester skirt and the background in the inspection video, while effectively dealing with the characteristic interference caused by changes in the coverage area of ​​the attached objects. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this embodiment of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is an exemplary flowchart of a machine-based video segmentation method according to the present application; Figure 2 This is a module structure diagram of a machine-based video segmentation system based on SAM provided in this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] This application provides a method and system for video segmentation based on SAM (Search Engine Analysis). The core of this method is to segment a pixel-level annotated map of the surge arrester skirt area from inspection videos collected by a drone; identify attachments in the skirt area of ​​the pixel-level annotated map to determine the coverage ratio of the attachments; when the coverage ratio exceeds a preset threshold, enhance the edge contour to obtain an edge feature map; separate the attachment background from the video frames in the inspection video using the edge feature map to obtain the interference distribution area of ​​the attachments; perform feature optimization and reconstruction on the interference distribution area based on a preset SAM model; when the reconstructed feature value is lower than the skirt matching threshold, perform local magnification to obtain an optimized feature reconstruction map; analyze the dynamic interference pattern caused by changes in the attachment area based on the feature reconstruction map to obtain a quantitative index map of the attachment area change; perform matching calibration on the quantitative index map and the pixel-level annotated map, and determine an evaluation report for the surge arrester skirt based on the matching calibration results; and update the metadata tags of the inspection video based on the evaluation report.

[0022] Example 1: To better understand the above technical solution, the following will provide a detailed description of the technical solution in conjunction with the accompanying drawings and specific implementation methods. (Refer to...) Figure 1 As shown in the figure, this is an exemplary flowchart of a machine-based video segmentation method according to this embodiment of the present application, which includes the following steps: In step S1, a pixel-level labeled map of the lightning arrester skirt area is segmented from the inspection video collected by the drone.

[0023] In this embodiment, the pixel-level annotation map of the lightning arrester skirt area segmented from the inspection video collected by the UAV can be achieved by the following steps: Extracting time-series continuous image sequences from inspection video data collected by drones; Based on the preset SAM model, the lightning arrester skirt region is segmented for each video frame in the time-series continuous image sequence to obtain the binary mask of the lightning arrester skirt region. The lightning arrester skirt area is pixel-level labeled according to the binary mask to obtain a pixel-level labeled map of the lightning arrester skirt area.

[0024] It should be noted that the temporally continuous image sequence mentioned in this application refers to a set of inspection images arranged in the order of video acquisition time and after inter-frame displacement compensation; the SAM model is the Segment Anything Model, which is a cue-based image segmentation model consisting of an image encoder, a cue encoder, and a mask decoder, used to segment the surge arrester skirt region from the input image under the guidance of point cues or box cues and output the corresponding segmentation mask; the mask confidence indicates the reliability of the segmentation mask output by the SAM model matching the real area of ​​the skirt; the binary mask represents the pixel category map used to distinguish the surge arrester skirt region from the background region; the pixel-level annotation map represents the annotation result after assigning target category labels to each pixel in the skirt region; the annotation confidence indicates the reliability of the current pixel annotation result matching the standard boundary features of the skirt.

[0025] In specific implementation, firstly, key frames are extracted from the drone inspection video at preset time intervals, and the inter-frame displacement vector is obtained based on the change of the center coordinates of the surge arrester device in adjacent key frames. When the inter-frame displacement vector exceeds a preset inter-frame displacement threshold, a linear interpolation method is used to supplement transition frames. The image set of key frames and transition frames sorted by acquisition time is used as a temporally continuous image sequence. Secondly, each video frame in the temporally continuous image sequence is subjected to size normalization and brightness normalization processing. The processed RGB image is input into a preset SAM model. The SAM model includes an image encoder, a cue encoder, and a mask decoder. The image encoder of the SAM model performs feature embedding on the processed RGB image to obtain image feature embedding. The geometric center point and bounding box of the surge arrester skirt area are used as the segmentation cue input to the SAM model. A prompt encoder of type A is used to obtain the prompt embedding. The image feature embedding and prompt embedding are fused and decoded by the mask decoder of the SAM model. The output is a skirt region segmentation mask of the same size as the input image and the corresponding mask confidence. Pixels with a mask confidence higher than a preset segmentation threshold are set as skirt regions, and the remaining pixels are set as background regions. The obtained binarization result is used as the binary mask of the surge arrester skirt region. Finally, the binary mask is labeled with connected components. Discrete noise connected components with an area less than 5% of the candidate skirt region are removed. The mask boundaries, circumscribed rectangles and central skeleton lines of the remaining connected components are extracted. The annotation results of the previous and next frames are mapped to the coordinate system of the current frame according to the device displacement vector of the adjacent frames. Low confidence pixels are weighted and fused for correction. The corrected pixel category correspondence is used as the pixel-level annotation map of the surge arrester skirt region.

[0026] For example, in one implementation, when a drone is used to inspect a surge arrester, a video acquisition device records the inspection process at a rate of 25 frames per second. The video data includes continuous images of the surge arrester at different angles and distances, where the surge arrester skirt, as a key monitoring component, needs to be accurately identified. Specifically, when extracting a temporally continuous image sequence, keyframes are extracted from the original video according to a preset 5-frame interval. The inter-frame displacement vector is obtained by calculating the pixel coordinate changes of the center point of the surge arrester in adjacent keyframes. When the magnitude of the displacement vector exceeds a preset threshold of 10% of the image width, a bilinear interpolation algorithm is used to generate a transition frame between the two keyframes. The interpolation weight is determined based on the linear relationship of the timestamps.

[0027] For example, the SAM model is a cue-based image segmentation model consisting of three parts: an image encoder, a cue encoder, and a mask decoder. The image encoder uses a visual Transformer structure to perform one-time feature embedding on the input image. The cue encoder encodes point cues and bounding box cues into cue embeddings. The mask decoder performs cross-attention fusion of the image feature embeddings and the cue embeddings, outputting a segmentation mask and its corresponding confidence score. The network input is a normalized 1024×1024 pixel RGB image, and the output is a segmentation mask of the same size for the umbrella-shaped region, where a mask value of 1 represents the umbrella-shaped region and a mask value of 0 represents the background region. When applying the aforementioned segmentation model, the image encoder first performs feature embedding on each video frame. Then, the geometric center point of the lightning arrester skirt region is used as the foreground cue, and the bounding box of the skirt region is used as the bounding box cue. The input to the cue encoder is then the mask decoder outputs the segmentation mask of the lightning arrester skirt region. When judging the integrity of the skirt boundary, the gradient magnitude of the mask edge is calculated by the Sobel operator. When the standard deviation of the gradient magnitude of consecutive edge pixels is less than a preset threshold, the boundary segment is considered to have good continuity, thus obtaining the preliminary segmentation result of the skirt region.

[0028] It should be noted that during the generation of pixel-level labeled maps, the label confidence is determined by calculating the Euclidean distance between the gradient value of each pixel at the segmentation boundary and the preset standard boundary gradient template of the umbrella skirt. The standard boundary gradient template is pre-established based on the ideal edge features of the uncontaminated umbrella skirt. When the label confidence of a pixel is lower than a preset threshold of 0.7, the labeling results of the corresponding spatial locations in the two frames before and after are extracted, and a weighted average is performed using the reciprocal of the time distance as the weight to obtain the corrected label value.

[0029] In step S2, the attachments in the umbrella-shaped region of the pixel-level annotation map are identified to determine the coverage ratio of the attachments. When the coverage ratio exceeds a preset threshold, the edge contour is enhanced to obtain an edge feature map.

[0030] In this embodiment, the identification of attachments in the umbrella-shaped area of ​​the pixel-level annotation map, and the determination of the coverage ratio of the attachments, can be achieved by the following steps: The attachments in the umbrella skirt region of the pixel-level annotation map are identified pixel by pixel, and the umbrella skirt body pixels and attachment pixels are distinguished to obtain the attachment pixel set and the umbrella skirt body pixel set. The coverage ratio of the attachment is determined based on the set of pixels of the attachment and the set of pixels of the umbrella skirt body.

[0031] It should be noted that, in this application, the candidate pixels of the attachments refer to pixels that deviate from the standard features of the clean umbrella skirt in terms of color, brightness, or texture and may belong to the surface attachments; the set of attachment pixels refers to the set of pixels used to characterize the coverage area of ​​the attachments after spatial connectivity filtering; the set of umbrella skirt body pixels refers to the set of body pixels in the umbrella skirt area that are not classified into the coverage area of ​​the attachments; and the coverage ratio refers to the proportion of the number of attachment pixels to the number of effective pixels in the umbrella skirt area.

[0032] In specific implementation, firstly, the umbrella skirt region in the pixel-level annotation map is scanned pixel by pixel to extract the RGB channel value, grayscale value, and local texture gradient of each pixel. This is then compared with the pre-established standard color range and standard texture range of the umbrella skirt sample without attachments. Pixels that exceed the standard range and are continuous with neighboring pixels are selected as candidate pixels for attachments. Secondly, an eight-neighbor connected region growing method is used to screen spatially continuous regions from the candidate pixels for attachments. Isolated regions with an area smaller than a preset noise area threshold are removed. The retained continuous pixel regions are used as the attachment pixel set, and the pixels in the umbrella skirt region that are not included in the attachment pixel set are used as the umbrella skirt body pixel set. Finally, a pixel counting method is used to count the number of pixels in the attachment pixel set and the total number of effective pixels in the umbrella skirt region. The proportion of the number of attachment pixels in the total number of effective pixels is used as the coverage ratio of the attachments.

[0033] In this embodiment, when the coverage ratio exceeds a preset threshold, enhancing the edge contour to obtain an edge feature map can be achieved through the following steps: When the coverage ratio exceeds the preset threshold, interference compensation is activated; The optimal response scale is determined based on the density distribution of edge points in each scale layer of the multi-scale edge point set in the umbrella skirt region. By connecting the edges of the skirt region using the optimal response scale, a continuous skirt outline is obtained; The umbrella skirt region is expanded based on the umbrella skirt outline to obtain an edge feature map.

[0034] It should be noted that, in this application, the interference compensation refers to the process of enhancing the edge of the umbrella skirt when the coverage of the attachment exceeds a preset threshold; the multi-scale edge point set refers to the set of candidate edge pixels of the umbrella skirt extracted at different smoothing scales; the optimal response scale refers to the edge extraction scale that can take into account both edge continuity and noise suppression effect; the umbrella skirt contour line refers to the outer edge of the umbrella skirt and the sheet boundary line formed by connecting effective edge points; and the edge feature map refers to the image containing the enhanced umbrella skirt contour response and used for subsequent attachment background separation.

[0035] In specific implementation, firstly, the coverage ratio of the attached material is compared with a preset coverage ratio threshold. The preset coverage ratio threshold can be set statistically based on the edge missegmentation rate of clean umbrella skirt samples and attached material occlusion samples. When the coverage ratio exceeds the preset threshold, the current umbrella skirt area is marked as an area requiring interference compensation. Secondly, smooth images of different scales are generated through continuous Gaussian filtering, and the difference image sequence is obtained by subtracting adjacent scale images. Pixels with edge response intensity exceeding the adaptive gradient threshold are extracted from the difference image sequence. The pixels extracted from different scale layers are used as a multi-scale edge point set, and the scale layer with uniform edge point density distribution and the fewest contour breakpoints is used as the optimal response scale. Next, the Canny edge detection algorithm is used on the optimal response scale to perform non-maximum suppression and double-threshold edge connection. Weak edge points connected to strong edge points are retained as valid edge points, and the continuously arranged valid edge points are used as continuous umbrella skirt contour lines. Finally, distance transformation and contour expansion are performed along both sides of the umbrella skirt contour lines, and median filtering is used to smooth the transition area of ​​the expanded contour. The image containing continuous umbrella skirt contours and enhanced edge responses is used as an edge feature map.

[0036] Specifically, the pixel classification process is based on the statistical distribution characteristics of the color space. By pre-collecting samples of lightning arrester skirts in a clean state, the distribution range of these skirts in the RGB color space is extracted to establish a standard color template. The standard color range is determined by calculating the mean and standard deviation of the sample pixels in the R, G, and B channels, forming a three-dimensional ellipsoidal color space. The center of the ellipsoid is the mean vector, and the radius is determined by a multiple of the standard deviation.

[0037] For example, when performing attachment pixel identification, the RGB three-channel values ​​of each pixel within the umbrella skirt region are extracted, and the distance from the pixel to the center of the standard color ellipsoid is calculated. When the distance exceeds the threshold corresponding to three times the standard deviation, the pixel is marked as an attachment candidate point. Based on the spatial distribution characteristics of the candidate points, the eight-neighborhood connectivity criterion is used for region growth. The specific process is as follows: starting from a candidate seed point, its eight surrounding neighboring pixels are checked. If neighboring pixels are also marked as candidate points and the color difference is within the allowable range, they are classified into the same attachment region. Through iterative expansion, until no new pixels are added, the growth of a connected region is completed. The above process is performed on all candidate points within the umbrella skirt region, ultimately forming multiple independent attachment regions, and the remaining pixels are classified into the umbrella skirt body pixel set.

[0038] In one possible implementation, multi-scale edge detection is achieved by constructing a Gaussian pyramid. First, the original umbrella-shaped image is smoothed using a Gaussian filter with a standard deviation of σ, resulting in the first smoothed image. Then, the standard deviation is increased to √2σ, 2σ, and 2√2σ, generating the second, third, and fourth smoothed images, respectively. Subtracting adjacent image layers yields a sequence of difference images, each reflecting edge information within a specific scale range. Edge points are extracted from the difference images by setting a gradient magnitude threshold. This threshold is adaptively determined based on the statistical characteristics of the overall gradient distribution of the image, typically taking the magnitude corresponding to 70% of the cumulative distribution function of the gradient histogram. Edge points extracted from different scale layers are aggregated to form a multi-scale edge point set. By calculating the spatial density distribution of edge points in each scale layer, the layer with the most uniform density distribution and a moderate number of edge points is selected as the optimal response scale.

[0039] Preferably, when applying the Canny edge detection algorithm at the optimal scale, the high threshold is set to the 80th percentile of the gradient magnitude, and the low threshold is 0.4 times the high threshold. Through this dual-threshold strategy, strong edge points are directly retained, while weak edge points are only retained if they are connected to strong edge points, thus achieving edge continuity. The algorithm refines edges through non-maximum suppression, comparing the gradient magnitudes of adjacent pixels along the gradient direction and retaining only local maxima points to obtain a continuous umbrella-shaped outline with a single pixel width.

[0040] Understandably, the distance transformation is achieved by calculating the Euclidean distance from each pixel within the umbrella-shaped region to the nearest contour point. A two-scan algorithm is employed: the first scan moves from the top left to the bottom right, updating the distance value for each pixel; the second scan moves from the bottom right to the top left, further optimizing the distance values. The gradient field is calculated based on the spatial distribution of the distance values, with the gradient direction pointing towards the nearest contour point.

[0041] In one embodiment, the contour expansion proceeds in the opposite direction of the gradient, with a fixed expansion step size of 2 pixels. The grayscale value of the expanded point is determined by linear interpolation, and the weighting coefficient is 0.8 of the original contour point's grayscale value plus 0.2 of the background grayscale value, achieving a gradual transition of the contour. Through the implementation of the above technical solution, even when heavily covered by attachments, the edge features of the umbrella skirt can still be accurately extracted, providing a reliable image basis for subsequent equipment status assessment.

[0042] In step S3, the background of the attachments is separated from the video frames in the inspection video by the edge feature map to obtain the interference distribution area of ​​the attachments.

[0043] In this embodiment, the process of separating the background of the attachments from the video frames in the inspection video using the edge feature map to obtain the interference distribution area of ​​the attachments can be achieved through the following steps: The edge feature map is used to perform brightness equalization fusion on the video frames in the inspection video to obtain a fusion result map. Based on the fusion result image, background separation is performed using color clustering to obtain the foreground umbrella skirt image; The interference pixels of the soiled umbrella skirt are extracted from the foreground umbrella skirt image to obtain the interference pixel distribution map of the attached material; The interference distribution area of ​​the attachment is determined based on the interference pixel distribution map.

[0044] It should be noted that the fusion result image mentioned in this application represents the image obtained after brightness equalization fusion of the edge feature map and the inspection video frame; the foreground umbrella skirt image represents the image retaining the umbrella skirt and its attached areas after removing the sky, vegetation and tower material background; the interference pixel distribution map represents the pixel location map that marks the interference of the attached objects on the umbrella skirt texture, brightness and edges; the interference distribution area represents the spatial distribution range of the attached objects formed by integrating the interference pixels through connected regions.

[0045] In specific implementation, firstly, edge feature maps are weighted and superimposed with corresponding frames in the original image sequence. The weight coefficients of the edge feature maps and the original image are summed to 1. The weight ratio is adaptively adjusted according to the edge intensity. A fused image is generated through pixel-level linear combination. Histogram equalization is then performed on the fused image to obtain a brightness-balanced fused result image. Secondly, background separation based on color clustering is performed on the brightness-balanced fused result image. The K-means clustering algorithm is used to group pixels according to color features. The number of clusters is preset to 3-5, which is suitable for three typical background scenarios: power line inspection sky, vegetation, and equipment. The average brightness of each cluster center is calculated, and the cluster with the highest average brightness is marked as the background region. Background region pixels are removed from the fused result image to obtain the foreground umbrella skirt image. Then, the foreground umbrella skirt image is compared with a preset umbrella skirt standard feature template image by image. For pixel comparison, the standard feature template is constructed by extracting the texture gradient direction and brightness distribution from clean umbrella skirt samples without any attached material. The Euclidean distance between the texture gradient direction angle difference and the brightness difference of each pixel is calculated as the comprehensive difference value. The preset comprehensive difference value threshold is 0.6, which is based on the statistical setting of feature differences between clean and dirty umbrella skirts. If the comprehensive difference value exceeds the preset threshold, it is marked as an interference pixel, resulting in an interference pixel distribution map. Finally, morphological closing operations are performed on the interference pixel distribution map. The preset structuring element size is 5×5, which can effectively fill the interference gaps without expanding the interference area. The preset structuring element is first expanded and then eroded to fill the interference pixel gaps. The processed connected regions are marked, and the area and centroid position of each connected region are calculated to determine the specific distribution area of ​​the attached material interference. The obtained specific distribution area is used as the interference distribution area of ​​the attached material.

[0046] For example, in one implementation, when the enhanced edge feature map is fused with the original image, an adaptive weight allocation mechanism is used to optimize the combination of image information. The fusion process enhances the edge contour features while preserving the original texture information.

[0047] Specifically, the adaptive adjustment of the weight coefficients is based on the statistical distribution of edge intensity. A global scan of the edge feature map is performed to statistically analyze the gradient magnitude distribution of edge pixels, and the richness of edge information is determined based on the variance of the gradient magnitude. When edge information is rich, the weight of the edge feature map increases to 0.4, and the weight of the original image decreases to 0.6 accordingly; when edge information is sparse, the weight of the edge feature map decreases to 0.2, and the weight of the original image increases to 0.8. This dynamic adjustment ensures that the fused image retains both details and highlights contours.

[0048] For example, the application of the K-means clustering algorithm in background separation includes three stages: initialization, iterative optimization, and convergence judgment. In the initialization stage, K pixels are randomly selected from the fused image as initial cluster centers. The value of K is preset according to the scene complexity; typically, 3 to 5 clusters are set for power inspection scenarios. In the iterative optimization stage, the Euclidean distance from each pixel to each cluster center in the RGB color space is calculated, and the pixel is assigned to the nearest cluster. Then, the centroid of each cluster is recalculated as the new cluster center. Convergence judgment is achieved by comparing the changes in the positions of the cluster centers in two adjacent iterations. The algorithm converges when the displacement of all cluster centers is less than a preset threshold. Background region identification is based on statistical analysis of cluster features. The average brightness value, pixel count ratio, and spatial distribution continuity of each cluster are calculated. Sky backgrounds typically exhibit high brightness, large area, and are concentrated in the upper part of the image. Based on this, clusters that meet the criteria are marked as background and removed from the image.

[0049] It should be noted that histogram equalization improves image contrast by redistributing pixel gray values, statistically merging the gray-level histogram of the image, calculating the cumulative distribution function, and mapping the original gray values ​​to new gray levels based on the cumulative distribution function, so that the histogram of the output image tends to be uniformly distributed.

[0050] In one possible implementation, the construction of the standard feature template for the umbrella skirt is based on statistical learning from multiple clean umbrella skirt samples. Images of umbrella skirts without attachments are acquired under different lighting conditions and shooting angles, and the texture gradient direction and brightness distribution features of each sample are extracted. The texture gradient direction is calculated using the Sobel operator to calculate the gradient components in the horizontal and vertical directions. The gradient direction angle of each pixel is obtained according to the arctangent function. The gradient direction distribution of pixels at the same position in all samples is statistically analyzed, and the mode is taken as the standard gradient direction at that position. The brightness distribution features are obtained by calculating the average brightness and brightness change rate of each region in the sample image. The folds of the umbrella skirt usually exhibit a regular alternation pattern of light and dark. This pattern is parameterized as a brightness change curve and stored in the template.

[0051] Preferably, the calculation of the comprehensive difference value adopts a weighted Euclidean distance metric. For each pixel position, its gradient direction angle and brightness value are extracted. The angle difference between the gradient direction angle and the template standard value is calculated and normalized to the range of 0 to 1. The difference between the brightness value and the template standard value is calculated and divided by the brightness dynamic range for normalization. The two normalized differences are multiplied by a preset weight coefficient, the sum of squares is taken, and the square root is taken to obtain the comprehensive difference value of the pixel.

[0052] In one embodiment, the morphological closing operation achieves connectivity of the interfering regions through a sequential operation of dilation followed by erosion. The dilation operation uses a structuring element of a preset size to convolve the distribution map of interfering pixels, marking all locations containing interfering pixels within the coverage area of ​​the structuring element as interference. The erosion operation uses the same structuring element in reverse processing, retaining only the interfering regions that can be completely covered by the structuring element.

[0053] Understandably, the connected region labeling employs a two-pass scanning algorithm. The first pass assigns a temporary label to each interfering pixel, while the second pass merges equivalent labels, ultimately resulting in a unique identifier for each independent connected region. The number of pixels and centroid coordinates for each region are then calculated. Through this processing flow, the spatial distribution of the adhering material on the umbrella skirt surface is accurately located, enabling precise identification and quantitative analysis of the interfering regions.

[0054] In step S4, the interference distribution area is reconstructed by feature optimization based on the preset SAM model. When the reconstructed feature value is lower than the umbrella skirt matching threshold, local magnification is performed to obtain the optimized feature reconstruction map.

[0055] In this embodiment, the interference distribution area is reconstructed and its features are optimized based on a preset SAM model. When the reconstructed feature value is lower than the skirt matching threshold, local magnification is performed to obtain the optimized feature reconstruction map. This can be achieved through the following steps: Based on the preset SAM model, the interference distribution area is reconstructed by feature optimization to obtain a reconstructed feature map; The reconstructed feature map is matched with a preset umbrella skirt standard feature template to obtain the global matching degree; Based on the global matching degree, the reconstructed feature map is marked with low-matching regions to obtain a low-matching region mask; The low-matching region mask is locally magnified, and the magnified region is edge-enhanced to obtain an optimized feature reconstruction map.

[0056] It should be noted that, in this application, the reconstructed feature map refers to the umbrella skirt feature representation map obtained after semantic feature recovery of the distribution area of ​​the attachment interference; the global matching degree refers to the overall similarity between the reconstructed feature map and the preset umbrella skirt standard feature template; the low matching region mask refers to the binary mask used to mark the region with insufficient feature recovery; and the optimized feature reconstruction map refers to the umbrella skirt feature map obtained after local magnification and edge enhancement of the low matching region.

[0057] In specific implementation, firstly, image patches of the interference distribution area are input into the image encoder of a pre-defined SAM model. The image encoder performs feature embedding to obtain the interference region feature embedding. The geometric center point of the candidate umbrella skirt region within the interference distribution area is used as the foreground cue, and the bounding box of the candidate umbrella skirt region is used as the bounding box cue. This is then input into the cue encoder of the SAM model to obtain the cue embedding. The mask decoder of the SAM model performs cross-attention fusion on the interference region feature embedding and the cue embedding, outputting a segmentation mask for the umbrella skirt body region and its corresponding mask confidence score. The mask confidence score is then used as a pixel-wise weight to weight the interference region feature embedding, suppressing the feature response of pixels containing attachments and enhancing the feature response of pixels containing the umbrella skirt body. The weighted features are used as reconstructed feature values. When the reconstructed feature value is lower than the umbrella skirt matching threshold, the cue point of the corresponding pixel is encrypted and re-input into the cue encoder. The mask decoder iteratively outputs the segmentation mask until the reconstructed feature value is not lower than the umbrella skirt matching threshold. The process involves several steps: First, obtaining a reconstructed feature map. Second, performing similarity matching between the reconstructed feature map and a preset standard feature template for the umbrella skirt. This involves extracting the gradient magnitude and direction vector of each pixel in the feature map, calculating the inner product of this vector and the corresponding position vector in the standard template, and dividing by the product of the vector magnitudes to obtain the cosine similarity. The average of all pixel similarity values ​​is then used to obtain the global matching degree. Next, a global matching degree threshold of 0.7 is preset. This threshold has been verified through feature reconstruction experiments; a value below 0.7 indicates insufficient feature restoration and requires enhancement. If the matching degree is below this preset threshold, it is marked as a low-matching region, resulting in a low-matching region mask. Finally, local magnification is performed based on the low-matching region mask. A bicubic interpolation algorithm is used to enhance the spatial resolution of the low-matching region. The interpolation factor is preset to 2-4 times, dynamically set according to the number of pixels in the region (4 times for small regions, 2 times for large regions) to balance accuracy and efficiency. The Laplacian operator is applied to the magnified region to enhance the edges, and Gaussian kernel convolution is used to smooth interpolation artifacts, resulting in a locally enhanced image. By replacing the corresponding content in the feature map with the locally enhanced image, a set of gradient constraint equations is constructed at the replacement boundary. The pixel value adjustment amount that makes the boundary gradient continuous is solved. The pixel value in the boundary transition area is corrected according to the adjustment amount to achieve seamless fusion and obtain the optimized feature reconstruction map.

[0058] For example, in one implementation, when reconstructing the features of the distribution area of ​​the attached object interference, a deep learning network is used to realize the semantic understanding and feature recovery of the interference area. The feature reconstruction process gradually improves the reconstruction quality through multiple iterations of optimization.

[0059] Specifically, the DeepLabV3+ semantic segmentation network adopts an encoder-decoder architecture. The encoder part is based on an improved ResNet-101 backbone network and includes a dilated spatial pyramid pooling module, which captures multi-scale contextual information through dilated convolutions with different dilation rates. The decoder part restores the feature map resolution through bilinear upsampling and fuses shallow features from the encoder to preserve detailed information.

[0060] For example, the iterative optimization process performs two phases in each iteration: forward propagation and backward propagation. In the forward propagation phase, a 256×256 pixel interference region image patch is input to the encoder. After continuous convolution, batch normalization, and activation function processing, semantic features at different levels are extracted. The dilated spatial pyramid pooling module uses dilated convolutions with dilation rates of 6, 12, and 18 to process the feature maps in parallel, capturing contextual information from different receptive fields. The decoder receives the high-level semantic features output from the encoder, restores them to one-quarter of the original resolution through a 4x upsampling, and simultaneously extracts low-level features from the second layer of the encoder. The two are then fused after adjusting the number of channels through a 1×1 convolution. The fused features are refined through a 3×3 convolution and then upsampled again by 4x to obtain a reconstructed feature map with the same resolution as the input. In the backward propagation phase, the gradient of each layer's parameters is calculated based on the loss value. The Adam optimizer is used to update the network weights, with the learning rate set to 0.001 and gradually decaying during training.

[0061] It should be noted that the loss function uses negative log-likelihood calculation. The cross-entropy between the predicted probability distribution and the true class label is calculated for each pixel location. The sum of the cross-entropies of all pixels is used as the total loss. The network parameters are optimized by minimizing the loss value.

[0062] In one possible implementation, cosine similarity matching achieves feature comparison through vectorization. Gradient information from a 3×3 neighborhood is extracted for each pixel in the reconstructed feature map, and the horizontal and vertical gradient components are calculated to form a two-dimensional gradient vector. A standard feature template stores the standard gradient vectors of the cleaning umbrella skirt at the same position. The cosine value is obtained by calculating the inner product of the two vectors and dividing by the product of their magnitudes; a cosine value closer to 1 indicates higher similarity. The global matching degree is obtained by taking the arithmetic mean of the cosine similarity across all pixel positions. When the global matching degree is below a preset threshold of 0.7, low-matching regions requiring further processing are identified. Low-matching regions are binarized to generate a mask; pixels with a value of 1 in the mask represent locations requiring local magnification and enhancement.

[0063] Preferably, the bicubic interpolation algorithm maintains image smoothness during local magnification. For each target pixel in a low-matching region, the algorithm uses 16 pixels (4×4 pixels) surrounding its corresponding position in the original image for interpolation calculation. The interpolation kernel function uses a cubic polynomial, and weight coefficients are calculated based on the distance from the target pixel to the source pixel; the closer the distance, the greater the weight. The interpolation factor is dynamically determined based on the number of pixels in the region; the smaller the region, the greater the factor, typically between 2 and 4 times.

[0064] In one embodiment, the Laplacian operator detects edges using its second derivative. A 3×3 Laplacian kernel is used to convolve the magnified image, with the kernel element being -4 and surrounding elements being 1. The convolution result reflects the degree of abrupt change in pixel grayscale. The Laplacian response is then superimposed onto the original image to achieve edge sharpening, with the sharpening intensity controlled by adjusting the superposition coefficient.

[0065] Understandably, the gradient constraint equations are solved using the least squares method. Gradient continuity constraints are established at each pixel of the replacement boundary, requiring the minimization of the difference between the replaced gradient field and the original gradient field. This forms a linear system of equations that are iteratively solved using the conjugate gradient method. Through the aforementioned feature reconstruction and optimization, accurate recovery of the interference region caused by the attachment material is achieved, improving the accuracy of umbrella skirt feature recognition.

[0066] In step S5, the dynamic interference pattern caused by the change in the area of ​​the attached material is analyzed based on the feature reconstruction map to obtain a quantitative index map of the change in the area of ​​the attached material.

[0067] In this embodiment, the following steps can be used to analyze the dynamic interference pattern caused by the change in the area of ​​the attached material based on the feature reconstruction map and obtain a quantitative index map of the change in the area of ​​the attached material: Extract the outline of the attachment region based on the reconstructed feature map; A temporal variation sequence of the area covered by the attachment region is constructed by the outline of the attachment region; The boundaries of the attachments in each frame of the time-series change sequence are tracked to obtain a boundary change trend map; A quantitative index map of the change in the area of ​​the attachment is determined based on the boundary change trend map.

[0068] It should be noted that, in this application, the outline of the attached area represents the closed outline formed by the outer boundary of the attached response area in the feature reconstruction image; the temporal change sequence represents the record of the change in the coverage area of ​​the attached material arranged according to the frame order of the inspection video; the boundary change trend map represents an image used to characterize the expansion, contraction and stable change state of the attached boundary; and the quantitative index map represents an index distribution map formed after writing the rate of change of the attached area, the expansion speed and the stability into the corresponding spatial location.

[0069] In specific implementation, firstly, adaptive threshold segmentation is performed on the response values ​​of interfering pixels in the feature reconstruction map. Isolated noise points are removed by morphological opening operations, and the circumscribed boundary and closed contour of each attachment region are extracted using a connected component labeling method. The circumscribed boundary and closed contour are then used as the attachment region contour. Secondly, feature reconstruction maps at different times are obtained from the continuous frame sequence of the inspection video. The connected component numbers and pixel counts are performed on the attachment region contour in each frame. The frame number, coverage area, and area change rate of each connected component in the continuous frames are arranged in chronological order, and the resulting temporal record table is used as the temporal change sequence of the attachment coverage area. Then, Sobel edge detection is used. The algorithm extracts the attachment boundary of each frame in the temporal change sequence and establishes a correspondence based on the spatial proximity and grayscale similarity of the boundary pixels in adjacent frames. The displacement vector between corresponding boundary pixels is used as the boundary expansion direction and expansion distance. The image composed of the boundary expansion direction, expansion distance, and stability marker in consecutive frames is used as the boundary change trend map. Finally, based on the boundary change trend map, the average expansion speed, contraction speed, stable expansion area ratio, and unstable change area ratio of each attachment region are statistically analyzed. The above indicators are written into the corresponding pixel positions, and different interference levels are identified by grayscale level or texture filling method. The obtained indicator distribution map is used as a quantitative indicator map of attachment area change.

[0070] For example, in one implementation, when performing time-series analysis on the optimized feature reconstruction map, the dynamic changes in coverage area are accurately monitored by extracting attachment distribution information at multiple time points. Specifically, inspection videos are continuously acquired at a rate of 25 frames per second, and a feature reconstruction map is extracted every 5 frames to form a time-series dataset. Connected component labeling is performed on the attachment region in each frame image, and the number of pixels in each connected component is counted as the coverage area value at that time. The difference in area between two adjacent time points is divided by a 0.2-second time interval to obtain the instantaneous rate of change.

[0071] For example, the Sobel edge detection algorithm achieves boundary tracking through convolutional kernels in two directions. The horizontal Sobel operator is a 3×3 matrix with weights of -1, -2, -1 in the left column, 0, 0, 0 in the middle column, and 1, 2, 1 in the right column; the vertical operator is its transpose. The two operators are convolved with the attached region image respectively to obtain the horizontal gradient Gx and the vertical gradient Gy. The motion characteristics of the boundary pixels are determined by calculating the gradient magnitude and direction angle. Boundary pixels in consecutive frames are matched, and correspondences are found based on pixel grayscale similarity and spatial proximity. The displacement vector between matched pixel pairs is calculated, reflecting the direction of boundary expansion or contraction.

[0072] It should be noted that the stability assessment is based on the directional consistency evaluation across multiple consecutive frames. The expansion direction angles of the same boundary segment are extracted from five consecutive frames, and the standard deviation of these angles is calculated. When the standard deviation is less than a preset threshold of 15 degrees, it is considered a stable expansion, indicating that the attachment is continuously growing in that direction; otherwise, it is considered a random fluctuation.

[0073] Preferably, the color mapping adopts a heatmap color scheme, linearly mapping the expansion velocity value to the color space. When the velocity is positive, it gradually changes from yellow to red, with the color becoming redder as the velocity increases; when the velocity is negative, it gradually changes from cyan to blue, with the color becoming bluer as the contraction speed increases; when the velocity is close to zero, it is displayed as green. Through the above processing, dynamic monitoring and visualization of the changes in the coverage area of ​​the attachments on the surface of the surge arrester skirt are achieved, providing an intuitive quantitative basis for equipment condition assessment.

[0074] In step S6, the quantitative index map and the pixel-level annotation map are matched and calibrated, and the evaluation report of the surge arrester skirt is determined based on the matching and calibration results.

[0075] In this embodiment, the following steps can be used to match and calibrate the quantitative index map and the pixel-level annotation map, and to determine the evaluation report of the surge arrester skirt based on the matching and calibration results: Spatial registration is performed between the quantized index map and the pixel-level annotation map to obtain the matching calibration result; Based on the matching calibration results, a decision tree-based classification of the contamination status is performed to obtain the contamination category identifier of the surge arrester skirt; An alarm is determined based on the soiling category identifier. When the soiling category identifier is a severe soiling category, a skirt soiling alarm signal is generated, and the trigger time of the skirt soiling alarm signal is recorded. Based on the pollution category identifier and the matching calibration results, the location number, pollution level, attachment coverage area ratio and attachment expansion rate of the surge arrester skirt are summarized to obtain the skirt pollution summary information. The arrester skirt contamination alarm signal, the trigger time, and the skirt contamination summary information are filled in according to a preset report template to obtain an assessment report of the arrester skirt.

[0076] It should be noted that, in this application, the matching calibration result refers to the pixel correspondence obtained after unifying the quantitative index map and the pixel-level annotation map into the same coordinate system; the contamination category identifier refers to the umbrella skirt contamination status category generated based on the coverage area, expansion rate, and distribution density; the umbrella skirt contamination alarm signal refers to the alarm flag triggered when the contamination category identifier reaches the severe contamination category, used to indicate that the umbrella skirt needs maintenance; the trigger time refers to the timestamp recorded when the umbrella skirt contamination alarm signal is generated; the umbrella skirt contamination summary information refers to the structured field set obtained by merging the location number, contamination level, attachment coverage area ratio, and attachment expansion rate on a per-surge arrester umbrella skirt basis; and the evaluation report refers to the structured report recording the umbrella skirt location, contamination level, attachment coverage changes, and treatment recommendations.

[0077] In specific implementation, firstly, the least squares method is used to calculate the affine transformation parameters to realize the coordinate system of feature points in the quantization index map and the pixel-level annotation map. The coverage change rate value of each pixel in the quantization index map is mapped to the corresponding category label in the pixel-level annotation map. The correlation coefficient between the rate value and the category is calculated as a measure of correlation. If the correlation coefficient is lower than a preset correlation threshold, the corner points of the umbrella skirt edge are reselected for secondary registration. If the correlation coefficient reaches the preset correlation threshold, the coverage change rate value, pixel category label, and spatial coordinates of the corresponding position are written into the calibration mapping table to obtain the matching calibration. The first step is to obtain the accurate result. Secondly, a decision tree-based state classification is performed on the matched calibration results, using the coverage area ratio, average expansion rate, and attachment distribution density as decision node features. A preset coverage area ratio threshold of 30% is set according to the power equipment pollution level classification standard. When the coverage area exceeds this preset threshold and the expansion rate is positive, it is judged as severe pollution; when the coverage area is below the threshold and the rate is negative, it is judged as mild pollution; and all other cases are judged as moderate pollution, thus obtaining the pollution category identifier of the surge arrester skirt. Next, an alarm judgment is performed based on the pollution category identifier. When the... When the contamination category is identified as severe contamination, a skirt contamination alarm signal is generated, and the trigger time of the alarm signal is recorded by the system clock. No alarm signal is generated when the contamination category is identified as moderate or light contamination. Subsequently, using the contamination category identifier as an index, the location number, contamination level, coverage area ratio of attached materials, and rate of expansion of attached materials are extracted from the matching calibration results and intermediate results of the preceding steps. Fields are aligned and merged according to the equipment dimension to obtain summary skirt contamination information. Finally, the skirt contamination information is... The alarm signal, the trigger time, and the summary information of the arrester skirt contamination are filled into the corresponding fields of the preset report template. The preset report template includes four parts: basic equipment information, test result data, attachment distribution image, and processing suggestions. The basic equipment information section is filled with the arrester skirt location number, the test result data section is filled with the contamination level, the attachment coverage area ratio, and the attachment expansion rate, the attachment distribution image section is filled with a pixel-level annotation image, and the processing suggestions section is filled with corresponding maintenance suggestions based on the arrester skirt contamination alarm signal and contamination level mapping, thereby outputting an evaluation report of the arrester skirt.

[0078] For example, in one implementation, a comprehensive assessment of the fouling status of the surge arrester skirt is achieved by precisely matching the quantitative index map with the annotation map. The assessment process integrates spatial information and temporal variation characteristics.

[0079] For example, spatial registration employs a coordinate system-one based on affine transformations of feature points. Corner features are extracted from both images, and the Harris corner detection algorithm is used to identify points with significant gradient changes. For each candidate corner, the gradient covariance matrix eigenvalues ​​within its 8×8 neighborhood are calculated; a strong corner is identified when both eigenvalues ​​are large. Intrapoints are filtered from matching point pairs using the RANSAC algorithm to eliminate false matches. The six parameters of the affine transformation, including translation, rotation, and scaling coefficients, are solved using the least squares method. The registered images are mapped pixel-wise, and the coverage change rate value in the quantization index map and the category label in the annotation map are extracted. The Pearson correlation coefficient is calculated to measure the degree of association between the two; a correlation coefficient closer to 1 indicates a stronger consistency between the dynamic change and the static category at that location.

[0080] Specifically, the decision tree classifier is constructed using the ID3 algorithm, with information gain as the feature selection criterion. At the root node, the information gain of three features—coverage ratio, mean expansion rate, and distribution density—is calculated, and the feature with the largest gain is selected as the splitting attribute. A coverage ratio exceeding 30% is used as the first-level decision condition, the sign of the expansion rate as the second-level condition, and the dispersion of the distribution density as the third-level condition. A complete decision tree is formed through recursive splitting.

[0081] It should be noted that the alarm triggering adopts a tiered response mechanism: severe pollution sends a real-time alarm immediately, moderate pollution records a warning message, and light pollution only records the status.

[0082] Preferably, the assessment report template includes four standardized sections: the basic equipment information section records the umbrella skirt number, inspection time, and geographical location; the inspection results section lists the coverage percentage, expansion rate, and contamination level; the image evidence section displays a comparison of the original image, labeled image, and quantitative indicator image; and the treatment recommendations section provides corresponding cleaning cycles and maintenance plans based on the contamination level. Through this comprehensive assessment process, a complete closed loop from image analysis to condition determination is achieved, providing a scientific basis for power equipment maintenance decisions.

[0083] In step S7, the metadata tags of the inspection video are updated according to the evaluation report.

[0084] In this embodiment, updating the metadata tags of the inspection video according to the evaluation report can be achieved through the following steps: Based on the contamination level, coverage ratio, expansion speed and umbrella skirt location number in the assessment report, metadata tags are generated for the inspection video, resulting in an inspection video file with metadata tags. The sampling interval of subsequent inspection video processing tasks is adjusted based on the metadata tags to obtain inspection video processing tags that match the state of dirtiness.

[0085] It should be noted that the metadata tags mentioned in this application refer to structured tags embedded in the inspection video file and used to describe the state, location, and processing priority of the umbrella skirt dirt; the inspection video file with metadata tags refers to a video file that can be retrieved and scheduled after the status field is written; the inspection video processing tags refer to the tag results that constrain the subsequent sampling interval, review frequency, and processing priority according to the dirt status.

[0086] In specific implementation, firstly, the contamination level, coverage area, expansion speed, umbrella skirt location number, inspection time, and equipment number are extracted from the assessment report. These fields are written into the auxiliary data segment of the video file in key-value pair format, and the written field set is used as the metadata tag for the inspection video. Secondly, the inspection video is status-marked according to the metadata tag. Umbrella skirt video segments with severe contamination are marked as high priority, umbrella skirt video segments with moderate contamination are marked as medium priority, and umbrella skirt video segments with mild contamination are marked as low priority. The video file with completed status marking is used as the inspection video file with metadata tag. Finally, the priority field in the inspection video file with metadata tag is read, and the sampling interval and review frequency of subsequent inspection video processing tasks are adjusted according to the priority field. The adjusted sampling interval, review frequency, and processing priority are written back to the metadata tag, and the written tag result is used as the inspection video processing tag matching the contamination status.

[0087] For example, in one implementation, intelligent management of inspection videos is achieved through metadata tags, which carry equipment status information to guide subsequent processing procedures.

[0088] Specifically, the metadata tags are stored using a key-value pair data structure, including a "dirt level" key corresponding to three levels: severe, moderate, or mild; a "coverage area" key storing a percentage value; and a "spread speed" key recording the amount of pixel change per second. These key-value pairs are embedded through auxiliary data segments in the video container format and written sequentially within the 2048-byte space reserved in the video file header.

[0089] For example, the processing queue employs a priority scheduling mechanism, assigning a priority value of 3 to video segments marked as severely polluted, 2 to moderately polluted, and 1 to lightly polluted. The queue is sorted according to priority values, with higher-priority videos processed first. The sampling interval is dynamically adjusted based on priority, with a baseline interval of 10 frames. This interval is shortened to 5 frames for severely polluted areas to increase monitoring density, and extended to 20 frames for lightly polluted areas to save computational resources. Through this adaptive processing mechanism, computational resources are rationally allocated, strengthening monitoring of severely polluted areas and reducing processing frequency for lightly polluted areas.

[0090] Therefore, in this application, the metadata tags of the inspection video can be updated according to the evaluation report. Firstly, by segmenting the pixel-level annotation map of the lightning arrester skirt area from the inspection video collected by the drone, the target area of ​​the skirt can be determined in complex video footage, providing an accurate spatial basis for attachment identification and background interference elimination, reducing the impact of irrelevant background on skirt edge recognition. Secondly, by identifying attachments in the skirt area of ​​the pixel-level annotation map and determining the attachment coverage ratio, the edge contour is enhanced when the coverage ratio exceeds a preset threshold. This adaptively strengthens the skirt boundary features according to the degree of change in attachment coverage area, avoiding local missegmentation caused by small-area attachments, while mitigating the problem of skirt contour confusion with background features caused by larger-area attachments. Furthermore, the attachment background is separated from the video frame using the edge feature map to obtain the interference distribution area of ​​the attachments, and the interference distribution area is reconstructed based on a preset SAM model. When the reconstructed feature value is low... By performing local magnification at the umbrella skirt matching threshold, targeted recovery of umbrella skirt features obscured or weakened by attached objects can be achieved, improving the continuous segmentation capability and feature expression stability of the umbrella skirt region under complex interference. Furthermore, by analyzing the dynamic interference patterns caused by changes in the area of ​​attached objects based on the feature reconstruction map, a quantitative index map of the changes in the area of ​​attached objects is obtained. This transforms the impact of changes in the coverage area of ​​attached objects on umbrella skirt recognition into a calibrable quantitative basis, thereby enhancing the adaptability of the segmentation process to changes in the dynamic environment. Finally, by matching and calibrating the quantitative index map and the pixel-level annotation map, and determining the evaluation report of the surge arrester umbrella skirt based on the matching and calibration results, and then updating the metadata tags of the inspection video, a closed-loop correction can be performed on the umbrella skirt segmentation results and the attached object interference analysis results. This reduces the probability of misjudgment when the attached object and umbrella skirt features are highly similar, thereby improving the overall accuracy of distinguishing the surge arrester umbrella skirt from the background in the inspection video, the reliability of the status assessment, and the traceability of subsequent video data management.

[0091] In summary, the technical solution adopted in this application can accurately distinguish between the lightning arrester skirt and the background in the inspection video, while effectively dealing with the characteristic interference caused by changes in the coverage area of ​​the attached objects.

[0092] Example 2: This application provides a machine-guided video segmentation system based on SAM, referring to... Figure 2 As shown in the figure, this is a module structure diagram of a machine-based video segmentation system according to this embodiment of the present application. The machine-based video segmentation system includes: The pixel annotation module 100 is used to segment out the pixel-level annotation map of the lightning arrester skirt area from the inspection video collected by the drone; The attachment identification and edge enhancement module 200 is used to identify attachments in the umbrella skirt area of ​​the pixel-level annotation map, thereby determining the coverage ratio of the attachments. When the coverage ratio exceeds a preset threshold, the edge contour is enhanced to obtain an edge feature map. The attachment background separation module 300 is used to separate the attachment background from the video frames in the inspection video using the edge feature map to obtain the interference distribution area of ​​the attachment. The feature reconstruction optimization module 400 is used to perform feature optimization and reconstruction on the interference distribution area based on the preset SAM model. When the reconstructed feature value is lower than the umbrella skirt matching threshold, local magnification is performed to obtain the optimized feature reconstruction map. The attachment area dynamic analysis module 500 analyzes the dynamic interference pattern caused by the change in attachment area based on the feature reconstruction map, and obtains a quantitative index map of the change in attachment area. The evaluation report generation module 600 is used to match and calibrate the quantitative index map and the pixel-level annotation map, and determine the evaluation report of the lightning arrester skirt based on the matching and calibration results. The inspection video update module 700 is used to update the metadata tags of the inspection video according to the evaluation report.

[0093] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0094] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-Erasable Programmable Read-Only Memory (EEPROM), compactdisc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0095] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

Claims

1. A method for video segmentation based on SAM (Search Engine Amplifier), characterized in that, The steps include the following: Pixel-level labeled images of the lightning arrester skirt area were segmented from inspection videos collected by drones; The umbrella-shaped region in the pixel-level annotation map is subjected to attachment identification, and the coverage ratio of the attachment is determined. When the coverage ratio exceeds a preset threshold, the edge contour is enhanced to obtain an edge feature map. The edge feature map is used to separate the background of the attachments in the video frames of the inspection video to obtain the interference distribution area of ​​the attachments. Based on the preset SAM model, the interference distribution area is reconstructed by feature optimization. When the reconstructed feature value is lower than the umbrella skirt matching threshold, local magnification is performed to obtain the optimized feature reconstruction map. Based on the feature reconstruction map, the dynamic interference pattern caused by the change in the area of ​​the attached material is analyzed, and a quantitative index map of the change in the area of ​​the attached material is obtained. The quantitative index map and the pixel-level annotation map are matched and calibrated, and the evaluation report of the surge arrester skirt is determined based on the matching and calibration results; Update the metadata tags of the inspection video according to the assessment report.

2. The SAM-based video segmentation method as described in claim 1, characterized in that, The pixel-level annotation map segmented from the inspection video collected by the drone specifically includes: Extracting time-series continuous image sequences from inspection video data collected by drones; Based on the preset SAM model, the lightning arrester skirt region is segmented for each video frame in the time-series continuous image sequence to obtain the binary mask of the lightning arrester skirt region. The lightning arrester skirt area is pixel-level labeled according to the binary mask to obtain a pixel-level labeled map of the lightning arrester skirt area.

3. The SAM-based video segmentation method as described in claim 1, characterized in that, The process of identifying attachments in the umbrella-shaped area of ​​the pixel-level annotated image and then determining the coverage ratio of the attachments specifically includes: The attachments in the umbrella skirt region of the pixel-level annotation map are identified pixel by pixel, and the umbrella skirt body pixels and attachment pixels are distinguished to obtain the attachment pixel set and the umbrella skirt body pixel set. The coverage ratio of the attachment is determined based on the set of pixels of the attachment and the set of pixels of the umbrella skirt body.

4. The SAM-based video segmentation method as described in claim 1, characterized in that, When the coverage ratio exceeds a preset threshold, enhancing the edge contour to obtain the edge feature map specifically includes: When the coverage ratio exceeds the preset threshold, interference compensation is activated; The optimal response scale is determined based on the density distribution of edge points in each scale layer of the multi-scale edge point set in the umbrella skirt region. By connecting the edges of the skirt region using the optimal response scale, a continuous skirt outline is obtained; The umbrella skirt region is expanded based on the umbrella skirt outline to obtain an edge feature map.

5. The SAM-based video segmentation method as described in claim 1, characterized in that, By performing attachment background separation on video frames in the inspection video using the edge feature map, the interference distribution area of ​​the attachment is obtained, specifically including: The edge feature map is used to perform brightness equalization fusion on the video frames in the inspection video to obtain a fusion result map. Based on the fusion result image, background separation is performed using color clustering to obtain the foreground umbrella skirt image; The interference pixels of the soiled umbrella skirt are extracted from the foreground umbrella skirt image to obtain the interference pixel distribution map of the attached material; The interference distribution area of ​​the attachment is determined based on the interference pixel distribution map.

6. The SAM-based video segmentation method as described in claim 1, characterized in that, Based on the preset SAM model, the interference distribution area is reconstructed by feature optimization. When the reconstructed feature value is lower than the skirt matching threshold, local magnification is performed to obtain the optimized feature reconstruction map, which specifically includes: Based on the preset SAM model, the interference distribution area is reconstructed by feature optimization to obtain a reconstructed feature map; The reconstructed feature map is matched with a preset umbrella skirt standard feature template to obtain the global matching degree; Based on the global matching degree, the reconstructed feature map is marked with low-matching regions to obtain a low-matching region mask; The low-matching region mask is locally magnified, and the magnified region is edge-enhanced to obtain an optimized feature reconstruction map.

7. The SAM-based video segmentation method as described in claim 1, characterized in that, Based on the analysis of the feature reconstruction map, the dynamic interference pattern caused by the change in the area of ​​the attached material is obtained, and the quantitative index map of the change in the area of ​​the attached material is specifically included: Extract the outline of the attachment region based on the reconstructed feature map; A temporal variation sequence of the area covered by the attachment region is constructed by the outline of the attachment region; The boundaries of the attachments in each frame of the time-series change sequence are tracked to obtain a boundary change trend map; A quantitative index map of the change in the area of ​​the attachment is determined based on the boundary change trend map.

8. The SAM-based video segmentation method as described in claim 1, characterized in that, The evaluation report for the surge arrester skirt, which involves matching and calibrating the quantitative index map and the pixel-level annotation map and determining the matching and calibration results, specifically includes: Spatial registration is performed between the quantized index map and the pixel-level annotation map to obtain the matching calibration result; Based on the matching calibration results, a decision tree-based classification of the contamination status is performed to obtain the contamination category identifier of the surge arrester skirt; An alarm is determined based on the soiling category identifier. When the soiling category identifier is a severe soiling category, a skirt soiling alarm signal is generated, and the trigger time of the skirt soiling alarm signal is recorded. Based on the pollution category identifier and the matching calibration results, the location number, pollution level, attachment coverage area ratio and attachment expansion rate of the surge arrester skirt are summarized to obtain the skirt pollution summary information. The arrester skirt contamination alarm signal, the trigger time, and the skirt contamination summary information are filled in according to a preset report template to obtain an assessment report of the arrester skirt.

9. The SAM-based video segmentation method as described in claim 1, characterized in that, Updating the metadata tags of the inspection video according to the assessment report specifically includes: Based on the contamination level, coverage ratio, expansion speed and umbrella skirt location number in the assessment report, metadata tags are generated for the inspection video, resulting in an inspection video file with metadata tags. The sampling interval of subsequent inspection video processing tasks is adjusted based on the metadata tags to obtain inspection video processing tags that match the state of dirtiness.

10. A SAM-based mobile video segmentation system, used to execute a SAM-based mobile video segmentation method as described in any one of claims 1 to 9, characterized in that, The machine patrol video segmentation system includes: The pixel annotation module is used to segment out pixel-level annotation maps of the lightning arrester skirt area from inspection videos collected by drones; The attachment recognition and edge enhancement module is used to identify attachments in the umbrella skirt area of ​​the pixel-level annotation map, and then determine the coverage ratio of the attachments. When the coverage ratio exceeds a preset threshold, the edge contour is enhanced to obtain an edge feature map. The attachment background separation module is used to separate the attachment background from the video frames in the inspection video using the edge feature map to obtain the interference distribution area of ​​the attachment. The feature reconstruction optimization module is used to perform feature optimization and reconstruction on the interference distribution area based on the preset SAM model. When the reconstructed feature value is lower than the umbrella skirt matching threshold, local magnification is performed to obtain the optimized feature reconstruction map. The attachment area dynamic analysis module is used to analyze the dynamic interference pattern caused by the change in attachment area based on the feature reconstruction map, and obtain a quantitative index map of the change in attachment area. The evaluation report generation module is used to match and calibrate the quantitative index map and the pixel-level annotation map, and determine the evaluation report of the surge arrester skirt based on the matching and calibration results; The inspection video update module is used to update the metadata tags of the inspection video according to the evaluation report.