Semantic segmentation method and device for distribution network line, medium and equipment

By using binocular camera imaging and depth information processing to remove obstructions, precise segmentation of power distribution network equipment was achieved. This solved the problem of recognition accuracy of two-dimensional image analysis methods under obstruction conditions, and improved the integrity and reliability of detection.

CN121640069APending Publication Date: 2026-03-10ANHUI JIYUAN SOFTWARE CO LTD +1
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-04
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing two-dimensional image analysis methods have difficulty accurately distinguishing the foreground and background of network equipment when dealing with occlusion problems, resulting in low recognition accuracy, especially when trees or buildings are obstructing the view.

Method used

By using a binocular camera to photograph the power distribution network lines, and through depth information extraction and layered processing, obstructions are removed, and region growing and pixel-level labeling are performed, accurate segmentation of power distribution network equipment is achieved.

Benefits of technology

It effectively removes obstructions, improves the integrity and reliability of power distribution equipment detection, and provides high-quality data support for equipment health status assessment and potential hazard investigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640069A_ABST
    Figure CN121640069A_ABST
Patent Text Reader

Abstract

The invention relates to a distribution network line semantic segmentation method and device, a medium and equipment, and relates to the technical field of semantic segmentation, and the method comprises the following steps: shooting a distribution network line through a binocular camera to obtain a distribution network scene stereo image, and carrying out the depth information extraction of the distribution network scene stereo image to obtain scene depth data; performing hierarchical processing on the scene depth data to obtain a scene hierarchical graph, and performing shielding object elimination based on the scene hierarchical graph to obtain a non-shielding scene graph; and carrying out region growing processing on the unshielded scene graph to obtain a distribution network equipment region graph, and carrying out pixel-level labeling on the distribution network equipment region graph to obtain a distribution network equipment segmentation result, thereby solving the problem that the existing two-dimensional image analysis method is difficult to accurately distinguish foreground and background objects. And the identification accuracy of the distribution network equipment is not high.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of semantic segmentation, in particular to a distribution network line semantic segmentation method, device, medium and equipment. BACKGROUND

[0002] The intelligent inspection and maintenance of distribution network lines has always been an important issue in the field of power system operation and maintenance. Traditional distribution network line inspection mainly relies on manual field investigation, which is not only time-consuming and labor-intensive, but also prone to missed detection in complex environments, especially when distribution network equipment is obscured by trees, buildings, etc. The accuracy and efficiency of manual identification will be greatly reduced. With the development of computer vision technology, automatic inspection schemes based on image processing have gradually become a research hotspot, but existing two-dimensional image analysis methods have obvious limitations in dealing with occlusion problems, making it difficult to accurately distinguish foreground and background objects, resulting in low recognition accuracy of distribution network equipment. SUMMARY

[0003] The present application aims to at least partially solve one of the problems in the prior art.

[0004] To achieve the above-mentioned purpose, the present application provides a distribution network line semantic segmentation method, comprising the following steps: A binocular camera is used to take pictures of the distribution network line to obtain a distribution network scene stereo image, and depth information is extracted from the distribution network scene stereo image to obtain scene depth data; The scene depth data is processed in layers to obtain a scene layered image, and an occlusion object is removed based on the scene layered image to obtain an unoccluded scene image; The unoccluded scene image is subjected to region growing processing to obtain a distribution network equipment region image, and the distribution network equipment region image is subjected to pixel-level labeling and segmentation to obtain a distribution network equipment segmentation result.

[0005] Further, the taking pictures of the distribution network line by the binocular camera to obtain the distribution network scene stereo image comprises: The left lens and the right lens of the binocular camera are calibrated to obtain calibrated lenses, and the distribution network line is synchronously photographed by the calibrated lenses to obtain a left photographed image and a right photographed image; The left photographed image and the right photographed image are time-stamped to obtain a time-synchronized image pair, and the time-synchronized image pair is positionally calibrated to obtain a calibrated image pair; The corresponding pixel points in the left photographed image and the right photographed image in the calibrated image pair are subjected to parallax calculation to obtain a parallax image, and the parallax image is subjected to image fusion to obtain the distribution network scene stereo image.

[0006] Further, the depth information extraction on the stereo image of the distribution network scene is performed to obtain scene depth data, comprising: pixel coordinate extraction is performed on the stereo image of the distribution network scene to obtain each pixel coordinate data, and disparity reverse calculation is performed on each pixel coordinate data based on binocular camera imaging principle to obtain each pixel disparity value; Based on each pixel disparity value, the stereo image of the distribution network scene is analyzed to obtain a disparity distribution map, and the disparity distribution map is subjected to smoothing filter processing to obtain a smoothed disparity map; The depth value conversion is performed on each pixel disparity value in the smoothed disparity map based on the focal length and baseline distance preset by the binocular camera to obtain each pixel depth value, and the scene depth map is constructed based on each pixel depth value.

[0007] Further, the scene depth data is subjected to hierarchical processing to obtain a scene hierarchical map, comprising: The depth value statistical sorting is performed on the scene depth data to obtain a depth value ordered sequence, and the depth value ordered sequence is subjected to interval division to obtain a depth hierarchical interval, wherein the depth hierarchical interval includes depth threshold features of different distances in the distribution network line scene, near, middle and far; Based on the depth hierarchical interval, the pixel classification is performed on the stereo image of the distribution network scene to obtain a depth classified pixel set, and the depth classified pixel set is subjected to region marking to obtain a depth marked image; The region contour regularization is performed on the depth marked image to obtain a regularized depth image, and the regularized depth image is subjected to hierarchical color rendering to obtain a scene hierarchical map.

[0008] Further, based on the depth hierarchical interval, the pixel classification is performed on the stereo image of the distribution network scene to obtain a depth classified pixel set, comprising: The depth value reading is performed on each pixel point in the stereo image of the distribution network scene to obtain a pixel depth value sequence, and the interval comparison and determination is performed on the pixel depth value sequence based on the depth hierarchical interval to obtain a pixel level label; Based on the pixel level label, the pixel points in the stereo image of the distribution network scene are subjected to same layer aggregation processing to obtain a depth classified pixel set.

[0009] Further, the occlusion removal is performed based on the scene hierarchical map to obtain an unoccluded scene map, comprising: Based on the scene hierarchical map, the occlusion area identification is performed to obtain an occlusion candidate area, and the contour extraction is performed on the occlusion candidate area to obtain an occlusion contour line; Based on the occlusion contour line, the occlusion area precise marking is performed on the occlusion candidate area to obtain an occlusion precise area map; The occlusion pixel is replaced in the occlusion precision area map to obtain an initial occlusion-free map, and the edge smoothing process is performed on the initial occlusion-free map to obtain an occlusion-free scene map.

[0010] Furthermore, the unobstructed scene map is subjected to region growing processing to obtain a distribution network equipment area map, including: The unobstructed scene image is filtered by pixel grayscale threshold to obtain candidate distribution network device pixels, and the candidate distribution network device pixels are marked with seed points to obtain distribution network device seed point information; Based on the seed point information of the power distribution equipment, the unobstructed scene map is initialized with regional growth to obtain the initial growth area; The initial growth region is merged to obtain a merged region, and small regions are removed from the merged region to obtain a filtered region. The selected area is regularized to obtain a regularized distribution network equipment area map, and the regularized distribution network equipment area map is marked with boundary enhancement to obtain a distribution network equipment area map.

[0011] Furthermore, based on the seed point information of the distribution network equipment, the unobstructed scene map is initialized with region growth to obtain an initial growth region, including: The coordinate positions of the seed points of the distribution network equipment are extracted to obtain a seed point location set, and the pixel grayscale values ​​of the unobstructed scene image are read based on the seed point location set to obtain the seed point grayscale values. Based on the seed point location set, a neighborhood search window is constructed for the unobstructed scene map to obtain the neighborhood search window, and grayscale values ​​are collected for the pixels within the neighborhood search window to obtain the neighborhood pixel grayscale set. The grayscale values ​​of the neighboring pixels and the seed point are calculated pixel by pixel to obtain the pixel grayscale difference. Pixels with a grayscale difference less than a preset growth threshold are marked with eight-neighbor connectivity to obtain the initial growth region.

[0012] Furthermore, the distribution network equipment area map is pixel-level labeled and segmented to obtain the distribution network equipment segmentation result, including: The pixel color features of the distribution network equipment area map are extracted to obtain a color feature vector. The color feature vector is then filtered by setting a color threshold range to obtain candidate pixels for device color. The candidate pixels for device color are then color-marked to obtain color-marked device information. Connectivity analysis is performed on adjacent pixels with the same color marking in the color-marked device information to obtain connected device regions, and the area of ​​the connected device regions is calculated to obtain region area data. By setting an area threshold, the area data of the region is filtered to obtain the effective equipment region. The boundary pixels of the effective equipment region are extracted to obtain the boundary pixel data. The boundary pixel data is then labeled with pixel values ​​to obtain a pixel-level labeled equipment map. Finally, the pixel-level labeled equipment map is segmented to obtain the distribution network equipment segmentation result.

[0013] The present invention also provides a semantic segmentation device for distribution network lines, comprising: The shooting module is used to shoot the power distribution network lines with a binocular camera to obtain a stereo image of the power distribution network scene, and to extract depth information from the stereo image of the power distribution network scene to obtain scene depth data; The culling module is used to perform layered processing on the scene depth data to obtain a scene layer map, and to remove occlusions based on the scene layer map to obtain an unobstructed scene map. The segmentation module is used to perform region growing processing on the unobstructed scene map to obtain a distribution network equipment area map, and to perform pixel-level annotation and segmentation on the distribution network equipment area map to obtain the distribution network equipment segmentation result.

[0014] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above methods.

[0015] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the above methods.

[0016] This invention provides a semantic segmentation method for distribution network lines, comprising the following steps: capturing images of the distribution network lines using a binocular camera to obtain a stereoscopic image of the distribution network scene; extracting depth information from the stereoscopic image to obtain scene depth data; performing layered processing on the scene depth data to obtain a scene layer map; removing occlusions based on the scene layer map to obtain an unobstructed scene map; performing region growing processing on the unobstructed scene map to obtain a distribution network equipment region map; and performing pixel-level annotation on the distribution network equipment region map to obtain the distribution network equipment segmentation result. This method solves the technical problem that existing two-dimensional image analysis methods struggle to accurately distinguish between foreground and background objects, resulting in low recognition accuracy of distribution network equipment. By layering the scene according to depth information, the system can clearly distinguish which objects are in the foreground and which are in the background, thereby achieving effective removal of occlusions. This approach directly solves the most common problem of tree and building obstruction in power distribution network inspections, allowing for accurate identification of obstructed equipment. This significantly improves the completeness and reliability of the inspection. Furthermore, by annotating the area map at the pixel level, the system can precisely outline the contours of each piece of equipment and even distinguish different components. This refined segmentation provides high-quality data support for applications such as equipment health status assessment and hazard identification, further enhancing the practical value of automated inspection. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the steps of a semantic segmentation method for distribution network lines in one embodiment of the present invention; Figure 2 This is a schematic diagram of a semantic segmentation device for distribution network lines in one embodiment of the present invention; Figure 3 This is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.

[0019] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] The embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0021] The following describes in detail, with reference to the accompanying drawings, a semantic segmentation method for distribution network lines according to an embodiment of the present invention.

[0022] Figure 1 This invention provides a semantic segmentation method for distribution network lines, comprising the following steps: Step S1: Take pictures of the power distribution network lines using a binocular camera to obtain a stereoscopic image of the power distribution network scene, and extract depth information from the stereoscopic image of the power distribution network scene to obtain scene depth data.

[0023] Specifically, in practical applications, binocular cameras need to be horizontally mounted at a fixed baseline distance (usually 6-12 cm) to create a human-like observation method. During shooting, both cameras simultaneously trigger their shutters to acquire two images of the same power distribution network scene, forming a stereoscopic image of the scene. The next step is to calculate the depth from these stereoscopic images. This is done by first matching feature points in the left and right images. For example, if a utility pole is located at position x1 in the left image and x2 in the right image, the difference is the parallax d. Then, the actual distance Z at that point is calculated using the formula Z = f × B / d, where f is the camera's focal length and B is the baseline distance. For instance, if the baseline is 10 cm, the focal length is 8 mm, and the measured parallax at a certain point is 4 pixels (pixel size 0.005 mm), then the depth is equal to 8 × 10 ÷ (4 × 0.005) = 4000 mm, or 4 meters. By performing this matching and calculation on every pixel of the entire image, a depth map can be obtained. The value at each location in the map represents the distance of the object from the camera. This is the scene depth data, which records the three-dimensional spatial information of the entire power distribution network scene.

[0024] Step S2: Perform layered processing on the scene depth data to obtain a scene layered map, and remove occlusions based on the scene layered map to obtain an unobstructed scene map.

[0025] Specifically, after obtaining the scene depth data, the image needs to be divided into several layers according to the depth values. In practice, the distribution range of the depth data is first calculated. Assuming the measured depth ranges from 2 meters to 20 meters, a layer can be created every 2 meters, resulting in depth ranges such as 2-4 meters, 4-6 meters, and 6-8 meters. Pixels with depth values ​​falling within the same range are grouped into one layer and assigned the same layer label, for example, the first layer is labeled L1, the second layer L2, etc. After traversing the entire depth data and completing the labeling, a scene layering map is formed. Taking a power distribution line as an example, leaves at a distance of 2-3 meters are labeled as layer L1, a utility pole at 5 meters belongs to layer L3, and a distant house at 15 meters is considered layer L7. This clearly shows the spatial relationships within the scene. Next, occlusions need to be removed from the scene layer map. The method is to first determine which layers are occlusions. Usually, areas that are close to the camera and have a large area are occlusion candidates. A distance threshold can be set, such as 5 meters. Any layer with a depth of less than 5 meters and a pixel connectivity area exceeding a certain area is considered an occlusion. Then, the pixel areas corresponding to these occlusion layers are deleted from the original image or marked as invalid. The remaining part is the occlusion-free scene map, which retains the distant view layer where the network equipment is located.

[0026] Step S3: Perform region growing processing on the unobstructed scene map to obtain the distribution network equipment region map, and perform pixel-level annotation and segmentation on the distribution network equipment region map to obtain the distribution network equipment segmentation result.

[0027] Specifically, after obtaining the unobstructed scene image, the specific locations of the distribution network equipment need to be identified on the image. This is done using the region growing method. First, several seed points are manually or automatically selected in the image. These points usually fall in relatively obvious locations of the equipment, such as the metal casing of a transformer or the surface of an insulator. After selecting the seed points, their eight adjacent pixels are checked. If the grayscale value or color of the adjacent pixels is not significantly different from that of the seed point (e.g., the grayscale difference is less than 15), this pixel is also included in the same region. Then, the process continues to check outwards from the newly added pixel as the center, expanding in this way until the surrounding pixels no longer meet the similarity condition. For example, if the grayscale of the seed point is 120, and the grayscale of the pixel to its right is 125, the difference of 5 is within the threshold and is therefore absorbed. However, the grayscale of the pixel to its left is 80, and the difference of 40 is too large and is discarded. Following this rule, pixels belonging to the same device are grouped together, ultimately forming several connected regions. This is the distribution network equipment region map. After obtaining the area map, fine segmentation is required. Specifically, each pixel in the area is labeled with a category. For example, the pixels of the transformer are labeled as category 1 and the pixels of the insulator are labeled as category 2. During labeling, the pixels are judged one by one along the boundary of the area. The boundary pixels are marked separately to facilitate the subsequent extraction of the contour. After all the pixels are labeled, the distribution network equipment segmentation result is obtained.

[0028] In a specific embodiment, the step of capturing images of the distribution network lines using a binocular camera to obtain a stereoscopic image of the distribution network scene includes: The parameters of the left and right lenses of the binocular camera are calibrated to obtain calibrated lenses. The distribution network line is then photographed synchronously using the calibrated lenses to obtain left and right images. The left and right images are timestamped to obtain time-synchronized image pairs, and the positions of the time-synchronized image pairs are calibrated to obtain calibrated image pairs. The disparity of corresponding pixels in the left and right images of the calibration image pair is calculated to obtain a disparity image. Based on the disparity image, image fusion is performed to obtain a stereoscopic image of the power distribution network scene.

[0029] Specifically, before using a binocular camera, the parameters of the left and right lenses must be calibrated. This calibration process requires a checkerboard calibration board. The calibration board is placed in front of the camera at different positions and angles, and 20-30 photos are taken. Then, the intrinsic parameter matrix (focal lengths fx, fy and optical center coordinates cx, cy) and distortion coefficients (radial distortion k1, k2 and tangential distortion p1, p2) of each lens are calculated using the Zhang Zhengyou calibration method. Simultaneously, the rotation matrix R and translation vector T between the two lenses must also be calculated. These parameters, once calculated, constitute the calibrated lenses. When using the calibrated camera to photograph power distribution lines, the shutters of the left and right lenses must be triggered strictly synchronously. In practice, hardware trigger signals can be used to control the exposure of both lenses at the same moment, ensuring that the scene is captured at the same instant. The left lens captures the left-side image, and the right lens captures the right-side image. Once you have these two images, immediately add a timestamp to them, recording the shooting time, such as 2025-12-25 14:32:15.234. The timestamps of the two images must match to form a pair, which is called a time-synchronized image pair.

[0030] Next, positional calibration is performed on the image pair. Even if the shutters are synchronized, the image planes of the two lenses are not on the same horizontal line. During calibration, the previously calculated rotation matrix R and translation vector T are used to perform epipolar correction on the images, ensuring that corresponding rows in the left and right images are strictly aligned. After calibration, the same object will only have a horizontal positional difference in the left and right images, without any vertical offset. This results in the calibrated image pair. Then, disparity is calculated. Specifically, a pixel is selected in the left image (e.g., coordinates x=100, y=50), and then the search continues to the left along the 50th row of the right image to find the position most similar to this point (let's assume it's at x=85). The disparity d is then equal to 100 minus 85, which equals 15 pixels. This calculation is performed for each point in the image to obtain the disparity image. Finally, the left, right, and disparity images are merged. During fusion, the left image is used as the base, and the disparity information is added as a third channel, or the disparity value is used to perform a weighted blending of the left and right images. This forms a stereoscopic image of the scene, which contains both texture and depth information of the scene.

[0031] In a specific embodiment, the step of extracting depth information from the 3D image of the power distribution network scene to obtain scene depth data includes: The pixel coordinates of the stereo image of the power distribution network scene are extracted to obtain the pixel coordinate data. Then, the disparity of each pixel coordinate data is calculated by back-calculating the disparity of each pixel using the binocular camera imaging principle to obtain the disparity value of each pixel. Based on the disparity values ​​of each pixel, the disparity value distribution of the stereo image of the power distribution network scene is analyzed to obtain a disparity distribution map. The disparity distribution map is then smoothed by filtering to obtain a smoothed disparity map. By using the preset focal length and baseline distance of the binocular camera, the disparity values ​​of each pixel in the smooth disparity map are transformed into depth values ​​to obtain the depth values ​​of each pixel, and a scene depth map is constructed based on the depth values ​​of each pixel.

[0032] Specifically, after obtaining the 3D image of the power distribution network scene, the first step is to record the coordinate position of each pixel in the image. If the image has a resolution of 1920×1080, there will be 1920 multiplied by 1080, which is 2,073,600 pixels. Each point has its horizontal coordinate x and vertical coordinate y. For example, the coordinates of the first point in the upper left corner are (0, 0), and the coordinates of the 10th point to the right are (9, 0). In this way, the coordinates of all points are extracted to form the pixel coordinate data. Next, we need to use the imaging principle of a binocular camera to inversely calculate the parallax. This process is actually stereo matching. Taking the point at coordinates (150, 200) in the left image as an example, we need to find the point that is most similar to it in the 200th row of the right image. When finding it, we usually use the block matching method. We take an 11×11 small window centered on this point, and then slide the window of the same size on the 200th row of the right image to compare. When comparing, we calculate the difference in pixel grayscale in the two windows. We can use the SAD method, which is to subtract the pixel grayscale of the corresponding position, take the absolute value, and then add them all up. We can also use the NCC method to calculate the normalized correlation coefficient. If the similarity calculated at coordinates (137, 200) in the right image is the highest, then the parallax is 150 minus 137, which equals 13 pixels. However, some points are difficult to match, such as areas with little texture like the sky. In these cases, the matching results are not very accurate. Therefore, a left-right consistency check is also needed. This involves performing the matching from the right image as well to see if the results match. If the disparity of 13 calculated for the left image point (150, 200) corresponds to the right image point (137, 200), then calculating the disparity of 13 from the right image point (137, 200) should also yield the disparity of 13 corresponding to the left image point (150, 200). Only when the results from both sides match is the disparity value considered reliable. Calculating the disparity value for each pixel in the image in this way yields the disparity value for each pixel.

[0033] After obtaining these disparity values, we need to examine their distribution. This is done by counting the frequency of each disparity value. Assuming the disparity range is 0 to 64, we can plot a histogram to see which disparity values ​​occur more frequently and which occur less frequently. For example, points with disparities between 10 and 15 are particularly dense, corresponding to a large plane, while points with disparities around 30 are sparser, representing distant objects. Representing this statistical result as an image is the disparity distribution map. Next, we need to smooth this distribution map because the calculated disparity values ​​inevitably contain some noise. For example, a point might have a disparity of 12 surrounding it but jump to 25. These isolated outliers need to be smoothed out. Median filtering or Gaussian filtering can be used. Median filtering involves sorting all disparity values ​​within a 5x5 radius and replacing the current point with the median value. Gaussian filtering uses a Gaussian kernel to apply a weighted average to the surrounding points, for example, a center weight of 0.4 and a perimeter weight of 0.15. The filtering radius is typically set to 3 to 5 pixels. After processing, the disparity map becomes smooth and continuous; this is called a smoothed disparity map.

[0034] Finally, parallax needs to be converted into actual depth. The conversion formula is Z = f multiplied by B and then divided by d. Here, Z is the depth distance, f is the camera focal length, B is the baseline distance between the two lenses, and d is the parallax value. For example, if the camera focal length f is 8 mm, which is approximately 1600 pixels (based on sensor size), the baseline distance B is 100 mm, and the parallax d measured at a certain point is 20 pixels, then the depth Z is equal to 1600 multiplied by 100 and divided by 20, which equals 8000 mm, or 8 meters. However, it is important to maintain consistent units. If the focal length is in pixels, the baseline distance should also be converted to pixels, or both should be in millimeters; in short, keep them consistent. This conversion is performed on every pixel in the smoothed parallax map. Points with a parallax of 0 indicate a failed match, so their depth is set to infinity or marked as invalid. After the calculation, each pixel has a corresponding depth value, which are the pixel depth values. Arrange these depth values ​​according to the coordinates of the original image to form a depth map of the same size as the original image. Instead of color, each position in the map stores distance data. For example, the position (100, 50) stores 6.5 meters, and the position (200, 100) stores 12.3 meters. This map is the scene depth map, also known as scene depth data. It completely records how far each point in the power distribution network scene is from the camera.

[0035] In a specific embodiment, the scene depth data is processed into layers to obtain a scene layer map, including: The depth data of the scene is statistically sorted by depth value to obtain an ordered sequence of depth values. The ordered sequence of depth values ​​is then divided into intervals to obtain depth layer intervals. The depth layer intervals include depth threshold features at different distances (near, medium, and far) in the distribution network line scene. The distribution network scene stereo image is classified into pixels based on the depth layering interval to obtain a depth-classified pixel set, and the depth-classified pixel set is region-marked to obtain a depth-marked image. The depth marker image is normalized to obtain a normalized depth image, and the normalized depth image is then rendered with layered colors to obtain a scene layering map.

[0036] Specifically, after obtaining the scene depth data, all depth values ​​need to be collected and statistically analyzed. For example, if the entire image has 2 million pixels, there will be 2 million depth values. These values ​​are arranged in ascending order to form an ordered sequence of depth values, such as 1.2 meters, 1.2 meters, 1.3 meters, 1.5 meters... up to the maximum of 25.8 meters. Next, this sequence needs to be divided into several segments. The segmentation is not a simple equal division, but rather takes into account the actual situation of the power distribution scenario. Usually, near distance refers to obstructions such as tree branches and weeds within 5 meters, medium distance is within 5 to 15 meters of utility poles or low-voltage lines, and far distance is often beyond 15 meters of high-voltage towers or background buildings. Therefore, three thresholds can be set: 5 meters, 15 meters, and 25 meters. This divides the data into four intervals: 0 to 5 meters, 5 to 15 meters, 15 to 25 meters, and greater than 25 meters. Each interval represents a depth level, and these intervals together form the depth layer intervals. The depth threshold feature mentioned here refers to these key dividing points, which reflect the distribution patterns of objects at different distances (near, medium, and far) in the distribution network scenario.

[0037] After defining the intervals, each pixel in the original image needs to be assigned to its corresponding layer. Specifically, this involves iterating through all pixels in the 3D image of the network distribution scene, reading the depth value of each pixel in the scene depth data, and then determining which interval this value falls into. For example, a pixel with a depth of 3.2 meters is less than 5 meters, so it's assigned to the first layer; another pixel with a depth of 12.5 meters is between 5 and 15 meters, so it's assigned to the second layer. This process categorizes all pixels in the image by depth value, placing pixels in the same layer into a set, resulting in several depth-categorized pixel sets. Next, these categorized pixels need to be labeled. The labeling method assigns the same layer number to pixels in the same layer: all pixels in the first layer are labeled 1, the second layer 2, the third layer 3, and the fourth layer 4. After labeling, a new image is formed. In this image, each location no longer stores RGB color or depth in meters, but rather a layer number; this is the depth-labeled image. For example, if the depth at coordinates (100, 50) in the original image is 3.2 meters, write 1 at position (100, 50) in the depth marker image; if the depth at coordinates (200, 100) is 12.5 meters, write 2.

[0038] After obtaining the depth-tagged image, you'll find that the boundaries between layers are somewhat rough. This is because the depth values ​​of adjacent pixels are just around the threshold, causing layer instability. For example, pixels with depths of 4.9 meters and 5.1 meters are adjacent, labeled as layer 1 and layer 2 respectively, resulting in a fragmented boundary. To smooth these boundaries, morphological operations are commonly used. First, an opening operation is performed to remove small imperfections. The opening operation involves erosion followed by dilation. During erosion, a 3×3 structuring element scans the image. If a pixel labeled 1 is surrounded by a pixel labeled 2, that pixel is changed to 2, thus smoothing out sharp corners. Then, dilation fills in the recessed areas, maintaining a relatively constant region size while smoothing the boundaries. Alternatively, a closing operation can be used to fill small holes. The closing operation involves dilation followed by erosion, filling in gaps within the layers. After these operations, the contours between layers become regular and continuous, forming a regular depth image.

[0039] Finally, this normalized image needs to be rendered into a color image that is easily recognizable by the human eye. During rendering, different colors are assigned to different layers. For example, the first layer, representing the nearest occlusion, is represented by red; the second layer by yellow; the third by green; and the fourth by blue. Specifically, the normalized depth image is traversed, and the layer number at each location is read. If it's 1, the corresponding position is filled with the RGB value (255, 0, 0), which is red; if it's 2, it's filled with (255, 255, 0), yellow; if it's 3, it's filled with (0, 255, 0), green; and if it's 4, it's filled with (0, 0, 255), blue. After all the rendering is complete, a clearly defined image is obtained, with different colors representing different depth layers. This is the scene layering map. Looking at this image, you can immediately see which areas are near-field occlusions, which are mid-field devices, and which are distant background elements; the layering relationships are very intuitive.

[0040] In a specific embodiment, the 3D image of the power distribution network scene is classified into pixels based on the depth layering interval to obtain a depth-classified pixel set, including: The depth value of each pixel in the 3D image of the power distribution network scene is read to obtain a pixel depth value sequence. The pixel depth value sequence is then compared and judged based on the depth layer interval to obtain a pixel level label. Based on the pixel-level labels, the pixels in the 3D image of the power distribution network scene are aggregated at the same layer to obtain a depth-classified pixel set.

[0041] Specifically, when processing 3D images of power distribution network scenes, the depth information needs to be read pixel by pixel. The operation involves establishing a double loop: the outer loop iterates through the image rows from 0 to the height minus 1, while the inner loop iterates through the columns from 0 to the width minus 1. If the image resolution is 1920×1080, the outer loop iterates 1080 times, and the inner loop iterates 1920 times each, ensuring that every pixel is accessed. During access, the scene depth data is retrieved using the coordinates (x, y), and the corresponding depth value is read. For example, the depth of point (100, 50) is 4.2 meters, and the depth of point (500, 300) is 11.8 meters. Arranging all the depth values ​​from left to right and top to bottom forms a pixel depth value sequence, the length of which is equal to the total number of pixels in the image.

[0042] After obtaining the pixel depth value sequence, it needs to be compared with the previously defined depth stratification intervals. Let's assume the intervals are set as follows: 0-5 meters for the near-field layer, 5-15 meters for the mid-field layer, 15-25 meters for the far-field layer, and greater than 25 meters for the ultra-far-field layer. The determination is based on which range each depth value falls into. For example, the value 4.2 meters we just read is used with an IF statement: if the depth value is less than 5, it's classified as near-field (L1); if the depth value is greater than or equal to 5 and less than 15, it's classified as mid-field (L2); if it's greater than or equal to 15 and less than 25, it's classified as far-field (L3); and if it's greater than or equal to 25, it's classified as ultra-far-field (L4). For the 4.2-meter point, because 4.2 is less than 5, it's labeled L1, while the 11.8-meter point, being between 5 and 15, is labeled L2. It's important to handle boundary cases here; a point exactly 5 meters is classified as mid-field, not near-field, to avoid duplicate classification. Every depth value in the entire sequence must be judged in this way. After the judgment is completed, each pixel obtains a level label. The labels of all pixels are combined to form the pixel level label.

[0043] After obtaining these labels, an aggregation operation is needed, which involves collecting pixels with the same label together. In actual programming, four empty sets can be created to correspond to the four levels L1 to L4 respectively. Then, all pixels in the 3D image of the network scene are traversed again, and the level label of each point is read. If the label is L1, the coordinates (x, y) of this point are added to the L1 set; if the label is L2, it is added to the L2 set, and so on. For example, when traversing to the point (100, 50), its label is found to be L1, so the operation L1_set.add((100, 50)) is executed to put this coordinate into the L1 set; when traversing to the point (500, 300), the label is L2, so L2_set.add((500, 300)) is executed to put it into the L2 set. After traversing the entire image in this way, set L1 stores the pixel coordinates of the foreground layer, set L2 stores the coordinates of the midground layer, set L3 stores the coordinates of the background layer, and set L4 stores the coordinates of the super-distant layer. The pixels in each set are spatially scattered throughout the image, but their depth values ​​belong to the same range.

[0044] However, coordinates alone are not enough. Sometimes it's also necessary to save other pixel attributes for later processing. Therefore, during aggregation, in addition to recording coordinates, the RGB color value and depth value of the pixel can also be stored together, making each element a quadruple (x, y, RGB, depth). For example, the color of point (100, 50) is (120, 150, 80), and its depth is 4.2 meters. When stored in set L1, it is recorded as (100, 50, (120, 150, 80), 4.2). This way, when this pixel is used later, all the information can be retrieved directly without having to look up the original image. After all pixels have been categorized, the four sets L1 to L4 are depth-categorized pixel sets, which organize the originally randomly distributed pixels according to depth levels. In some implementations, to improve access efficiency, a spatial index is also created for each set, such as using a quadtree to organize pixels at the same level according to their spatial location. This allows for quick location of pixels within a certain area without traversing the entire set. In addition, in practical applications, if a pixel with an invalid depth value is encountered, such as a point that fails stereo matching due to occlusion or reflection, its depth is marked as negative or infinite. Such points need to be handled separately during the determination. Usually, they are classified into a special set of invalid pixels or skipped directly and not included in the layering.

[0045] In a specific embodiment, the step of removing occlusions based on the scene layer map to obtain an unobstructed scene map includes: Based on the scene layer map, occlusion region identification is performed to obtain occlusion candidate regions, and contour extraction is performed on the occlusion candidate regions to obtain occlusion contour lines. Based on the occlusion contour line, the occlusion candidate region is accurately marked to obtain an accurate occlusion region map; The occlusion pixel is replaced in the occlusion precision area map to obtain an initial occlusion-free map, and the edge smoothing process is performed on the initial occlusion-free map to obtain an occlusion-free scene map.

[0046] Specifically, after obtaining the scene layer map, the first step is to identify which regions are occluders. The criteria for this are mainly twofold: first, the depth level, typically the near-field layer marked L1, which is most likely to be occluders; and second, the area of ​​the region, as areas that are too small are simply noise and not considered occlusions. In practice, first, perform connected component analysis on the scene layer map to extract all pixels marked L1. Then, use an 8-connectivity algorithm to group adjacent L1 pixels into connected regions. For example, a group of L1 pixels in the upper left corner of the image forms region A, and another group of L1 pixels in the lower right corner forms region B. Next, calculate the number of pixels in each region, setting an area threshold, such as 500 pixels. If a region contains more than this threshold, it is considered a true occluder rather than isolated noise. Collecting these qualifying regions yields the occlusion candidate regions. For example, region A has 1200 pixels, exceeding the threshold, so it is selected, while region B, with only 80 pixels, is too small and is excluded.

[0047] After finding these candidate regions, their boundary contours need to be delineated. Boundary tracking algorithms are used to extract the contours, starting from a certain edge pixel of the region and traversing the boundary, recording the points passed. Specifically, the image is first scanned to find the first boundary point of the occlusion candidate region. The judgment method is to see if this point itself belongs to the occlusion region, but its upper adjacent pixels do not belong to the occlusion region. After finding this starting point, tracking begins. Each time, the neighbors in the eight directions around the current point are checked, and the next boundary point is found in a clockwise or counterclockwise order. The criteria for the next boundary point are that it must belong to the occlusion region and at least one neighbor does not belong to the occlusion region. This process continues until the starting point is returned, forming a closed curve. For example, starting from point (100, 50), if the point to its right (101, 50) is also a boundary, the process moves there. Then, starting from (101, 50), if the point to its lower right (102, 51) is found to be a boundary, the process continues. In this way, the boundary points of the entire region are connected to form the occlusion contour line. If there are multiple disconnected blocks in the occlusion candidate region, the contour of each block must be extracted separately, and the final occlusion contour line contains multiple curves.

[0048] After obtaining the outline, the occluded areas need to be marked more precisely. Previously, the L1 layer was only roughly considered as an occluder, but it contains some elements that shouldn't be removed, such as nearby network equipment components. Precise marking requires further filtering based on the shape characteristics of the outline. Occluders are usually irregular in shape; their roundness or rectangularity can be calculated. Roundness is defined as 4π times the area divided by the square of the perimeter. If this value is close to 1, the outline is very round and regular, indicating equipment components rather than occluders. If the value is small, the outline is fragmented and irregular, likely indicating occluders like branches or weeds. Additionally, the texture characteristics of the area should be checked. Occluders, like leaves, often have messy textures. The standard deviation of pixel grayscale within the area can be calculated, or an LBP texture descriptor can be used. A large standard deviation indicates messy textures. By combining these features, only areas that are in the foreground layer, have a large enough area, irregular shape, and messy texture are finally identified as occlusions. The identified occlusion areas are marked on the original image with special values, such as changing the label of these pixels to -1 to indicate occlusion. This gives us a precise map of the occlusion area.

[0049] Next, the pixels in the occluded area need to be replaced. There are several replacement methods, the simplest being filling in with the surrounding background color. Specifically, this involves iterating through the precise occlusion area map to find all pixels marked as -1. For each occluded pixel, the average color of its surrounding non-occluded pixels is taken and used as the filler. For example, if point (100, 50) is occluded, and its 8 neighbors include 6 non-occluded pixels with RGB colors of (120, 130, 110), (125, 128, 115), etc., the average of the R, G, and B channels of these 6 colors is calculated to obtain (122, 129, 112). This color is then used to replace the original value of point (100, 50). However, this simple averaging method is ineffective when the occlusion area is large. A better approach is to use image inpainting algorithms, such as patch-based inpainting, which finds similar texture patches around the occluded area to fill the gaps. During the repair process, first find a pixel to be filled at the boundary of the occluded area. Take a small window, such as 9×9, centered on it. Half of this window is in the known area and the other half is in the occluded area. Then, search for the window that is most similar to the known half in the non-occluded part of the image. After finding it, use its center pixel value to fill the current point. After filling all the occluded pixels one by one, the initial unoccluded image is obtained.

[0050] Finally, edge smoothing is needed on the replaced image because color abrupt changes or seam marks appear at the boundary between the occluded area and the surrounding background. During smoothing, a transition zone is formed by extending a few pixels inwards and outwards along the original occlusion outline, for example, extending 3 pixels inwards and outwards. Within this transition zone, pixels are weighted and blended, with points closer to the occluded area receiving a higher weight and using the restored color, and points closer to the original background receiving a higher weight and using the original color. Points in the middle are blended using distance interpolation. Specific algorithms can use bilateral filtering or guided filtering. Bilateral filtering preserves edge details while smoothing without blurring the entire image. The filter radius is set to 5 to 7 pixels, and the intensity parameter is adjusted according to the size of the occluded area. After processing the transition zone, the connection between the occluded area and the background becomes natural, with no obvious stitching marks, resulting in an unoccluded scene image.

[0051] In a specific embodiment, the unobstructed scene map is subjected to region growing processing to obtain a distribution network equipment area map, including: The unobstructed scene image is filtered by pixel grayscale threshold to obtain candidate distribution network device pixels, and the candidate distribution network device pixels are marked with seed points to obtain distribution network device seed point information; Based on the seed point information of the power distribution equipment, the unobstructed scene map is initialized with regional growth to obtain the initial growth area; The initial growth region is merged to obtain a merged region, and small regions are removed from the merged region to obtain a filtered region. The selected area is regularized to obtain a regularized distribution network equipment area map, and the regularized distribution network equipment area map is marked with boundary enhancement to obtain a distribution network equipment area map.

[0052] Specifically, after obtaining the unobstructed scene image, the first step is to filter the pixels of the power distribution equipment from the background. Power distribution equipment is usually made of metal or ceramic, and it appears in a specific grayscale range in the image. For example, transformer casings are mostly grayish-white with a grayscale value between 180 and 230, while insulators are white or light gray with a grayscale value between 200 and 250. During the filtering process, the unobstructed scene image is first converted to a grayscale image. For each pixel, its grayscale value is checked to see if it falls within the set range. Assuming the set threshold is a lower limit of 170 and an upper limit of 240, then a sky pixel with a grayscale of 120 does not meet the criteria and is excluded, while a device pixel with a grayscale of 210 meets the criteria and is retained. Collecting all pixels that meet the criteria yields the candidate power distribution equipment pixels. However, this screening method may miss some equipment components because some equipment surfaces may have rust or shadows that cause the grayscale to be low. Therefore, it is necessary to combine color information for supplementary judgment, convert the RGB three-channel values ​​to the HSV color space, and determine whether the hue H is in the gray or white range and whether the saturation S is relatively low (the saturation of gray objects is usually less than 0.3). Pixels that meet both grayscale and color conditions are more reliable.

[0053] After selecting candidate pixels, seed points need to be chosen as the starting point for region growing. Seed points cannot be chosen arbitrarily; they must be selected at locations where device features are clearly defined. Specifically, gradient calculations are performed on the candidate distribution network device pixels. Large gradients indicate drastic grayscale changes, typically at edges, making them unsuitable as seed points. Small gradients indicate relatively flat areas, making them suitable as seeds. The Sobel operator is used to calculate the gradients in the x and y directions, respectively. The gradient magnitude G is then calculated as sqrt(Gx² + Gy²). All candidate pixels are iterated through to find points with gradients less than a certain threshold, such as 20. These points are then clustered according to their grayscale values, grouping points with similar grayscale values ​​into one group. A representative point is selected from each group as a seed point. This ensures that different devices or different parts of the same device have seed point coverage. For example, three seeds are selected from candidate points near a grayscale value of 210 on the transformer body, and two seeds are selected from candidate points near a grayscale value of 230 on the insulator. The coordinates and grayscale values ​​of these seed points are recorded to form the distribution network device seed point information.

[0054] With seed points, region growing can begin. Initially, a blank marker map of the same size as the unoccluded scene map is created. The seed point is marked as 1 on the marker map, indicating it has been visited. The seed point is then pushed into a queue. The first point popped from the queue, for example, with coordinates (100, 50) and a grayscale value of 210, is checked. The grayscale values ​​of its eight neighbors are examined. If a neighbor with coordinates (101, 50) has a grayscale value of 208, and the difference between it and the seed point is 2 (less than a set similarity threshold, such as 15), this neighbor is considered to belong to the same device region. It is then marked as 1 on the marker map and pushed into the queue. The next point is popped from the queue, and the process is repeated. This continues expanding outwards until the queue is empty. At this point, all pixels marked as 1 form a connected region, which is grown from the first seed. The same operation is then repeated with the second seed, and the newly grown region is marked as 2 on the marker map. The region grown from the third seed is marked as 3, and so on. After processing all seeds, the initial grown region is obtained. Different numbers on the marker map represent different device regions.

[0055] After the initial growth is complete, some regions that should be connected are found to be separated. For example, the upper and lower parts of a transformer may have a grayscale jump due to a weld seam in the middle, resulting in two regions labeled 1 and 2 growing from two different seeds, even though they are actually the same device. At this point, region merging is required. The criterion is whether the boundaries of two regions are very close and the grayscale difference is small. Specifically, the marker map is traversed to find adjacent different markers. For example, if a point is labeled 1 and its right neighbor is labeled 2, then region 1 and region 2 are adjacent. The average grayscale of these two regions is calculated. If the difference is less than a threshold, such as 10, they are considered to be merged. All pixels labeled 2 are changed to label 1, thus merging the two regions into one. The gaps between regions also need to be checked. If two regions are not directly adjacent but are separated by only a few pixels, morphological closing operations can be used to fill the gaps and connect them. The closing operation parameters are set to a structuring element size of 5×5 and performed 3 iterations. After processing all adjacent regions, the merged region is obtained.

[0056] After merging, some small fragments will remain. These small areas, only a few dozen pixels in size, are noise or debris and not actual network equipment; they need to be removed. During removal, the number of pixels corresponding to each marker value on the marker map is counted. For example, marker 1 has 5800 pixels, marker 3 has 45 pixels, and marker 5 has 12000 pixels. An area threshold is set, such as 300 pixels. Areas with fewer pixels than this threshold are deleted. Specifically, the markers of these small areas are changed to 0 to represent the background. For example, marker 3 has only 45 pixels, less than 300, so all points marked 3 are changed to 0. This removes the small fragments, leaving only the large enough device areas—this is the filtering area.

[0057] After filtering, the boundaries of the area may still have some rough edges or depressions. To smooth the boundaries, we need to refine the outline. The method is similar to handling occluded areas, using morphological operations. First, we perform an opening operation to remove protruding sharp corners. The structuring element is a circle with a radius of 3 pixels. The opening operation involves erosion followed by dilation. During erosion, if a device pixel is surrounded by background pixels, it is temporarily replaced with background pixels, thus smoothing out protruding points on the boundary. Then, dilation fills in the recessed boundary. Next, we perform a closing operation to fill in the small depressions on the boundary. The closing operation involves dilation followed by erosion. During dilation, the device area expands outward to fill in the depressions, and then erosion recedes back to maintain the overall size. After the opening and closing operations, the area outline becomes rounded and regular, resulting in a regularized distribution network device area map.

[0058] Finally, the boundaries of the device areas need to be highlighted for easier subsequent processing. This is done by drawing another ring around the perimeter of the device area on the marked map, scanning the map to find all device pixels with non-zero markers, and checking their eight neighbors. If a neighbor has a marker of 0, it means the current point is on the boundary. This point is then marked with a special symbol, such as -1 to represent a boundary point, or by adding 1000 to the original marker value for differentiation. For example, a device internal point originally marked as 1 remains unchanged, while a boundary point is changed to 1001. After all markings are completed, the internal areas and boundary outlines of the devices are clearly separated. This is the distribution network device area map, which indicates which pixels belong to the devices and which belong to the background, and also marks the precise boundary position of each device.

[0059] In a specific embodiment, the unobstructed scene map is initialized with region growth based on the seed point information of the distribution network equipment to obtain an initial growth region, including: The coordinate positions of the seed points of the distribution network equipment are extracted to obtain a seed point location set, and the pixel grayscale values ​​of the unobstructed scene image are read based on the seed point location set to obtain the seed point grayscale values. Based on the seed point location set, a neighborhood search window is constructed for the unobstructed scene map to obtain the neighborhood search window, and grayscale values ​​are collected for the pixels within the neighborhood search window to obtain the neighborhood pixel grayscale set. The grayscale values ​​of the neighboring pixels and the seed point are calculated pixel by pixel to obtain the pixel grayscale difference. Pixels with a grayscale difference less than a preset growth threshold are marked with eight-neighbor connectivity to obtain the initial growth region.

[0060] Specifically, after obtaining the seed point information of the distribution network equipment, the first step is to extract the coordinate data. The seed point information is usually stored as a list or array, with each element containing the x-coordinate, y-coordinate, and other attributes of the seed point. During extraction, the list is traversed to extract the (x, y) coordinate pairs of each seed and store them separately. For example, the first seed coordinate is (120, 85), the second is (340, 150), and the third is (580, 220). These coordinates are collected to form the seed point location set. After obtaining the location set, we need to look up the grayscale values ​​of these points in the unoccluded scene image. This is because grayscale is used as a standard when determining similarity in region growing. When reading, we first convert the unoccluded scene image into a grayscale image. If the original image is in color, we use the formula Gray = 0.299 multiplied by R + 0.587 multiplied by G + 0.114 multiplied by B to calculate the grayscale value. After conversion, we read the grayscale value based on the coordinates of the seed points. For example, the grayscale value corresponding to point (120, 85) is 205, the grayscale value of point (340, 150) is 218, and the grayscale value of point (580, 220) is 192. By pairing and recording the coordinates of each seed point with its corresponding grayscale value, we obtain the grayscale value of the seed point.

[0061] Next, starting from each seed point, we search for pixels belonging to the same device in the surrounding area. This search isn't aimless but rather within a defined range, known as the neighborhood search window. The window is constructed by expanding outwards from the seed point by a fixed distance. For example, if the window size is 7x7, then for the seed point (120, 85), the top-left corner coordinates are (120 - 3, 85 - 3) = (117, 82), and the bottom-right corner is (120 + 3, 85 + 3) = (123, 88). This defines a rectangular area containing 49 pixels as the neighborhood search window. However, we must pay attention to boundary conditions. If the seed point is close to the image edge, the window may exceed the image's boundaries. In this case, cropping is necessary. For example, if the seed point is at (5, 10), the left boundary of a 7x7 window should be 5 - 3 = 2, but the top boundary is 10 - 3 = 7, both within the image. The key is to ensure the window doesn't exceed these boundaries. A window is constructed for each seed point; overlapping windows between different seeds are acceptable, as they are processed independently.

[0062] After the window is created, the grayscale values ​​of all pixels within the window need to be collected. This is done by using a double loop to traverse the window range. The outer loop scans from the beginning row to the end row, and the inner loop scans from the beginning column to the end column. For each point scanned, its grayscale value is read from the grayscale image. For example, for a 7×7 window with a seed point (120, 85), the outer loop loops from row 82 to 88, and the inner loop loops from column 117 to 123. The grayscale value of point (117, 82) is 200, the grayscale value of point (118, 82) is 203, and the grayscale value of point (119, 82) is 198. After reading the grayscale values ​​of all 49 points within the window, these values ​​are stored in an array, which is the neighborhood pixel grayscale set. Each seed point has its own corresponding neighborhood pixel grayscale set; for example, seed 1 has a set of 49 grayscale values, and seed 2 also has a set of 49 values. These sets are independent of each other.

[0063] After collecting the grayscale values ​​of the neighboring pixels, it's necessary to determine which pixels are sufficiently similar to the seed point to be grouped into the same region. This is done by calculating the grayscale difference between the neighboring pixels and the seed point. For example, the seed point (120, 85) has a grayscale value of 205, while the point within the window (117, 82) has a grayscale value of 200. The difference is 205 minus the absolute value of 200, which equals 5. The point (118, 82) has a grayscale value of 203, so the difference is 2. The point (119, 82) has a grayscale value of 198, so the difference is 7. This process of subtracting the seed point's grayscale value from each value in the neighboring pixel set and taking the absolute value gives the pixel grayscale difference. It's important to maintain consistency in the calculation order: always subtract the neighboring pixel's grayscale value from the seed point's grayscale value and then take the absolute value, or vice versa. The key is to keep this consistent and remove the sign when taking the absolute value. After calculating all the neighboring points, a set of difference data will be obtained, for example, the difference set is {5, 2, 7, 3, 1, 12, 8, 4, ...}.

[0064] After obtaining the difference, it needs to be compared with the preset growth threshold. The threshold is usually set between 10 and 20, and the specific value is adjusted according to the image contrast. For images with high contrast, the threshold can be set smaller, such as 10, and for images with low contrast, a larger threshold can be set, such as 18. Assuming the threshold is set to 12, then the difference 5 calculated earlier is less than 12, so point (117, 82) meets the condition. The difference 2 is less than 12, so point (118, 82) also meets the condition. The difference 7 meets the condition, but the difference 12 is exactly equal to the threshold, which does not meet the condition according to the less-than rule. If it is changed to less than or equal to, it meets the condition. This depends on the specific implementation. Generally, the less-than sign is more strict. All pixels in the neighboring pixel grayscale set whose difference is less than the threshold are selected. These pixels are considered to belong to the same device region as the seed point.

[0065] The selected pixels undergo connectivity checks because isolated points, despite similar gray levels, are spatially disconnected and should not be included in the region. The connectivity check uses the eight-neighbor rule: it checks if at least one point in each of the current pixel's eight surrounding directions (up, down, left, right, and four diagonal angles) has been marked as part of the growth region. If so, the current pixel is connected to the region; otherwise, it's isolated and temporarily excluded. In practice, a labeling matrix of the same size as the image is created, initially all values ​​are 0. Seed points are marked as 1 at their corresponding positions in the matrix. Then, the previously selected similar pixels are checked. For each similar pixel, its eight neighbors are checked. If a neighbor's value in the labeling matrix is ​​1, the pixel is considered connected, and it is also marked as 1 and added to a processing queue. This process is repeated incrementally. For example, the seed point (120, 85) is first marked as 1. Its neighbor (121, 85) has a matching grayscale difference and is horizontally adjacent to the seed point, so it is also marked as 1 because it is connected. Next, the neighbor (122, 85) of (121, 85) is checked. It is adjacent to (121, 85) and also has a matching grayscale difference, so it is marked as 1. This process continues outwards in circles. After processing the entire neighborhood window corresponding to the seed point, all pixels marked as 1 on the labeling matrix form a connected region. This is the initial growth region grown from this seed.

[0066] If there are multiple seed points, repeat the above process. The second seed is marked with a 2 in the labeling matrix to indicate its growth area, the third seed with a 3, thus distinguishing the different devices grown from different seeds. After all seeds have been processed, the number 0 in the labeling matrix represents the background, the number 1 represents the first device area, the number 2 represents the second device area, and so on. The entire labeling matrix is ​​a complete representation of the initial growth area. It is important to note the situation where two regions from different seeds meet during the growth process. For example, if the first seed grows from the left and the second from the right, and they meet in the middle, it is necessary to determine whether merging is allowed. If the gray levels of the seed points in the two regions are very similar, for example, both between 200 and 210, they can be merged into one region; otherwise, they remain independent.

[0067] In a specific embodiment, the distribution network equipment area map is pixel-level labeled and segmented to obtain the distribution network equipment segmentation result, including: The pixel color features of the distribution network equipment area map are extracted to obtain a color feature vector. The color feature vector is then filtered by setting a color threshold range to obtain candidate pixels for device color. The candidate pixels for device color are then color-marked to obtain color-marked device information. Connectivity analysis is performed on adjacent pixels with the same color marking in the color-marked device information to obtain connected device regions, and the area of ​​the connected device regions is calculated to obtain region area data. By setting an area threshold, the area data of the region is filtered to obtain the effective equipment region. The boundary pixels of the effective equipment region are extracted to obtain the boundary pixel data. The boundary pixel data is then labeled with pixel values ​​to obtain a pixel-level labeled equipment map. Finally, the pixel-level labeled equipment map is segmented to obtain the distribution network equipment segmentation result.

[0068] Specifically, after obtaining the distribution network equipment area map, color features need to be extracted from each pixel. During extraction, the original unobstructed scene image is first retrieved to read the RGB three-channel values. For example, if a pixel's R is 180, G is 185, and B is 175, these three values ​​are combined to form a three-dimensional vector (180, 185, 175). However, the RGB color space is not intuitive enough for describing color, so it is usually converted to the HSV color space. In the conversion formula, H represents the color angle (0 to 360 degrees), S represents the saturation range (0 to 1), and V represents the lightness (also 0 to 1). After conversion, a vector like (210°, 0.15, 0.73) is obtained. This conversion is performed on each pixel marked as a device in the distribution network equipment area map. Collecting the HSV values ​​of all pixels yields a color feature vector, which is a large array where each element is a triplet.

[0069] Next, threshold ranges need to be set based on the color characteristics of the distribution network equipment to filter out the true device pixels. Common colors of distribution network equipment include gray for transformers, white for insulators, and black for wires. Multiple thresholds can be set to correspond to different equipment types. For example, the threshold range for gray equipment is H between 180 and 240 degrees, S less than 0.3, and V between 0.5 and 0.9; for white equipment, H is unlimited, S less than 0.2, and V greater than 0.8; and for black equipment, H is unlimited, S less than 0.3, and V less than 0.4. During the filtering process, each element in the color feature vector is traversed to determine whether it falls within a certain threshold range. If a pixel's HSV is (210°, 0.15, 0.73), its H is 210 (within the range of 180 to 240), its S is 0.15 (less than 0.3), and its V is 0.73 (within the range of 0.5 to 0.9). All three conditions are met, so this pixel matches the characteristics of gray equipment and is selected. All pixels that meet any set of thresholds are collected to form candidate pixels for device colors.

[0070] After filtering out candidate pixels, they need to be labeled with color categories. The labeling method is determined by which threshold the pixel matches. Pixels matching the gray threshold are labeled as category 1, those matching the white threshold as category 2, and those matching the black threshold as category 3. In practice, a labeling matrix of the same size as the original image is created, with all initial values ​​of 0. The coordinates of the candidate pixels and their corresponding categories are recorded. For example, if the point (120, 85) matches gray and is labeled as 1, then 1 is written at position (120, 85) in the matrix; if the point (340, 150) matches white and is labeled as 2, then 2 is written. After processing all candidate pixels, this labeling matrix is ​​the color-labeled device information.

[0071] After obtaining the color markers, the goal is to identify which pixels are connected to form complete devices. This is done using a connected component analysis algorithm. The algorithm scans the color marker device information to find the first pixel with a non-zero marker as the starting point. It checks its eight neighbors; if a neighbor's marker value is the same, the pixel is considered connected. This neighbor is added to the current region, and the process continues recursively until no more pixels with the same marker are found. This forms a connected component. The process continues, scanning for the next unprocessed non-zero pixel and repeating the above steps. Finally, all spatially connected pixels with the same marker are grouped together; each group is a connected device region. For example, in the upper left corner of an image, there might be a patch of pixels marked 1 forming a 500-pixel area, and in the lower right corner, another patch of pixels marked 1 forming an 800-pixel area. Although both are marked 1, they are not connected and are therefore two independent connected device regions. Adding the regions marked 2 and 3, a total of more than ten connected components are obtained.

[0072] After obtaining the connected components, we need to calculate how many pixels each region contains. The calculation method is simple: count the number of pixels in each connected component. This count can be accumulated during the connectivity analysis process. For example, for the first connected component, increment the counter by 1 for each connected pixel found. After traversing this component, the counter value is its area. For the second connected component, start counting again from 0. Record the areas of all connected components to form the region area data. Suppose the obtained data is a sequence like {500, 800, 45, 6800, 120, 3200, 28, ...}, where each number represents the number of pixels in a connected component.

[0073] Next, an area threshold is used to filter out areas that are too small. These small areas are simply noise or clutter, not actual network equipment. The threshold is typically set to 300 to 500 pixels; let's say 400. Then, for each value in the area data, keep the area if it's 500 > 400, keep the area if it's 800 > 400, delete the area if it's 45 < 400, keep the area if it's 6800 > 400, delete the area if it's 120 < 400, delete the area if it's 28 < 400, and keep the area if it's 3200 > 400. The remaining areas after filtering are the valid equipment areas. During filtering, the color-coded device information should be updated simultaneously. The pixels corresponding to the deleted small areas should be changed back to 0 in the label matrix to represent background, while the retained areas maintain their original labels.

[0074] After identifying the valid device regions, the boundary contours of each region need to be extracted. When extracting boundaries, the marker matrix is ​​scanned to find pixels belonging to the device region. The method to determine if a pixel is a boundary pixel is to check its eight neighbors. If at least one of the eight neighbors is a background pixel (marked as 0), the current point is on the boundary. If all eight neighbors are of the same type, the current point is inside the region and not on the boundary. For example, point (120, 85) is marked as 1. Its upper neighbor (120, 84) is also marked as 1, and its right neighbor (121, 85) is marked as 1, but its left neighbor (119, 85) is marked as 0 (background). Therefore, point (120, 85) is a boundary pixel. The boundary pixel coordinates of all valid device regions are collected to form boundary pixel data.

[0075] After extracting the boundaries, each pixel in the device region needs to be labeled in detail. This labeling not only records which device region a pixel belongs to, but also distinguishes whether it's on the boundary or inside. Specifically, an additional labeling layer is built on top of the labeling matrix. Pixels inside a region are labeled "Device Region X - Inside", and pixels on the boundary are labeled "Device Region X - Boundary", where X is the region number. Alternatively, numerical encoding can be used. For example, pixels inside region 1 are labeled 101, and pixels on the boundary are labeled 1001; pixels inside region 2 are labeled 102, and pixels on the boundary are labeled 1002. This way, the pixel's affixation and location attribute can be determined through numerical values. After labeling all pixels in the entire image, a pixel-level labeled device map is obtained, where each pixel has a clear semantic label.

[0076] Finally, the different device regions need to be segmented based on the annotation information. Segmentation is based on the region number, extracting all pixels (including interior and boundary pixels) labeled as belonging to the same region and storing them separately. For example, all pixels in region 1 are saved as one independent image block, and pixels in region 2 are saved as another image block. If the original image has 10 valid device regions, it will be segmented into 10 image blocks. Each segmented image block contains only the pixels of the corresponding device, with other areas filled with transparent or white backgrounds. This achieves pixel-level precise segmentation. The collection of all segmented image blocks constitutes the distribution network device segmentation result, which preserves the complete shape of each device while also indicating the device category and boundary information.

[0077] The above describes a semantic segmentation method for distribution network lines according to an embodiment of the present invention. The following describes a semantic segmentation device for distribution network lines according to an embodiment of the present invention. Please refer to [link / reference]. Figure 2 One embodiment of the semantic segmentation device for distribution network lines in this invention includes: The shooting module 21 is used to shoot the power distribution line with a binocular camera to obtain a stereo image of the power distribution scene, and to extract depth information from the stereo image of the power distribution scene to obtain scene depth data. The culling module 22 is used to perform layered processing on the scene depth data to obtain a scene layer map, and to remove occlusions based on the scene layer map to obtain an unobstructed scene map. The segmentation module 23 is used to perform region growing processing on the unobstructed scene map to obtain a distribution network equipment area map, and to perform pixel-level annotation and segmentation on the distribution network equipment area map to obtain the distribution network equipment segmentation result.

[0078] In this embodiment, the specific implementation of each module in the above device embodiment is described in the above method embodiment, and will not be repeated here.

[0079] like Figure 3 As shown in the diagram, this embodiment of the invention provides a structural schematic block diagram of a computer device, including: At least one processor; At least one memory for storing at least one program; When at least one program is executed by at least one processor, the at least one processor implements the above-described fault diagnosis method for current transformers.

[0080] It is evident that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented in this device embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0081] Furthermore, this application also discloses a computer program product or computer program stored in a computer-readable storage medium. A processor of a computer device can read the computer program from the computer-readable storage medium and execute the computer program, causing the computer device to perform the aforementioned fault diagnosis method for current transformers. Similarly, the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0082] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A method for semantic segmentation of a network configuration line, characterized in that, The method comprises the following steps: The power distribution line is photographed by a binocular camera to obtain a power distribution scene stereo image, and depth information of the power distribution scene stereo image is extracted to obtain scene depth data; The scene depth data is processed in layers to obtain a scene layered image, and occlusion is removed based on the scene layered image to obtain an unoccluded scene image; The unoccluded scene image is processed by region growing to obtain a power distribution equipment region image, and pixel-level labeling and segmentation are performed on the power distribution equipment region image to obtain a power distribution equipment segmentation result.

2. The network configuration line semantic segmentation method according to claim 1, characterized in that, The power distribution line is photographed by a binocular camera to obtain a power distribution scene stereo image, and depth information of the power distribution scene stereo image is extracted to obtain scene depth data; The left lens and the right lens of the binocular camera are calibrated to obtain calibrated lenses, and the power distribution line is synchronously photographed by the calibrated lenses to obtain left and right photographed images; The left and right photographed images are time-stamped to obtain time-synchronized image pairs, and the time-synchronized image pairs are positionally calibrated to obtain calibrated image pairs; The corresponding pixel points in the left and right photographed images in the calibrated image pairs are calculated for parallax to obtain a parallax image, and image fusion is performed based on the parallax image to obtain the power distribution scene stereo image.

3. The network configuration line semantic segmentation method according to claim 2, characterized in that, The power distribution scene stereo image is processed for depth information extraction to obtain scene depth data, comprising: Pixel coordinate data of the power distribution scene stereo image is extracted to obtain each pixel coordinate data, and parallax reverse calculation is performed on each pixel coordinate data based on the binocular camera imaging principle to obtain each pixel parallax value; Based on each pixel parallax value, parallax value distribution analysis is performed on the power distribution scene stereo image to obtain a parallax distribution map, and the parallax distribution map is subjected to smoothing filter processing to obtain a smoothed parallax map; The depth value conversion is performed on each pixel parallax value in the smoothed parallax map based on the focal length and baseline distance of the binocular camera to obtain each pixel depth value, and a scene depth map is constructed based on each pixel depth value.

4. The network configuration line semantic segmentation method according to claim 1, characterized in that, The scene depth data is processed in layers to obtain a scene layered image, comprising: The scene depth data is subjected to depth value statistical sorting to obtain an ordered sequence of depth values, and the ordered sequence of depth values is subjected to interval division to obtain depth layered intervals, wherein the depth layered intervals include depth threshold features of different distances in the power distribution line scene, near, medium, and far; Based on the depth layered intervals, the pixel set of the power distribution scene stereo image is classified to obtain a depth classified pixel set, and the depth classified pixel set is marked by region to obtain a depth marked image; The depth marked image is subjected to region contour regularization to obtain a regularized depth image, and the regularized depth image is subjected to layered color rendering to obtain a scene layered image.

5. The network configuration line semantic segmentation method according to claim 1, characterized in that, Based on the scene layered image, occlusion is removed to obtain an unoccluded scene image, comprising: Based on the scene layered image, an occlusion candidate region is identified to obtain an occlusion candidate region, and an occlusion contour line is extracted from the occlusion candidate region; Based on the occlusion contour line, the occlusion candidate region is precisely marked by region to obtain an occlusion precise region image; The occluded pixel replacement is performed on the occlusion accurate region map to obtain an initial unoccluded map, and edge smoothing is performed on the initial unoccluded map to obtain an unoccluded scene map.

6. The network configuration line semantic segmentation method according to claim 1, characterized in that, The unoccluded scene map is subjected to region growing processing to obtain a network equipment region map, including: The unoccluded scene map is subjected to pixel gray threshold screening to obtain candidate network equipment pixels, and the candidate network equipment pixels are subjected to seed point marking to obtain network equipment seed point information; Based on the network equipment seed point information, the unoccluded scene map is subjected to region growing initialization to obtain an initial growing region; The initial growing region is subjected to growing region merging to obtain a merged region, and the merged region is subjected to small region elimination to obtain a screened region; The screened region is subjected to region contour regularization to obtain a regularized network equipment region map, and the regularized network equipment region map is subjected to boundary strengthening marking to obtain a network equipment region map.

7. The network configuration line semantic segmentation method according to claim 1, characterized in that, The network equipment region map is subjected to pixel-level labeling and segmentation to obtain a network equipment segmentation result, including: The network equipment region map is subjected to pixel color feature extraction to obtain a color feature vector, and the color feature vector is screened by setting a color threshold range to obtain device color candidate pixels, and the device color candidate pixels are subjected to color labeling to obtain color-labeled device information; Adjacent and color-labeled same pixels in the color-labeled device information are subjected to connected region analysis to obtain connected device regions, and region area calculation is performed on the connected device regions to obtain region area data; By setting an area threshold, the region area data is screened to obtain an effective device region, and boundary pixel data is extracted from the effective device region, and pixel value labeling is performed on the boundary pixel data to obtain a pixel-level labeled device map, and the pixel-level labeled device map is subjected to region segmentation to obtain a network equipment segmentation result.

8. An apparatus for network configuration line semantic segmentation, comprising: A device for performing the network line semantic segmentation method of any one of claims 1 to 7, comprising: A shooting module configured to shoot a network line by a binocular camera to obtain a network scene stereo image, and extract depth information from the network scene stereo image to obtain scene depth data; A removal module configured to perform layered processing on the scene depth data to obtain a scene layered map, and remove an occlusion object based on the scene layered map to obtain an unoccluded scene map; A segmentation module configured to perform region growing processing on the unoccluded scene map to obtain a network equipment region map, and perform pixel-level labeling and segmentation on the network equipment region map to obtain a network equipment segmentation result. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. A processor executes a computer program to implement the steps of the method of any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, A computer program is executed by a processor to implement the steps of the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Virtual terrain rendering method and device, equipment and medium

    CN112245926A

  • Building model generation system based on depth map analysis

    CN113808262A

  • Mass container batch rendering method and device

    CN115294251A

  • Monocular depth estimation method based on super-pixel processing occlusion

    CN115330874A

  • Role model rendering method and device, computer equipment and readable storage medium

    CN117745905A