Cable image segmentation method and device, terminal equipment and storage medium
By constructing an image pyramid and introducing an attention mechanism for cable image segmentation, the accuracy problem of cable image segmentation in complex backgrounds is solved, and accurate segmentation of cable inspection images in different complex backgrounds is achieved.
Patent Information
- Application Number
- CN202510742561.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-05
AI Technical Summary
Existing cable image segmentation technologies have low feature extraction accuracy under complex backgrounds, resulting in low accuracy of cable image segmentation.
By constructing an image pyramid, using convolution kernel parameters to extract cable features and perform upsampling alignment, combining channel attention and spatial attention enhancement, cable detail features are enhanced, and finally a cable segmentation mask map is generated.
The accuracy of cable image segmentation and the generalization ability of the model are improved, and the cable area can be accurately segmented in complex backgrounds, reducing mis-segmentation and missed segmentation.
Smart Images

Figure CN120599264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power image processing, and in particular to a cable image segmentation method, device, terminal equipment and storage medium. Background Art
[0002] As the scale of power systems continues to expand, cables, as the core carrier of power transmission, require accurate monitoring of their operating status to ensure grid security. In recent years, automated inspection technology based on image recognition has been gradually applied in the field of power inspection. By segmenting and processing cable images, it can achieve intelligent analysis of the cable itself and its surrounding environment, providing key data support for defect detection and fault location.
[0003] Existing cable image segmentation technologies are primarily based on traditional image processing algorithms and deep learning models. Traditional algorithms rely on artificially designed features, such as color and texture. In scenes with uneven lighting and complex backgrounds, feature extraction capabilities decline, easily leading to mis-segmentation or missed segmentation of cable areas. Deep learning models are also less adaptable to complex backgrounds. Cables are often laid in complex outdoor environments, and existing deep learning models lack adaptive processing mechanisms for background information, leading to confusion between background noise and cable features. Consequently, existing cable image segmentation technologies suffer from low feature extraction accuracy in complex backgrounds, resulting in low cable image segmentation accuracy. Summary of the Invention
[0004] The embodiments of the present invention provide a cable image segmentation method, apparatus, terminal device and storage medium, which can effectively solve the problem of low accuracy of feature extraction in complex backgrounds in the prior art, thereby leading to low accuracy of cable image segmentation.
[0005] An embodiment of the present invention provides a cable image segmentation method, comprising:
[0006] Obtain a cable inspection image of the area where the cables to be segmented are located;
[0007] Inputting the cable inspection image into a preset image segmentation model for recognition to obtain a cable segmentation image;
[0008] The training of the image segmentation model includes:
[0009] Obtain several cable image samples and initial model parameters;
[0010] Marking the cable area on the cable image sample to obtain a cable label sample;
[0011] Training the image segmentation model to be trained according to the cable label samples and current model parameters to obtain a cable prediction image;
[0012] Calculating a loss value of a preset loss function according to the cable label sample and the cable prediction image;
[0013] When the loss value converges, a trained image segmentation model is obtained;
[0014] When the loss value does not converge, the current model parameters are updated according to the loss value, and the updated current model parameters are used as the current model parameters for the next training; wherein, the current model parameters during the first training are the initial model parameters.
[0015] Furthermore, the image segmentation model to be trained is trained according to the cable label samples and the current model parameters to obtain a cable prediction image, including:
[0016] Inputting the cable label sample into the image segmentation model to be trained for slicing to obtain a slice image sample;
[0017] generating an image hierarchy according to the slice image samples and constructing an image pyramid;
[0018] Perform cable feature extraction and upsampling alignment according to the image pyramid and preset convolution kernel parameters to obtain a cable fusion feature map;
[0019] Performing cable detail feature enhancement according to the cable fusion feature map to obtain a cable segmentation mask map;
[0020] The cable prediction image is obtained by fusing the depth feature map corresponding to the cable label sample and the cable segmentation mask map.
[0021] Furthermore, generating an image hierarchy according to the slice image samples and constructing an image pyramid includes:
[0022] Determining the number of pyramid layers according to the resolution corresponding to the slice image sample;
[0023] Performing Gaussian smoothing and downsampling operations on the slice image samples until the number of pyramid levels is reached, thereby generating a plurality of layers of image samples;
[0024] According to the resolution corresponding to each layer of image samples, a pyramid structure is generated in descending order of resolution, and an image pyramid is constructed according to the image samples of each layer and the pyramid structure.
[0025] Furthermore, cable feature extraction and upsampling alignment are performed according to the image pyramid and preset convolution kernel parameters to obtain a cable fusion feature map, including:
[0026] Traversing the image pyramid layer by layer, and performing a convolution operation on each layer of the image pyramid according to preset convolution kernel parameters to extract appearance features and structural features of the cable, thereby obtaining a feature map of each layer of the image pyramid;
[0027] Perform upsampling alignment based on the feature maps of each layer of the image pyramid, and perform jump connections on the feature maps of each layer of the image pyramid to obtain a multi-scale fusion feature map;
[0028] Channel attention enhancement and spatial attention enhancement are performed according to the multi-scale fusion feature map to obtain a cable fusion feature map.
[0029] Furthermore, channel attention enhancement and spatial attention enhancement are performed according to the multi-scale fusion feature map to obtain a cable fusion feature map, including:
[0030] Performing global average pooling on the multi-scale fusion feature map, extracting global information of each channel of the multi-scale feature fusion map based on a channel attention mechanism, and obtaining a first feature vector after global average pooling;
[0031] Obtaining a channel weight according to the first feature vector mapping, and weighting the channel weight with the multi-scale fusion feature map to obtain a first feature map after channel attention enhancement;
[0032] Perform average pooling and maximum pooling on the first feature map, and concatenate the average pooling results and the maximum pooling results to obtain a spatial attention weight map;
[0033] The first feature map is spatially weighted according to the spatial attention weight map to obtain a cable fusion feature map.
[0034] Furthermore, cable detail features are enhanced according to the cable fusion feature map to obtain a cable segmentation mask map, including:
[0035] Performing a deformable convolution operation, a dilated convolution operation, and an edge detection on the cable fusion feature map to obtain a feature map after the deformable convolution operation, a feature map after the dilated convolution operation, and a feature map after the edge detection operation, respectively;
[0036] The feature map after the deformable convolution operation, the feature map after the hole convolution, and the feature map after the edge detection are spliced and normalized to obtain a normalized feature map;
[0037] Perform pixel-by-pixel classification according to the normalized feature map to obtain a cable segmentation mask map.
[0038] Furthermore, a cable prediction image is obtained by fusing the depth feature map corresponding to the cable label sample and the cable segmentation mask map, including:
[0039] Generating a corresponding depth feature map according to the cable label sample;
[0040] Performing preliminary fusion according to the depth feature map and the cable segmentation mask map to obtain a preliminary fusion feature map;
[0041] The preliminary fusion feature map is converted into a binary mask to obtain a binary image, and the binary image is converted into an RGB image to obtain a final cable prediction image.
[0042] As an improvement to the above solution, another embodiment of the present invention provides a cable image segmentation device, comprising:
[0043] The cable image acquisition module is used to obtain the cable inspection image of the area where the cable to be segmented is located;
[0044] The cable image segmentation module is used to input the cable inspection image into a preset image segmentation model for recognition to obtain a cable segmentation image;
[0045] Wherein, it also includes: a segmentation model training module; the segmentation model training module is used to train the image segmentation model, including:
[0046] A sample data acquisition unit, used for acquiring a number of cable image samples and initial model parameters;
[0047] a sample labeling unit, configured to label the cable region of the cable image sample to obtain a cable label sample;
[0048] a cable prediction image generating unit, configured to train the image segmentation model to be trained according to the cable label samples and current model parameters to obtain a cable prediction image;
[0049] A loss value calculation unit, configured to calculate a loss value of a preset loss function based on the cable label sample and the cable prediction image;
[0050] A first determining unit is configured to obtain a trained image segmentation model when the loss value converges;
[0051] The second determination unit is used to update the current model parameters according to the loss value when the loss value has not converged, and use the updated current model parameters as the current model parameters for the next training; wherein the current model parameters during the first training are the initial model parameters.
[0052] Another embodiment of the present invention provides a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a cable image segmentation method as described in the above embodiment is implemented.
[0053] Another embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the cable image segmentation method described in the above embodiment.
[0054] By implementing the present invention, at least the following beneficial effects are achieved:
[0055] The present invention provides a cable image segmentation method, apparatus, terminal device, and storage medium. The method can obtain a cable inspection image of an area where a cable to be segmented is located; input the cable inspection image into a preset image segmentation model for recognition to obtain a cable segmentation image; wherein the image segmentation model is trained, including: obtaining a plurality of cable image samples and initial model parameters; labeling the cable area of the cable image samples to obtain cable label samples; training the image segmentation model to be trained based on the cable label samples and current model parameters to obtain a cable prediction image; calculating a loss value of a preset loss function based on the cable label samples and the cable prediction image; when the loss value converges, obtaining a trained image segmentation model; when the loss value does not converge, updating the current model parameters based on the loss value, and using the updated current model parameters as the current model parameters for the next training; wherein the current model parameters during the first training are the initial model parameters. The image segmentation model learns cable features through a plurality of cable image samples, without the need for manually designed features. The model can adapt to various complex backgrounds and can accurately segment cable inspection images under different complex backgrounds, thereby improving the generalization ability of the model. Cable label samples provide a clear learning objective for the image segmentation model, enabling it to adjust model parameters so that the predicted cable image more closely resembles the actual cable area. This improves feature extraction, avoids mis-segmentation or missed segmentation of cable areas, and enhances cable image segmentation accuracy. During training, the model calculates the loss value of a preset loss function based on the cable label samples and the predicted cable image, and continuously adjusts the model parameters. The model gradually learns to distinguish between cables and background, suppressing the interference of background noise on cable features. If the loss value fails to converge, the model parameters are updated based on the loss value, allowing the image segmentation model to optimize the image segmentation task during the next training cycle, gradually improving feature extraction accuracy and cable image segmentation precision. When the loss value converges, the image segmentation model has learned sufficient feature information to adapt to cable feature extraction in complex backgrounds, improving cable feature extraction accuracy. For input cable inspection images, it can accurately obtain cable segmentation images, thereby improving cable image segmentation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 This is a flow chart of a cable image segmentation method provided by one embodiment of the present invention;
[0057] Figure 2 This is a schematic diagram of cable image segmentation visualization provided by an embodiment of the present invention;
[0058] Figure 3 This is a schematic diagram of a cable segmentation image provided by an embodiment of the present invention;
[0059] Figure 4 It is a structural diagram of a cable image segmentation device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0061] See also Figure 1 To solve the problem in the prior art of low accuracy in feature extraction under complex backgrounds, which leads to low accuracy in cable image segmentation, an embodiment of the present invention provides a flow chart of a cable image segmentation method, including:
[0062] S1. Obtain a cable inspection image of the area where the cable to be segmented is located;
[0063] Specifically, cable inspection images are sourced from drones and handheld devices, with image resolutions better than 0.1m, ensuring clear visibility of the cable and its connections. The captured images are optimized for shooting angles, selecting vertical or optimal pitch angles to enhance cable outline clarity and segmentation feasibility.
[0064] In a preferred embodiment of the present invention, the grayscale normalization processing of the cable inspection image is also included to unify the grayscale value range of the image and improve the image consistency. The histogram equalization algorithm is used to enhance the contrast of the cable inspection image, making the cable more clear and prominent against the background. The median filtering algorithm is used to remove Gaussian noise and improve the clarity and details of the cable inspection image. The detection and repair processing of cable inspection image artifacts, lens stains and shadow areas are increased to ensure the quality of image data. The low-light enhancement algorithm is used for the cable inspection images collected at night to improve the visibility of the cable under low-light conditions and obtain the final cable inspection image.
[0065] S2. Inputting the cable inspection image into a preset image segmentation model for recognition to obtain a cable segmentation image;
[0066] Specifically, the collected cable inspection images are input into a preset image segmentation model to identify the cable area, which can segment the cable from the complex background and obtain a cable segmentation image, which is of great significance for cable detection.
[0067] In a preferred embodiment of the present invention, after obtaining the cable segmentation image, a three-dimensional cable model is constructed, and the spacing between adjacent cable routings is calculated based on the cable three-dimensional model and cable parameters, and the rationality of the cable arrangement is evaluated, specifically including: extracting the center lines of adjacent cables through the cable three-dimensional model and calculating their minimum Euclidean distance; setting a reasonable spacing range. When the minimum Euclidean distance is less than the minimum spacing range, it may cause heat accumulation or short circuit risks. When the minimum Euclidean distance is greater than the maximum spacing range, the space utilization rate is low.
[0068] In a preferred embodiment of the present invention, based on the three-dimensional model of the cable and the geometric information of the cable, a cable space model that introduces mechanical analysis, environmental factors, etc. is used to analyze the suspension point load and predict the risk of high stress nodes. Specifically, the cable suspension point is extracted, and the tension at the point is calculated, and a stress threshold is set. If the tension at a certain suspension point is higher than the stress threshold, it is considered that the point has a high stress risk.
[0069] In a preferred embodiment of the present invention, a 3D spatial distribution map of the cable is generated based on the cable segmentation image, enabling high-precision visualization of the cable image. Based on cable structural attribute information, including numbering, spatial coordinates, and insulation status, a cable inspection analysis report is generated, covering equipment status, abnormality alerts, and maintenance recommendations. The cable inspection analysis is integrated into a GIS platform to display cable spatial information and enable data sharing. The GIS platform integrates 3D modeling and spatial measurement capabilities, supporting interactive manipulation of cable models. It supports comparative analysis of multi-time series inspection data, demonstrating cable status trends. It also features foreign object intrusion detection, enabling timely identification of external obstacles posing threats to cable safety. It supports cable insulation aging detection and early warning, assisting in health assessment. It supports drone inspection task scheduling, enabling autonomous flight inspections and image acquisition. It also supports 3D cable heat map display, visually reflecting risk level distribution. Data encryption, permission management, and backup and recovery mechanisms ensure data security. It supports multi-user collaboration and remote inspection management, improving task execution efficiency. Therefore, cable segmentation images are of great significance for subsequent cable inspections and power system monitoring.
[0070] The training of the image segmentation model includes:
[0071] Obtain several cable image samples and initial model parameters;
[0072] Marking the cable area on the cable image sample to obtain a cable label sample;
[0073] Training the image segmentation model to be trained according to the cable label samples and current model parameters to obtain a cable prediction image;
[0074] Calculating a loss value of a preset loss function according to the cable label sample and the cable prediction image;
[0075] When the loss value converges, a trained image segmentation model is obtained;
[0076] When the loss value does not converge, the current model parameters are updated according to the loss value, and the updated current model parameters are used as the current model parameters for the next training; wherein, the current model parameters during the first training are the initial model parameters.
[0077] Specifically, the cable image samples represent the samples used in the image segmentation model; the cable label samples represent the image samples after the cable image samples are labeled with cable areas; and the cable prediction images represent the predicted cable segmentation images output during the model training process.
[0078] In a preferred embodiment of the present invention, conventional algorithms for traditional cable image segmentation rely on manually designed features such as color and texture, which are prone to failure in scenes with uneven lighting and complex backgrounds. This embodiment collects a large number of cable image samples, allowing the model to automatically learn cable features during training. For example, in outdoor scenes with complex lighting, the model can learn cable features that are not affected by lighting, such as geometric shape and outline, rather than being limited to features such as color that are easily affected by lighting, thereby improving the accuracy of feature extraction in complex scenes. A large number of cable image samples from different environments (such as mountainous areas, urban areas, and different weather conditions) allows the model to learn richer and more representative features, enhancing the model's generalization ability. The model can adapt to various complex backgrounds, such as those with interference from vegetation, buildings, and utility poles, reducing the interference of background noise on cable feature extraction. Cable image samples are labeled with cable regions to obtain cable label samples, providing clear supervision information for model training. Based on the labeled samples, the model can clearly determine which areas are cables and which are background areas, and then adjust its parameters in a targeted manner to make the predicted cable areas closer to the actual situation, improving segmentation accuracy. Each training iteration updates the model parameters based on the previous one, allowing the model to gradually accumulate experience and continuously improve its feature extraction and segmentation capabilities. As training progresses, the model's understanding of cable characteristics deepens and its adaptability to complex backgrounds improves. In real-world applications, cable images may constantly change. Through iterative training, the model can continuously adapt to new data features and maintain excellent segmentation performance. Even when encountering new lighting conditions, background interference, and other conditions, the model can adjust and optimize through continuous training.
[0079] Specifically, a number of cable image samples are obtained, including cable image samples covering different environments and lighting conditions, to construct a diversified training data set. The types of cables in the cable image samples are expanded, including cable image samples with different materials, colors, and insulation layer characteristics, to improve the model's generalization ability for complex scenes. Data enhancement operations are performed on the cable image samples, including rotation, scaling, mirror flipping, brightness contrast adjustment, random occlusion, etc., to enhance the robustness of the model. Random Gaussian blur and halo simulation are added to improve the model's recognition ability under complex shooting conditions. At the same time, the cable image samples are divided into training set and validation set at 80% and 20%, and stratified sampling is used to ensure balanced sample distribution. Cable image samples collected under extreme conditions are optimized separately to ensure full coverage of data samples.
[0080] Preferably, training the image segmentation model to be trained according to the cable label samples and current model parameters to obtain the cable prediction image includes:
[0081] Inputting the cable label sample into the image segmentation model to be trained for slicing to obtain a slice image sample;
[0082] generating an image hierarchy according to the slice image samples and constructing an image pyramid;
[0083] Perform cable feature extraction and upsampling alignment according to the image pyramid and preset convolution kernel parameters to obtain a cable fusion feature map;
[0084] Performing cable detail feature enhancement according to the cable fusion feature map to obtain a cable segmentation mask map;
[0085] The cable prediction image is obtained by fusing the depth feature map corresponding to the cable label sample and the cable segmentation mask map.
[0086] Specifically, the cable label sample is divided into multiple small slice images. In real-world scenarios, cable images may be large, and direct processing consumes a large amount of computing resources and makes it difficult to capture local details. Slicing allows large images to be broken down into a series of smaller slice image samples, facilitating subsequent processing. An image pyramid consists of layers of images with varying resolutions. The slice image samples are downsampled to generate a series of images with progressively lower resolutions, which together form the image pyramid. Each layer contains information at different scales, ranging from high-resolution details to low-resolution global information. Convolution operations are performed on each layer of the image pyramid using preset convolution kernel parameters to extract cable features. Convolution automatically learns characteristic patterns in the image, such as edges and textures. Subsequently, an upsampling operation is performed to resize the feature maps extracted from different layers to the same resolution, and then fused to produce a fused cable feature map. The fused cable feature map is further processed to enhance cable details. This may involve using techniques such as edge detection and texture enhancement to highlight cable details such as edges and textures. Then, a cable segmentation mask is generated based on the enhanced feature map. The value of each pixel in the mask represents the probability that the pixel belongs to the cable area. The depth feature map reflects the depth information of different objects in the cable image. The depth feature map and the cable segmentation mask are fused, combining color, texture, and depth information to produce the final cable prediction image.
[0087] In a preferred embodiment of the present invention, a sliding window method is used to slice cable label samples, with window sizes set to 512x512 pixels and 1024x1024 pixels to ensure target integrity. An adaptive slicing algorithm is introduced to dynamically adjust the slice size to accommodate the complexity of varying cable scenarios. The overlap ratio of the sliced areas is optimized to avoid target information loss and reduce redundant computations, ultimately resulting in sliced image samples.
[0088] Preferably, generating an image hierarchy according to the slice image samples and constructing an image pyramid comprises:
[0089] Determining the number of pyramid layers according to the resolution corresponding to the slice image sample;
[0090] Performing Gaussian smoothing and downsampling operations on the slice image samples until the number of pyramid levels is reached, thereby generating a plurality of layers of image samples;
[0091] According to the resolution corresponding to each layer of image samples, a pyramid structure is generated in descending order of resolution, and an image pyramid is constructed according to the image samples of each layer and the pyramid structure.
[0092] In a preferred embodiment of the present invention, the number of layers of the image pyramid is determined by the original resolution of the sliced image. For example, a high-resolution image (such as 2048×2048) may require more layers (such as 5-6 layers), while a low-resolution image (such as 512×512) may only require 3-4 layers. The determination of the number of layers usually follows empirical rules (such as halving the resolution of each layer until the minimum effective size is reached, such as 16×16 or 32×32), ensuring that features of different scales can be captured while avoiding excessive computational complexity. Cables may appear different in size in the image due to different shooting distances and viewing angles, such as clear details of close-up cables and blurred or slender cables in the distance. The dynamic determination of the number of pyramid layers ensures that the model can handle cable features of different scales and avoid loss of details or insufficient global structural information at a single resolution.
[0093] A Gaussian filter is used to blur the image and reduce high-frequency noise, such as sensor noise and leaf texture interference, to prevent the noise from being amplified in subsequent downsampling. The image resolution is halved through downsampling to generate the next level of image. This process is repeated until the preset number of pyramid layers is reached, resulting in multi-layer image samples from high resolution (bottom layer) to low resolution (top layer). For example, the bottom layer is the original image, the upper layer is 1 / 2 resolution, the upper layer is 1 / 4 resolution, and so on. Gaussian smoothing weakens the effects of complex textures (such as grass and rocks) and uneven lighting in the background, making the contrast between the cable and the background more prominent, and alleviating the problem of missegmentation caused by noise interference with artificially designed features (such as color and texture) in traditional algorithms.
[0094] The generated multi-layer images are arranged from high to low resolution, from bottom to top, forming a pyramid structure. The bottom layer is the original slice image (highest resolution, richest detail), and the top layer is the minimum resolution image (most pronounced global structure). Spatial correspondence is preserved across each layer, enabling accurate feature alignment during subsequent upsampling. This image pyramid provides multi-scale input for subsequent convolutional neural networks (such as U-Net and DeepLab). The model extracts features at different levels. The bottom layer's high-resolution features are used to precisely locate cable edges, while the top layer's low-resolution features are used to distinguish the overall semantics of the cable from the background. This addresses the problem of traditional deep learning models confusing background noise with cable features due to their lack of multi-scale mechanisms. In field environments, cables may be partially obscured, such as by foliage, or have a similar color to the background. The pyramid structure enables the model to eliminate local noise interference at low-resolution layers and capture the overall continuity of the cable. It also inpaints occluded details at high-resolution layers, improving segmentation accuracy in complex scenes and reducing missed segmentations, such as partially occluded cable segments, and missegments, such as misidentifying a cable as a background with a similar texture.
[0095] Specifically, cable feature extraction and upsampling alignment are performed according to the image pyramid and preset convolution kernel parameters to obtain a cable fusion feature map, including:
[0096] Traversing the image pyramid layer by layer, and performing a convolution operation on each layer of the image pyramid according to preset convolution kernel parameters to extract appearance features and structural features of the cable, thereby obtaining a feature map of each layer of the image pyramid;
[0097] Perform upsampling alignment based on the feature maps of each layer of the image pyramid, and perform jump connections on the feature maps of each layer of the image pyramid to obtain a multi-scale fusion feature map;
[0098] Channel attention enhancement and spatial attention enhancement are performed according to the multi-scale fusion feature map to obtain a cable fusion feature map.
[0099] In a preferred embodiment of the present invention, the image pyramid contains multiple layers of images of different resolutions, such as the original high-resolution image at the bottom layer and the gradually downsampled low-resolution image at the top layer, and each layer of images is processed layer by layer. Convolution kernels of different sizes or types (such as 3×3, 5×5 convolution kernels, or learnable dynamic convolution kernels) are used to convolve each layer of images to extract the appearance features (such as color, texture, edges) and structural features (such as long strip structure, spatial continuity) of the cable, and obtain the feature map of each layer. After each layer of convolution, a corresponding feature map is generated. The low-level feature map (high resolution) retains the detailed information of the cable (such as local texture), and the high-level feature map (low resolution) captures the global structure of the cable (such as the overall direction, long strip outline). During the feature extraction process, each layer of the image pyramid is subjected to multiple convolution layers to extract local features. The convolution layer can use a standard 3×3 convolution kernel combined with the ReLU activation function to enhance the nonlinear expression capability. At the same time, batch normalization is introduced to normalize the features to improve the stability of the model. A feature pyramid network (FPN) is used to fuse features at different scales. High-level features are upsampled and skip-connected with low-level features, preserving detailed information in low-level features in high-level representations, resulting in a multi-scale fused feature map. Channel attention (SE module) and spatial attention (CBAM module) are introduced on the multi-scale fused feature map. Feature weights are calculated through global pooling to enhance the response of key areas while suppressing background noise interference, resulting in a cable fusion feature map. A dynamic convolution kernel resizing mechanism is introduced into the model to automatically switch convolution kernel parameters based on the size of the input image, improving feature adaptability.
[0100] Specifically, channel attention enhancement and spatial attention enhancement are performed according to the multi-scale fusion feature map to obtain a cable fusion feature map, including:
[0101] Performing global average pooling on the multi-scale fusion feature map, extracting global information of each channel of the multi-scale feature fusion map based on a channel attention mechanism, and obtaining a first feature vector after global average pooling;
[0102] Obtaining a channel weight according to the first feature vector mapping, and weighting the channel weight with the multi-scale fusion feature map to obtain a first feature map after channel attention enhancement;
[0103] Perform average pooling and maximum pooling on the first feature map, and concatenate the average pooling results and the maximum pooling results to obtain a spatial attention weight map;
[0104] The first feature map is spatially weighted according to the spatial attention weight map to obtain a cable fusion feature map.
[0105] In a preferred embodiment of the present invention, global average pooling is performed on the multi-scale fusion feature map, that is, the global average is calculated for the feature map of each channel, and the two-dimensional features of each channel are compressed into a one-dimensional scalar value to obtain a first feature vector (the dimension is the number of channels). The essence is to extract the global statistical information of each channel to reflect the overall response strength of the channel in the entire image range. The first feature vector is subjected to a nonlinear transformation (usually through a fully connected layer, an activation function, etc.) to generate a channel weight for each channel (the value range is usually [0,1]). The higher the weight, the more important the channel is to cable segmentation. The channel weight is multiplied by the original multi-scale fusion feature map to obtain the first feature map after channel attention enhancement, that is, the response to important channels is enhanced and the response to irrelevant channels (such as channels dominated by background noise) is suppressed. The first feature map after channel attention enhancement is subjected to average pooling and maximum pooling (both are pooling of spatial dimensions) to obtain two two-dimensional feature maps reflecting the importance of spatial positions. The two pooling results are spliced along the channel dimension through 3×3 convolution, and a spatial attention weight map is generated through convolution operation (the dimension is the same as the input feature map). Figure 1 The spatial attention weight map is multiplied by the first feature map after channel attention enhancement to obtain the final cable fusion feature map. The spatial weights are used to weight each pixel in the feature map, enhancing the characteristic response of the cable area. This allows the model to focus on learning details of cable edges and damaged areas while suppressing interference from irrelevant background areas.
[0106] The combined channel-spatial attention mechanism enables the model to simultaneously "focus on key feature channels" and "locate the spatial position of cables," achieving dual enhancement of cable features: at the channel level, it filters channels associated with background noise (such as vegetation and rock texture channels in complex environments); at the spatial level, it locks onto the actual distribution area of cables (such as irregular cable paths in complex field scenes). Compared to traditional deep learning models, this mechanism dynamically adjusts feature weights through an adaptive attention mechanism, rather than relying on a fixed network structure or manually designed features. This significantly improves the model's robustness to complex backgrounds and reduces both mis-segmentation and missed segmentation.
[0107] Specifically, the cable detail feature is enhanced according to the cable fusion feature map to obtain a cable segmentation mask map, including:
[0108] Performing a deformable convolution operation, a dilated convolution operation, and an edge detection on the cable fusion feature map to obtain a feature map after the deformable convolution operation, a feature map after the dilated convolution operation, and a feature map after the edge detection operation, respectively;
[0109] The feature map after the deformable convolution operation, the feature map after the hole convolution, and the feature map after the edge detection are spliced and normalized to obtain a normalized feature map;
[0110] Perform pixel-by-pixel classification according to the normalized feature map to obtain a cable segmentation mask map.
[0111] In a preferred embodiment of the present invention, the sampling points of traditional convolution are fixed grids (such as 3×3 squares), while deformable convolution enables the sampling points to adaptively adjust their positions (such as irregular shapes) by learning offsets, thereby capturing irregular morphological features such as bending and twisting of the cable. Input the cable fusion feature map; output the feature map after deformable convolution (retaining the irregular geometric features of the cable). Insert "holes" into the traditional convolution kernel to expand the receptive field without reducing the spatial resolution, thereby capturing the long-distance structural information of the cable (such as the overall outline of a long strip of cable) without losing details. Input the cable fusion feature map; output the feature map after hole convolution (retaining multi-scale contextual information to avoid loss of details caused by downsampling). Use an edge detection algorithm (such as the Canny operator or a convolution-based edge extraction layer) to directly extract the boundary contour information between the cable and the background and enhance the edge details of the cable. Input the cable fusion feature map; output the feature map after edge detection (highlighting the contour edge of the cable). The deformable convolution feature map, the hole convolution feature map, and the edge detection feature map are spliced together (for example, assuming the number of channels of the three are C1, C2, and C3 respectively, the number of channels after splicing is C1+C2+C3), and the features extracted by different operations are fused. The spliced feature map is normalized (such as Batch Normalization or LayerNormalization) to adjust the data distribution to a similar range to avoid training instability caused by excessive numerical differences. Each pixel of the normalized feature map is classified through a convolutional layer or a fully convolutional network (FCN), and the probability value of the pixel belonging to the cable (foreground) or background is output. Finally, a binary mask map (pixel value is 0 or 1) is generated through threshold segmentation. Input the normalized feature map (size is H×W×C); output the cable segmentation mask map (size is H×W×1, each pixel value indicates whether it is a cable). Convolution kernels of different sizes (such as 3×3 and 5×5) are used to extract local texture, edge and other detail information; dilated convolutions with different expansion rates are used to expand the receptive field and obtain richer structural features; multi-scale information is cascaded or weightedly fused to improve the resolution of cable edges and detail areas.
[0112] Specifically, the cable prediction image is obtained by fusing the depth feature map corresponding to the cable label sample and the cable segmentation mask map, including:
[0113] Generating a corresponding depth feature map according to the cable label sample;
[0114] Performing preliminary fusion according to the depth feature map and the cable segmentation mask map to obtain a preliminary fusion feature map;
[0115] The preliminary fusion feature map is converted into a binary mask to obtain a binary image, and the binary image is converted into an RGB image to obtain a final cable prediction image.
[0116] In a preferred embodiment of the present invention, the deep feature map contains the global semantic information of the label (such as the overall shape and category attributes of the cable) and local details (such as edge positions), providing semantic guidance for subsequent fusion and preventing the segmentation results from deviating from the semantic constraints of the real label. It complements the shallow spatial features of the segmentation mask map and improves the semantic consistency of the prediction results. The deep feature map provides the prior semantics of the label (such as the shape and structure that the cable should have), and the segmentation mask map provides the cable area details learned by the model. After fusion, the integrity of the features can be enhanced. Avoid errors caused by local noise or incomplete feature extraction in the segmentation mask map, and use the global semantic constraints of the label to converge the prediction results to the real cable area. The binarization operation converts the continuous probability distribution or feature value into a clear cable outline, removes fuzzy predictions, and makes the segmentation results more decision-making valuable. For example, when automatically identifying whether the cable is damaged, clear boundaries are crucial. RGB images conform to human visual habits and general image formats, making them convenient for superimposition and display with the original cable image, or input into downstream tasks (such as defect detection and length measurement), thereby improving the practicality of the model output.
[0117] Preferably, a preset loss function composed of a cross entropy loss function and a boundary perception loss function is used to jointly optimize the image segmentation model to improve the accuracy of the segmentation boundary. The specific process is as follows: Forward propagation calculation loss: input cable label samples, calculate the cable prediction image through the RSAM-Seg model (image segmentation model), and use the category segmentation loss function to evaluate the pixel-level classification error, and use the edge preservation loss function to evaluate the edge detail retention. Gradient calculation and back propagation: Calculate the gradient of the segmentation loss Lseg and the edge loss Ledge with respect to the network weight (the loss value of the loss function) respectively, combine the gradient information of the two, and update the parameter model. Adopt Adam optimizer or SGD optimizer, adjust the network weight according to the loss gradient, update the model parameters, so that the model can enhance the clarity and integrity of the cable edge while ensuring the overall segmentation accuracy. Repeat the above process until the loss value of the loss function converges, and finally obtain an image segmentation model with high recognition ability for the cable area.
[0118] Specifically, the image segmentation model is the RSAM-Seg model, which mainly includes multi-scale feature extraction (image pyramid construction), convolutional feature extraction and multi-scale fusion (obtaining a multi-scale fusion feature map), channel and spatial attention enhancement (obtaining a cable fusion feature map) and detail enhancement (obtaining a cable segmentation mask map). Figure 2 As shown in , the cable inspection image is collected, and then the cable inspection image is preprocessed and input into the image segmentation model. The image is sliced first, and then the cable image is segmented to obtain the cable segmentation image, as shown in Figure 3As shown, three-dimensional point cloud reconstruction is performed based on the cable segmentation image, and the cable visualization results are displayed.
[0119] In a preferred embodiment of the present invention, a class balance factor is introduced into the preset loss function to improve the model's learning ability for small sample classes (such as cable cracks and joint anomalies). In the training data set, the number of pixels in different classes is counted, and the class weights are calculated to reduce the impact of class imbalance. Based on the standard cross entropy loss function Lce, the class balance factor is introduced, so that the loss function becomes Among them, y c is the true category, is the predicted probability, w c By adjusting the class weights, the disadvantage of minority classes during training is reduced. Backpropagation is used to calculate the gradient of the loss with respect to the model weights, and optimization algorithms (such as SGD) are used to update the model parameters to improve the model's ability to recognize minority sample classes. A dynamic learning rate strategy is applied, with an initial learning rate of 0.001, which is periodically decreased during training. The learning rate is dynamically adjusted every 10 training iterations with an adjustment coefficient of 0.1 to improve the model convergence speed. A cosine annealing scheduling strategy is added to further smooth the learning rate changes and enhance the model training effect. Mini-batch gradient descent (Mini-Batch SGD) is used to optimize the model parameters with a batch size of 16. Batch normalization technology is introduced to standardize the feature distribution and improve training stability. After training is completed, the optimal model weights are saved as one of the model parameters of the trained image segmentation model.
[0120] By implementing this embodiment, the image segmentation model learns cable features through a number of cable image samples. There is no need for artificially designed features, and the model can adapt to various complex backgrounds. Cable inspection images under different complex backgrounds can be accurately segmented, thereby improving the generalization ability of the model. The cable label samples provide a clear learning goal for the image segmentation model, so that the image segmentation model adjusts the model parameters so that the cable prediction image is closer to the actual cable area, thereby improving the feature extraction capability, avoiding mis-segmentation or missed segmentation of the cable area, and improving the accuracy of cable image segmentation. By calculating the loss value of the preset loss function based on the cable label samples and the cable prediction image during the training process, and continuously adjusting the model parameters, the model gradually learns to distinguish between cables and backgrounds, and suppresses the interference of background noise on cable features. When the loss value has not converged, the model parameters are updated according to the loss value so that the image segmentation model can be optimized on the image segmentation task in the next training process, and the accuracy of feature extraction and the precision of cable image segmentation can be gradually improved; when the loss value converges, the image segmentation model has learned enough feature information and can adapt to the cable feature extraction under complex backgrounds, thereby improving the accuracy of cable feature extraction. For the input cable inspection image, the cable segmentation image can be accurately obtained, thereby improving the accuracy of cable image segmentation.
[0121] See also Figure 4 , is a schematic structural diagram of a cable image segmentation device provided by one embodiment of the present invention, comprising:
[0122] The cable image acquisition module is used to obtain the cable inspection image of the area where the cable to be segmented is located;
[0123] The cable image segmentation module is used to input the cable inspection image into a preset image segmentation model for recognition to obtain a cable segmentation image;
[0124] Wherein, it also includes: a segmentation model training module; the segmentation model training module is used to train the image segmentation model, including:
[0125] A sample data acquisition unit, used for acquiring a number of cable image samples and initial model parameters;
[0126] a sample labeling unit, configured to label the cable region of the cable image sample to obtain a cable label sample;
[0127] a cable prediction image generating unit, configured to train the image segmentation model to be trained according to the cable label samples and current model parameters to obtain a cable prediction image;
[0128] A loss value calculation unit, configured to calculate a loss value of a preset loss function based on the cable label sample and the cable prediction image;
[0129] A first determining unit is configured to obtain a trained image segmentation model when the loss value converges;
[0130] The second determination unit is used to update the current model parameters according to the loss value when the loss value has not converged, and use the updated current model parameters as the current model parameters for the next training; wherein the current model parameters during the first training are the initial model parameters.
[0131] The present invention provides a cable image segmentation device, which obtains a cable inspection image of the area where the cable to be segmented is located according to a cable image acquisition module; in the cable image segmentation module, the cable inspection image is input into a preset image segmentation model for recognition to obtain a cable segmentation image; wherein, it also includes: a segmentation model training module; in the segmentation model training module, the image segmentation model is trained, and a number of cable image samples and initial model parameters are obtained according to a sample data acquisition unit; according to a sample labeling unit, the cable image samples are labeled with cable areas to obtain cable label samples; in a cable prediction image generation unit, the image segmentation model to be trained is trained according to the cable label samples and current model parameters to obtain a cable prediction image; in a loss value calculation unit, the loss value of a preset loss function is calculated according to the cable label samples and the cable prediction image; through a first determination unit, when the loss value converges, a trained image segmentation model is obtained; through a second determination unit, when the loss value does not converge, the current model parameters are updated according to the loss value, and the updated current model parameters are used as the current model parameters for the next training; wherein, the current model parameters during the first training are the initial model parameters.
[0132] The image segmentation model learns cable features from a number of cable image samples, eliminating the need for manual feature design. The model can adapt to various complex backgrounds and accurately segment cable inspection images under varying conditions, improving model generalization. Cable label samples provide a clear learning objective for the image segmentation model, enabling it to adjust model parameters so that the predicted cable image more closely resembles the actual cable region. This improves feature extraction, avoids mis-segmentation or under-segmentation of the cable region, and enhances cable image segmentation accuracy. During training, a pre-set loss function is calculated based on the cable label samples and the predicted cable image, and model parameters are continuously adjusted. The model gradually learns to distinguish between cables and background, suppressing the interference of background noise on cable features. If the loss value fails to converge, the model parameters are updated based on the loss value, allowing the image segmentation model to optimize its performance on the image segmentation task during the next training cycle, gradually improving feature extraction accuracy and cable image segmentation precision. When the loss value converges, the image segmentation model has learned sufficient feature information to adapt to cable feature extraction in complex backgrounds, improving the accuracy of cable feature extraction. For input cable inspection images, it can accurately generate cable segmentation images, thereby improving cable image segmentation accuracy.
[0133] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.
[0134] Those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0135] Another embodiment of the present invention provides a terminal device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the cable image segmentation method described in the above embodiment. The terminal device can be a computing device such as a desktop computer, a laptop, a PDA, or a cloud server. The terminal device can include, but is not limited to, a processor and a memory.
[0136] The processor may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device using various interfaces and lines.
[0137] The memory can be used to store the computer program, and the processor implements various functions of the terminal device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device or other volatile solid-state storage device.
[0138] Another embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the cable image segmentation method described in the above embodiment.
[0139] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. The computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium.
[0140] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A cable image segmentation method, characterized in that: include: Obtain a cable inspection image of the area where the cables to be segmented are located; Inputting the cable inspection image into a preset image segmentation model for recognition to obtain a cable segmentation image; The training of the image segmentation model includes: Obtain several cable image samples and initial model parameters; Marking the cable area on the cable image sample to obtain a cable label sample; Training the image segmentation model to be trained according to the cable label samples and current model parameters to obtain a cable prediction image; Calculating a loss value of a preset loss function according to the cable label sample and the cable prediction image; When the loss value converges, a trained image segmentation model is obtained; When the loss value does not converge, the current model parameters are updated according to the loss value, and the updated current model parameters are used as the current model parameters for the next training; wherein, the current model parameters during the first training are the initial model parameters.
2. A cable image segmentation method according to claim 1, characterized in that: The image segmentation model to be trained is trained according to the cable label sample and the current model parameters to obtain a cable prediction image, including: Inputting the cable label sample into the image segmentation model to be trained for slicing to obtain a slice image sample; generating an image hierarchy according to the slice image samples and constructing an image pyramid; Perform cable feature extraction and upsampling alignment according to the image pyramid and preset convolution kernel parameters to obtain a cable fusion feature map; Performing cable detail feature enhancement according to the cable fusion feature map to obtain a cable segmentation mask map; The cable prediction image is obtained by fusing the depth feature map corresponding to the cable label sample and the cable segmentation mask map.
3. A cable image segmentation method according to claim 2, characterized in that: Generating an image hierarchy according to the slice image samples and constructing an image pyramid includes: Determining the number of pyramid layers according to the resolution corresponding to the slice image sample; Performing Gaussian smoothing and downsampling operations on the slice image samples until the number of pyramid levels is reached, thereby generating a plurality of layers of image samples; According to the resolution corresponding to each layer of image samples, a pyramid structure is generated in descending order of resolution, and an image pyramid is constructed according to the image samples of each layer and the pyramid structure.
4. A cable image segmentation method according to claim 2, characterized in that: Cable feature extraction and upsampling alignment are performed according to the image pyramid and preset convolution kernel parameters to obtain a cable fusion feature map, including: Traversing the image pyramid layer by layer, and performing a convolution operation on each layer of the image pyramid according to preset convolution kernel parameters to extract appearance features and structural features of the cable, thereby obtaining a feature map of each layer of the image pyramid; Perform upsampling alignment based on the feature maps of each layer of the image pyramid, and perform jump connections on the feature maps of each layer of the image pyramid to obtain a multi-scale fusion feature map; Channel attention enhancement and spatial attention enhancement are performed according to the multi-scale fusion feature map to obtain a cable fusion feature map.
5. A cable image segmentation method according to claim 4, characterized in that: Channel attention enhancement and spatial attention enhancement are performed according to the multi-scale fusion feature map to obtain a cable fusion feature map, including: Performing global average pooling on the multi-scale fusion feature map, extracting global information of each channel of the multi-scale feature fusion map based on a channel attention mechanism, and obtaining a first feature vector after global average pooling; Obtaining a channel weight according to the first feature vector mapping, and weighting the channel weight with the multi-scale fusion feature map to obtain a first feature map after channel attention enhancement; Perform average pooling and maximum pooling on the first feature map, and concatenate the average pooling results and the maximum pooling results to obtain a spatial attention weight map; The first feature map is spatially weighted according to the spatial attention weight map to obtain a cable fusion feature map.
6. A cable image segmentation method according to claim 2, characterized in that: The cable detail feature is enhanced according to the cable fusion feature map to obtain a cable segmentation mask map, including: Performing a deformable convolution operation, a dilated convolution operation, and an edge detection on the cable fusion feature map to obtain a feature map after the deformable convolution operation, a feature map after the dilated convolution operation, and a feature map after the edge detection operation, respectively; The feature map after the deformable convolution operation, the feature map after the hole convolution, and the feature map after the edge detection are spliced and normalized to obtain a normalized feature map; Perform pixel-by-pixel classification according to the normalized feature map to obtain a cable segmentation mask map.
7. A cable image segmentation method according to claim 2, characterized in that: The cable prediction image is obtained by fusing the depth feature map corresponding to the cable label sample and the cable segmentation mask map, including: Generating a corresponding depth feature map according to the cable label sample; Performing preliminary fusion according to the depth feature map and the cable segmentation mask map to obtain a preliminary fusion feature map; The preliminary fusion feature map is converted into a binary mask to obtain a binary image, and the binary image is converted into an RGB image to obtain a final cable prediction image.
8. A cable image segmentation device, characterized in that: include: The cable image acquisition module is used to obtain the cable inspection image of the area where the cable to be segmented is located; The cable image segmentation module is used to input the cable inspection image into a preset image segmentation model for recognition to obtain a cable segmentation image; Wherein, it also includes: a segmentation model training module; the segmentation model training module is used to train the image segmentation model, including: A sample data acquisition unit, used for acquiring a number of cable image samples and initial model parameters; a sample labeling unit, configured to label the cable region of the cable image sample to obtain a cable label sample; a cable prediction image generating unit, configured to train the image segmentation model to be trained according to the cable label samples and current model parameters to obtain a cable prediction image; A loss value calculation unit, configured to calculate a loss value of a preset loss function based on the cable label sample and the cable prediction image; A first determining unit is configured to obtain a trained image segmentation model when the loss value converges; The second determination unit is used to update the current model parameters according to the loss value when the loss value has not converged, and use the updated current model parameters as the current model parameters for the next training; wherein the current model parameters during the first training are the initial model parameters.
9. A terminal device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, a cable image segmentation method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the cable image segmentation method according to any one of claims 1 to 7.
Citation Information
Cited By
Prejudgment method and system for reset of charging gun cable, readable storage medium and equipment
CN121236064A