Method and device for removing components in transmission line inspection image
By combining feature decoupling and cross-reconstruction networks with optical flow methods, the problem of redundant shooting and complex scenes in transmission line inspection images is solved, achieving accurate image deduplication, reducing storage and analysis costs, and adapting to the complex environment of transmission line inspection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies for power transmission line inspection suffer from data redundancy due to repeated shooting, insufficient adaptability of deduplication methods, and severe interference in complex scenarios, leading to increased storage and analysis costs. Furthermore, existing methods cannot effectively handle changes under different angles, lighting conditions, and occlusions.
A feature decoupling and cross-reconstruction network is adopted. Through multi-loss collaborative optimization such as mask-guided consistency, cross-reconstruction consistency and cross-path subject consistency, combined with optical flow method and cosine similarity algorithm, feature extraction, decoupling and image reconstruction of adjacent images are realized. Global motion compensation and multi-source consistency fusion judgment are performed to identify and remove duplicate parts.
It achieves accurate, efficient, and lightweight image deduplication, reduces redundant data, lowers storage and analysis costs, adapts to the complex scene characteristics of inspection images, and improves the accuracy and robustness of component identification.
Smart Images

Figure CN121259615B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the present application relates to the technical field of computer vision and digital image processing, and particularly relates to a component deduplication method and device for transmission line section inspection images. BACKGROUND
[0002] With the wide application of unmanned aerial vehicles, high-altitude cameras and other equipment in transmission line inspection, a large amount of sequence image data is generated. These images usually take large-format pictures of line section components, and the following technical pain points exist:
[0003] (1) Redundant data caused by repeated shooting: In order to ensure coverage, there are a large number of transmission line components that are repeatedly shot between adjacent images, resulting in the same component being captured multiple times and generating a large amount of redundant data, which greatly increases the burden and cost of storage, transmission and subsequent intelligent analysis.
[0004] (2) Existing deduplication methods are not adaptable: Traditional image deduplication methods are sensitive to content and cannot adapt to changes in the same component under different angles, lighting and slight occlusion. General object tracking algorithms are too heavy and are not suitable for jumping, non-continuous video frames, i.e. sequences composed of a series of independently shot high-definition pictures.
[0005] (3) Serious interference in complex scenes: The background of the inspection image is complex, including buildings, vegetation, mountains, etc., and the camera moves and shakes, resulting in changes in the apparent features and positions of the same component in different images, further increasing the difficulty of accurate deduplication.
[0006] Therefore, there is an urgent need for a precise, efficient, lightweight and special-purpose deduplication technology that can adapt to the characteristics of inspection images to solve the above problems. SUMMARY
[0007] The embodiment of the present application provides a component deduplication method for transmission line section inspection images, which provides a precise, efficient, lightweight and special-purpose deduplication technology that can adapt to the characteristics of inspection images, reduces redundant data in transmission line section inspection images, and reduces costs. The method comprises:
[0008] Obtaining a transmission line section inspection image;
[0009] According to the name of the transmission line section inspection image, constructing an adjacent image pair;
[0010] input the adjacent image pair into a pre-trained recognition network, and output a recognition result; the recognition result includes a component region recognition result and a background region recognition result; the recognition network is obtained by training a convolutional neural network using a global image carrying a component region mask and a component region image corresponding to the global image; and the recognition network is used for feature extraction, feature decoupling and image reconstruction on the adjacent image pair;
[0011] The global motion compensation module is configured to perform global motion compensation on the adjacent image pair by using an optical flow method, and generate a global motion compensated adjacent image pair;
[0012] The similarity calculation module is configured to calculate a component similarity, a background similarity and a global motion compensated full image similarity according to the recognition result and the global motion compensated adjacent image pair by using a cosine similarity algorithm;
[0013] The similarity calculation module is configured to calculate a component similarity, a background similarity and a global motion compensated full image similarity according to the recognition result and the global motion compensated adjacent image pair by using a cosine similarity algorithm;
[0014] According to the repeated component, the adjacent image pair is de-duplicated.
[0015] Another aspect of the present application also provides a device for removing components in power line inspection images, which provides a precise, efficient, lightweight and adaptive de-duplication technology for inspection images, reduces redundant data in power line inspection images, and reduces costs. The device comprises:
[0016] An image acquisition module is configured to acquire a power line inspection image;
[0017] An adjacent image pair construction module is configured to construct an adjacent image pair according to the name of the power line inspection image;
[0018] A region recognition module is configured to input the adjacent image pair into a pre-trained recognition network, and output a recognition result; the recognition result includes a component region recognition result and a background region recognition result; the recognition network is obtained by training a convolutional neural network using a global image carrying a component region mask and a component region image corresponding to the global image; and the recognition network is used for feature extraction, feature decoupling and image reconstruction on the adjacent image pair;
[0019] A global motion compensation module is configured to perform global motion compensation on the adjacent image pair by using an optical flow method, and generate a global motion compensated adjacent image pair;
[0020] A similarity calculation module is configured to calculate a component similarity, a background similarity and a global motion compensated full image similarity according to the recognition result and the global motion compensated adjacent image pair by using a cosine similarity algorithm;
[0021] The repeated component determination module is configured to determine the repeated components in the adjacent image pair according to the component similarity, the background similarity, the global motion compensated full image similarity, and the preset threshold value.
[0022] The deduplication module is configured to deduplicate the adjacent image pair according to the repeated components.
[0023] The embodiment of the present application also provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the power transmission line section inspection image component deduplication method when executing the computer program.
[0024] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the power transmission line section inspection image component deduplication method.
[0025] The embodiment of the present application also provides a computer program product, which comprises a computer program, and the computer program is executable on a processor to implement the power transmission line section inspection image component deduplication method.
[0026] Compared with the prior art, the embodiment of the present application acquires the power transmission line section inspection image, constructs an adjacent image pair according to the name of the power transmission line section inspection image, inputs the adjacent image pair into a pre-trained recognition network, and outputs a recognition result; the recognition result comprises a component region recognition result and a background region recognition result; the recognition network is obtained by training a convolutional neural network using a global image carrying a component region mask and a component region image corresponding to the global image; the recognition network is used for feature extraction, feature decoupling, and image reconstruction of the adjacent image pair; a global motion compensation is performed on the adjacent image pair by using an optical flow method to generate a global motion compensated adjacent image pair; a cosine similarity algorithm is used to calculate a component similarity, a background similarity, and a global motion compensated full image similarity according to the recognition result and the global motion compensated adjacent image pair; a hierarchical and step-by-step judgment is performed on the components in the adjacent image pair according to the component similarity, the background similarity, the global motion compensated full image similarity, and a preset threshold value to determine the repeated components in the adjacent image pair; and the adjacent image pair is deduplicated according to the repeated components, so that a precise, efficient, lightweight, and adaptive section inspection image deduplication technology can be provided, the redundant data in the power transmission line section inspection image can be reduced, and the cost can be reduced. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to make the technical solutions in the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from these drawings without creative effort. In the drawings:
[0028] Figure 1 The flow chart of the component deduplication method in the transmission line section inspection image in the embodiment of the present application;
[0029] Figure 2 The flow chart of the specific example of the component deduplication method in the transmission line section inspection image in the embodiment of the present application;
[0030] Figure 3 The flow chart of the specific example of the component deduplication method in the transmission line section inspection image in the embodiment of the present application;
[0031] Figure 4 The structural diagram of the component deduplication device in the transmission line section inspection image in the embodiment of the present application;
[0032] Figure 5 The structural diagram of the computer device in the embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the technical solutions in the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can be obtained from these drawings without creative effort. In the drawings:
[0034] The transmission line inspection image is divided into two parts of section and tower, the section is the part of overhead line between two towers, and the embodiment of the present application designs the component deduplication method in the transmission line section inspection image according to the characteristics of the section image.
[0035] In recent years, with the wide application of unmanned aerial vehicles, high-altitude cameras and other equipment in power transmission line inspection, a large amount of sequence image data is generated. These images are usually taken in large format to shoot the equipment in the line section, and there are several technical pain points. First, repeated shooting leads to data redundancy. In order to ensure coverage, there is a lot of overlap between adjacent images, resulting in the same component being captured multiple times, generating a large amount of redundant data, greatly increasing the burden and cost of storage, transmission and subsequent intelligent analysis. Second, the existing deduplication method is not adaptable. Traditional image deduplication methods are sensitive to content and cannot adapt to changes in the same component under different angles, lighting and slight occlusion. General target tracking algorithms are too heavy and are not suitable for jumping, non-continuous video frames, i.e. sequences composed of a series of independently shot high-definition pictures. Finally, the interference in complex scenes is serious. The background of the inspection image is complex, including buildings, vegetation, mountains, etc., and the camera may move or shake, causing the apparent features and positions of the same component in different images to change, further increasing the difficulty of accurate deduplication. Therefore, there is an urgent need for a specialized deduplication technology that is accurate, efficient, lightweight and can adapt to the characteristics of inspection images to solve the above problems.
[0036] The application belongs to the technical field of computer vision and digital image processing, and relates to intelligent processing of large-format time sequence images in a power system transmission line section inspection scene, in particular to a component repeated detection and data deduplication method and device, an electronic equipment and a storage medium. The application takes feature decoupling and cross-reconstruction network as the core: in the training stage, based on the data set of the applicant self-built and manually labeled subject mask, mask guided consistency, cross-reconstruction consistency, cross-path subject consistency and semantic supervision are used for multi-loss collaborative optimization to obtain explicit decoupling representation of the subject and the background; in the inference stage, without external mask, stable subject and context feature representation can be automatically formed. At the same time, for time sequence jitter and view angle change, global motion estimation and compensation based on feature points are introduced to realize robust alignment of adjacent images; on this basis, a multi-source consistency fusion and deduplication determination process is constructed, which comprehensively considers subject consistency, context cross matching and whole image similarity to complete reliable determination of the same component and effective elimination of repeated samples.
[0037] The purpose of the present application is to provide a large-format file-oriented feature decoupling and cross-reconstruction network for image component de-duplication of transmission line files in the inspection image, by introducing dynamic scene perception and compensation technology, and multi-source consistency fusion judgment, the precise identification and de-duplication of the target component are realized. The method first extracts the main features and background features of the component through the feature decoupling and cross-reconstruction network, and decouples the target area, so as to avoid the interference of background information on component identification; then, the global motion compensation is carried out on the adjacent images by using the optical flow method, the image difference caused by camera motion and object displacement is eliminated, and the component area is aligned; finally, the multi-source consistency fusion judgment strategy is adopted, the main features, context background features and global similarity are calculated, and the target is accurately judged whether it is the same component, which significantly improves the accuracy and robustness of the component de-duplication. The present application can be widely applied to image de-duplication, target detection and other fields, and is especially suitable for complex inspection environment, solving the problem of repeated component judgment caused by changes such as light, angle and shielding.
[0038] Figure 1 The flow chart of the component de-duplication method in the transmission line file inspection image in the embodiment of the present application is shown in Figure 1 The method comprises the following steps:
[0039] Step 101, acquiring the transmission line file inspection image;
[0040] Step 102, constructing adjacent image pairs according to the names of the transmission line file inspection images;
[0041] Step 103, inputting the adjacent image pairs into the pre-trained identification network to output the identification results; the identification results include component region identification results and background region identification results; the identification network is obtained by training the convolutional neural network using the global image carrying the component region mask and the component region image corresponding to the global image; the identification network is used for feature extraction, feature decoupling and image reconstruction of the adjacent image pairs;
[0042] Step 104, using the optical flow method to perform global motion compensation on the adjacent image pairs to generate the global motion compensated adjacent image pairs;
[0043] Step 105, using the cosine similarity algorithm to calculate the component similarity, the background similarity and the global motion compensated global image similarity according to the identification results and the global motion compensated adjacent image pairs;
[0044] Step 106, according to the component similarity, the background similarity, the global motion compensated global image similarity and the preset threshold, the components in the adjacent image pairs are judged layer by layer, and the repeated components in the adjacent image pairs are determined;
[0045] Step 107: Based on the duplicate components, remove duplicates from adjacent image pairs.
[0046] Depend on Figure 1 As shown in the flowchart, compared with the image deduplication technology in the prior art, the embodiments of the present invention obtain inspection images in the transmission line arches; construct adjacent image pairs according to the names of the inspection images in the transmission line arches; input the adjacent image pairs into a pre-trained recognition network, and output the recognition results; the recognition results include component region recognition results and background region recognition results; the recognition network is obtained by training a convolutional neural network using a global image with component region mask annotations and the component region images corresponding to the global image; the recognition network is used to perform feature extraction, feature decoupling, and image reconstruction on adjacent image pairs; and optical flow is used to perform global image reconstruction on adjacent image pairs. Motion compensation generates adjacent image pairs after global motion compensation. A cosine similarity algorithm is used to calculate component similarity, background similarity, and overall image similarity after global motion compensation based on the recognition results and the adjacent image pairs. Based on component similarity, background similarity, overall image similarity after global motion compensation, and a preset threshold, components in adjacent image pairs are judged hierarchically and level by level, identifying duplicate components in adjacent image pairs. Based on the duplicate components, duplicate components are removed from adjacent image pairs. This provides a precise, efficient, lightweight deduplication technique that adapts to the characteristics of inspection images, reducing redundant data in inspection images of transmission line spans and lowering costs.
[0047] Figure 2 This is a flowchart illustrating a specific example of a method for deduplicating components in inspection images of transmission line files according to an embodiment of the present invention, such as... Figure 2 As shown, the component deduplication method in the inspection images of transmission line sections in this embodiment of the invention may include: Input: data captured by an inspection line; intelligent matching of temporally adjacent frames; feature decoupling and cross-reconstruction network: decoupling of background and subject features, multi-level feature reconstruction; dynamic scene perception and compensation: calculating optical flow between images, estimating the global motion matrix using RANSAC (random sample consensus algorithm), and motion compensation; multi-source consistency fusion and deduplication determination: subject feature matching, context feature matching, and global feature matching after motion compensation; Output: encoding mapping table for the same component.
[0048] Figure 3 This is a flowchart illustrating a specific example of a method for deduplicating components in inspection images of transmission line files according to an embodiment of the present invention, such as... Figure 3As shown, in one embodiment, in step 102, constructing the adjacent image pair according to the name of the inspection image in the power line file can include: step 201, parsing the name of the inspection image in the power line file to extract the serial number in the name; step 202, sorting the inspection images in the power line file according to the serial number; and step 203, constructing the adjacent image pair according to the sorted inspection images in the power line file.
[0049] In this embodiment, after obtaining the inspection images in the power line file, the time-sequential adjacent frame intelligent pairing is performed. By parsing the file name of the image to be processed, the digital serial number in the picture shooting number is extracted, and the adjacent image pair is automatically constructed according to the continuity of the serial number, so as to ensure that the comparison only occurs between the time-sequential adjacent images, greatly improving the image processing efficiency.
[0050] In one embodiment, the recognition network can include a feature extraction module, a feature decoupling module, a feature reconstruction module, and a loss function; the feature extraction module is a backbone network constructed based on a convolutional neural network, and is used to extract features from the adjacent image pair; the feature decoupling module is used to decouple the features into component features and background features; the feature reconstruction module is used to reconstruct the image through a decoder according to the features, the component features, and the background features, to generate a reconstructed image; and the loss function is a result of weighted summation of a plurality of sub-loss functions; and the plurality of sub-loss functions are used to optimize the loss function of each part in the recognition network.
[0051] In this embodiment, the recognition network is a feature decoupling and cross-reconstruction network for large-format file images, which aims to decouple the subject features and background features in the input image and restore the complete image through the reconstruction module. The feature decoupling and cross-reconstruction network for large-format file images is used to automatically predict the subject component region and the context background region from the input shooting image. The core innovation lies in that, first, the explicit decoupling representation of the subject features and the background features is simultaneously learned in the same framework, and through the cross-reconstruction mechanism of the shared background and the double-path subject, the subject representation is constrained to remain stable and consistent under different viewing angles, scales, and brightness conditions; second, a multi-loss collaborative optimization system composed of mask consistency, cross-reconstruction consistency, cross-path consistency, and semantic supervision is constructed, so that the decoupled subject features can not only accurately restore the global content, but also retain the component semantics and discriminability, thereby providing a robust and interpretable feature basis for subsequent similarity fusion and deduplication judgment.
[0052] The recognition network includes multiple modules: an input layer, a feature extraction module, a feature decoupling module, a feature reconstruction module, and a loss function. During training, the input of the recognition network includes an original image with component main body region mask annotation. The mask is used to mark the region where the component is located in the image. Through this mask information, the recognition network can concentrate on processing the features of the target component and its surrounding background region; a trained convolutional neural network is used as the backbone network for feature extraction, and the features extracted from the original image are decoupled through the main body feature encoder and the background feature encoder, respectively, to obtain the background features and the main body features. Through backward propagation learning, the decoupled features are aligned with the actual image to restore the details of the image; the network introduces multiple loss functions to ensure that the network learns effective component feature and background feature decoupling and strengthens the reconstruction quality of the image.
[0053] In the training phase, based on the self-built and self-labeled data set, the input is composed of a global original image with a main body region mask annotation and a corresponding main body region image. The main body mask is used as a strong supervision signal for end-to-end optimization of multiple losses, which is used to accurately constrain the feature decoupling and cross-reconstruction of the main body or the background, so that the model learns a restorable and distinguishable main body representation under the condition of sharing the background. In the inference phase, without providing any external mask, the input is only the original image to be processed. The network relies on the decoupling ability formed in the training to automatically form the main body component representation and the context background representation, and outputs stable multi-source features for subsequent fusion and de-duplication judgment. This process can maintain robustness and generalization to different viewing angles, scales, and lighting conditions without relying on external annotations, providing more accurate and interpretable feature basis for component de-duplication tasks.
[0054] The input data of the input layer includes a global image I global , which contains the input of the entire image; a main body region mask M mask , which marks the position of the target component in the image, usually a binary mask, where the target component region is 1 and the remaining region is 0; and a main body component image I component , which is a main body component region image cropped from the global image. The input image and its mask are processed by a feature extraction network, and the mask is used to constrain the feature learning of the network in a specific region.
[0055] The feature extraction module is a backbone network based on a convolutional neural network (CNN, Convolutional Neural Networks), which is responsible for extracting high-dimensional features from the global image I global . The backbone network uses a pre-trained ResNet-50 model, and removes the last fully connected layer to obtain high-dimensional convolutional features. The global image I global and the main body region mask M mask output the extracted feature map F globalwith size H x W x C, where H and W are the height and width of the image, respectively, and C is the number of channels of the convolution output.
[0056] In one embodiment, the feature decoupling module includes a component feature encoder and a background feature encoder; the component feature encoder includes a max-pooling layer, a convolution layer, a full connection layer, and a feature extractor; the component feature encoder is configured to decouple the feature through the max-pooling layer, the convolution layer, and the full connection layer to obtain a first component feature, and extract a second component feature from the feature through the feature extractor; the background feature encoder includes a max-pooling layer, a convolution layer, and a full connection layer; the background feature encoder is configured to decouple the feature to obtain a background feature.
[0057] The feature decoupling module decouples the feature of the global image into two parts: component feature and background feature. The module consists of two independent encoders: component feature encoder E component , which receives the subject region mask M mask of the global image, combines it with the feature map F global , and extracts the first component feature F component1 of the subject component; background feature encoder E background , which separates the background region in the global image through M mask , and extracts the background feature F background ; the component feature encoder uses four convolution layers to extract the feature of the target component. The first layer has a convolution kernel size of 3x3, a convolution kernel number of 64, and a stride of 1, followed by a max-pooling layer (2x2, stride 2). The next three convolution layers are: 3x3 convolution kernel, convolution kernel number is 128, 256, 512 in turn, stride is 1, and each layer is followed by a max-pooling layer. Finally, a 1024-dimensional feature vector is output through a full connection layer, representing the global feature of the target component. The structure of the background feature encoder is similar to that of the component feature encoder, but the focus is on the extraction of the background region. It also contains four convolution layers, with a convolution kernel size of 3x3, a convolution kernel number starting from 64, and then 128, 256, 512, a stride of 1, and the same pooling layer settings as the component feature encoder. After these convolution layers and pooling layers, a 1024-dimensional feature vector is output through a full connection layer, representing the feature of the background region around the target component. It should be noted that the subject component image I component directly obtains the second component feature F component2 through the feature extractor, which will participate in feature reconstruction and loss function learning later.
[0058] In one embodiment, the feature reconstruction module is used to reconstruct a first reconstructed image by means of a decoder based on the features of a first component and background features; the feature reconstruction module is also used to reconstruct a second reconstructed image by means of a decoder based on the features of a second component and background features; the decoder includes a deconvolution layer, a batch normalization function and an activation function, and the deconvolution layer is used to restore the spatial resolution of the image.
[0059] The feature reconstruction module performs a feature reconstruction on the first decoupled component, F. component1 and background features F background and main component diagram I component The second component feature F is obtained directly through the feature extractor. component2 Reconstruction is performed using a separate decoder module to generate a reconstructed image. Where F... component1 and F background Reconstructed as the first reconstructed image I construct1 F component2 and F background Reconstructed into a second reconstructed image I construct2 The decoder structure includes deconvolutional layers to progressively restore the spatial resolution of the image; batch normalization functions and ReLU activation functions to increase the network's stability and non-linear expressive power. The reconstructed image is subsequently compared with the original image to calculate the loss function.
[0060] The recognition network comprises six sub-loss functions, each optimizing a different part of the network to ensure effective decoupling of part features from background features and high-quality image reconstruction. Ultimately, these six sub-loss functions are merged into a single loss function to optimize the entire network training process, as follows:
[0061] (1) Subject consistency loss L caused by masking mask This constraint aligns the subject features extracted by the subject encoder with the semantics of the masked region, avoiding the aliasing of subject and background information and ensuring that the network learns accurate subject features within the subject region (the region marked by the mask). This sub-loss function constrains the network's feature learning within the part region by calculating the difference between the subject features in the masked region and the subject region features in the original image. The specific formula is as follows:
[0062] ;
[0063] Where M mask [i] is the mask region; F global [i] represents the features of the global image; F component1 [i] represents the characteristics of the main component after decoupling.
[0064] (2) Self-reconstruction consistency loss L construct1, the constraint loss of reconstructing the first image (one of the two adjacent image pairs) and the global image, to verify F component1 and F background The reconstructability of the global image content promotes the complementarity of the decoupled subject feature and the background feature. It ensures that the decoupled features of the network can be reconstructed into an image consistent with the original image. During the training process, the network generates a reconstructed image I construct1 , which is compared with the original image to optimize the recovery effect of the image. The specific formula is as follows:
[0065] ;
[0066] where I construct1 [i] is the reconstructed image of the first image, and I global [i] is the global image.
[0067] (3) Cross-reconstruction consistency loss L construct2 , the constraint loss of reconstructing the second image (the other image in the adjacent image pair) and the global image, to ensure that the background decoupled features of the network can be reconstructed into an image consistent with the original image with the subject features obtained by the feature extractor. During the training process, the network generates a reconstructed image I construct2 , which is compared with the original image to optimize the recovery effect of the image. The specific formula is as follows:
[0068] ;
[0069] where I construct2 [i] is the reconstructed image of the second image.
[0070] (4) Cross-path subject consistency loss L component , which constrains the subject features of the two paths to remain subject invariance under different perspectives or scales by calculating the difference between the subject feature 1 and the subject feature 2, avoiding excessive decoupling. The specific formula is as follows:
[0071] ;
[0072] where F conponent1 [i] is the subject feature 1, and F conponent2 [i] is the subject feature 2.
[0073] (5) Subject semantic supervision loss L component1_class , the subject feature 1 classification loss, used to optimize the classification accuracy of the subject feature 1, and the network learns the category information of the components through the cross-entropy loss. The specific formula is as follows:
[0074] ;
[0075] where y[i] is the true label, is the predicted class probability.
[0076] (6) Main semantic supervision loss L component2_class , the main feature 2 classification loss, is used to optimize the classification accuracy of the main feature 2, and the network learns the class information of the component through the cross-entropy loss. The specific formula is as follows:
[0077] ;
[0078] where y[i] is the true label, is the predicted class probability.
[0079] Finally, the loss function L total is the result of weighting and merging the above six sub-loss functions. By assigning a weight coefficient a i to each sub-loss function, the contribution of each part is adjusted according to the requirements of the task. The expression of the total loss function is as follows:
[0080] .
[0081] The training process is trained end-to-end through the above loss function by inputting the image with the main body region mask and the extracted component region map. During the training process, these loss functions are optimized to gradually improve the accuracy of feature decoupling and the quality of the reconstructed image. In the inference stage, the original image is input, and the network can automatically predict the main body region feature and the background region feature in the image, and reconstruct the target component and the background image through the decoder module, and finally output the main body component region and the context background region in the original image.
[0082] In one embodiment, in step 104, the global motion compensation of the adjacent image pair is generated by using an optical flow method, which can include: using an ORB (Oriented FAST and Rotated BRIEF) feature point detector based on the ORB algorithm to extract feature points in the adjacent image pair; using a brute force matcher to match the feature points by calculating the Hamming distance between the descriptors of the feature points to determine the matching feature points; using a random sample consensus algorithm to calculate the geometric relationship between the matching feature points to obtain an affine transformation matrix; and compensating the coordinates of one image in the adjacent image pair according to the affine transformation matrix, transforming the coordinates of the image into the coordinate system of the other image in the adjacent image pair, and then aligning the coordinates; and determining the adjacent image pair after coordinate alignment as the globally motion compensated adjacent image pair.
[0083] In this embodiment, dynamic scene perception and compensation, by estimating the motion between adjacent images accurately through a feature point based global motion estimation method, so as to make motion compensation in the image contrast process, eliminate the image difference caused by camera motion or background change. First, enough local feature points are extracted from the input adjacent images for subsequent matching. The ORB algorithm is used for feature point detection and description; the BFMatcher (brute force matcher) is used to calculate the feature point matching in the image according to the ORB descriptor. This step matches by calculating the Hamming distance between the feature points; the transformation matrix between the matching points is estimated by RANSAC (random sample consensus algorithm), and the affine transformation or similarity transformation matrix is obtained. The RANSAC algorithm can effectively eliminate the mismatched points and calculate the transformation matrix that can truly reflect the global motion between the images; through the obtained global transformation matrix, the coordinates of the target in the image are compensated, the offset caused by the camera motion or background change is compensated by mapping the coordinates of the target from the second image to the coordinate system of the first image, and the alignment of the target in the two images is ensured. In this way, the subsequent similarity calculation can be carried out in the same coordinate system, so as to improve the accuracy of image matching.
[0084] Dynamic scene perception and compensation are used to accurately estimate the motion between adjacent images, so as to make motion compensation in the image contrast process, eliminate the image difference caused by camera motion or background change. Feature point detection and matching using optical flow method, this step uses optical flow method for global motion estimation, by detecting and matching feature points in adjacent images, to calculate the motion between images. The matching of feature points can help to identify the common part in the image, so as to provide data support for subsequent motion compensation.
[0085] Feature point detection uses ORB feature point detector to extract feature points in the image. ORB algorithm can maintain robustness under rotation, scale transformation and illumination change, and has high calculation efficiency, and the specific formula is as follows:
[0086] kp1, des1 = ORB (I1);
[0087] kp2, des2 = ORB (I2);
[0088] Where I1 and I2 are the input adjacent image pair, kp1 and kp2 are the feature point coordinates of the first image and the second image, and des1 and des2 are the descriptors corresponding to the feature point coordinates of the first image and the second image.
[0089] Feature point matching uses BFMatcher (brute force matcher) to match feature points by calculating the Hamming distance between the descriptors. Through the cross-validation strategy, it is ensured that the matching points are consistent in both directions, and the specific formula is as follows:
[0090] matches = BFMMatcher(desl, des2);
[0091] where desl and des2 are the descriptors corresponding to the feature point coordinates of the first and second images, and matches is the set of matched points.
[0092] Global motion estimation and transformation matrix calculation, the global motion between images is calculated by RANSAC algorithm, that is, the affine transformation matrix M, which can describe the translation, rotation and scaling of the second image relative to the first image. RANSAC estimates the transformation matrix: by calculating the geometric relationship between the matched feature points, the affine transformation matrix M is estimated using the RANSAC algorithm:
[0093] ;
[0094] where a, b, c, d are the rotation and scale parameters of the affine transformation, tx and ty are the translation parameters. The RANSAC algorithm calculates the optimal transformation matrix by randomly sampling and removing outliers, as follows:
[0095] M = RANSAC(matches).
[0096] Coordinate transformation and motion compensation, the target coordinates in the second image are transformed into the coordinate system of the first image using the affine transformation matrix M, motion compensation is performed to ensure that the target regions of the images are aligned. Given a target point p2=(x2, y2) in the second image, use the affine transformation matrix M to transform it to get the new coordinates of the target in the first image coordinate system p1=(x1, y1), the transformation process is as follows:
[0097] ;
[0098] Apply the affine transformation matrix M to the coordinates of each target point in the second image to compensate for the coordinates consistent with the first image, and apply it to the subsequent target comparison and similarity calculation.
[0099] In one embodiment, the component region identification result can include component feature vectors of the adjacent image pair; the background region identification result can include context background feature vectors of each component of the adjacent image pair; the context background feature vectors are feature vectors of adjacent regions located above, below, left and right of the component.
[0100] In the embodiment, in step 105, the cosine similarity algorithm is adopted to calculate the part similarity, the background similarity and the global motion compensated full image similarity according to the recognition result and the adjacent image pair after global motion compensation, which can include: adopting the cosine similarity algorithm to calculate the cosine similarity between the part feature vectors of the two images in the adjacent image pair, and determining the cosine similarity between the part feature vectors of the two images in the adjacent image pair as the part similarity; calculating the cosine similarity between the context background feature vectors of one image and the context background feature vectors of the other image in the adjacent image pair, and determining the maximum value of the calculation result as the background similarity; calculating the cosine similarity between the two images in the adjacent image pair after global motion compensation, and determining the cosine similarity between the two images in the adjacent image pair after global motion compensation as the global motion compensated full image similarity.
[0101] First, the part similarity sim component of the first image and the second image is calculated, which measures the similarity between the part main feature vectors extracted by the deep feature decoupling and reconstruction network. The formula is as follows:
[0102] ;
[0103] Where F component1 is the first part feature; F component2 is the second part feature.
[0104] Secondly, the background similarity (context background feature similarity) is determined:
[0105] The background feature is directly obtained by the deep feature decoupling network prediction, and the background information extracted in the four extended directions (up, down, left and right) of the part position. Due to the forward and upward movement of the camera in the two actual adjacent repeated shooting pictures, the context information of the part in the front and back two pictures changes relative to the position of the part. The same context information with obvious features around the part in the previous shooting picture may appear in the upper, lower, left and right positions of the part in the next shooting picture. Therefore, the context background feature matching is no longer limited to point-to-point matching, but the background features of the same part in the two pictures are matched two by two in the upper, lower, left and right four regions, the maximum value of the matching result is taken, and a high determination threshold is set. The specific formula is as follows:
[0106] ;
[0107] Where the background feature of the first image is F i1 , the background feature of the second image is F j2 , i, j ∈ {up, down, left, right}, represents the combination of direction pair, for example, sim up,leftsim is the similarity between the feature above the part in the first image and the feature left of the part in the second image. The final similarity of each direction is selected from all the calculated similarities, which is the maximum similarity value of each direction matching with other directions:
[0108] sim contex_max_i =max(sim i,j )。
[0109] For each direction, the maximum similarity of the corresponding direction with other directions is selected, and finally the maximum similarity value of the four directions is obtained, which is the background similarity:
[0110] sim contex_max =max(sim contex_max_i )。
[0111] The global motion compensated full image similarity formula is as follows:
[0112] ;
[0113] The image pair comparison after global motion compensation is to further verify whether the targets in the two images belong to the same physical part by calculating the similarity between the first image and the second image after global motion compensation. The cosine similarity is used to measure the similarity between the features of the images after motion compensation. Wherein, I compensated_1 and I compensated_2 are the first image and the second image after global motion compensation. The similarity value obtained by calculation is compared with the preset threshold value, if the similarity is greater than the threshold value, it is considered that the targets of the two images belong to the same physical part. Through the global motion estimation of the optical flow method, the accuracy of the image comparison can be improved, especially in the case of camera movement.
[0114] In one embodiment, in step 106, according to the part similarity, the background similarity, the global motion compensated full image similarity and the preset threshold, the parts in the adjacent image pair are judged layer by layer, and the determination of the repeated parts in the adjacent image pair can include: when the part similarity is not less than the first preset threshold, the part is determined as a repeated part; when the part similarity is less than the first preset threshold, and the background similarity is not less than the second preset threshold, the part is determined as a repeated part; when the part similarity is less than the first preset threshold, the background similarity is less than the second preset threshold, and the global motion compensated full image similarity is not less than the third preset threshold, the part is determined as a repeated part.
[0115] In this embodiment, multi-source consistency fusion judgment is used to comprehensively integrate various feature similarity information, and different thresholds are set to accurately determine whether the targets in the image pair are the same part. The subject features and context background features obtained through the feature decoupling and cross-reconstruction network are used to determine whether the targets in adjacent images are the same physical part step by step. The multi-level feature similarity collaborative decision module finally combines multiple similarity measures (subject features, context features, full-image features, etc.) to make a part deduplication decision through layer-by-layer screening. This process determines through three main stages. The step-by-step determination process is as follows:
[0116] First layer determination: part similarity feature judgment. If sim component ≥ the first preset threshold T1, it is directly determined as a repeated part.
[0117] Second layer determination: background similarity determination. If the part feature similarity does not reach the threshold, the background similarity sim context_max around the target part is calculated, and according to the background similarity and the second preset threshold T2, it is determined whether the targets in the two images are the same part.
[0118] Third layer determination: global similarity judgment after global motion compensation. If the similarity of the part features and the background features does not reach the preset threshold, the full-image similarity sim global after global motion compensation is used to further evaluate the similarity.
[0119] Determination rules and threshold settings: The multi-level feature similarity collaborative decision module uses three similarity judgment stages, each of which sets different thresholds: the first preset threshold T1 is set to 0.90, which is used to determine whether the part features are similar enough. If the part similarity sim component of the two images exceeds this threshold, it is determined as the same part; the second preset threshold T2 is set to 0.85, which is used to determine the background similarity of the background region around the part with significant distinguishable features; the third preset threshold T3 is set to 0.7, which is used to assist in judgment. When the part feature and background feature similarity cannot pass the threshold, the full-image similarity is used for supplementary judgment.
[0120] Through the hierarchical and step-by-step determination logic, different feature similarities can be gradually screened, improving the accuracy of the decision. At the same time, it can efficiently process large-scale data, avoid redundant calculation, improve the efficiency of part deduplication, and effectively reduce false positives and omissions. By setting adjustable thresholds, the decision engine can flexibly adapt to different image features and is suitable for complex situations such as changes in lighting, occlusion, and angle differences.
[0121] The step synthesizes the main body consistency, context matching and global similarity and the like multi-source evidences, adopts the process of 'hierarchical threshold + fusion judgment': firstly, taking the main body consistency as the main criterion, setting the pass or reject double threshold, quickly confirming the same components or obviously different samples; for the image pairs in the gray area, further combining the context and global similarity to form a comprehensive score, and making the final judgment according to the calibrated threshold. For the still doubtful samples, the multi-scale and expanded context review is started. The mechanism improves the determination robustness under different viewing angles, scales and illumination conditions under the constraint of multi-source information, realizes the accurate identification of the same components and the effective elimination of the repeated samples.
[0122] The embodiment of the present application also provides a component deduplication device for transmission line patrol images, as described in the following embodiment. Since the device solves the problem by the same principle as the component deduplication method for transmission line patrol images, the implementation of the device can be referred to the implementation of the component deduplication method for transmission line patrol images, and the repeated parts will not be described again.
[0123] Figure 4 The structural diagram of the component deduplication device for transmission line patrol images in the embodiment of the present application is shown in FIG. 1, and the device comprises: Figure 4
[0124] An image acquisition module 401 is configured to acquire the transmission line patrol images.
[0125] An adjacent image pair construction module 402 is configured to construct adjacent image pairs according to the names of the transmission line patrol images.
[0126] A region identification module 403 is configured to input the adjacent image pairs into a pre-trained identification network to output identification results. The identification results include component region identification results and background region identification results. The identification network is obtained by training a convolutional neural network using a global image carrying a component region mask and a component region image corresponding to the global image. The identification network is used for feature extraction, feature decoupling and image reconstruction of the adjacent image pairs.
[0127] A global motion compensation module 404 is configured to perform global motion compensation on the adjacent image pairs by using an optical flow method to generate globally motion compensated adjacent image pairs.
[0128] A similarity calculation module 405 is configured to calculate the component similarity, the background similarity and the globally motion compensated global similarity by using a cosine similarity algorithm according to the identification results and the globally motion compensated adjacent image pairs.
[0129] The repeated component determination module 406 is configured to determine the repeated components in the adjacent image pair according to the component similarity, the background similarity, the global motion compensated full image similarity, and a preset threshold value, and determine the repeated components in the adjacent image pair in a hierarchical and step-by-step manner.
[0130] The deduplication module 407 is configured to deduplicate the adjacent image pair according to the repeated components.
[0131] In an embodiment, the adjacent image pair construction module 402 is specifically configured to:
[0132] Parse the name of the transmission line inspection image in the transmission line file, and extract the serial number in the name;
[0133] Sort the transmission line inspection images in the transmission line file according to the serial numbers;
[0134] Construct the adjacent image pair according to the sorted transmission line inspection images in the transmission line file.
[0135] In an embodiment, the recognition network includes a feature extraction module, a feature decoupling module, a feature reconstruction module, and a loss function.
[0136] The feature extraction module is a backbone network constructed based on a convolutional neural network, and is configured to extract features from the adjacent image pair;
[0137] The feature decoupling module is configured to decouple the features into component features and background features;
[0138] The feature reconstruction module is configured to reconstruct an image through a decoder according to the features, the component features, and the background features, to generate a reconstructed image;
[0139] The loss function is a result of weighted summation of a plurality of sub-loss functions; the plurality of sub-loss functions are configured to optimize the loss function of each part in the recognition network.
[0140] In an embodiment, the feature decoupling module includes a component feature encoder and a background feature encoder.
[0141] The component feature encoder includes a max-pooling layer, a convolutional layer, a fully connected layer, and a feature extractor; the component feature encoder is configured to decouple the features through the max-pooling layer, the convolutional layer, and the fully connected layer to obtain first component features, and extract second component features from the features through the feature extractor;
[0142] The background feature encoder includes a max-pooling layer, a convolutional layer, and a fully connected layer; the background feature encoder is configured to decouple the features to obtain background features.
[0143] In an embodiment, the feature reconstruction module is configured to reconstruct through the decoder according to the first component features and the background features, to generate a first reconstructed image.
[0144] The feature reconstruction module is further configured to reconstruct, by the decoder, according to the second component feature and the background feature, to generate a second reconstructed image.
[0145] The decoder comprises a deconvolution layer, a batch normalization function and an activation function, and the deconvolution layer is configured to restore the spatial resolution of the image.
[0146] In one embodiment, the global motion compensation module 404 is specifically configured to:
[0147] extract feature points in the adjacent image pair using an ORB feature point detector adopting an ORB algorithm;
[0148] determine the matching feature points by matching the feature points by calculating the Hamming distance between the descriptors of the feature points using a brute-force matcher;
[0149] calculate the geometric relationship between the matching feature points using a random sample consensus algorithm to obtain an affine transformation matrix;
[0150] compensate for the coordinates of one image in the adjacent image pair according to the affine transformation matrix, transform the coordinates of the image into the coordinate system of the other image in the adjacent image pair, and then align the coordinates;
[0151] determine the adjacent image pair after the coordinate alignment as the adjacent image pair after the global motion compensation.
[0152] In one embodiment, the component region recognition result comprises a component feature vector of the adjacent image pair, and the background region recognition result comprises a context background feature vector of each component in the adjacent image pair; the context background feature vector is a feature vector of a neighboring region located above, below, left and right of the component;
[0153] The similarity calculation module 405 is specifically configured to:
[0154] calculate the cosine similarity between the component feature vectors of the two images in the adjacent image pair using a cosine similarity algorithm, and determine the cosine similarity between the component feature vectors of the two images in the adjacent image pair as the component similarity;
[0155] calculate the cosine similarity between the context background feature vector of one image and the context background feature vector of the other image in the adjacent image pair, and determine the maximum value of the calculation result as the background similarity;
[0156] calculate the cosine similarity between the two images in the adjacent image pair after the global motion compensation, and determine the cosine similarity between the two images in the adjacent image pair after the global motion compensation as the global motion compensation full image similarity.
[0157] In one embodiment, the repeating component determining module 406 is specifically configured to:
[0158] when the component similarity is not less than the first preset threshold, determining that the component is a repeating component;
[0159] when the component similarity is less than the first preset threshold and the background similarity is not less than the second preset threshold, determining that the component is a repeating component;
[0160] when the component similarity is less than the first preset threshold, the background similarity is less than the second preset threshold, and the global motion compensated global image similarity is not less than the third preset threshold, determining that the component is a repeating component.
[0161] Based on the foregoing inventive concept, the present application further provides a computer device 500, as shown in Figure 5 The present application further provides a computer device 500, which comprises a memory 510, a processor 520, and a computer program 530 stored in the memory 510 and capable of running on the processor 520, wherein the processor 520 implements the foregoing power transmission line corridor inspection image component deduplication method when executing the computer program 530.
[0162] Based on the foregoing inventive concept, the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the foregoing power transmission line corridor inspection image component deduplication method.
[0163] Based on the foregoing inventive concept, the present application provides a computer program product, which comprises a computer program, wherein the computer program is executed by a processor to implement a power transmission line corridor inspection image component deduplication method.
[0164] In the technical solution of the present application, the acquisition, storage, use, processing, etc. of data all comply with relevant regulations.
[0165] Compared with the image deduplication technical solution in the prior art, the embodiment of the present application can obtain the inspection image in the power transmission line file; construct an adjacent image pair according to the name of the inspection image in the power transmission line file; input the adjacent image pair into a pre-trained recognition network to output a recognition result; the recognition result includes a component region recognition result and a background region recognition result; the recognition network is obtained by training a convolutional neural network using a global image carrying a component region mask and a component region image corresponding to the global image; the recognition network is used for feature extraction, feature decoupling and image reconstruction of the adjacent image pair; the optical flow method is used to perform global motion compensation on the adjacent image pair to generate a globally motion-compensated adjacent image pair; the cosine similarity algorithm is used to calculate the component similarity, the background similarity and the globally motion-compensated full image similarity according to the recognition result and the globally motion-compensated adjacent image pair; and the component in the adjacent image pair is determined to be a duplicate component through hierarchical and step-by-step judgment according to the component similarity, the background similarity, the globally motion-compensated full image similarity and a preset threshold; and the adjacent image pair is deduplicated according to the duplicate component, so as to provide a deduplication technology which is accurate, efficient, lightweight and can adapt to the characteristics of the inspection image, reduce the redundant data in the inspection image in the power transmission line file and reduce the cost.
[0166] The method of the present application can significantly improve the accuracy and efficiency of the inspection image interval deduplication in the power transmission line file. Its core advantages are:
[0167] The application of feature decoupling and cross-reconstruction network can effectively decouple the main features of the target component from the background features, thereby avoiding the interference of background information and ensuring the accurate identification of the target component. Through this decoupling technology, the network can extract clearer and more representative component features, avoiding the negative impact of background changes or irrelevant regions on the recognition result. This technology greatly improves the robustness and accuracy of target component identification, especially in complex backgrounds and dynamic scenes.
[0168] In addition, the introduction of dynamic scene perception and compensation technology makes the differences between adjacent images due to camera motion or object displacement can be effectively compensated, thereby ensuring the accurate alignment of the target component in the image and eliminating the errors caused by motion and perspective changes. Through global motion estimation, the application of the optical flow method further improves the accuracy of dynamic compensation and avoids the motion mismatch problem in traditional methods.
[0169] Finally, the multi-source consistency fusion judgment strategy comprehensively considers information such as subject features, contextual background features, and overall image similarity to progressively filter and accurately determine whether targets belong to the same component. The advantage of this engine lies in its hierarchical judgment mechanism, which ensures that even if the similarity at a certain layer is insufficient, subsequent supplementary judgments can improve the accuracy and reliability of the final decision, effectively solving the problem of misjudgment caused by complex backgrounds and different perspectives. Overall, this invention, combining deep feature decoupling and reconstruction networks, dynamic scene perception and compensation technology, and a multi-level feature similarity collaborative decision engine, significantly improves the accuracy and robustness of component deduplication. Especially in complex and dynamic inspection environments, it effectively solves the problem of duplicate component judgments caused by factors such as lighting, perspective, and occlusion, improving the recognition accuracy and recall rate of target components and providing reliable support for subsequent component matching and defect detection.
[0170] This method has broad application prospects, and is particularly suitable for fields such as transmission line inspection and equipment testing. It is of great significance for the intelligent operation and maintenance of power systems.
[0171] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0172] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0173] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1the function specified in the one or more blocks.
[0174] These computer program instructions can also be loaded into computer or other programmable data processing devices, so that a series of operational steps are performed on the computer or other programmable data processing devices to generate computer-implemented processes, thus the instructions executed on the computer or other programmable data processing devices provide processes for implementing the flows Figure 1 the flows or the flows and / or blocks Figure 1 the steps of the function specified in the one or more blocks.
[0175] The above-described specific embodiments, the purpose, technical solutions and beneficial effects of the present application are further described in detail, it should be understood that the above-described is only a specific embodiment of the present application, and is not used to limit the protection scope of the present application, any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for removing components from inspection images in power line posts, characterized by, The method comprises the following steps: acquiring a transmission line inspection image; constructing a pair of adjacent images according to the name of the transmission line inspection image; inputting the adjacent image pair into a pre-trained recognition network, and outputting a recognition result; the recognition result comprises a component region recognition result and a background region recognition result; the recognition network is obtained by training a convolutional neural network using a global image carrying a component region mask and a component region image corresponding to the global image; the recognition network is used for feature extraction, feature decoupling and image reconstruction on the adjacent image pair; the recognition network comprises a loss function; the loss function is a result of weighted summation of a plurality of sub-loss functions; the plurality of sub-loss functions are used for optimizing the loss function of each part in the recognition network; the plurality of sub-loss functions comprise a mask-guided subject consistency loss L mask , and the formula is as follows: ; where M mask [i] is a mask region; F global [i] is a feature of the global image; F component1 [i] is a feature of the decoupled subject component; reconstruction consistency loss L construct1 , as follows: ; where I construct1 [i] is a reconstructed image of the first image, I global [i] is a global image; cross-reconstruction consistency loss L construct2 The formula is as follows: ; where I construct2 [i] is a reconstructed image of the second image; Cross-path subject consistency loss L component The formula is as follows: ; where F conponent1 [i] is the subject feature 1, F conponent2 [i] is the subject feature 2; The main body semantic supervision loss L component1_class , for optimizing the classification accuracy of the main body feature 1, the formula is as follows: ; where y[i] is the true label, is the predicted class probability; The main body semantic supervision loss L component2_class , for optimizing the classification accuracy of the main body feature 2, the formula is as follows: ; where y[i] is the true label, is the predicted class probability; performing global motion compensation on the adjacent image pair by using an optical flow method to generate a globally motion compensated adjacent image pair; using a cosine similarity algorithm to calculate a part similarity, a background similarity and a global motion compensated full image similarity according to the recognition result and the pair of adjacent images after global motion compensation; judging the parts in the pair of adjacent images in layers according to the part similarity, the background similarity, the global motion compensated full image similarity and a preset threshold value, and determining repeated parts in the pair of adjacent images; de-duplicating the pair of adjacent images according to the repeated parts; the background area recognition result comprises a context background feature vector of each part in the pair of adjacent images; the context background feature vector is a feature vector of a neighboring area located above, below, left and right of the part; using a cosine similarity algorithm to calculate a part similarity, a background similarity and a global motion compensated full image similarity according to the recognition result and the pair of adjacent images after global motion compensation, comprising: calculating the background similarity according to the following formula: ; where F is the background feature of the first image in the pair of adjacent images i1 F is the background feature of the second image in the pair of adjacent images j2 i, j e {up, down, left, right}, represents a combination of directions, i represents a neighboring region of the first image located above, below, left or right of the component, and j represents a neighboring region of the second image located above, below, left or right of the component; selecting the maximum similarity value of each direction matched with other directions from all calculated similarities according to the following formula to obtain the final similarity of each direction: sim contex_max_i = max(sim i,j ); calculating the background similarity according to the following formula: sim contex_max = max(sim contex_max_i ).
2. The method of claim 1, wherein, The recognition network further comprises a feature extraction module, a feature decoupling module and a feature reconstruction module; the feature extraction module is a backbone network constructed based on a convolutional neural network, and is used for extracting features from the pair of adjacent images; the feature decoupling module is used for decoupling the features into part features and background features; the feature reconstruction module is used for generating a reconstructed image through a decoder according to the features, the part features and the background features.
3. The method of claim 2, wherein, The feature decoupling module comprises a part feature encoder and a background feature encoder; the part feature encoder comprises a max-pooling layer, a convolutional layer, a fully connected layer and a feature extractor; the part feature encoder is used for decoupling the features through the max-pooling layer, the convolutional layer and the fully connected layer to obtain first part features, and extracting second part features from the features through the feature extractor; the background feature encoder comprises a max-pooling layer, a convolutional layer and a fully connected layer; the background feature encoder is used for decoupling the features to obtain background features.
4. The method of claim 3, wherein, The feature reconstruction module is used for generating a first reconstructed image through a decoder according to the first part features and the background features; the feature reconstruction module is further used for generating a second reconstructed image through a decoder according to the second part features and the background features; The decoder comprises an inverse convolutional layer, a batch normalization function and an activation function, and the inverse convolutional layer is used for restoring the spatial resolution of the image.
5. The method of claim 1, wherein, The part area recognition result comprises a part feature vector of the pair of adjacent images; using a cosine similarity algorithm to calculate a part similarity, a background similarity and a global motion compensated full image similarity according to the recognition result and the pair of adjacent images after global motion compensation, further comprising: The cosine similarity algorithm is used to calculate the cosine similarity between the part feature vectors of the two images in the adjacent image pair, and the cosine similarity between the part feature vectors of the two images in the adjacent image pair is determined as the part similarity; The cosine similarity between the two images in the adjacent image pair after global motion compensation is calculated, and the cosine similarity between the two images in the adjacent image pair after global motion compensation is determined as the global motion compensation full image similarity.
6. The method of claim 1, wherein, According to the part similarity, the background similarity, the global motion compensation full image similarity and the preset threshold, the parts in the adjacent image pair are hierarchically and progressively judged, and the repeated parts in the adjacent image pair are determined, including: When the part similarity is not less than the first preset threshold, the part is determined as a repeated part; When the part similarity is less than the first preset threshold, and the background similarity is not less than the second preset threshold, the part is determined as a repeated part; When the part similarity is less than the first preset threshold, the background similarity is less than the second preset threshold, and the global motion compensation full image similarity is not less than the third preset threshold, the part is determined as a repeated part.
7. A device for removing duplicates of components in images of power line poles, characterized by Including: An image acquisition module is configured to acquire transmission line inspection images; An adjacent image pair construction module is configured to construct adjacent image pairs according to the names of the transmission line inspection images; The region identification module is configured to input the adjacent image pair into a pre-trained identification network, and output an identification result. The identification result includes a component region identification result and a background region identification result. The identification network is obtained by training a convolutional neural network using a global image carrying a component region mask and a component region image corresponding to the global image. The identification network is configured to perform feature extraction, feature decoupling, and image reconstruction on the adjacent image pair. The identification network includes a loss function. The loss function is a result of weighted summation of a plurality of sub-loss functions. The plurality of sub-loss functions are configured to optimize the loss function of each part in the identification network. The plurality of sub-loss functions include a mask-guided subject consistency loss L mask , and the formula is as follows: ; where M mask [i] is a mask region; F global [i] is a feature of the global image; F component1 [i] is a feature of the decoupled subject component; reconstruction consistency loss L construct1 , as follows: ; where I construct1 [i] is a reconstructed image of the first image, I global [i] is a global image; cross-reconstruction consistency loss L construct2 The formula is as follows: ; where I construct2 [i] is a reconstructed image of the second image; Cross-path subject consistency loss L component The formula is as follows: ; where F conponent1 [i] is the subject feature 1, F conponent2 [i] is the subject feature 2; The main body semantic supervision loss L component1_class , for optimizing the classification accuracy of the main body feature 1, the formula is as follows: ; where y[i] is the true label, is the predicted class probability; The main body semantic supervision loss L component2_class , for optimizing the classification accuracy of the main body feature 2, the formula is as follows: ; where y[i] is the true label, is the predicted class probability; A global motion compensation module is configured to perform global motion compensation on the adjacent image pairs using an optical flow method to generate adjacent image pairs after global motion compensation; A similarity calculation module is configured to calculate part similarity, background similarity and global motion compensation full image similarity according to the recognition results and the adjacent image pairs after global motion compensation using a cosine similarity algorithm; A repeated part determination module is configured to hierarchically and progressively judge the parts in the adjacent image pairs according to the part similarity, the background similarity, the global motion compensation full image similarity and the preset threshold, and determine the repeated parts in the adjacent image pairs; A deduplication module is configured to deduplicate the adjacent image pairs according to the repeated parts. The background area recognition result includes a contextual background feature vector of each part in the adjacent image pair; the contextual background feature vector is a feature vector of a neighboring area located above, below, left and right of the part; The similarity calculation module is configured to: The background similarity is calculated according to the following formula: ; where F is the background feature of the first image in the pair of adjacent images i1 F is the background feature of the second image in the pair of adjacent images j2 i,j e {up, down, left, right}, representing a combination of directions, i represents a neighboring region of the first image located above, below, left or right of the component, and j represents a neighboring region of the second image located above, below, left or right of the component; The final similarity of each direction is obtained by selecting the maximum similarity value matching each direction and other directions from all calculated similarities according to the following formula: sim contex_max_i =max(sim i,j ) The background similarity is obtained according to the following formula: sim contex_max =max(sim contex_max_i ).
8. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the method of any one of claims 1-6 when executing the computer program.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-6.
10. A computer program product, characterised in that, The computer program product includes a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-6.
Citation Information
Patent Citations
Transmission tower inspection image de-weighting method and system under visible light
CN115689928A
Power transmission line multi-mode warning system and expelling method
CN120472598A