Electric power field semi-automatic labeling method based on target detection

By introducing multi-scale feature extraction and adaptive weighted residual connection structure, combined with rendering engine and adaptive segmentation algorithm, the accuracy and robustness issues of power equipment component identification in complex scenarios are solved, and efficient data acquisition and annotation are achieved.

CN120913205APending Publication Date: 2025-11-07STATE GRID HEBEI ELECTRIC POWER RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510936851.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing power equipment component identification technologies suffer from low accuracy and poor robustness in complex scenarios, as well as low data acquisition and annotation efficiency. They are also ill-suited to problems such as multiple component occlusions and uneven lighting, and exhibit discontinuous inter-frame prediction in video sequence processing.

Method used

We employ multi-scale feature extraction based on deep convolutional neural networks and adaptive weighted residual connection structures, combined with a rendering engine to generate multi-angle images. We use adaptive thresholding and region growing algorithms for segmentation, and introduce spatiotemporal attention feature aggregation and adaptive Kalman filtering algorithms for motion prediction.

Benefits of technology

It improves the accuracy and robustness of power equipment component identification, enhances adaptability to complex scenarios, reduces computational complexity, and improves the efficiency of data acquisition and annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913205A_ABST
    Figure CN120913205A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power field semi-automatic labeling method based on target detection, and belongs to the technical field of target detection, and the method comprises the steps: importing collected three-dimensional modeling data of an electric power equipment combination part into a rendering engine, so as to obtain a multi-angle rendering image in a preset angle range; a deep convolutional neural network is adopted to extract contour information of the rendered image, and features are fused and trained through a multi-scale feature extraction module and an adaptive weighted residual connection structure; performing perspective transformation and adaptive threshold segmentation on the combined part detected after training, completing part segmentation by adopting a region growing algorithm, and establishing a hierarchical relationship between the combined part and an independent part; and extracting video frames of the power equipment according to a preset frame rate, performing preprocessing, and performing semi-automatic fine tuning on a model detection result obtained based on a preprocessed image to obtain annotation data. Through the scheme of the invention, the identification performance of the power equipment component can be improved, and the method is suitable for a complex practical application scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of target detection, and particularly relates to a method for semi-automatic labeling in the field of power based on target detection. BACKGROUND

[0002] In the important field of power equipment operation and maintenance, component recognition technology as a key means to ensure the safe and stable operation of equipment has attracted much attention in its development and application. Looking back at the development history of existing technology, traditional component recognition methods mainly rely on extracting and matching image features under a single perspective. This technical route can achieve basic recognition function under ideal conditions. However, in the actual power equipment inspection scene, due to the spatial limitations of equipment installation location and the fixedness of inspection route, etc., the collected image data often has serious problems such as complex perspective changes, uneven illumination, and mutual occlusion of components. In view of these challenges, the existing technology usually adopts the method of expanding the training sample set to improve the model performance, but this scheme needs to invest a lot of human resources for data collection and labeling work, and the efficiency is low. Especially when dealing with combined structures with multiple components, due to the phenomenon of spatial overlap and mutual occlusion between components, the feature extraction is incomplete, which seriously affects the recognition accuracy. At the same time, when processing continuous video sequences, due to the failure to fully utilize the rich temporal information contained in adjacent frames, there are often technical problems such as discontinuous inter-frame prediction and unstable detection results, which greatly limits the popularization of this technology in practical application.

[0003] To solve the above technical problems, researchers have disclosed a variety of improvement schemes. In the feature extraction stage, the existing technology introduces a residual network structure to enhance the feature extraction capability, but this structure adopts a unified residual learning strategy for different input features, without fully considering the difference in the discrimination of features, which makes it difficult for the network to adaptively adjust the intensity of feature extraction according to the importance of different features. In the process of multi-scale feature fusion, the current method generally uses simple methods such as feature splicing or weighted averaging for feature fusion, which ignores the importance difference between different scale features and fails to effectively utilize the complementary information contained in each scale feature. In the aspect of motion target tracking in video sequences, the traditional Kalman filter algorithm uses a fixed state transition model and observation model for target tracking, but it shows poor robustness when dealing with component occlusion and other abnormal situations. In addition, in the recognition method based on deep learning, the geometric deformation problem caused by perspective transformation and occlusion and other factors makes it difficult for the neural network to extract stable feature representation, which seriously reduces the generalization ability of the model. Although these improvement schemes improve the system performance to some extent, there are still many challenges in practical application, and key problems such as how to further improve the recognition accuracy, enhance the system robustness, and reduce the computational complexity still need to be solved.

[0004] Therefore, there is an urgent need for a technical solution to improve the performance of power equipment component identification and adapt to complex practical application scenarios. SUMMARY

[0005] To solve the problems of the prior art, the embodiments of the present application disclose a method for semi-automatic labeling in the field of power based on target detection. Through improvements in view diversity, adaptive segmentation, and tracking robustness, the present application solves the technical problems of poor robustness in the prior art.

[0006] In a first aspect, the embodiments of the present application disclose a method for semi-automatic labeling in the field of power based on target detection, comprising: importing collected three-dimensional modeling data of power equipment combined components into a rendering engine to obtain multi-angle rendering images within a preset angle range; extracting contour information of the rendering images using a deep convolutional neural network, and fusing and training features through a multi-scale feature extraction module and an adaptive weighted residual connection structure; performing perspective transformation and adaptive threshold segmentation on the combined components detected after training, completing component segmentation using a region growing algorithm, and establishing a hierarchical relationship between the combined components and independent components; extracting power equipment video frames at a preset frame rate and performing preprocessing, and performing semi-automatic fine-tuning on model detection results obtained based on the preprocessed images to obtain labeling data.

[0007] In an implementation, the three-dimensional modeling data of the power equipment combined components is imported into the rendering engine to obtain multi-angle rendering images within a preset angle range, comprising: collecting three-dimensional modeling data of power equipment combined components, and importing the three-dimensional modeling data into the rendering engine for parameter setting; taking the combined components as the center in a spherical coordinate system, rendering at a preset elevation angle range and azimuth angle range respectively to obtain multi-angle rendering images; setting diffuse reflection coefficient, specular reflection coefficient, and roughness for the rendering images, and setting corresponding metallicity parameters according to the material type; setting a main light source and an auxiliary light source based on ambient light, wherein the brightness of the main light source is greater than that of the ambient light, and the brightness of the auxiliary light source is less than that of the main light source; performing brightness adjustment, contrast adjustment, saturation adjustment, and geometric transformation processing on the generated images.

[0008] In an implementation manner, the contour information of the rendered image is extracted by using a deep convolutional neural network, and features are fused and trained by a multi-scale feature extraction module and an adaptive weighted residual connection structure, including: contour information of the combined component is extracted by an image segmentation algorithm, and the segmented combined component region is labeled; a deep learning training environment is configured, and a deep convolutional neural network architecture including a multi-scale feature extraction module is constructed; a deformable convolution module is set in the network to process the geometric transformation of the target, and a candidate box is generated by a region proposal network; an adaptive weighted residual connection structure is used to calculate weight coefficients for the residual branch and the identity mapping branch respectively; a multi-scale feature aggregation method based on an attention mechanism is used to weight and fuse feature maps of different scales.

[0009] In an implementation manner, perspective transformation and adaptive threshold segmentation are performed on the combined component detected after training, a region growing algorithm is used to complete component segmentation, and a hierarchical relationship between the combined component and the independent component is established, including: perspective transformation is performed on the detected combined component region to convert the inclined view angle to an orthographic view; adaptive histogram equalization processing is performed on the orthographic view, and an edge detection operator is used to extract the component contour; an improved Otsu algorithm is used to perform adaptive threshold segmentation on the combined component region; a region growing algorithm is used to segment the components, and the relative position relationship of each independent component is recorded; the independent components obtained by segmentation are labeled, and a hierarchical relationship between the combined component and the independent components is established.

[0010] In an implementation manner, video frames of the power equipment are extracted at a preset frame rate and preprocessed, and a model detection result obtained based on the preprocessed image is semi-automatically fine-tuned to obtain labeled data, including: video frames are extracted at a preset frame rate, and selected frames are preprocessed using Gaussian blur; the motion of the target component is tracked by feature point matching, and a ratio test method is used to filter mismatched points; a feature aggregation method based on spatiotemporal attention is used to process the video frame sequence; a multi-target tracking algorithm based on Kalman filtering is used for motion prediction; the detection result is visually displayed, and the processing result is saved as labeled data containing a time stamp.

[0011] In an implementation manner, the adaptive threshold segmentation of the combined component region is performed by using an improved Otsu algorithm, including: dividing the combined component region image into a plurality of sub-blocks, calculating a local optimal threshold for each of the sub-blocks; dividing pixels in the sub-blocks into foreground pixels and background pixels according to the local optimal threshold, and calculating foreground pixel proportion and background pixel proportion respectively; calculating average gray values of the foreground pixels and the background pixels, and determining an optimal segmentation threshold based on the average gray values; performing median filtering processing on the local optimal thresholds of the sub-blocks, and the filtering uses a preset size of a filtering window; and performing smooth transition on the thresholds of adjacent sub-blocks by an interpolation method, and the interpolation weight is inversely proportional to the cube of the distance.

[0012] In an implementation manner, the video frame sequence is processed by using a feature aggregation method based on space-time attention, including: extracting a feature map sequence of a current frame and a preset number of frames before and after the current frame, and constructing an inter-frame correlation matrix of the feature maps; performing feature transformation on the feature map sequence, and obtaining a transformed feature by convolution operation; calculating an attention weight in a time dimension based on the correlation matrix, and adjusting the weight by a preset temperature parameter; and performing weighted calculation on the attention weight and the transformed feature to obtain a feature aggregation result.

[0013] In an implementation manner, the motion prediction is performed by using a multi-target tracking algorithm based on Kalman filtering, including: constructing a state vector containing a target center coordinate, an aspect ratio, an area and a speed component; setting a state transition matrix and an observation matrix, wherein the state transition matrix contains a unit matrix and a time interval; updating a process noise covariance matrix on line by a maximum likelihood estimation method, and calculating based on a state transition bias; updating an observation noise covariance matrix on line by a maximum likelihood estimation method, and calculating based on a bias between an observation value and a state estimation; and increasing diagonal elements of the observation noise covariance matrix when a target is occluded, and enhancing a prediction weight of a state transition model.

[0014] In a second aspect, the application also provides a semi-automatic labeling device, including: a processor, a memory, a system bus; the processor and the memory are connected through the system bus; the memory is used for storing one or more programs, the one or more programs include instructions, the instructions make the processor execute any of the above-mentioned semi-automatic labeling methods in the field of power based on target detection when executed by the processor.

[0015] In a third aspect, the application also provides a computer program product, when the computer program product runs on a terminal device, makes the terminal device execute any of the above-mentioned semi-automatic labeling methods in the field of power based on target detection.

[0016] In a target detection-based semi-automatic labeling method in the field of power, the embodiment of the application introduces a physical-based rendering technology and an adaptive data enhancement method to improve the diversity of training data while ensuring the authenticity of the rendered image. By designing an adaptive weighted residual structure and an attention mechanism-based feature fusion method, the network's expression ability for different features is enhanced. In the image segmentation stage, an improved Otsu algorithm is used to realize adaptive threshold segmentation, and a monocular depth estimation method based on deep learning is used to improve the accuracy of combined component decomposition. In video sequence processing, a spatiotemporal attention-based feature aggregation mechanism and an adaptive Kalman filtering algorithm are designed to realize stable tracking and prediction of the motion state of the component. These technical improvements improve the performance of power equipment component recognition from different levels, making the method suitable for complex practical application scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0018] Figure 1 A flowchart of a target detection-based semi-automatic labeling method in the field of power disclosed by the embodiment of the present application;

[0019] Figure 2 A flowchart of a component segmentation method disclosed by the embodiment of the present application;

[0020] Figure 3 A flowchart of a power equipment video frame preprocessing method disclosed by the embodiment of the present application. DETAILED DESCRIPTION

[0021] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. Note that the relative arrangement, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.

[0022] Those skilled in the art can understand that the terms "first", "second" and the like in the embodiments of the disclosure are only used to distinguish different steps, devices or modules and the like, and do not represent any specific technical meaning, nor indicate their logical order. It should also be understood that in the embodiments of the disclosure, "multiple" can mean two or more, and "at least one" can mean one, two or more. It should also be understood that for any component, data or structure mentioned in the embodiments of the disclosure, without explicit limitation or in the context of the preceding and following, it can be understood as one or more. In addition, the term "and / or" in the disclosure only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B together, and the existence of B alone. In addition, the character " / " in the disclosure generally represents an "or" relationship between the front and rear associated objects. It should also be understood that the description of various embodiments of the disclosure emphasizes the differences between various embodiments, and the same or similar parts can be referred to each other, and for the sake of brevity, they will not be repeated.

[0023] At the same time, it should be understood that in order to facilitate the description, the size of each part shown in the drawings is not drawn according to the actual proportional relationship. The following description of at least one exemplary embodiment is actually only illustrative, but not as any limitation on the disclosure and its application or use. The technology, methods and devices known to those skilled in the relevant art can not be discussed in detail, but in appropriate cases, the technology, methods and devices should be considered as part of the specification. It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.

[0024] In order to make the purposes, technical solutions and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0025] Figure 1 A flowchart of a target detection-based semi-automatic labeling method in the power field disclosed by the embodiments of the present application.

[0026] It needs to be understood that the current semi-automatic labeling technology in the field of electric power mainly includes power equipment detection method based on ResNet-101, transmission line insulator defect recognition system based on Fast R-CNN, and overhead transmission line image defect detection technology based on YOLOv5. These conventional technologies have obvious technical limitations in practical application. The power equipment detection method based on ResNet-101 uses a fixed weight residual connection structure, which cannot dynamically adjust the feature extraction strategy according to the nature of the input features when processing image features of different scales and complexities. The simple channel splicing method used for feature fusion cannot reasonably allocate the importance of different level features, resulting in a significant performance decline when processing images with severe occlusion or uneven lighting. When rendering training samples, this method only considers basic lighting and material parameters, and the generated images lack realism and differ significantly from actual inspection scenes.

[0027] The transmission line insulator defect recognition system based on Fast R-CNN uses anchor boxes of fixed size and aspect ratio to generate candidate regions during target detection, which is difficult to adapt to the diversity of power equipment component sizes and shapes. Its region proposal network uses a single threshold non-maximum suppression algorithm to process overlapping detection boxes, which can easily miss detection in dense component areas. When processing video sequences, this system uses the basic frame difference method for motion detection, which does not fully utilize the temporal information of the video, resulting in discontinuous detection results. Its deep feature extraction module uses pre-trained weights and does not optimize the feature distribution for power equipment, limiting the feature expression capability.

[0028] The overhead transmission line image defect detection technology based on YOLOv5 uses the global histogram equalization method in the image preprocessing stage, which cannot handle local area uneven lighting problems. Its network structure uses a fixed feature pyramid for multi-scale feature extraction, lacking an adaptive evaluation mechanism for the importance of different scale features. When dealing with component occlusion, it is difficult to accurately divide the boundaries of overlapping components due to the lack of effective depth information estimation and three-dimensional structure understanding. When performing multi-target tracking, this technology uses the standard Kalman filter, whose state transition model and observation model parameters are fixed, and cannot adapt to the dynamic changes of target motion states.

[0029] In contrast, the adaptive weighting residual structure adopted by the embodiments of the present application dynamically calculates the weight coefficients of the residual branch and the identity mapping branch, so that the network can adaptively adjust the feature extraction strategy according to the characteristics of the input features. In the feature fusion stage, a multi-scale feature aggregation method based on attention mechanism is introduced, which learns the importance weights of different scale feature maps to realize adaptive fusion of features. By designing a monocular depth estimation network based on deep learning, combined with edge-aware depth smoothing constraints, more accurate three-dimensional structure information is disclosed, which provides an effective basis for the segmentation of overlapping parts. When processing video sequences, a spatiotemporal attention feature aggregation mechanism and an adaptive Kalman filtering algorithm are used to realize accurate modeling and prediction of the motion state of the parts. In addition, the embodiments of the present application use a physics-based rendering technique when generating training samples, which significantly improves the realism of the rendered images by setting fine material parameters and simulating lighting, enabling the model to better migrate to actual application scenarios.

[0030] As shown in Figure 1 At step S101, the collected three-dimensional modeling data of the power equipment combined components is imported into the rendering engine to obtain multi-angle rendering images within a preset angle range. This includes: collecting three-dimensional modeling data of the power equipment combined components and importing the three-dimensional modeling data into the rendering engine for parameter setting; taking a viewing angle in the preset elevation angle range and azimuth angle range in the spherical coordinate system with the combined components as the center to perform rendering and obtain multi-angle rendering images; setting the diffuse reflection coefficient, specular reflection coefficient, and roughness of the rendering images, and setting the corresponding metallicity parameter according to the material type; setting the main light source and the auxiliary light source based on the ambient light, wherein the brightness of the main light source is greater than that of the ambient light, and the brightness of the auxiliary light source is less than that of the main light source; performing brightness adjustment, contrast adjustment, saturation adjustment, and geometric transformation processing on the generated images.

[0031] Specifically, first, the three-dimensional modeling data of the power equipment combined components is collected, and multi-angle rendering images of the combined components are generated by three-dimensional modeling software. Specifically, the three-dimensional modeling data is imported into the rendering engine, and the rendering parameters are set, including an ambient light intensity of 2000-3500 lux, a shadow softening radius of 8.7-12.3 pixels, and a material reflectivity of 0.15-0.85. Subsequently, in the spherical coordinate system, taking the combined components as the center, an angle is taken every 4.2 degrees in the elevation angle range of 37.8 degrees to 142.2 degrees, and an angle is taken every 6.3 degrees in the azimuth angle range of 0 degrees to 360 degrees, and each angle is rendered to obtain a set of multi-angle combined component rendering images.

[0032] Preferably, in one embodiment, for the rendering process of the combined component, a physically-based rendering technique is adopted, with a diffuse reflectance coefficient of 0.67-0.89, a specular reflectance coefficient of 0.11-0.33, and a roughness of 0.23-0.78. For metal material components, the metal degree is set to 0.82-0.97; for non-metal material components, the metal degree is set to 0.03-0.18. When rendering, a bidirectional path tracing algorithm is used, with the number of samples per pixel set to 1024-2048, and the maximum number of ray bounces set to 6-8. To enhance the realism of the rendered image, three area lights are added in addition to the ambient light, with the color temperature of the main light set to 5600K, the brightness set to 2.3-2.8 times that of the ambient light, the color temperature of the auxiliary lights set to 4200K and 6500K, and the brightness set to 0.42-0.57 times and 0.31-0.46 times that of the main light, respectively.

[0033] The rendered image is subjected to data augmentation processing, including randomly adjusting the brightness range to 0.82-1.23 times the original value, the contrast range to 0.87-1.18 times the original value, and the saturation range to 0.91-1.12 times the original value. At the same time, the image is subjected to geometric transformation, including randomly scaling within 92%-108% of the original size, randomly rotating within ±7.8 degrees, and randomly translating in the horizontal and vertical directions by no more than 4.7% of the image size. To simulate the blur effect in actual shooting scenes, motion blur is added to some images, with the blur kernel size set to 3-7 pixels and the blur direction at an angle of 27.5-152.5 degrees with the image gradient direction.

[0034] At step S102, the contour information of the rendered image is extracted using a deep convolutional neural network, and the features are fused and trained through a multi-scale feature extraction module and an adaptive weighted residual connection structure. This includes: extracting the contour information of the combined component through an image segmentation algorithm, and labeling the segmented combined component region; configuring a deep learning training environment and building a deep convolutional neural network architecture containing a multi-scale feature extraction module; setting a deformable convolution module in the network to handle the geometric transformation of the target, and generating candidate boxes through a region proposal network; using an adaptive weighted residual connection structure to calculate weight coefficients for the residual branch and the identity mapping branch, respectively; using a multi-scale feature aggregation method based on attention mechanism to weight and fuse feature maps of different scales.

[0035] Specifically, for each rendered image, the contour information of the combined components is extracted by an image segmentation algorithm, wherein the segmentation algorithm adopts an instance segmentation network based on deep learning, uses cross entropy as the loss function, sets the learning rate to 0.0023, sets the batch size to 16, and sets the training rounds to 235. In the segmentation process, the image is divided into 256x256 size tiles, each tile has an overlap rate of 23.7%, and each tile is processed separately before being spliced. The contour information obtained by segmentation is saved in the form of a binary image at the pixel level, and the relative coordinate information of each point in the contour is recorded. For each combined component region after segmentation, the component class is manually labeled, including insulators, fuses, and surge arresters, to form an initial training data set. The labeled data is divided into a training set and a validation set in a ratio of 7:3, the training set is used for model training, and the validation set is used for model performance evaluation.

[0036] Further, in the instance segmentation, an improved Mask R-CNN network structure is adopted, and an attention mechanism is introduced into the feature pyramid network to adaptively weight the feature maps of different scales. The attention module is composed of spatial attention and channel attention, wherein the spatial attention calculates a spatial weight map through a two-layer convolutional network, the first layer of convolution kernel size is 7x7, and the second layer is 1x1, and a GELU activation function is added in the middle; the channel attention extracts channel descriptors through global average pooling and maximum pooling, and obtains channel weights through a shared multi-layer perceptron, the number of hidden layer neurons of the multi-layer perceptron is 0.37 times the number of input channels. Finally, the outputs of spatial attention and channel attention are fused by element-wise multiplication. In the training process, a cosine annealing learning rate scheduling strategy is adopted, the initial learning rate is 0.002, the minimum learning rate is 0.00008 times the initial value, and the annealing period is 67 epochs. To prevent overfitting, a Dropout layer is added after the convolution layer, and the inactivation probability is set to 0.3.

[0037] In one implementation scenario, based on the obtained combined component annotation data, a deep learning training environment is configured, and a deep convolutional neural network architecture suitable for an instance segmentation task is selected. The network input end adopts a multi-scale feature extraction module, which contains 3 feature maps of different scales, with sizes of 1 / 4, 1 / 8 and 1 / 16 of the input image respectively. Each feature map extracts features through 4 consecutive convolution layers, with a convolution kernel size of 3x3, a step size of 1 and a padding mode of SAME. The feature-extracted feature maps are fused through an upsampling module. An inverse convolution operation is used to upsample the low-resolution feature maps to the same size as the high-resolution feature maps, with an upsampling factor of 2. Then, the feature maps of the same size are spliced in the channel dimension. At the decoding end of the network, a deformable convolution module is used to process the geometric transformation of the target. The offset field of the deformable convolution is generated by a separate convolution branch, and the learning rate of the offset is 0.1 times the base learning rate. Then, a region proposal network is used to generate candidate boxes. Five different scale anchor boxes are set, with aspect ratios including 1:2, 1:1 and 2:1. The step size of the sliding window on the feature map is 16 pixels. For each candidate box, the intersection over union (IOU) with the true annotation box is calculated. When the IOU is greater than 0.7, it is marked as a positive sample, and when the IOU is less than 0.3, it is marked as a negative sample. Samples between the two are not involved in training. During training, an online hard example mining strategy is used. The negative samples in each mini-batch are arranged in descending order of loss value, and the top K negative samples are selected for backpropagation, where K is 3 times the number of positive samples in the current batch. Finally, the detection results are post-processed by a non-maximum suppression algorithm, and the threshold for non-maximum suppression is set to 0.45.

[0038] In one embodiment, in the backbone network of the neural network, an improved residual structure is used, and the traditional residual connection is replaced by an adaptive weighted residual connection. Specifically, weight coefficients a and b are calculated for the residual branch and the identity mapping branch respectively, and the calculation formula is: a = s(W1 P + b1), b = s(W2 Q + b2), where P and Q are feature vectors of the residual branch and the identity mapping branch respectively, obtained through global average pooling, W1 and W2 are learnable weight matrices, and b1 and b2 are bias terms. The final output feature map F is calculated as follows: F = a T(X) + b X, where T(X) is the output of the residual branch and X is the input feature map. By introducing adaptive weights, the network can dynamically adjust the degree of residual learning according to different input features.

[0039] In the feature fusion stage, a multi-scale feature aggregation method based on attention mechanism is used. For a feature map Fs of scale s, first calculate its attention weight w s = softmax(MLP(GAP(F swhere GAP represents a global average pooling operation, and MLP is a two-layer neural network with 0.43 times the input dimension of neurons in the first layer and an output dimension of 1 in the second layer. Then the feature maps of different scales are fused by weighting: F = å(w s • Up(F s ), where Up represents an operation of upsampling the feature map to the same size using a bicubic interpolation method. In the upsampling process, to reduce information loss, guide points are additionally inserted at the four corner positions of the feature map, and the values of the guide points are calculated by weighted average of adjacent feature points, with the weight being inversely proportional to the distance.

[0040] At step S103, perspective transformation and adaptive threshold segmentation are performed on the combined component detected after training, a region growing algorithm is used to complete component segmentation, and a hierarchical relationship between the combined component and the independent component is established. As shown in Figure 2 : perspective transformation is performed on the detected combined component region to convert the inclined view angle to an orthographic view; adaptive histogram equalization processing is performed on the orthographic view, and an edge detection operator is used to extract the component contour; an improved Otsu algorithm is used to perform adaptive threshold segmentation on the combined component region; a region growing algorithm is used to segment the components, and the relative positional relationship of each independent component is recorded; the independent components obtained by segmentation are labeled, and a hierarchical relationship between the combined component and the independent component is established.

[0041] At step S104, video frames of the power equipment are extracted at a preset frame rate and preprocessed, and the model detection result obtained based on the preprocessed image is semi-automatically fine-tuned to obtain labeled data. As shown in Figure 3 : video frames are extracted at a preset frame rate, and the selected frames are preprocessed using Gaussian blur; the motion of the target component is tracked by feature point matching, and a ratio test method is used to filter the mismatched points; a feature aggregation method based on spatiotemporal attention is used to process the video frame sequence; a multi-target tracking algorithm based on Kalman filtering is used for motion prediction; the detection result is visualized and displayed, and the processing result is saved as labeled data containing a timestamp.

[0042] Figure 2 A flowchart of a component segmentation method disclosed in an embodiment of the present application.

[0043] As shown in Figure 2 : at step S201, perspective transformation is performed on the detected combined component region to convert the inclined view angle to an orthographic view.

[0044] The trained combined part detection model can be used to further decompose the combined parts in the inspection image. First, perspective transformation is performed on each detected combined part region to convert the part image under the tilted view to an orthographic view, and the transformation matrix is calculated from the corresponding relationship of the four corner points.

[0045] At step S202, adaptive histogram equalization processing is performed on the orthographic view, and a part contour is extracted using an edge detection operator.

[0046] In the orthographic view, the image quality is improved by brightness equalization processing, and an adaptive histogram equalization algorithm is used to divide the image into 32x32 local regions, and histogram equalization is performed on each local region. Bilateral interpolation is used for smooth transition between adjacent regions. Then, a part contour is extracted using an edge detection operator, and the Canny operator is selected, with the low threshold set to 85, the high threshold set to 255, and the Gaussian kernel size set to 5x5.

[0047] At step S203, an improved Otsu algorithm is used for adaptive threshold segmentation of the combined part region. This includes: dividing the combined part region image into multiple sub-blocks, calculating the local optimal threshold for each sub-block; dividing the pixels in the sub-blocks into foreground pixels and background pixels according to the local optimal threshold, and calculating the foreground pixel ratio and background pixel ratio, respectively; calculating the average gray values of the foreground pixels and the background pixels, and determining the optimal segmentation threshold based on the average gray values; performing median filtering on the local optimal threshold of the sub-blocks, and the filtering uses a filter window of a predetermined size; and smoothing the threshold values of adjacent sub-blocks by interpolation, with the interpolation weight being inversely proportional to the cube of the distance.

[0048] Preferably, in one embodiment, for the detected combined part region, adaptive threshold segmentation is first performed, and an improved Otsu algorithm is used to calculate the optimal threshold. The image is divided into MxN sub-blocks, and the local optimal threshold is calculated for each sub-block where ω1(t) and ω2(t) are the proportions of foreground and background pixels at threshold t, and μ1(t) and μ2(t) are the corresponding average gray values. To improve robustness, the calculated local threshold is median filtered with a filter window size of 3x3. The threshold values of adjacent sub-blocks are smoothly transitioned by bicubic interpolation, with the interpolation weight being inversely proportional to the cube of the distance.

[0049] In the perspective transformation, a monocular depth estimation network based on deep learning is used to predict the three-dimensional structure information of the combined component. The network uses an encoder-decoder architecture, and the encoder consists of 6 residual blocks, each containing three layers of convolution with kernel sizes of 1x1, 3x3, and 1x1. In the decoder, a skip connection structure is used to concatenate the feature maps of the corresponding layers of the encoder with the upsampled feature maps, and then fuse the features through 1x1 convolution. The loss function for depth estimation uses a weighted sum of depth smooth L1 loss and edge perception term: where D and D* are the predicted depth map and the true depth map, respectively, and I is the input image, represents the gradient operator, and λ1, λ2, and β are weight coefficients, set to 0.85, 0.37, and 0.063, respectively.

[0050] At step S204, the region growing algorithm is used to segment the components, and the relative positional relationship of each independent component is recorded.

[0051] Specifically, based on the contour after region segmentation, the region growing algorithm is used to segment the components, and the local maximum point inside the contour is selected as the seed point, and the similarity threshold in the growing process is 1.8 times the local region gray scale standard deviation. For each independent component obtained by segmentation, the relative positional relationship in the combined component is recorded, including the center point coordinates, the rotation angle, the occlusion relationship, etc.

[0052] At step S205, the independent components obtained by segmentation are labeled, and the hierarchical relationship between the combined component and the independent component is established.

[0053] In one embodiment, these independent components can be labeled, and the labeling information includes component type and component state. The labeling information is associated with the previous combined component detection result to establish the hierarchical relationship between the combined component and the independent component, forming structured labeling data. Finally, the accuracy of the labeling is reviewed, and the corrected result after review is fed back to the training data set for continuous optimization of the model.

[0054] Figure 3 A flowchart of a power equipment video frame preprocessing method disclosed in an embodiment of the present application.

[0055] As shown in Figure 3 At step S301, video frames are extracted at a preset frame rate, and Gaussian blur is used for preprocessing on the selected frames.

[0056] In the actual application process, when processing continuous inspection videos, video frames are extracted by a video decoder at a rate of 25 frames per second. In order to reduce redundant calculation, one frame is selected every 3 frames for processing, and Gaussian blur is used for preprocessing of the selected frame, with a kernel size of 3x3 and a standard deviation of 0.8 to reduce noise influence.

[0057] At step S302, the motion of the target component is tracked by feature point matching, and a ratio test method is used to filter mis-matching points.

[0058] At step S303, a feature aggregation method based on spatiotemporal attention is used to process the video frame sequence. This includes: extracting the feature map sequence of the current frame and a preset number of frames before and after it, and constructing the inter-frame correlation matrix of the feature map; performing feature transformation on the feature map sequence, and obtaining the transformed feature through convolution operation; calculating the attention weight in the time dimension based on the correlation matrix, and adjusting the weight through a preset temperature parameter; and performing weighted calculation on the attention weight and the transformed feature to obtain the feature aggregation result.

[0059] At step S304, a multi-target tracking algorithm based on Kalman filtering is used for motion prediction. This includes: constructing a state vector containing the coordinates of the target center, the aspect ratio, the area, and the speed components; setting a state transition matrix and an observation matrix, wherein the state transition matrix contains a unit matrix and a time interval; updating the process noise covariance matrix online through maximum likelihood estimation method, based on state transition bias calculation; updating the observation noise covariance matrix online through maximum likelihood estimation method, based on observation value and state estimation bias calculation; and increasing the diagonal elements of the observation noise covariance matrix when the target is occluded, to enhance the prediction weight of the state transition model.

[0060] Specifically, between adjacent frames, the motion of the target component is tracked by feature point matching, SIFT feature descriptors are used to extract feature points, and a 128-dimensional description vector is extracted for each feature point. K-Nearest Neighbor algorithm is used for feature matching, with K value set to 2, and a ratio test method is used to filter mis-matching points, with a ratio threshold set to 0.75. Based on the spatial relationship between the matching point pairs, a RANSAC algorithm is used to estimate the inter-frame transformation matrix, with the number of iterations set to 1000 and the inlier judgment threshold set to 2.5 pixels. The detected component regions are associated between consecutive frames, and if the detection probability of a certain component is greater than 0.85 in more than 5 consecutive frames, the detection result is considered reliable. For regions with large fluctuations in detection probability, multi-frame fusion is performed between adjacent frames, and the weight of each frame is proportional to its detection probability. When occlusion of the component is detected, the information of the occluded region is compensated according to motion prediction, with a prediction window size set to 3 frames before and after.

[0061] Preferably, in one embodiment, a spatio-temporal attention based feature aggregation method is adopted in the video frame sequence processing. For the current frame ftand the feature map sequence {ft-N,..., ft+N} of N frames before and after it, first calculate the correlation matrix where and ψ are two independent feature transformation functions, realized by 1x1 convolution, d is the feature dimension. Then calculate the attention weight in the time dimension based on the correlation matrix: where τ is the temperature parameter, set to 0.182. The final feature aggregation result F is obtained by weighted summation: F = ∑(A i ·g(f i )), where g is the feature transformation function, also realized by 1x1 convolution.

[0062] To improve the accuracy of motion prediction, a multi-target tracking algorithm based on Kalman filtering can be used. The state vector contains the center coordinates, aspect ratio, area and corresponding velocity components of the target. The state transition matrix F and the observation matrix H are respectively: H = [I6 0], where I6 is a 6x6 identity matrix, Δt is the time interval between adjacent frames. The process noise covariance matrix Q and the observation noise covariance matrix R are updated online by maximum likelihood estimation: Q = E[(x-Fx')(x-Fx') T ], R = E[(z-Hx)(z-Hx) T ], where x and x' are the current state and the last state respectively, z is the observation value. To deal with the case where the part is occluded, when the observation value is missing, increase the diagonal elements of the observation noise covariance matrix R to 2.7 times the normal value, so that the prediction result depends more on the state transition model.

[0063] At step S305, the detection result is visualized and the processing result is saved as labeled data containing a timestamp.

[0064] In one embodiment, the final processing result is visualized, supporting manual fine-tuning of the detection result, and the adjustment range includes the position of the bounding box, the part category, etc. The fine-tuned result is saved in XML format, containing timestamp, part position, category information and other labeled data.

[0065] In summary, the embodiments of the present application introduce the physical-based rendering technology and the adaptive data enhancement method, which improves the diversity of the training data while ensuring the authenticity of the rendered image. By designing the adaptive weighted residual structure and the attention mechanism-based feature fusion method, the network's expression ability for different features is enhanced. In the image segmentation stage, the improved Otsu algorithm is used to realize adaptive threshold segmentation, and combined with the monocular depth estimation method based on deep learning, the accuracy of the combined component decomposition is improved. In the video sequence processing, the feature aggregation mechanism based on spatiotemporal attention and the adaptive Kalman filtering algorithm are designed to realize stable tracking and prediction of the component motion state. These technical improvements improve the performance of power equipment component recognition from different levels, making the method adapt to complex practical application scenarios.

[0066] Further, the embodiments of the present application also disclose a semi-automatic marking device, comprising a processor, a memory and a system bus; the processor and the memory are connected through the system bus; the memory is used for storing one or more programs, the one or more programs comprising instructions which, when executed by the processor, cause the processor to perform any of the above methods.

[0067] Further, the embodiments of the present application also disclose a computer program product, which, when running on a terminal device, causes the terminal device to perform any of the above methods.

[0068] From the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the above-mentioned embodiment methods can be implemented by means of software and the necessary universal hardware platforms. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) execute the methods described in the various embodiments or some parts of the embodiments.

[0069] It should be noted that the embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the related parts can be referred to the method part.

[0070] It should also be noted that, in the embodiments of the present application, the terms such as first and second, etc. are used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between such entities or operations. Moreover, the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements in the list, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0071] The above description of disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined in the embodiments of the present application can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown in the embodiments of the present application, but will conform to the widest scope consistent with the principles and novel features disclosed in the embodiments of the present application.

Claims

1. A method for semi-automatic labeling in the field of power based on target detection, characterized in that, The method comprises the steps of: Importing the collected three-dimensional modeling data of the combined components of the power equipment into a rendering engine to obtain multi-angle rendering images within a preset angle range; Using a deep convolutional neural network to extract the contour information of the rendering images, and fusing and training the features through a multi-scale feature extraction module and an adaptive weighted residual connection structure; Performing perspective transformation and adaptive threshold segmentation on the detected combined components after training, using a region growing algorithm to complete component segmentation, and establishing a hierarchical relationship between the combined components and the independent components; Extracting power equipment video frames at a preset frame rate and performing preprocessing, and performing semi-automatic fine-tuning on the model detection results obtained based on the preprocessed images to obtain labeled data.

2. The method of claim 1, wherein, In the step of importing the collected three-dimensional modeling data of the combined components of the power equipment into a rendering engine to obtain multi-angle rendering images within a preset angle range, the method comprises the steps of: Collecting three-dimensional modeling data of the combined components of the power equipment, and importing the three-dimensional modeling data into a rendering engine for parameter setting; In a spherical coordinate system, taking the combined components as the center, rendering is performed within a preset elevation angle range and azimuth angle range respectively to obtain multi-angle rendering images; Setting the diffuse reflection coefficient, the specular reflection coefficient and the roughness of the rendering images, and setting the corresponding metallicity parameters according to the material type; Setting a main light source and an auxiliary light source based on the ambient light, wherein the brightness of the main light source is greater than that of the ambient light, and the brightness of the auxiliary light source is less than that of the main light source; Performing brightness adjustment, contrast adjustment, saturation adjustment and geometric transformation processing on the rendering generated images. In the step of using a deep convolutional neural network to extract the contour information of the rendering images, and fusing and training the features through a multi-scale feature extraction module and an adaptive weighted residual connection structure, the method comprises the steps of:

3. The method of claim 1, wherein, Extracting the contour information of the combined components through an image segmentation algorithm, and labeling the segmented combined component regions; Configuring a deep learning training environment, and constructing a deep convolutional neural network architecture containing a multi-scale feature extraction module; Setting a deformable convolution module in the network to process the geometric transformation of the target, and generating a candidate box through a region proposal network; Using an adaptive weighted residual connection structure to calculate weight coefficients for the residual branch and the identity mapping branch respectively; Using a multi-scale feature aggregation method based on an attention mechanism to perform weighted fusion on feature maps of different scales. In the step of performing perspective transformation and adaptive threshold segmentation on the detected combined components after training, using a region growing algorithm to complete component segmentation, and establishing a hierarchical relationship between the combined components and the independent components, the method comprises the steps of: Performing perspective transformation on the detected combined component regions to convert the inclined viewing angle into an orthographic view; 4. The method of claim 1, wherein, Performing adaptive histogram equalization processing on the orthographic view, and using an edge detection operator to extract the component contour; Using an improved Otsu algorithm to perform adaptive threshold segmentation on the combined component regions; Using a region growing algorithm to segment the components, and recording the relative position relationship of each independent component; Labeling the independent components obtained by segmentation, and establishing a hierarchical relationship between the combined components and the independent components. In the step of performing perspective transformation and adaptive threshold segmentation on the detected combined components after training, using a region growing algorithm to complete component segmentation, and establishing a hierarchical relationship between the combined components and the independent components, the method comprises the steps of: ​ ​ 5. The method of claim 1, wherein, ​ The power equipment video frames are extracted and preprocessed according to a preset frame rate, and the model detection results obtained based on the preprocessed images are semi-automatically fine-tuned to obtain labeled data, including: Video frames are extracted at a preset frame rate, and selected frames are preprocessed using Gaussian blur; The motion of the target component is tracked through feature point matching, and a ratio test method is used to filter mis-matched points; A spatiotemporal attention-based feature aggregation method is used to process the video frame sequence; A multi-target tracking algorithm based on Kalman filtering is used for motion prediction; The detection results are visualized and saved as labeled data containing timestamps.

6. The method of claim 4, wherein, Among them, An improved Otsu algorithm is used for adaptive threshold segmentation of the combined component region, including: The combined component region image is divided into multiple sub-blocks, and the local optimal threshold value of each sub-block is calculated; The pixels in the sub-blocks are divided into foreground pixels and background pixels according to the local optimal threshold value, and the foreground pixel ratio and background pixel ratio are calculated respectively; The average gray values of the foreground pixels and background pixels are calculated, and the optimal segmentation threshold value is determined based on the average gray values; The local optimal threshold values of the sub-blocks are median filtered, and the filtering uses a preset size of the filtering window; The threshold values of adjacent sub-blocks are smoothly transitioned through an interpolation method, and the interpolation weight is inversely proportional to the cube of the distance.

7. The method of claim 5, wherein, Among them, A spatiotemporal attention-based feature aggregation method is used to process the video frame sequence, including: The feature map sequence of the current frame and a preset number of frames before and after it is extracted, and an inter-frame correlation matrix of the feature map is constructed; Feature transformation is performed on the feature map sequence, and the transformed features are obtained through convolution operation; The attention weight in the time dimension is calculated based on the correlation matrix, and the weight is adjusted through a preset temperature parameter; The attention weight and the transformed features are weighted to obtain the feature aggregation result.

8. The method of claim 5, wherein, Among them, A multi-target tracking algorithm based on Kalman filtering is used for motion prediction, including: A state vector containing the target center coordinates, aspect ratio, area, and velocity components is constructed; The state transition matrix and the observation matrix are set, wherein the state transition matrix contains the unit matrix and the time interval; The process noise covariance matrix is updated online through the maximum likelihood estimation method, and the state transition bias is calculated; The observation noise covariance matrix is updated online through the maximum likelihood estimation method, and the observation value and state estimation bias are calculated; When the target is occluded, the diagonal elements of the observation noise covariance matrix are increased to enhance the prediction weight of the state transition model.

9. A semi-automatic marking apparatus comprising: Processor, memory, system bus; The processor and the memory are connected through the system bus; The memory is used to store one or more programs, and the one or more programs include instructions, which, when executed by the processor, cause the processor to perform any of the above-mentioned target detection-based semi-automatic labeling methods in the power field.

10. A computer program product, characterised in that, The computer program product runs on a terminal device, and causes the terminal device to perform any of the above-mentioned target detection-based semi-automatic labeling methods in the power field. The computer program product runs on a terminal device, and causes the terminal device to perform any of the above-mentioned target detection-based semi-automatic labeling methods in the power field.