A high-resolution power component defect image generation method

By combining the YOLO or DETR model with the CycleGAN style transfer network to generate large-resolution defect images of power components, the problem of generating high-resolution defect images of power components is solved, and the accuracy and generalization performance of the defect detection model are improved.

CN117315226BActive Publication Date: 2025-10-17WUHAN HUAZHONG KUANGTENG OPTICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311180514.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-13
Publication Date
2025-10-17
Estimated Expiration
2043-09-13

AI Technical Summary

Technical Problem

Existing technologies have difficulty in efficiently generating high-resolution defect images of electrical components, which affects the performance of defect detection models.

Method used

The target detection based on the YOLO series or DETR series models is combined with the CycleGAN style transfer network to generate large-resolution defect images of power components. The component position is determined through target detection, and the style transfer model is used to convert normal images into defective images.

Benefits of technology

The sample set diversity and breadth of the power component defect detection model are improved, and the accuracy and generalization performance of the detection model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315226B_ABST
    Figure CN117315226B_ABST
Patent Text Reader

Abstract

The application discloses a kind of high-resolution power component defect image generation methods, input a normal power component image, first using power component target detection model determines the position of each power component in image, then it is cut out, the rest part is as background map, next using style transfer network converts normal power component into abnormal power component, finally abnormal power component is put back to background map, so high-resolution power component defect image is obtained;The defect image generation method of the patent can efficiently generate a variety of types of power component defect map, improve the diversity and universality of sample set, improve the accuracy and generalization performance of subsequent defect detection model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of image recognition, and particularly relates to a large-resolution power component defect image generation method. BACKGROUND

[0002] A large amount of high-quality data is the basis of a neural network algorithm model, and so is defect detection of power components. However, power defects are subdivided into many categories, and the forms of defects are various. Compared with normal sample images, defect samples are very rare, which seriously affects the performance of a defect detection model.

[0003] Using a deep neural network to synthesize defect samples is a solution, however, the resolution of images shot by a UAV inspection is often large, up to 4096*2160 or even higher, and how to generate such high-resolution power component defect images becomes a difficult problem. SUMMARY

[0004] In view of the above defects of the prior art, the application provides a large-resolution power component defect image generation method.

[0005] The technical scheme adopted by the application to solve the technical problem is: a large-resolution power component defect image generation method, comprising the following steps:

[0006] S1, collecting power component sample images including normal images and defect images shot by a UAV inspection, constructing a data set that can be trained by a deep neural network, constructing and training a power component target detection model based on a YOLO series model or a DETR series model, and outputting the category and position of an image target;

[0007] S2, inputting a power component normal image img into the power component target detection model, accurately identifying the power component in the image by the training of the data set and marking the position of the power component, outputting the category of the power component, marking the left upper corner coordinates (x1, y1) and the right lower corner coordinates (x2, y2) of the rectangular frame of the power component in the normal image img, thereby determining the position of the power component, and taking the surrounding area after the rectangular frame is cut as a background image;

[0008] S3, constructing a style transfer model based on CycleGAN or other models, wherein the CycleGAN comprises a style transfer network composed of two generators G1 and G2 and two discriminators D1 and D2, the generator G1 converts an input image x of style X into an image G1(x) of style Y, the image G1(x) is input into the generator G2 again, and is restored into an image x' of style X, the discriminator D1 judges whether the input is an image x of style X, and the discriminator D2 judges whether the input is an image y of style Y;

[0009] S4, determine the size W of the input image of the style transfer model, train the style transfer network from the normal power component to the defective power component: if x2-x1>W or y2-y1>W, the power component cannot be converted from normal to abnormal, return to step S2, otherwise, take the rectangular frame where the power component is located as the center, crop a W*W image area noted as img_crop, and use it for subsequent style transfer, and record the center coordinates (x c , y c ) of the cropped area;

[0010] S5, input the cropped image img_crop into the style transfer network to obtain a defective power component image img_crop';

[0011] S6, according to the center coordinates (x c , y c ), replace the cropped image img_crop area in the original image with the defective power component image img_crop', and return to the background image to obtain a new cropped image img', which is a large-resolution power component defect image.

[0012] Further, the YOLO model is composed of a feature extractor and a detection head, the feature extractor extracts deep semantic features of the input image using a deep convolutional network, and the detection head is responsible for determining and outputting the category and position of the image target.

[0013] Further, the DETR series model uses a Transformer instead of a convolutional neural network, and detects the target through a multi-head attention mechanism to output the category and position of the image target.

[0014] The large-resolution power component defect image generation method, the power component includes a tower, a tower base, an insulator on the tower, a wire, a screw, and a pin.

[0015] The large-resolution power component defect image generation method, the normal image and the defective image are the entire image containing the power component and its surrounding environment.

[0016] The beneficial effects of the present application are that the defect image generation method of the present patent can efficiently generate various types of power component defect images, improve the diversity and universality of the sample set, and improve the accuracy and generalization performance of the subsequent defect detection model. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 An example normal image img of the tower foundation is accurately labeled by the detection model of the present application.

[0018] Figure 2 The rectangular image region img crop tailored for the present application is shown in the figure;

[0019] Figure 3 The tailored image img crop' generated by the style transfer network is shown in the figure;

[0020] Figure 4 The large resolution ground subsidence map generated by the embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to make the above features and advantages of the present application more obvious and easy to understand, the following specific examples are described in detail below with the help of the accompanying drawings.

[0022] The present application discloses a large resolution power component defect image generation method based on target detection and style transfer. Taking power tower base normal image and subsidence image as samples, the specific implementation process of the embodiment is described, and the steps are as follows.

[0023] S1, collect power component sample images including normal images and defect images taken by unmanned aerial vehicle inspection, construct a data set that can be trained by deep neural network, construct and train a power component target detection model based on YOLO series model, wherein the power components include tower, tower base (ground), insulator on tower, wire, screw, pin, etc., and the normal image and defect image are the whole image containing the power component and its surrounding environment, the defect image is the image of the power component facing safety hazards due to changes in the surrounding environment or its own state, such as broken insulator, tower base subsidence (sinking, subsidence), wire wrapped with plastic bag or cloth and other foreign matters, at this time the tower base and wire itself have no problem but there is a hidden danger in future normal operation.

[0024] The YOLO v5 target detection model used in the embodiment is composed of a feature extractor and a detection head. The feature extractor uses a deep convolutional network to extract deep semantic features of the input image, and the detection head is responsible for determining and outputting the category and position of the image target.

[0025] S2, input a power component normal image img into the power component target detection model, so that the target detection model can accurately identify the power component in the image and output the category of the power component through the training of the data set, mark the left upper corner coordinates of the rectangular frame where the power component is located in the normal image img as (x1, y1), and the right lower corner coordinates as (x2, y2), thereby determining the position of the power component, and taking the surrounding area remaining after the rectangular frame is cut as the background image.

[0026] Taking the tower base as an example, Figure 1The result example of the normal image input detection model of the iron tower base: the normal image img size is 2048x1536, it is difficult for the neural network to generate such a large resolution image finely, but the power component image cropped by the power component target detection model has the upper left corner coordinates of (802, 825) and the lower right corner coordinates of (1248, 1150), and the box is drawn according to the coordinates, and it can be seen that the power component target detection model successfully determines the position of the iron tower base.

[0027] S3, construct a CycleGAN-based style transfer model. Style transfer is a technology for converting a normal image into other image styles, such as converting a photo style into a cartoon style or converting a Van Gogh style into a Picasso style. In this embodiment, the normal image and the defect image of the power component are regarded as two different styles, and the neural network is trained to enable conversion between different styles, so that a defect image can be generated from a normal power component image.

[0028] The normal foundation and the sunken foundation are regarded as two different styles of images, and the CycleGAN model is used for style transfer. CycleGAN is a pioneering work of unpaired image style transfer method proposed by researchers at the University of California, Berkeley. It consists of two generators G1 and G2 and two discriminators D1 and D2. Generator G1 converts an input image x of style X into an image G1(x) of style Y, and G1(x) is input into generator G2 again to restore the image x' of style X. Discriminator D1 judges whether the input is an image x of style X, and discriminator D2 judges whether the input is an image y of style Y. Through training, the CycleGAN model can transfer between various illustration styles, such as conversion between zebra and horse, conversion between orange and apple, etc.

[0029] S4, prepare to input the iron tower foundation and the sunken iron tower foundation as two different styles of images into the style transfer network for style transfer: determine the size W of the input image of the style transfer model, train the style transfer network from the normal power component to the defective power component, in this embodiment, W is set to 512, that is, the input size of the style transfer network is 512x512: if x2-x1>W or y2-y1>W, the power component cannot be converted from normal to abnormal, return to step S2, otherwise, crop the WxW image region centered on the rectangular frame of the power component as img_crop, which is used for subsequent style transfer, and record the center coordinates (x c , y c ) of the cropped region; the cropped img_crop in this embodiment is shown in Figure 2 , and the center coordinates of the cropped region are (1025, 988).

[0030] S5, input the cropped image img crop into the style transfer network to obtain a power component image with defects img crop', and the power component image with defects img crop' has the same size as the cropped image img crop. The sunken tower foundation image generated by the style transfer model in this embodiment is as shown in Figure 3 It can be seen that the ground around the base has obviously sunk, and the style transfer model has completed the image transformation from a normal foundation to a sunken foundation.

[0031] S6, according to the center coordinates (x c , y c ), replace the cropped image img crop in the original image with the power component image with defects img crop' to obtain a new cropped image img', and return to the background image. The new cropped image img' is the high-resolution power component defect image generated in this embodiment.

[0032] The high-resolution foundation settlement image generated in this embodiment is as shown in Figure 4 . Figure 4 The sunken ground replaces the normal ground in Figure 1 , and is very realistic. Through this method, a large number of rich high-resolution power tower base settlement image samples can be automatically generated, the diversity and universality of the sample set are improved, and the accuracy and generalization performance of the subsequent defect detection model are improved.

[0033] As another embodiment, the power component target detection model based on the DETR model can also be constructed and trained. The DETR series model uses the Transformer to replace the convolutional neural network, detects the target through the multi-head attention mechanism, and outputs the category and position of the image target.

[0034] To sum up, the above is only a preferred embodiment of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for generating a high-resolution defect image of an electric component, characterized by: Includes the following steps S1 collects sample images of power components, including normal and defective images taken by drone inspections, to build a dataset suitable for deep neural network training. It then builds and trains a power component target detection model based on the YOLO model or the DETR model, outputting the type and location of the image targets. S2: Input a normal image img of an electric component into the electric component target detection model. The electric component target detection model is trained to accurately identify the electric component in the image and output the type of the electric component. The upper left corner coordinates of the rectangular box containing the electric component in the image img are marked as (x1, y1), and the lower right corner coordinates are marked as (x2, y2). The area around the rectangular box is used as the background image. S3: Build a style transfer model based on CycleGAN. CycleGAN includes a style transfer network consisting of two generators G1 and G2 and two discriminators D1 and D2. Generator G1 transforms an input image x with style X into an image G1(x) with style Y. G1(x) is then fed back into generator G2 to be restored to an image x′ with style X. Discriminator D1 determines whether the input image x is of style X, and discriminator D2 determines whether the input image y is of style Y. S4, determine the size W of the input image, and train the style transfer network from the normal power component to the defective power component: if x2-x1>W or y2-y1>W, return to step S2, otherwise crop the W×W image area centered on the rectangular box where the power component is located, record it as img_crop, and record the center coordinates of the cropped area (x c ,y c ); S5, inputting the cropped image img_crop into the style transfer network to obtain an image img_crop′ of the power component with defects of the same size; S6, according to the center coordinates (x c ,y c ), replace the cropped image img_crop with the defective power component image img_crop′, put the background image back, and obtain the new cropped image img′, which is the power component defect image with a large resolution.

2. The method for generating a high-resolution defect image of an electric component according to claim 1, characterized in that: The YOLO model consists of a feature extractor and a detection head. The feature extractor uses a deep convolutional network to extract deep semantic features of the input image, and the detection head is responsible for determining and outputting the type and location of the image target.

3. The method for generating a high-resolution defect image of an electric component according to claim 1, wherein: The DETR model uses Transformer instead of convolutional neural network, detects targets through a multi-head attention mechanism, and outputs the type and location of image targets.

4. A method for generating a high-resolution defect image of an electric component according to claim 2 or 3, characterized in that: The power components include an iron tower, an iron tower base, insulators on the iron tower, wires, screws and pins.

5. The method for generating a high-resolution defect image of an electric component according to claim 2 or 3, characterized in that: The normal image and defect image are entire images including the electrical component and its surrounding environment.

Citation Information

Patent Citations

  • Defect sample generation method and device, electronic equipment and storage medium

    CN111681162A

  • Defective insulator sample generation method and system based on style migration method

    CN112884758A