Method for augmenting small target data set in power transmission line
The small object data set is processed through masking algorithm and random erase algorithm, combined with the improved SRGAN model and higher-order degradation model, the problems of low detection accuracy of small object in transmission lines and unbalanced samples are solved, and the accuracy and resolution of small object detection are improved.
Patent Information
- Application Number
- CN202510226270.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, small target detection of transmission lines has problems such as low detection accuracy, unbalanced sample and difficulty in labeling, which makes it difficult for the model to accurately detect small targets.
Small object data sets are processed through masking algorithms and random erase algorithms, small object adversarial generation model is built, and high-resolution small object images are generated using improved SRGAN models and higher-order degradation models to enhance data set diversity and detection capabilities.
It improves the accuracy of small object detection under complex backgrounds and insufficient light conditions, enhances the resolution and realism of small object images, and improves the model's recognition effect of small objects.
Smart Images

Figure CN120375112A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data augmentation, and particularly relates to a method for augmenting a small target data set in a transmission line. Background Art
[0002] There are many types of small target detection objects in transmission lines. Insulators are one of the key target objects with the largest quantity, the most complex types, and are more easily affected by environmental light. According to the structural form of insulators, they can be divided into post insulators, suspension insulators, long rod insulators, pin insulators, etc.; according to the material structure of insulators, they can be divided into porcelain insulators, glass insulators, composite insulators, etc. In addition to insulator type targets, the inspection of transmission lines also needs to pay attention to important detection targets such as wire clamps and connecting fittings. Most transmission lines are in inaccessible areas. After observing a large number of transmission line pictures, it can be summarized that the image background can be generalized into multiple types such as tower conductor type, sky type, and dense vegetation type. Due to the influence of the shooting distance, pitch angle, aperture size, shooting focal length, and even shooting weather of the inspection UAV, insulator images present problems such as complex background, different light environments, and difficulty in determining the instance size threshold. According to the detection experience of transmission lines, small target detection has an obvious guiding role in long-distance positioning of detection targets, autonomous flight navigation of UAVs, rotation of the on-board gimbal of the flying UAV, and adjustment of the camera focal length. Therefore, the first step in small target detection of transmission lines is to construct a small target detection data set with diverse backgrounds and sufficient positive and negative examples.
[0003] However, existing general object detection data sets focus more on medium and large-sized targets, pay less attention to small target types, the distribution of small target instances in the same data set is uneven, and the annotation difficulty is relatively large. At the same time, small targets are more sensitive to the deviation caused by annotation errors, and it is difficult to introduce prior knowledge for learning. In the detection and positioning process, the general method is to determine whether the generated anchor box belongs to a positive sample or a negative sample through a set threshold. Therefore, if there is a large error between the calibrated anchor box and the true boundary of the small target, then for the same threshold of the same object detector, more small target prediction boxes will be discarded, resulting in the number of positive samples of small targets being much smaller than that of medium and large-sized positive samples, making the model pay more attention to the detection of medium and large-sized targets. This makes it difficult to accurately detect small targets in the detection of transmission lines. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for augmenting a small target data set in a transmission line to overcome the above-mentioned defects existing in the prior art.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] On the one hand, the present invention provides an augmentation method for small target datasets in transmission lines, comprising the following steps:
[0007] Step S1: Obtain the basic dataset of the transmission line. According to the proportion of the target to be detected in the transmission line in the basic dataset on the image of the basic dataset, divide the basic dataset into a training sample and a small target dataset;
[0008] Step S2: Process the small target dataset through a masking algorithm and a random erasing algorithm to obtain an intermediate dataset;
[0009] Step S3: Construct a small target adversarial generation model, train the small target adversarial generation model with the training sample, and use the trained small target adversarial generation model to reconstruct and augment the images in the intermediate dataset to obtain an augmented dataset.
[0010] Further, the basic dataset includes high-definition inspection images generated by the autonomous inspection of transmission lines by drones, and the targets to be detected include composite insulators, glass insulators, connecting fittings, and wire clips.
[0011] Further, the division of the basic dataset into a training sample and a small target dataset includes:
[0012] Preprocess the high-definition inspection images of the basic dataset, input the preprocessed images into a pre-trained object detection model, output the bounding box information of the targets to be detected in each image, calculate the proportion of the targets to be detected in each image relative to the image according to the bounding box information of the targets to be detected in each image, obtain the proportion of the targets to be detected in the transmission line in the basic dataset on the image of the basic dataset, and make a judgment according to the proportion of each high-definition inspection image. If the proportion is less than or equal to the first preset value, divide the corresponding high-definition inspection image into the small target dataset. If the proportion is greater than the first preset value, divide the corresponding high-definition inspection image into the training sample.
[0013] Further, the calculation of the proportion of the targets to be detected in each image relative to the image has the following formula:
[0014]
[0015] where Per i is the proportion of the target to be detected in the i-th high-definition inspection image relative to the image, n i is the number of targets to be detected in the i-th image, is the upper left coordinate of the bounding box of the j-th target to be detected in the i-th high-definition inspection image, is the lower right coordinate of the bounding box of the j-th target to be detected in the i-th high-definition inspection image, (x i ,yi ) is the lower right corner coordinate of the i-th high-definition inspection image, where the upper left corner coordinate of the high-definition inspection image is (0, 0).
[0016] Furthermore, the small target dataset is processed by the mask algorithm and the random erasing algorithm, specifically including:[[]]
[0017] Each image in the small target dataset is processed using the mask algorithm. The specific steps are as follows:[[]]
[0018] Randomly select the unit length d = random(d min , d max ), where d is the length size of a mask unit, d min , d max is the upper and lower limits of the value of the preset unit length d. Randomly select r = random(radio, 1 - radio), where r is the retention rate of the input image. Randomly select δ x (δ y ) = random(0, d - 1), where δ x and δ y are the horizontal and vertical distances between the upper left corner of the mask area and the image edge respectively. Determine the four corner coordinates (x x , x y ), y min , y max ), y min , y max ) of the mask area according to (r, d, δ x , δ y ). Crop the original image of the mask area from the input image and fill it with random pixels to obtain the image processed by the mask algorithm;
[0019] Process the image processed by the mask algorithm in the small target dataset using the random erasing algorithm. The specific steps are as follows:[[]]
[0020] Select the image I processed by the mask algorithm in the small target dataset and obtain the width W, height H, and image area S of the image;
[0021] Randomly select the area S of the erasing area e ← Rand(s l , s h ) × S, where s l , s h are the upper and lower thresholds of the rectangular area of the preset random erasing area;
[0022] Randomly select the aspect ratio r of the erasing area e ← Rand(r1, r2), where r1 and r2 are the upper and lower thresholds of the aspect ratio of the preset random erasing area;
[0023] According to the area S of the erasure region e and the aspect ratio r e , calculate the height H of the erasure region e and the width W e :
[0024] According to the height H of the erasure region e and the width W e Obtain the position (x e , y e ) of the randomly selected erasure region in the image e x ← Rand(0, W - W e ), y e ← Rand(0, H - H e );
[0025] If the erasure region is completely within the image, i.e., x e + W e ≤ W and y e + H a ≤ H, set the pixel values of the region (x e , y e , x e + W e , y e + H e ) in the image I to random values Rand(0, 255), and output the image processed by the random erasure algorithm;
[0026] Use the image processed by the random erasure algorithm as the intermediate data set.
[0027] Furthermore, the small target adversarial generation model is an improved SRGAN model, including a generator and a discriminator. The improvements include: removing all batch normalization BN layers in the generator, replacing the original basic block in SRGAN with a residual dense block RRDB that fuses multi-level residual networks and dense connections in the generator, introducing a residual scaling coefficient β in the generator, using a pixel cancellation buffer operation in the generator to reduce the network space size and expand the channel size, replacing the standard discriminator D with a relative average discriminator RaD in the discriminator, adopting a U-Net type network design with skip connections in the discriminator, and introducing spectral normalization in the discriminator; the U-Net type network provides real value feedback for each pixel of the generator and generates accurate gradient feedback.
[0028] Furthermore, training the small target adversarial generation model with the training samples specifically includes:
[0029] Processing the images in the training samples through a high-order degradation model to generate low-resolution training images;
[0030] Using the generated low-resolution training images as the input and the original images in the training samples as the target, train the small object adversarial generation model.
[0031] Furthermore, the high-order degradation model is a second-order degradation model, including a first degradation process and a second degradation process;
[0032] The processing of the images in the training samples by the high-order degradation model specifically includes:
[0033] Convolve the images in the training samples using a generalized Gaussian blur kernel to obtain blurred images, downsample the blurred images to obtain low-resolution images, add Poisson noise to the downsampled images to obtain low-resolution images with noise, and perform JPEG compression on the low-resolution images with noise to obtain the first degraded images; Process the first degraded images through a sinc filter to simulate ringing effects and overshoot artifacts; Process the images processed by the sinc filter through further blur processing, downsampling operations, noise addition, and JPEG compression to generate low-resolution training images.
[0034] Furthermore, the training of the small object adversarial generation model by using the generated low-resolution training images as the input and the original images in the training samples as the target specifically includes:
[0035] Input the low-resolution training images into the generator of the small object adversarial generation model. The generator extracts and reconstructs features through a multi-level residual dense block (RRDB) and outputs high-resolution training images x f ;
[0036] Input the high-resolution training images x f and the original images x r in the training samples into the relative average discriminator (RaD). The relative average discriminator RaD discriminates between the images output by the generator and the real images, and generates the probability that the original images x r in the training samples are more real than the generated high-resolution training images x f and the probability that the high-resolution training images x f are more fake than the original images x r in the training samples. The formula is:
[0037]
[0038] where D Ra (x r , x f ) is the probability that the original images x r in the training samples are more real than the generated high-resolution training images x f , and DRa (x f , x r ) is the probability that the high-resolution training image x f is more fake than the original image x in the training samples r . C(x r ) and C(x f ) are the original outputs of the discriminator for the original image x in the training samples r and the high-resolution training image x f respectively. σ is the Sigmoid activation function represents the mean of the output of the high-resolution training image x f ; represents the mean of the output of the original image x in the training samples r ;
[0039] Use the pre-trained network VGG to process the original image x in the training samples r and the high-resolution training image x f to extract the features of the original image x in the training samples r and the high-resolution training image x f ;
[0040] According to the D Ra (x r , x f ) and D Ra (x f , x r ) output by the Relative Average Discriminator RaD and the features of the original image x in the extracted training samples r and the high-resolution training image x f , train the small target adversarial generation model through the loss function of the discriminator and the loss function of the generator
[0041] Furthermore, the loss function of the discriminator is as follows
[0042]
[0043] where is the loss function of the discriminator and respectively represent the expected values of the original image x in the training samples r and the high-resolution training image x f ;
[0044] The loss function of the generator is as follows
[0045]
[0046] where L Gis the loss function of the generator, L percep is the perceptual loss, is the discriminator-based adversarial loss, L1 is the content loss, λ and η are balance coefficients, φ i (x r ) is the original image x in the training samples r is the feature map of the i-th layer in the pre-trained network VGG, φ i (x f ) is the high-resolution training image x f is the feature map of the i-th layer in the pre-trained network VGG, is the Euclidean distance, represents the expected value of the high-resolution training image x f , T represents the set of small target positions in the image, P(x f ) ij represents the pixel value of the high-resolution training image x f at the position (i, j), P(x r ) ij represents the pixel value of the original image x in the training samples r at the position (i, j), ||.||1 is the L1 norm.
[0047] Compared with the prior art, the present invention has the following advantages:
[0048] (1) The present invention processes the small target dataset through the mask algorithm and the random erasing algorithm, simulates complex backgrounds and lighting changes, and enhances the diversity of the dataset. In this way, the model can more effectively cope with changing environmental conditions during training, improving the detection ability of small targets in complex backgrounds. Especially in the case of insufficient light or large environmental interference, it can still maintain high detection accuracy.
[0049] (2) The present invention adopts an improved SRGAN model. Through the adversarial training of the generator and the discriminator, the generated small target images are more realistic, improving the recognition effect of the target detection algorithm on small targets. The generator uses the design of multi-level residual dense blocks (RRDB) and the residual scaling coefficient β, greatly improving the ability to restore image details, effectively enhancing the resolution of small target images, and ensuring the accurate presentation of target details.
[0050] (3) The present invention simulates the degradation processes such as image denoising, blurring, and compression through a high-order degradation model, making the generated small target images closer to the low-resolution images in the real world. The images processed by the high-order degradation model not only enhance the model's learning ability for low-resolution images but also improve the model's reconstruction ability for small targets, overcoming the disadvantages of traditional methods in dealing with small targets in low-resolution environments.
[0051] (4) The present invention uses a Relative Average Discriminator (RaD) to replace the traditional discriminator, and optimizes the performance of the discriminator through the U-Net type network design and spectral normalization. The Relative Average Discriminator can better distinguish the differences between the generated image and the real image, improve the image generation quality of the generator, make the generated small target images have higher realism and resolution, and effectively improve the performance of small target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a flowchart of a method according to an embodiment of the present application;
[0053] Figure 2 is a statistical schematic diagram of the positions, lengths, and widths of targets of various sizes in a dataset according to an embodiment of the present application;
[0054] Figure 3 is a schematic diagram of an insulator after being processed by a masking algorithm according to an embodiment of the present application;
[0055] Figure 4 is a schematic diagram of an insulator after being processed by a random erasing algorithm according to an embodiment of the present application;
[0056] Figure 5 is a schematic diagram of the processing of a second-order degradation model according to an embodiment of the present application;
[0057] Figure 6 is a schematic diagram of the main structure of the generator G according to an embodiment of the present application;
[0058] Figure 7 is a schematic diagram of the structure of the Residual-in-Residual Dense Block (RRDB) with dense connections according to an embodiment of the present application;
[0059] Figure 8 is a schematic diagram of the discriminator D of the U-Net structure with spectral normalization introduced according to an embodiment of the present application;
[0060] Figure 9 is a schematic diagram of the change of the EL-ESRNet network Loss during training and its fitting curve according to an embodiment of the present application;
[0061] Figure 10 is a schematic diagram of the change of the generator G-Loss of the EL-ESRGAN network during training and its fitting curve according to an embodiment of the present application;
[0062] Figure 11 is a schematic diagram of the change of the discriminator D-Loss of the EL-ESRGAN network during training and its fitting curve according to an embodiment of the present application; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0064] Embodiment 1:
[0065] This embodiment provides an augmentation method for small target datasets in transmission lines, as Figure 1 shown, including the following steps:
[0066] Step S1: Obtain the basic dataset of the transmission line, and divide the basic dataset into a training sample and a small target dataset according to the proportion of the target to be detected in the transmission line in the basic dataset on the image of the basic dataset;
[0067] Step S2: Process the small target dataset through a mask algorithm and a random erasing algorithm to obtain an intermediate dataset;
[0068] Step S3: Construct a small target adversarial generation model, train the small target adversarial generation model with the training samples, and use the trained small target adversarial generation model to reconstruct and augment the images in the intermediate dataset to obtain an augmented dataset.
[0069] Further, the basic dataset includes high-definition inspection images generated by drones autonomously inspecting transmission lines, and the targets to be detected include composite insulators, glass insulators, connecting fittings, and wire clamps.
[0070] Further, the division of the basic dataset into a training sample and a small target dataset includes:
[0071] Preprocess the high-definition inspection images of the basic dataset, input the preprocessed images into a pre-trained object detection model, output the bounding box information of the targets to be detected in each image, calculate the proportion of the targets to be detected in each image relative to the image according to the bounding box information of the targets to be detected in each image, obtain the proportion of the targets to be detected in the transmission line in the basic dataset on the image of the basic dataset, and make a judgment according to the proportion of each high-definition inspection image. If the proportion is less than or equal to the first preset value, the corresponding high-definition inspection image is divided into the small target dataset. If the proportion is greater than the first preset value, the corresponding high-definition inspection image is divided into the training sample.
[0072] Further, the calculation formula for the proportion of the targets to be detected in each image relative to the image is:
[0073]
[0074] Among them, Per i is the ratio of the target to be detected in the i-th high-definition inspection image to the image, and n i is the number of targets to be detected in the i-th image. is the upper left corner coordinate of the bounding box of the j-th target to be detected in the i-th high-definition inspection image, is the lower right corner coordinate of the bounding box of the j-th target to be detected in the i-th high-definition inspection image, (x i , y i ) is the lower right corner coordinate of the i-th high-definition inspection image, where the upper left corner coordinate of the high-definition inspection image is (0, 0).
[0075] Furthermore, the small target dataset is processed by the mask algorithm and the random erasing algorithm, which specifically includes:
[0076] Each image in the small target dataset is processed using the mask algorithm. The specific steps are as follows:
[0077] Randomly select a unit length d = random(d min , d max ), where d is the length size of a mask unit, and d min , d max are the upper and lower limits of the preset value of the unit length d. Randomly select r = random(radio, 1 - radio), where r is the retention rate of the input image. Randomly select δ x (δ y ) = random(0, d - 1), where δ x and δ y are the horizontal and vertical distances between the upper left corner of the mask area and the image edge respectively. Determine the four corner coordinates (x x , x y , y min , y max , y min , y max ) of the mask area according to (r, d, δ x , δ y ). Crop the original image of the mask area from the input image and fill it with random pixels to obtain the image after being processed by the mask algorithm;
[0078] Process the image after being processed by the mask algorithm in the small target dataset using the random erasing algorithm. The specific steps are as follows:
[0079] Select the image I after being processed by the mask algorithm in the small target dataset, and obtain the width W, height H, and image area S of the image;
[0080] Randomly select the area S of the erasing area e ← Rand(sl , s h ) × S, where s l , s h is the upper and lower threshold of the rectangular area of the preset random erasure area;
[0081] Randomly select the aspect ratio r of the erasure area e ← Rand(r1, r2), where r1 and r2 are the upper and lower thresholds of the aspect ratio of the preset random erasure area;
[0082] According to the area S of the erasure area e and the aspect ratio r e , calculate the height H of the erasure area e and the width W e :
[0083] According to the height H of the erasure area e and the width W e Obtain the position (x e , y e ) of the randomly selected erasure area in the image e ← Rand(0, W - W e ), y e ← Rand(0, H - H e );
[0084] If the erasure area is completely within the image, that is, x e + W e ≤ W and y e + H a ≤ H, set the pixel values of the area (x e , y e , x e + W e , y e + H e ) in the image I to random values Rand(0, 255), and output the image processed by the random erasure algorithm;
[0085] Use the image processed by the random erasure algorithm as the intermediate data set.
[0086] This application uses the mask algorithm and the random erasure algorithm for the small target data set to occlude some feature points in the small target data set, so as to simulate the problem that small targets are easily occluded and feature information is lost in the actual detection task, in order to improve the data quality and data abundance of the small target detection data set. And this application performs super-resolution reconstruction on the small target data set through the generative adversarial network, greatly enhancing the image clarity of the small target detection data set of the transmission line and the interpretability of the detection instances, and at the same time providing sufficient training samples for the small target detection under complex and extreme backgrounds.
[0087] The basic dataset used in this application is 4,935 high-definition inspection images generated by the autonomous inspection of a 500 kV transmission line by an unmanned aerial vehicle. The image size is 5,472 * 3,078 pixels. According to the previous data cleaning, data annotation and verification, 1,209 composite insulators, 2,946 glass insulators, 2,983 connecting fittings and 2,052 clamps are sorted out. The above examples constitute the basic dataset of this experiment.
[0088] The size distribution of the objects in the basic dataset is as Figure 2 shown, and the length and width distribution of the annotation boxes of the data is as Figure 2 shown in the upper right corner. It can be found that the main detection target of this dataset is the vertical long strip annotation box. The coordinate distribution and size distribution of the objects in the basic dataset are as Figure 2 shown in the lower left corner and the lower right corner respectively. In terms of the dataset distribution, the number of composite insulator instances is the least, and there is an uneven distribution of clamps and fittings as important detection objects for small target detection.
[0089] Referring to the definition of the relative size of small target detection, objects with an image area below 0.12% of the region can be called relatively small targets. The position and length-width statistics of each size target in the dataset are as Figure 3 shown. The objects within the red frame are the targets that need to be augmented for relatively small targets. Among the 9,190 target instances, 1,969 instances belong to the category of relatively small targets.
[0090] In the process of constructing the basic dataset, mainly 1,969 small-size detection instances are used as the small target dataset, and the remaining instances are used as the training samples of the adversarial generation algorithm. After appropriate image augmentation means, they participate in the masking algorithm and random erasing as the data supplement of the small target dataset.
[0091] In the actual scenario of small target detection of transmission lines, there will be occlusion problems at different spatial positions between insulators and targets such as towers and conductors. Therefore, the masking algorithm and random erasing algorithm are used for data enhancement, effectively realizing the deletion of some existing information, that is, artificially creating factors such as spatial occlusion and light reflection in the actual space.
[0092] The masking algorithm adds a regularization term to the image detection network, which can effectively avoid network overfitting. According to the task requirements of small target detection of transmission lines, it is required to retain as much as possible the main body of the detection target and its context information, avoid the detector placing the detection target on irrelevant factors such as the background, and at the same time avoid over-retaining the main body information, in order to improve the robustness of the target detection and classification system and realize the artificial function of random target occlusion. Therefore, a balance between the deleted area and the retained area should be achieved in information deletion, and the masking algorithm can perfectly fit the task objectives of this application.
[0093] In practical applications, the masking algorithm removes the discontinuous regions of a pixel set. The removed pixel region M is determined by four parameters, namely (r, d, δ x , δ y ). The red line part represents a masking unit. r is defined as the ratio of the shorter gray edge in a masking unit to the unit length. δ x and δ y are respectively defined as the distances from a masking unit to the left side and the top of the image. d is defined as the length size of a masking unit.
[0094] r determines the retention rate of the input image. First, define the retention rate k of the mask M as
[0095]
[0096] where sum(M) is the total number of pixels retained in the mask. H is the height of the input image. W is the width of the input image.
[0097] The retention rate k of the mask M represents the area ratio between the retained image and the input image. The retention rate is a very important parameter for controlling the masking algorithm. Ignoring the incomplete units in the mask, we can get
[0098] k = 1 - (1 - r) 2 = 2r - r 2
[0099] In this application, r can take 0.3.
[0100] The length of the unit length d does not affect the retention rate k, but it determines the size of a mask M. When r is fixed, the relationship between the side length l of the mask M and d is
[0101] l = r × d
[0102] Since the size of the mask should be randomly generated to obtain sufficient images, the unit length d is randomly generated within a certain range, and its relationship is:
[0103] d = random(d min , d max )
[0104] where d min , d max : The upper and lower limits of the value of the unit length d. random represents the random function. A smaller unit length d can generate a smaller mask, but for small target recognition, the smaller the mask, the less information it blocks, which is ineffective for convolution operations. Therefore, an appropriate unit length d needs to be selected to ensure the effectiveness of the mask for data augmentation.
[0105] δx and δ y is the distance between the first complete mask unit and the image boundary. To ensure that the mask covers all possible cases, δ x and δ y should be random numbers within a certain range:
[0106] δ x (δ y ) = random(0, d - 1)
[0107] The pseudo - code of the mask algorithm is shown in Table 1 below.
[0108] Table 1 Pseudo - code of the mask algorithm
[0109]
[0110]
[0111] The insulator image after being processed by the mask algorithm is as Figure 3 shown.
[0112] The random erasing algorithm randomly selects a rectangular area during training and randomly replaces the pixels within the area with random colors. Compared with the mask algorithm, the random erasing algorithm can better simulate the occlusion of a single - region large target, while the mask algorithm is more suitable for the occlusion of multiple - region small - scale areas. The two operations together can realize the construction of an intermediate data set with different regional occlusion levels. The implementation of the random erasing algorithm is shown in Table 2:
[0113] Table 2 Pseudo - code of the random erasing algorithm
[0114]
[0115] The insulator image after being processed by the random erasing algorithm is as Figure 4 shown.
[0116] Furthermore, the small - target adversarial generation model is an improved SRGAN model, including a generator and a discriminator. The improvements include: removing all batch normalization (BN) layers in the generator; replacing the original basic block in SRGAN with a residual dense block (RRDB) that integrates a multi - level residual network and dense connections in the generator; introducing a residual scaling coefficient β in the generator; using a pixel cancellation buffer operation in the generator to reduce the network spatial size and expand the channel size; replacing the standard discriminator D with a relative average discriminator (RaD) in the discriminator; adopting a U - Net - type network design with skip connections in the discriminator; introducing spectral normalization in the discriminator; the U - Net - type network provides real - value feedback for each pixel of the generator and generates accurate gradient feedback.
[0117] Further, training the small target adversarial generation model using the training samples specifically includes:
[0118] Processing the images in the training samples through a high-order degradation model to generate low-resolution training images;
[0119] Using the generated low-resolution training images as the input and the original images in the training samples as the target to train the small target adversarial generation model.
[0120] Further, the high-order degradation model is a second-order degradation model, including a first degradation process and a second degradation process;
[0121] Processing the images in the training samples through the high-order degradation model specifically includes:
[0122] Convolving the images in the training samples using a generalized Gaussian blur kernel to obtain a blurred image, downsampling the blurred image to obtain a low-resolution image, adding Poisson noise to the downsampled image to obtain a low-resolution image with noise, and performing JPEG compression on the low-resolution image with noise to obtain a first-degraded image; Processing the first-degraded image through a sinc filter to simulate ringing effects and overshoot artifacts; Processing the image after passing through the sinc filter through blur processing, downsampling operations, noise addition, and JPEG compression again to generate low-resolution training images.
[0123] Further, using the generated low-resolution training images as the input and the original images in the training samples as the target to train the small target adversarial generation model specifically includes:
[0124] Inputting the low-resolution training images into the generator of the small target adversarial generation model. The generator extracts and reconstructs features through a multi-level residual dense block RRDB and outputs a high-resolution training image x f ;
[0125] Inputting the high-resolution training image x f and the original image x in the training samples r into the relative average discriminator RaD. The relative average discriminator RaD discriminates between the image output by the generator and the real image, generating the probability that the original image x in the training samples r is more real than the generated high-resolution training image x f and the probability that the high-resolution training image x f is more fake than the original image x in the training samples r . The formula is:
[0126]
[0127] where DRa (x r , x f ) is the original image x in the training samples r is more real than the generated high-resolution training image x f probability, D Re (x f , x r ) is the high-resolution training image x f is more fake than the original image x in the training samples r probability, C(x r ), C(x f ) are the original outputs of the discriminator for the original image x in the training samples r and the high-resolution training image x f respectively, σ is the Sigmoid activation function, represents the mean of the output for the high-resolution training image x f ; represents the mean of the output for the original image x in the training samples r ;
[0128] The pre-trained network VGG is used to process the original image x in the training samples r and the high-resolution training image x f to extract the features of the original image x in the training samples r and the high-resolution training image x f ;
[0129] According to the D Ra (x r , x f ) and D Ra (x f , x r ) output by the relative average discriminator RaD and the features of the original image x in the extracted training samples r and the high-resolution training image x f are used to train the small target adversarial generation model through the loss function of the discriminator and the loss function of the generator.
[0130] Furthermore, the loss function of the discriminator is:
[0131]
[0132] Among them, is the loss function of the discriminator, and respectively represent the expected values for the original image x in the training samples r and the high-resolution training image x f ;
[0133] The loss function of the generator is as follows:
[0134]
[0135] where L G is the loss function of the generator, L percep is the perceptual loss, is the adversarial loss based on the discriminator, L1 is the content loss, λ and η are balance coefficients, and φ i (x r ) is the original image x in the training sample r at the i-th layer feature map in the pre-trained network VGG, and φ i (x f ) is the high-resolution training image x f at the i-th layer feature map in the pre-trained network VGG, is the Euclidean distance, denotes the expected value of the high-resolution training image x f , T is the set representing the positions of small objects in the image, and P(x f ) ij denotes the pixel value of the high-resolution training image x f at the position (i, j), and P(x r ) ij denotes the pixel value of the original image x in the training sample r at the position (i, j), and ||.||1 is the L1 norm.
[0136] In view of the characteristics of small target images, such as low resolution, small pixel proportion, and limited shallow information, the idea of the super-resolution generative adversarial network (SRGAN) can be introduced for image reconstruction and dataset expansion, so that while improving the resolution of small target images, realistic texture details are retained. Using single-image super-resolution (SISR) as the deconstruction method for underlying vision problems, a single low-resolution image is restored to a high-resolution image.
[0137] To achieve a good balance between simplicity and effectiveness, this study adopted a second-order degradation process. Compared with traditional image degradation methods, high-order degradation modeling is more flexible and attempts to simulate the real degradation generation process. In the further synthesis process, a sinc filter is added to simulate the ringing and overshoot artifact noises commonly found in the field of super-resolution. For a specific network for small target recognition in transmission lines, due to a larger degradation space, the discriminator needs to be more powerful in distinguishing complex training outputs from real images, and at the same time, more accurate gradient feedback is required to enhance the local details of the adversarial generated images by the GAN network. However, the U-Net type network structure and complex multi-dimensional degradation will increase the instability of training. Therefore, spectral normalization (SN) of the GAN network is adopted to enhance the dynamic stability of training.
[0138] Therefore, when designing a small target adversarial generation algorithm based on adversarial generation and scene fusion, in this application:
[0139] (1) A high-order degradation process is proposed to simulate the actual image degradation process, simulate the actual degradation process of high-definition transmission line targets into low-resolution images, and add a filter to simulate signal noises such as common overshoot artifacts;
[0140] (2) Use a spectral-normalized U-Net discriminator as the discriminator for adversarial generation of the GAN network to improve the discriminator's ability while stabilizing the training dynamics;
[0141] (3) Based on the GAN network trained with pure synthetic data simulating low clarity, restore small target images with low clarity in the real world to achieve a directional mapping from low-clarity images to high-clarity images.
[0142] The high-order degradation model is improved based on the classical degradation model. In the classical degradation model, first, the reference real image y is convolved with the blur kernel k, then the downsampling operation is performed using the scale factor r, then the signal noise n is added, and the JPEG compression operation is performed to obtain the actual low-resolution image x.
[0143] x = D(y) = [(y # k) ↓ r + n] JPEG
[0144] where D(y) represents the entire degradation process, including blurring, downsampling, adding noise, and JPEG compression. y # k represents the convolution calculation of the reference real image y and the blur kernel k, and ↓ r represents the downsampling operation. n represents the signal noise, and JPEG represents the JPEG compression operation.
[0145] The isotropic or anisotropic Gaussian filter is usually used to model the blurring degradation process, so this linear blurring filter is also called the blurring convolution kernel. For a Gaussian blurring kernel k with a convolution kernel size of 2t + 1, (i, j) ∈ [-t, t] conforms to the Gaussian distribution form as follows:
[0146]
[0147] In the formula, k(i, j) represents the value of the blurring kernel k at the position (i, j), T represents the transpose of the matrix, Σ is the covariance matrix; C is the spatial coordinate; N is the normalization constant. The covariance matrix Σ can be expanded as follows:
[0148]
[0149] Where σ1 and σ2 are the eigenvalues of the covariance matrix, and θ is the rotation degree. R represents the rotation matrix. When σ1 = σ2, k is an isotropic Gaussian blurring kernel, otherwise k is an anisotropic Gaussian blurring kernel.
[0150] Although Gaussian blurring kernels are widely used in modeling blurring degradation, they may not be able to well approximate the real image blurring. To include more different kernel shapes, this application further adopts generalized Gaussian blurring kernels and flat-top distribution blurring kernels. The probability density function Pdf of the generalized Gaussian blurring kernel is:
[0151]
[0152] The probability density function of the flat-top distribution blurring kernel is:
[0153]
[0154] Where β is the shape parameter, and these blurring kernels can produce clearer output images than the narrow-sense Gaussian blurring kernels.
[0155] Downsampling is a basic operation for synthesizing low-resolution images. Usually, downsampling and upsampling are considered simultaneously, and its physical meaning is to adjust the image size. Common algorithms for upsampling and downsampling include nearest neighbor interpolation, area interpolation, bilinear interpolation, and bicubic interpolation algorithms, etc. Since nearest neighbor interpolation introduces misalignment problems, only area interpolation, bilinear interpolation, and bicubic interpolation operations are considered.
[0156] In terms of noise design, additive white Gaussian noise with a noise probability density function belonging to a Gaussian distribution and the noise intensity controlled by the standard deviation of the Gaussian distribution is used to replace traditional Gaussian noise. When there is independent sampling noise in each channel of the RGB image, the synthesized noise is color noise. Poisson noise follows a Poisson distribution and is usually used to approximately simulate statistical sensor noise, i.e., the variation in the number of photons sensed by the sensor at a given exposure level. The intensity of Poisson noise is proportional to the image intensity, and the noise at different pixels is independent.
[0157] JPEG compression is a commonly used lossy digital image compression technique. It first converts the image into the YCbCr color space and downsamples the chrominance channels. Then the image is divided into 8×8 blocks, each block is transformed using a two-dimensional discrete cosine transform (DCT), and then the DCT coefficients are quantized. JPEG compression is very likely to cause blocky artifacts, and the quality of its compressed image is determined by the quality factor q ∈ [0, 100], where the lower q indicates a higher compression ratio and worse image quality.
[0158] In this application, considering the algorithm complexity and operation time cost, a second-order degradation model is adopted to replace the traditional degradation model. This model consists of 2 repeated degradation processes, and each degradation process is a classical degradation model with the same data processing process but adjusted hyperparameters. The processing schematic diagram of the second-order degradation model is as Figure 5 shown.
[0159] This application uses a sinc filter to synthesize ringing and overshoot artifacts. The sinc filter kernel can be expressed as:
[0160]
[0161] where (i,j) are the filter kernel coordinates.
[0162] In the last step of the blurring process and synthesis, this application uses a sinc filter as a simulator for ringing and overshoot artifacts. At the same time, in order to cover a larger degradation space, the order of the last sinc filter and JPEG compression is randomly swapped to achieve that some images are sharpened first and then JPEG compressed, while the remaining images are JPEG compressed first and then sharpened.
[0163] In this application, in order to further improve the quality of the restored image of the GAN network and achieve a more realistic restoration of small insulator targets, as Figure 6 shown, two modifications are made to the structure of the generator G: (1) Remove all BN layers; (2) Replace the original basic block in SRGAN with a residual dense block (RRDB) that integrates a multi-level residual network and dense connections.
[0164] The architecture of generator G is similar to SRResNet. Different from SRResNet, the basic block is adjusted to a densely connected residual dense block (RRDB), and its structure is as Figure 7 shown.
[0165] In the tasks of super-resolution construction and deblurring, since the original BN layer depends on the mean and variance parameters of batch data when regularizing features, but mainly depends on the estimated mean and variance parameters of the test set during the test process. When the data difference between the two is large, the BN layer will introduce artifact interference and limit the generalization ability of the model. Removing the BN layer can obtain a stable and consistent training effect, improve the generalization ability, and reduce the computational complexity and the dependence of the GAN network on training memory.
[0166] The overall structure of the GAN network retains the architecture design of SRGAN, uses RRDB as the basic block, and improves the performance through more hidden layers and network connections. RRDB has a deeper and more complex structure compared to the original residual block in SRGAN, and its network capacity is more abundant due to dense connections. In addition, the residual scaling coefficient β is also introduced into the GAN network. Multiplying the residual by a constant between 0 and 1 before feeding it back to the forward function can effectively prevent the non-steady oscillation of the GAN network. At the same time, introducing initial parameters with smaller variances can also make the residual network more likely to tend to a steady state. Since ESRGAN is a heavy network, first use pixel unshuffle (the inverse operation of pixel shuffle) to reduce the network spatial size and expand the channel size, and then downsample the image and input it into the main ESRGAN architecture. Therefore, most calculations are performed in a small-resolution space, which can significantly reduce the consumption of GPU video memory and computing resources.
[0167] In the discriminator D part of the GAN network, an effective improvement is made by referring to Relativistic GAN. The difference in the discriminator output between the two is that the output value of the discriminator in SRGAN is the probability that the image generated by the generator is consistent with the real image, while the output value of the discriminator D in Relativistic GAN is the probability that the real image x r is more likely to be real than the fake image x f is.
[0168] To form real local textures, the VGG-style discriminator is further improved to a U-Net type network design with skip connections to generate accurate gradient feedback. As Figure 8As shown, U-Net outputs the true value of each pixel and can provide detailed per-pixel feedback to the generator. At the same time, the U-Net network structure and the complex degradation process also greatly increase the instability of training. Therefore, spectral normalization is introduced to stabilize the training dynamics. In addition, spectral normalization also helps to alleviate the boundary over-sharpness and artifacts brought by GAN training. Through these adjustments, a good balance is achieved between local detail enhancement and artifact suppression.
[0169] Experimental analysis: The training process of the GAN network is divided into two stages. First, a second-order image degradation model for peak signal-to-noise ratio based on the L1 loss function was trained, and the resulting model was named Electric-ESRNet (hereinafter referred to as EL-ESRNet). Its function is to degrade high-definition transmission line target images into low-definition target instances, thus forming paired transmission line target image pairs. Then, the EL-ESRNet model was used as the initialization of the generator, and combined with the L1 loss function, perceptual loss, and GAN loss to train the adversarial generative network Electric-ESRGAN (hereinafter referred to as EL-ESRGAN) suitable for small target detection of transmission lines.
[0170] This application uses the non-small target sequences in the dataset to train EL-ESRNet. The model was trained using 3 NVIDIA GeForce RTX 3090 graphics cards (with 24G of video memory). The system is Ubuntu 20.04 LTS 64-bit system, and the processor is Xeon(R) Gold 6242R CPU@3.10GHz×20, with 128G of memory, Python version 3.7.0, and PyTorch version 1.7.1.
[0171] The total batch size of training was set to 18, and the Adam optimizer was used as the optimizer to accelerate the convergence speed. EL-ESRNet was trained by iterating 500,000 times at a learning rate of 2×10-4. After that, EL-ESRGAN was trained by iterating 300,000 times at a learning rate of 1×10-4. The exponential moving average method (EMA) was used during training to maintain more stable training and better training performance. The training of EL-ESRGAN combines the L1 loss function, perceptual loss function, and GAN loss function, with weights of {1, 1, 0.1} respectively. In the perceptual loss function layer, it mainly reflects the loss of content and style, and the {conv1,..., conv5} feature maps (weights are {0.1, 0.1, 1, 1, 1}) before activation in the pre-trained VGG19 network are used as the perceptual loss.
[0172] The second-order degradation model's blurring kernels adopt Gaussian blurring kernels, generalized Gaussian blurring kernels, and plateau-distribution blurring kernels, with probabilities {0.7, 0.15, 0.15}. The size of the blurring kernel is randomly selected from {7, 9, … 21}, and the blurring standard deviation σ is sampled from [0.2, 3] (for the second degradation process, it is [0.2, 1.5]). For the generalized Gaussian blurring kernel and the plateau-distribution blurring kernel, the shape parameter β is randomly sampled from [0.5, 4] and [1, 2] respectively. Additionally, there is a 0.1 probability in the experiment of using the sinc function to increase ringing artifacts and overshoot artifacts, and a 0.2 probability of skipping the second blurring degradation. Regarding noise, Gaussian noise and Poisson noise are used with probabilities {0.5, 0.5}. The Sigma range of Gaussian noise and the scale of Poisson noise are set to [1, 30] and [0.05, 3] respectively (parameters [1, 25] and [0.05, 5] are used respectively during the second degradation process), and the generation probability of grayscale noise is set to 0.4. The JPEG compression quality factor p is set to [30, 95], and meanwhile, the application probability of the sinc filter after JPEG compression is 0.8. The detailed training parameters are shown in Table 3 and Table 4:
[0173] Table 3 Training Parameters of EL-ESRNet
[0174]
[0175] Table 4 Parameter Settings of EL-ESRGAN
[0176]
[0177]
[0178] To improve the training efficiency of EL-ESRNET and EL-ESRGAN, all degradation processes are accelerated by CUDA in PyTorch, and training pairs are dynamically synthesized. This application uses a training pair pool to increase the degradation diversity of batches. In each iteration, training samples are randomly selected from the training pairs to form a training batch. The training queue size is set to 120, and certain sharpening operations are performed on the target images during training.
[0179] Analysis of EL-ESRGAN training results: The transmission line insulator detection dataset is imported into the training model. Based on the pre-trained model, after 500,000 iterations in 75 hours with 500 epochs of augmented training for small targets on the transmission line, the loss function of the EL-ESRNet network gradually stabilizes, and the Loss value stabilizes at around 0.05, indicating that it has reached the training steady state, as Figure 9 shown.
[0180] Taking the above-trained EL-ESRNet network as a pre-trained model, after 72 hours of 407 epochs with a total of 300,000 iterations, the EL-ESRGAN network is obtained. Its loss function has shown convergence;
[0181] The three sub-loss values of the generator G and the generator G-Loss value remain stable after training. The training process is as Figure 10 shown.
[0182] Since GANLoss represents the ability of the generator to pass off fakes as real, that is, the closer the output of the fake image is to 1, the better. Eventually, GANLoss gradually tends to stabilize at 0.6. From Figure 10 it can be seen that the generator G-Loss has reached a steady state and the training of EL-ESRGAN can be carried out.
[0183] In EL-ESRGAN, the Loss value of the discriminator D is divided into the Loss value Real-Loss for the real images generated by the generator and the Loss value Fake-Loss for the fake images generated by the generator. Both values are stable at around 0.01. The training process is as Figure 11 shown.
[0184] Using the trained EL-ESRGAN to augment and magnify the intermediate dataset, the magnification results are as follows. The display effects all use the 4-fold magnification effect. In the construction of the small target dataset, augmented images with different magnification factors of 2 times, 3 times, and 4 times are actually generated.
[0185] The trained EL-ESRGAN is much higher than the original small target dataset of transmission lines in terms of image clarity. At the same time, it meets the requirements of object detection in terms of background fusion and edge sharpening of the target main body. Considering the problems of background fusion and feature extraction in small target detection of transmission lines, how to use the augmented method to ensure that the image quality is not distorted while realizing edge sharpening for easy feature extraction for the small target detection data where overexposure occurs due to direct sunlight on sunny days, and the glass insulator and the background space both show light color systems and are difficult to distinguish. EL-ESRGAN can also achieve the corresponding work task goals.
[0186] Under extreme conditions such as overexposure, the background of the original image is extremely easy to erode with the glass insulator, making the boundary of its small target blurred. After the adversarial generation by EL-ESRGAN, its edge features are more obvious, which has a necessary positive induction for small target detection in complex backgrounds.
[0187] However, due to the small size and uneven distribution of the dataset used for small target adversarial generation, there are certain deficiencies in the graphic features of small targets such as insulators, which may lead to extreme cases of distortion. According to manual review, such situations occur in the adversarial generation super-resolution of extremely small targets and account for a very small proportion in the overall augmented images, having little impact on the overall quality of the small target dataset.
[0188] There are many types and large quantities of small target detection objects in transmission lines, and they are easily affected by environmental illumination. After observing a large number of transmission line pictures, it can be summarized that the image background can be classified into multiple types such as tower conductor type, sky type, and dense vegetation type. Due to the influence of factors such as the shooting distance, pitch angle, aperture size, shooting focal length, and even shooting weather of the inspection drone, insulator images present problems such as complex backgrounds, different light environments, and difficulty in determining the instance size threshold. It is necessary to balance the distribution of the small target dataset through data augmentation means and expand the information abundance of the images after small target detection.
[0189] In this application, the EL-ESRGAN adversarial generation network is used to perform super-resolution reconstruction on the small target dataset, greatly enhancing the image clarity of the transmission line small target detection dataset and the interpretability of detection instances. At the same time, it provides sufficient training samples for small target detection in complex and extreme backgrounds. After the clarity of small targets is improved by the adversarial generation network, this application uses augmentation algorithms such as the mask algorithm and the random erasing algorithm on the small target dataset to artificially occlude some feature points in the small target dataset to simulate the problem that small targets are easily occluded and feature information is lost in actual detection tasks, so as to improve the data quality and data abundance of the small target detection dataset.
[0190] After the augmentation and reconstruction of small target data, the small target dataset with originally less than two thousand samples is expanded to a target dataset with nearly six thousand different image sizes and different occlusion degrees, providing a good data foundation for subsequent small target detection. At the same time, it also has an obvious guiding role in long-distance positioning of components in transmission line unmanned inspection, autonomous flight navigation of drones, and focusing of the on-board cameras of flying drones. This application constructs a small target detection dataset with diverse backgrounds and sufficient positive and negative examples, and also provides a super-resolution amplification tool for subsequent post-processing of clear detection targets.
[0191] Embodiment 2:
[0192] This embodiment provides an electronic device, including a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, it implements a method for augmenting a small target dataset in a transmission line as described in any one of the above.
[0193] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements an augmentation method for small target data sets in a transmission line as described in any one of the above.
[0194] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0195] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for augmenting a small target data set in a transmission line, characterized in that Including the following steps: Step S1: Obtain the basic dataset of the transmission line. According to the proportion of the target to be detected in the transmission line in the basic dataset on the image of the basic dataset, divide the basic dataset into a training sample and a small target dataset; Step S2: Process the small target dataset through a masking algorithm and a random erasing algorithm to obtain an intermediate dataset; Step S3: Construct a small target adversarial generation model, train the small target adversarial generation model with the training sample, and use the trained small target adversarial generation model to reconstruct and augment the images in the intermediate dataset to obtain an augmented dataset.
2. The augmentation method for small target data sets in transmission lines according to claim 1, characterized in that, The basic dataset includes high-definition inspection images generated by the autonomous inspection of the transmission line by an unmanned aerial vehicle. The targets to be detected include composite insulators, glass insulators, connecting fittings, and wire clamps.
3. The method for augmenting a small target data set in a transmission line according to claim 1 or 2, characterized in that, The division of the basic dataset into a training sample and a small target dataset includes: Preprocess the high-definition inspection images in the basic dataset, input the preprocessed images into a pre-trained object detection model, output the bounding box information of the targets to be detected in each image, calculate the proportion of the targets to be detected in each image relative to the image according to the bounding box information of the targets to be detected in each image, obtain the proportion of the targets to be detected in the transmission line in the basic dataset on the image of the basic dataset, make a judgment according to the proportion of each high-definition inspection image. If the proportion is less than or equal to the first preset value, divide the corresponding high-definition inspection image into the small target dataset. If the proportion is greater than the first preset value, divide the corresponding high-definition inspection image into the training sample.
4. A method for augmenting a small target data set in a transmission line, according to claim 3, characterized in that The calculation formula for the proportion of the targets to be detected in each image relative to the image is: Among them, Per i is the ratio of the target to be detected in the i-th high-definition inspection image relative to the image, and n i is the number of targets to be detected in the i-th image. is the upper left corner coordinate of the bounding box of the j-th target to be detected in the i-th high-definition inspection image, is the lower right corner coordinate of the bounding box of the j-th target to be detected in the i-th high-definition inspection image, and (x i , y i ) is the lower right corner coordinate of the i-th high-definition inspection image, where the upper left corner coordinate of the high-definition inspection image is (0, 0).
5. A method for augmenting a small target data set in a transmission line, according to claim 1, wherein The processing of the small target dataset through the masking algorithm and the random erasing algorithm specifically includes: Process each image in the small target dataset using the masking algorithm. The specific steps are as follows: Randomly select a unit length \(d = random(d min ,d max ), where \(d\) is the length of a mask unit, and \(d min ,d max are the upper and lower limits of the preset value of the unit length \(d\). Randomly select \(r = random(radio, 1 - radio)\), where \(r\) is the retention rate of the input image. Randomly select \(\delta x (\delta y ) = random(0, d - 1)\), where \(\delta x and \(\delta y are the horizontal and vertical distances between the upper left corner of the mask area and the image edge respectively. Determine the four corner coordinates \((x x ,x y ,y min ,y max ,y min ,y max ) of the mask area according to \((r, d, \delta x ,\delta y ). Crop the original image of the mask area from the input image and fill it with random pixels to obtain the image processed by the mask algorithm; Process the images in the small target dataset that have been processed by the masking algorithm using the random erasing algorithm. The specific steps are as follows: Select an image I in the small target dataset that has been processed by the masking algorithm, and obtain the width W, height H, and image area S of the image; Randomly select the area S of the erasure region e ←Rand(s l ,s h )×S, where s l ,s h is the upper and lower thresholds of the area of the preset random erasure region rectangle; Randomly select the aspect ratio r of the erasure area e ←Rand(r1, r2), where r1 and r2 are the upper and lower thresholds of the aspect ratio of the preset random erasure area; According to the area S of the erasure region e and the aspect ratio r e , calculate the height H e and width W e of the erasure region: According to the height H of the erasure area e and the width W e Obtain the position (x e , y e ) of the randomly selected erasure area in the image, where x e ← Rand(0, W - W e ), and y e ← Rand(0, H - H e ); If the erasure area is completely within the image, i.e., x e +W e ≤Wandy e +H e ≤H, set the pixel values of the area (x e ,y e ,x e +W e ,y e +H e ) in the image I to random values Rand(0, 255), and output the image processed by the random erasure algorithm; Use the image processed by the random erasing algorithm as the intermediate dataset.
6. The augmentation method of the small target data set in the transmission line according to claim 1, characterized in that, The small target adversarial generation model is an improved SRGAN model, including a generator and a discriminator. The improvements include: removing all batch normalization BN layers in the generator, replacing the original basic block in SRGAN in the generator with a residual dense block RRDB that integrates a multi-level residual network and dense connections, introducing a residual scaling coefficient β in the generator, using a pixel cancellation buffer operation in the generator to reduce the network spatial size and expand the channel size, replacing the standard discriminator D with a relative average discriminator RaD in the discriminator, adopting a U-Net type network design with skip connections in the discriminator, and introducing spectral normalization in the discriminator; the U-Net type network provides the true value feedback of each pixel for the generator and generates accurate gradient feedback.
7. A method for augmenting a small target data set in a transmission line, according to claim 1, characterized in that The training of the small target adversarial generation model with the training sample specifically includes: Process the images in the training samples through a high-order degradation model to generate low-resolution training images; Use the generated low-resolution training images as the input and the original images in the training samples as the targets to train the small object adversarial generation model.
8. A method for augmenting a small target data set in a transmission line, according to claim 7, characterized in that The high-order degradation model is a second-order degradation model, including a first degradation process and a second degradation process; The process of processing the images in the training samples through the high-order degradation model specifically includes: Convolve the images in the training samples using a generalized Gaussian blur kernel to obtain a blurred image, downsample the blurred image to obtain a low-resolution image, add Poisson noise to the downsampled image to obtain a low-resolution image with noise, and perform JPEG compression on the low-resolution image with noise to obtain a first degraded image; Process the first degraded image through a sinc filter to simulate ringing effects and overshoot artifacts; Process the image processed by the sinc filter through blur processing, downsampling operations, noise addition, and JPEG compression again to generate low-resolution training images.
9. A method for augmenting a small target data set in a transmission line, according to claim 7, characterized in that The process of using the generated low-resolution training images as the input and the original images in the training samples as the targets to train the small object adversarial generation model specifically includes: Input the low-resolution training image into the generator of the small target adversarial generation model. The generator extracts and reconstructs features through the multi-level residual dense block (RRDB) and outputs the high-resolution training image x f ; Input the high-resolution training image x f and the original image x in the training sample r into the relative average discriminator RaD. The relative average discriminator RaD discriminates between the image output by the generator and the real image, and generates the original image x in the training sample r The probability that it is more real than the generated high-resolution training image x f and the probability that it is more fake than the original image x in the training sample f are given by the formula: r where, D Ra (x r , x f ) is the probability that the original image x r in the training sample is more real than the generated high-resolution training image x f , and D Ra (x f , x r ) is the probability that the high-resolution training image x f is more fake than the original image x r in the training sample. C(x r ) and C(x f ) are the original outputs of the discriminator for the original image x r in the training sample and the high-resolution training image x f respectively. σ is the Sigmoid activation function, represents the mean of the output for the high-resolution training image x f , and represents the mean of the output for the original image x r in the training sample; Process the original image x in the training samples using the pre-trained network VGG r and the high-resolution training image x f to extract the features of the original image x in the training samples r and the high-resolution training image x f ; According to D output by the Relative Average Discriminator RaD Ra (x r , x f ), and D Ra (x f , x r ), and the original image x in the extracted training samples r and the high-resolution training image x f The features of are used to train the small target adversarial generation model through the loss function of the discriminator and the loss function of the generator.
10. A method for augmenting a small target data set in a transmission line, characterized in that, according to claim 9 The loss function of the discriminator is: Among them, is the loss function of the discriminator, and respectively represent the expected values of the original image x r and the high-resolution training image x f in the training samples; The loss function of the generator is: Among them, L G is the loss function of the generator, L percep is the perceptual loss, is the discriminator-based adversarial loss, L1 is the content loss, λ and η are balance coefficients, and φ i (x r ) is the original image x r in the i-th layer feature map of the pre-trained network VGG, φ i (x f ) is the high-resolution training image x f in the i-th layer feature map of the pre-trained network VGG, is the Euclidean distance, represents the expected value of the high-resolution training image x f , T is the set representing the positions of small targets in the image, P(x f ) ij represents the pixel value of the high-resolution training image x f at the position (i, j), P(x r ) ij represents the pixel value of the original image x r in the training sample at the position (i, j), and ∥.∥1 is the L1 norm.