Image data augmentation method, device, equipment, medium and program product

Through the combination of feature encoding and decoding networks, high-quality and diverse augmented images of power equipment are generated, solving the problem of limited effects of traditional image data augmentation methods and improving the accuracy of power equipment image recognition and fault detection.

CN120374412APending Publication Date: 2025-07-25TIANJIN UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510250407.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The traditional image data augmentation method has limited effect on the augmentation of small sample data, and it is difficult to synthesize more effective augmented images while ensuring image quality. Especially in the complex operation and maintenance environment of power equipment, image recognition and fault detection face challenges.

Method used

By acquiring the original image and augmented guide image of the power equipment, using the feature encoding network to perform multi-scale feature extraction and feature fusion, combining the feature decoding network to generate augmented images, fusing the key information of the original image and augmented guide image, and generating high-quality and diverse augmented images.

Benefits of technology

It improves the generation efficiency and generation quality of augmented images, enhances the diversity of data sets, and improves the accuracy of image recognition and fault detection of power equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374412A_ABST
    Figure CN120374412A_ABST
Patent Text Reader

Abstract

The invention relates to an image data augmentation method and device, equipment, a medium and a program product. The image data augmentation method comprises the following steps: acquiring a to-be-processed image of power equipment; the to-be-processed image comprises an original image and an augmented guide image; image backgrounds of the augmented guide image and the original image are at least partially different; for each to-be-processed image, performing feature coding processing on the to-be-processed image based on the feature coding network to obtain a coding feature corresponding to the to-be-processed image; and based on a feature decoding network, performing feature decoding processing on the coding features corresponding to the original image and the coding features corresponding to the augmented guide image to obtain an augmented image of the power equipment. By adopting the steps, the generation efficiency and the generation quality of the augmented image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data augmentation, and particularly to an image data augmentation method, apparatus, device, medium, and program product. Background Art

[0002] With the continuous development of the power industry, the construction scale of power infrastructure has been expanding. In the actual operation and maintenance process of power equipment, especially in image analysis and image recognition of power equipment, many challenges are being faced. Due to the complex acquisition environment of operation and maintenance images, operation and maintenance images often contain different types of noise and have low image quality, which brings great difficulties to automatic image recognition and fault detection. At the same time, the fault types of power equipment are complex and diverse, and some faults often exhibit similar visual characteristics. Traditional image classification methods often face problems such as overfitting due to insufficient sample quantity. Based on this, image data augmentation technology has emerged. By augmenting image data, the scale and diversity of the training set can be expanded without actually increasing the data collection cost.

[0003] In traditional image data augmentation methods, image augmentation of the original image mostly relies on means such as image rotation, flipping, and scaling. However, the augmentation effect of these methods on small sample data is limited. How to synthesize more effective augmented images while ensuring image quality has become an urgent problem to be solved. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide an image data augmentation method, apparatus, device, medium, and program product, so as to improve the generation efficiency and generation quality of augmented images.

[0005] In a first aspect, this application provides an image data augmentation method, including:

[0006] Obtain a to-be-processed image of a power equipment; the to-be-processed image includes an original image and an augmentation guidance image; the augmentation guidance image is at least partially different from the image background of the original image;

[0007] For each to-be-processed image, perform feature encoding processing on the to-be-processed image based on a feature encoding network to obtain an encoded feature corresponding to the to-be-processed image;

[0008] Based on a feature decoding network, perform feature decoding processing on the encoded feature corresponding to the original image and the encoded feature corresponding to the augmentation guidance image to obtain an augmented image of the power equipment.

[0009] In one embodiment, feature encoding is performed on the image to be processed to obtain the encoded features corresponding to the image to be processed, including: performing multi-scale feature extraction on the image to be processed to obtain the low-level features and corresponding high-level features of the image to be processed; and fusing the low-level features and corresponding high-level features of the image to be processed to obtain the encoded features of the image to be processed.

[0010] In one embodiment, fusing the low-level features and corresponding high-level features of the image to be processed to obtain the encoded features of the image to be processed includes: extracting the key semantic information in the high-level features of the image to be processed; fusing the key semantic information with the low-level features of the image to be processed to obtain the corrected features of the image to be processed; and performing feature encoding on the corrected features and the key semantic information to obtain the encoded features of the image to be processed.

[0011] In one embodiment, performing feature encoding on the corrected features and the key semantic information to obtain the encoded features of the image to be processed includes: extracting the target space features in the corrected features; and enhancing the attention of the target space features according to the key semantic information to obtain the encoded features of the image to be processed.

[0012] In one embodiment, extracting the target space features in the corrected features includes: performing feature redrawing on the corrected features to obtain the initial space features; performing max pooling on the initial space features to obtain the first space features, and performing average pooling on the initial space features to obtain the second space features; and fusing the first space features and the second space features to obtain the target space features.

[0013] In one embodiment, the method further includes: performing feature extraction on the augmented image to obtain augmented features; performing feature extraction on the original image to obtain original features; generating a semantic perception loss according to the difference between the augmented features and the original features; and adjusting the network parameters of the feature encoding network and the feature decoding network according to the semantic perception loss.

[0014] In a second aspect, the present application further provides an image data augmentation device, including:

[0015] An acquisition module, configured to acquire the image to be processed of the power device; the image to be processed includes the original image and the augmented guiding image; the augmented guiding image is at least partially different from the image background of the original image;

[0016] An encoding module, configured to perform feature encoding processing on each image to be processed based on the feature encoding network to obtain the encoded features corresponding to the image to be processed;

[0017] A decoding module, configured to perform feature decoding processing on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guidance image based on a feature decoding network, to obtain an augmented image of the power equipment.

[0018] In a third aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0019] Obtain a to-be-processed image of the power equipment; the to-be-processed image includes an original image and an augmented guidance image; the augmented guidance image is at least partially different from the image background of the original image;

[0020] For each to-be-processed image, perform feature encoding processing on the to-be-processed image based on a feature encoding network, to obtain encoded features corresponding to the to-be-processed image;

[0021] Based on a feature decoding network, perform feature decoding processing on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guidance image, to obtain an augmented image of the power equipment.

[0022] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0023] Obtain a to-be-processed image of the power equipment; the to-be-processed image includes an original image and an augmented guidance image; the augmented guidance image is at least partially different from the image background of the original image;

[0024] For each to-be-processed image, perform feature encoding processing on the to-be-processed image based on a feature encoding network, to obtain encoded features corresponding to the to-be-processed image;

[0025] Based on a feature decoding network, perform feature decoding processing on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guidance image, to obtain an augmented image of the power equipment.

[0026] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0027] Obtain a to-be-processed image of the power equipment; the to-be-processed image includes an original image and an augmented guidance image; the augmented guidance image is at least partially different from the image background of the original image;

[0028] For each to-be-processed image, perform feature encoding processing on the to-be-processed image based on a feature encoding network, to obtain encoded features corresponding to the to-be-processed image;

[0029] Based on the feature decoding network, perform feature decoding processing on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guidance image to obtain the augmented image of the power equipment.

[0030] The above image data augmentation method, device, equipment, medium and program product obtain the original image of the power equipment and at least partially different augmented guidance images of the image background, so as to increase the diversity of the data set. By performing feature encoding processing on the image to be processed based on the feature encoding network, the key information in the original image and the augmented guidance image can be extracted. Since the information of the original image and the augmented guidance image is incorporated in the encoding and decoding processes, the generated augmented image not only retains the key features of the original image but also increases new diversity. At the same time, by introducing the encoding feature network and the decoding feature network, it helps to improve the generation efficiency and generation quality of the augmented image. Description of the Drawings

[0031] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0032] Figure 1 It is a schematic flowchart of the image data augmentation method in one embodiment;

[0033] Figure 2A It is a schematic flowchart of obtaining the encoded features of the image to be processed in one embodiment;

[0034] Figure 2B It is a schematic flowchart of the fusion process of the cross-fusion attention network in one embodiment;

[0035] Figure 3 It is a schematic flowchart of the adjustment steps of the network parameters in one embodiment;

[0036] Figure 4 It is a schematic structural diagram of the image data augmentation model in one embodiment;

[0037] Figure 5 It is a schematic diagram of the augmented image in one embodiment;

[0038] Figure 6 It is a schematic flowchart of the image data augmentation method in another embodiment;

[0039] Figure 7 It is a structural block diagram of the image data augmentation device in one embodiment;

[0040] Figure 8 Internal structure diagram of a computer device in an embodiment. Specific implementation manner

[0041] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0042] In one embodiment, as Figure 1 shown, an image data augmentation method is provided. In this embodiment, this method is exemplified by being applied to a terminal. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0043] S110. Obtain the image to be processed of the power equipment; the image to be processed includes the original image and the augmentation guidance image; the augmentation guidance image is at least partially different from the image background of the original image.

[0044] Optionally, the power equipment may include at least one of power generation equipment, power transmission equipment, power transformation equipment, power distribution equipment, and power consumption equipment, etc. Exemplarily, the power equipment may include at least one of poles, insulators, and wires, etc. The present application does not make any limitation on the specific type of the power equipment.

[0045] Among them, the original image, that is, the source semantic image, is a captured image with the power equipment as the source target. The augmentation guidance image is used to guide the original image for diversity augmentation, so as to generate the augmented image. The augmentation guidance image is at least partially different from the image background of the original image, and the source target of the augmentation guidance image and the original image may be the same or different.

[0046] Optionally, the original image and the augmentation guidance image of the power equipment can be obtained by a patrol device during the patrol process. Among them, the patrol device may include at least one of an unmanned aerial vehicle and a camera, etc. The present application does not make any limitation on the acquisition method of the image to be processed.

[0047] S120. For each image to be processed, perform feature encoding processing on the image to be processed based on the feature encoding network to obtain the encoded feature corresponding to the image to be processed.

[0048] Among them, the feature encoding network can be understood as a network for performing feature encoding on the image to be processed, and among them, the encoder can be composed of the feature encoding network.

[0049] In an alternative embodiment, the feature encoding network may include an original image feature encoding network and a guided image feature encoding network. Among them, the original image feature encoding network is used to perform feature encoding processing on the original image to obtain the encoded features corresponding to the original image; the guided image feature encoding network is used to perform feature encoding processing on the augmented guided image to obtain the encoded features corresponding to the augmented guided image.

[0050] In another alternative embodiment, the original image and the augmented guided image may share a feature encoding network, and through the feature encoding network, the encoded features corresponding to the original image and the encoded features corresponding to the augmented guided image are obtained respectively.

[0051] In an alternative embodiment, the feature encoding network may be a trained VGG16 (Visual Geometry Group 16) network. Among them, the VGG16 network is a deep convolutional neural network model. Exemplarily, before training the VGG16, the VGG16 can be pre-trained to facilitate reducing the training difficulty and improving the training efficiency. The present application does not make any limitation on the specific network structure and type of the feature encoding network.

[0052] In an alternative embodiment, the feature encoding network may include a cross-fusion attention network (FusionNet) for performing context encoding on features at different levels in the image to be processed, and increasing the differences between image features. Among them, the cross-fusion attention network models the relationship between features at different levels, reduces the information loss in the feature processing process, effectively extracts the background information in different dimensions, is conducive to improving the training and inference speed, and facilitates generating more realistic images.

[0053] Exemplarily, the feature encoding network may include an original image feature encoding network and a guided image feature encoding network, and the cross-fusion attention network is applied to the original image feature encoding network and the guided image feature encoding network, so as to ensure the consistency of the original image feature encoding network and the guided image feature encoding network and improve the system performance.

[0054] In an alternative embodiment, multi-scale feature extraction may be performed on the image to be processed to obtain the low-level features and the corresponding high-level features of the image to be processed; the low-level features and the corresponding high-level features of the image to be processed are fused to obtain the encoded features of the image to be processed. Specifically, the low-level features and the corresponding high-level features of the original image may be fused to obtain the encoded features of the original image; the low-level features and the corresponding high-level features of the augmented guided image may be fused to obtain the encoded features of the augmented guided image.

[0055] Among them, the low-level features can be understood as the basic visual attributes of the image to be processed, and can include at least one of image contours, image edges, image colors, image textures, and image shapes. The high-level features can be understood as semantic information or abstract information parsed or extracted from the image to be processed.

[0056] Among them, multi-scale feature extraction can be performed on the image to be processed based on the feature encoding network to obtain the low-level features and corresponding high-level features of the image to be processed.

[0057] It should be noted that by performing multi-scale feature extraction on the image to be processed, the diversity of image features can be improved, which is beneficial to improving the authenticity and augmentation effect of the generated augmented image in the subsequent augmented image generation process.

[0058] S130. Based on the feature decoding network, perform feature decoding processing on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guidance image to obtain the augmented image of the power equipment.

[0059] Among them, the feature decoding network can be understood as a network for performing feature decoding on the image to be processed. Among them, the decoder can be composed of the feature decoding network.

[0060] In an optional embodiment, the encoded features corresponding to the original image and the encoded features corresponding to the augmented guidance image can be feature-fused to obtain the augmented image of the power equipment.

[0061] In an optional embodiment, the augmented image can also be discriminated based on the image discrimination network and the augmented guidance image to obtain the discrimination result of the augmented image.

[0062] Among them, the image discrimination network can be understood as a network for discriminating the augmented image. Among them, the discriminator can be composed of the image discrimination network, and the discrimination result can include the authenticity discrimination data of the augmented image.

[0063] Exemplarily, the adversarial loss can be determined according to the discrimination result of the augmented image ; according to the adversarial loss , adjust the network parameters of networks such as the feature encoding network and the feature decoding network, so that the image quality and augmentation effect of the output augmented image are better.

[0064] Exemplarily, the semantic perception loss can be determined according to the augmented image and the original image , according to the semantic perception loss , adjust the network parameters of networks such as the feature encoding network and the feature decoding network, so that the image quality and augmentation effect of the output augmented image are better.

[0065] The above image data augmentation method obtains the original image of the power equipment and at least partially different augmented guiding images of the image background, so as to increase the diversity of the data set. By performing feature encoding processing on the image to be processed based on the feature encoding network, the key information in the original image and the augmented guiding image can be extracted. Since the information of the original image and the augmented guiding image is incorporated into the encoding and decoding processes, the generated augmented image not only retains the key features of the original image but also increases new diversity. At the same time, by introducing the encoding feature network and the decoding feature network, it helps to improve the generation efficiency and generation quality of the augmented image.

[0066] Based on the technical solutions of the above embodiments, the present application also provides an alternative embodiment. In this alternative embodiment, the encoding feature acquisition step of the image to be processed is refined.

[0067] See Figure 2A The encoding feature acquisition step of the image to be processed shown, includes:

[0068] S210. Extract the key semantic information in the high-level features of the image to be processed.

[0069] In an alternative embodiment, global average pooling can be performed on the high-level features of the image to be processed to eliminate redundant information in the high-level features, thereby reducing the network calculation amount. Then, convolution and non-linear normalization are sequentially performed on the high-level features to obtain the key semantic information in the high-level features.

[0070] Refer to Figure 2B The schematic diagram of the fusion process of the cross-fusion attention network shown. The key semantic information of the high-level features can be obtained by sequentially performing global average pooling, convolution (Convolution, Conv) on the high-level features, and non-linearly normalizing the convolved high-level features using the Sigmoid function. Among them, the Sigmoid function is used to normalize to the range of [0, 1].

[0071] It should be noted that the high-level features can obtain more important semantic information through cross-channel attention. Among them, the influence of each channel on the network can be determined, so that the network adaptively focuses on the key semantic information.

[0072] S220. Perform feature fusion on the key semantic information and the low-level features of the image to be processed to obtain the corrected features of the image to be processed.

[0073] Continue to refer to Figure 2B , global average pooling can be performed on the low-level features of the image to be processed to eliminate redundant information in the low-level features, thereby reducing the network calculation amount. Then, the low-level features and the key semantic information are subjected to feature fusion to obtain the corrected features of the image to be processed.

[0074] Exemplarily, the correction feature can be determined according to the following formula :

[0075] ;

[0076] wherein represents the Sigmoid function; represents the high-level feature; represents the low-level feature; represents convolution.

[0077] It should be noted that by fusing the key semantic information with the low-level features of the image to be processed, information interaction can be enabled between the low-level features and the high-level features, compensating for the loss of details in the extraction process.

[0078] S230. Feature-encode the correction feature and the key semantic information to obtain the encoded feature of the image to be processed.

[0079] In an optional embodiment, the target spatial feature in the correction feature can be extracted; the target spatial feature is attention-enhanced according to the key semantic information to obtain the encoded feature of the image to be processed.

[0080] Optionally, the correction feature can be feature-redrawn to obtain the initial spatial feature; the initial spatial feature is max-pooled to obtain the first spatial feature, and the initial spatial feature is average-pooled to obtain the second spatial feature; the first spatial feature and the second spatial feature are feature-fused to obtain the target spatial feature.

[0081] Continuing to refer to Figure 2B , the feature-redrawing of the correction feature can be achieved by performing Conv (Convolution), Bn (Batch Normalization), and ReLU (Rectified Linear Unit) processing on the correction feature to obtain the initial spatial feature.

[0082] Wherein, the first spatial feature is obtained by max-pooling the initial spatial feature; the second spatial feature is obtained by average-pooling the initial spatial feature; the first spatial feature and the second spatial feature are feature-fused and processed through the Sigmoid function to obtain the target spatial feature.

[0083] Continuing to refer to Figure 2B , the key semantic information can be multiplied by the target spatial feature to achieve attention enhancement and obtain the encoded feature of the image to be processed.

[0084] In an optional embodiment, for the feature encoding network corresponding to the original image and the feature encoding network corresponding to the augmented guidance image, the network parameters of the two can be independent of each other or shared with each other.

[0085] It can be understood that high-level features often contain a large amount of semantic information, and each pass can correspond to specific semantic information, but often more details are lost in the semantic information, while low-level features retain a large amount of texture and details in the image to be processed. Based on this, by extracting the key semantic information in the high-level features of the image to be processed and fusing the key semantic information with the low-level features of the image to be processed, information interaction can be carried out between the low-level features and the high-level features, compensating for the loss of details in the extraction process. Feature encoding is performed on the corrected features and key semantic information obtained by feature fusion, so that the fused features retain both the high-level semantic information of the image to be processed and the low-level texture detail information, and a good balance can be achieved between the two, which is beneficial to improving the generation effect of the subsequent generated augmented image.

[0086] Based on the technical solutions of the above embodiments, the present application also provides an optional embodiment, in which a step of adjusting network parameters is added.

[0087] See Figure 3 The steps of adjusting the network parameters shown include:

[0088] S310. Extract features from the augmented image to obtain augmented features; and extract features from the original image to obtain original features.

[0089] In an optional embodiment, the augmented features of the augmented image and the original features of the original image can be extracted based on a VGG16 feature extractor.

[0090] Exemplarily, the original features and the augmented features can be obtained according to the following formula:

[0091] ;

[0092] where represents the original image, represents the augmented image, and i represents the index of the image.

[0093] Exemplarily, the original features and the augmented features satisfy the following conditions:

[0094] ;

[0095] Among them, \(R\) represents the set of real numbers; \(N\) represents the number of characteristic data; \(C\) represents the number of channels; \(H\) represents the height of the feature map; \(W\) represents the width of the feature map.

[0096] S320. Generate a semantic perception loss according to the difference between the augmented feature and the original feature.

[0097] In an optional embodiment, the original feature can be subjected to feature mapping to obtain an original mapped feature; the augmented feature can be subjected to feature mapping to obtain an augmented mapped feature; according to the original mapped feature, a first channel autocorrelation matrix corresponding to the original image is determined; according to the augmented mapped feature, a second channel autocorrelation matrix corresponding to the augmented image is determined; according to the first channel autocorrelation matrix and the second channel autocorrelation matrix, the difference between the augmented feature and the original feature is determined.

[0098] Optionally, can be mapped to to obtain the original mapped feature and the augmented mapped feature . Among them, , represents the product of the height and width of the feature map. It should be noted that by performing feature mapping on the original feature and the augmented feature, the amount of calculation is reduced, and it is convenient to determine the difference between the augmented feature and the original feature.

[0099] Optionally, the first channel autocorrelation matrix and the second channel autocorrelation matrix can be determined based on the Ragham matrix operation. Exemplarily, the first channel autocorrelation matrix and the second channel autocorrelation matrix can be determined by the following formula:

[0100] ;

[0101] ;

[0102] Among them, represents the Ragham matrix operation, represents the original mapped feature, represents the augmented mapped feature.

[0103] Optionally, the semantic perception loss can be generated by the following formula:

[0104]

[0105] Among them, \(n\) represents the number of features; represents the first channel autocorrelation matrix; represents the second channel autocorrelation matrix.

[0106] S330. Adjust the network parameters of the feature encoding network and the feature decoding network according to the semantic perception loss.

[0107] Among them, when training the feature encoding network and the feature decoding network, the network parameters of the feature encoding network and the feature decoding network can be adjusted according to the semantic perception loss.

[0108] In an optional embodiment, the network parameters of the feature encoding network and the feature decoding network can be adjusted according to the semantic perception loss and / or the adversarial loss .

[0109] Optionally, the adversarial loss can be determined according to the following formula:

[0110] #timg# #timg#

[0111] Among them, represents the augmented image; represents the guided augmented image; represents the output probability that the discriminator D judges the augmented image to be real; represents the output probability that the discriminator D judges the guided augmented image to be real; represents the guided augmented image distribution and the guided augmented image sampled from it ; represents the augmented image sampled from the guided augmented image distribution ; .

[0112] Among them, the adversarial loss can be used to optimize the discriminator to improve its discrimination ability; the adversarial loss can be used to optimize the generator to minimize the probability that the augmented image is judged to be fake.

[0113] In one embodiment, as Figure 4 shown, the present application also provides an image data augmentation model, including a generator and a discriminator. Among them, the generator includes an encoder and a decoder.

[0114] Continuing to refer to Figure 4 , the encoder includes two VGG16 networks, which are respectively used for feature extraction of the original image X and the augmented guidance image Y. The encoder is also embedded with a cross-fusion attention network (FusionNet), which is used for feature fusion of the low-level features and high-level features in the original image X to obtain the encoded features of the original image; and feature fusion of the low-level features and high-level features in the augmented guidance image Y to obtain the encoded features of the augmented guidance image.

[0115] Among them, the decoder is used to perform feature decoding on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guidance image to obtain the augmented image of the power equipment. 。

[0116] Among them, the discriminator is used to determine the discrimination result of the augmented image and the augmented guidance image Y.

[0117] Continue to refer to Figure 4 , the image data augmentation model can determine the adversarial loss according to the discrimination result output by the discriminator; determine the semantic perception loss according to the augmented image and the original image ; according to the adversarial loss and the semantic perception loss , adjust the model parameters of the image data augmentation model.

[0118] Exemplarily, the original images and augmented guidance images of three categories of tower poles, insulators, and wires can be collected. 30 pairs of images are selected from each category as the training set; 20 pairs of images are selected from each category as the test set. Among them, each pair of images consists of the original image and the corresponding augmented guidance image. Optionally, the images in the training set can be standardized to facilitate the model's standardized learning. Exemplarily, the original image and the augmented guidance image can be cropped into fixed-size images of 256*256.

[0119] Exemplarily, the initial parameters of the model can be preset, and the initial parameters can include at least one of the number of iterations, the size of the convolutional kernel, etc. Put the training set samples and the corresponding label files into the specified folder and train the model. After the training reaches the convergence condition or the preset number of steps, the training is completed. Save the preprocessed and postprocessed pictures to the corresponding folders. Conduct a comparative analysis of the test data. Among them, the evaluation and analysis can be carried out by adopting the subjective evaluation of the augmented image, determining the objective index IS (Inception Score) and the classification result ACC (Accuracy) of the augmented image.

[0120] Based on the technical solutions of the above embodiments, the present application also provides a verification embodiment to verify the above image data augmentation method.

[0121] Exemplarily, in this embodiment, the Adam (Adaptive Moment) optimizer is used to train the above image data augmentation model, and the initial learning rate is 1*10 -4 , and it is reduced to 1*10 through the cosine annealing strategy -6。The image data augmentation model is trained on 256*256 patches, with a Batch Size of 4 and a total of 300 training rounds, and is tested under three categories: tower poles, insulators, and conductors. See Figure 5 The schematic diagram of the augmented images is shown, which respectively shows the augmented images corresponding to insulators, conductors, and tower poles. It can be seen that the augmented images generated based on the original images and the augmented guidance images have good authenticity. In addition, the objective index IS (Inception Score) reaches 1.0777, and the ACC (classification accuracy) reaches 96.45. Therefore, by using the image data augmentation method provided in this application, the obtained augmented images have both the authenticity seen by the human eye and the robustness for downstream tasks, presenting a good data augmentation effect. The real-time inference speed of the image data augmentation method designed in this application is relatively fast, and it has a high image deblurring efficiency.

[0122] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment, in which the image data augmentation method is described in detail.

[0123] See Figure 6 The image data augmentation method shown includes:

[0124] S601. Obtain the image to be processed of the power equipment; the image to be processed includes the original image and the augmented guidance image; at least part of the image background of the augmented guidance image is different from that of the original image.

[0125] S602. For each image to be processed, based on the feature encoding network, perform multi-scale feature extraction on the image to be processed to obtain the low-level features and corresponding high-level features of the image to be processed.

[0126] S603. Extract the key semantic information in the high-level features of the image to be processed.

[0127] S604. Perform feature fusion on the key semantic information and the low-level features of the image to be processed to obtain the corrected features of the image to be processed.

[0128] S605. Perform feature redrawing on the corrected features to obtain the initial spatial features.

[0129] S606. Perform max pooling on the initial spatial features to obtain the first spatial features, and perform average pooling on the initial spatial features to obtain the second spatial features.

[0130] S607. Perform feature fusion on the first spatial features and the second spatial features to obtain the target spatial features.

[0131] S608. Enhance the attention of the target spatial features according to the key semantic information to obtain the encoded features of the image to be processed.

[0132] S609. Based on the feature decoding network, perform feature decoding processing on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guidance image to obtain the augmented image of the power equipment.

[0133] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0134] Based on the same inventive concept, an embodiment of the present application also provides an image data augmentation device for implementing the above-mentioned image data augmentation method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following image data augmentation device can refer to the limitations on the image data augmentation method in the above text, and will not be repeated here.

[0135] In an exemplary embodiment, as Figure 7 shown, an image data augmentation device is provided, including: an acquisition module 710, an encoding module 720, and a decoding module 730, where:

[0136] The acquisition module 710 is used to acquire the image to be processed of the power equipment; the image to be processed includes the original image and the augmented guidance image; the augmented guidance image is at least partially different from the image background of the original image;

[0137] The encoding module 720 is used to perform feature encoding processing on each image to be processed based on the feature encoding network to obtain the encoded features corresponding to the image to be processed;

[0138] The decoding module 730 is used to perform feature decoding processing on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guidance image based on the feature decoding network to obtain the augmented image of the power equipment.

[0139] In one embodiment, the encoding module 720 includes: a first extraction unit configured to perform multi-scale feature extraction on the image to be processed to obtain low-level features and corresponding high-level features of the image to be processed; and a first fusion unit configured to perform feature fusion on the low-level features and the corresponding high-level features of the image to be processed to obtain encoded features of the image to be processed.

[0140] In one embodiment, the first fusion unit includes: a first extraction subunit configured to extract key semantic information from the high-level features of the image to be processed; a first fusion subunit configured to perform feature fusion on the key semantic information and the low-level features of the image to be processed to obtain corrected features of the image to be processed; and an encoding subunit configured to perform feature encoding on the corrected features and the key semantic information to obtain encoded features of the image to be processed.

[0141] In one embodiment, the encoding subunit is specifically configured to: extract target spatial features from the corrected features; and perform attention enhancement on the target spatial features according to the key semantic information to obtain encoded features of the image to be processed.

[0142] In one embodiment, the encoding subunit is further specifically configured to: perform feature redrawing on the corrected features to obtain initial spatial features; perform max pooling on the initial spatial features to obtain first spatial features, and perform average pooling on the initial spatial features to obtain second spatial features; and perform feature fusion on the first spatial features and the second spatial features to obtain target spatial features.

[0143] In one embodiment, it further includes: a first extraction module configured to perform feature extraction on the augmented image to obtain augmented features; a second extraction module configured to perform feature extraction on the original image to obtain original features; a generation module configured to generate a semantic perception loss according to the difference between the augmented features and the original features; and an adjustment module configured to adjust network parameters of the feature encoding network and the feature decoding network according to the semantic perception loss.

[0144] Each module in the above image data augmentation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form so that the processor can call and execute operations corresponding to the above modules.

[0145] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an image data augmentation method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0146] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0147] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0148] Obtain the image to be processed of the power device; the image to be processed includes an original image and an augmentation guidance image; the augmentation guidance image is at least partially different from the image background of the original image;

[0149] For each image to be processed, based on a feature encoding network, perform feature encoding processing on the image to be processed to obtain the encoded feature corresponding to the image to be processed;

[0150] Based on a feature decoding network, perform feature decoding processing on the encoded feature corresponding to the original image and the encoded feature corresponding to the augmentation guidance image to obtain the augmented image of the power device.

[0151] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing multi-scale feature extraction on the image to be processed to obtain low-level features and corresponding high-level features of the image to be processed; performing feature fusion on the low-level features and the corresponding high-level features of the image to be processed to obtain encoded features of the image to be processed.

[0152] In one embodiment, when the processor executes the computer program, the following steps are further implemented: extracting key semantic information in the high-level features of the image to be processed; performing feature fusion on the key semantic information and the low-level features of the image to be processed to obtain corrected features of the image to be processed; performing feature encoding on the corrected features and the key semantic information to obtain encoded features of the image to be processed.

[0153] In one embodiment, when the processor executes the computer program, the following steps are further implemented: extracting target space features in the corrected features; performing attention enhancement on the target space features according to the key semantic information to obtain encoded features of the image to be processed.

[0154] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing feature redrawing on the corrected features to obtain initial space features; performing max pooling on the initial space features to obtain first space features, and performing average pooling on the initial space features to obtain second space features; performing feature fusion on the first space features and the second space features to obtain target space features.

[0155] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing feature extraction on the augmented image to obtain augmented features; and performing feature extraction on the original image to obtain original features; generating a semantic perception loss according to the difference between the augmented features and the original features; adjusting the network parameters of the feature encoding network and the feature decoding network according to the semantic perception loss.

[0156] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0157] Obtaining an image to be processed of an electrical device; the image to be processed includes an original image and an augmented guiding image; at least part of the image background of the augmented guiding image is different from that of the original image;

[0158] For each image to be processed, performing feature encoding processing on the image to be processed based on a feature encoding network to obtain encoded features corresponding to the image to be processed;

[0159] Based on a feature decoding network, performing feature decoding processing on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guiding image to obtain an augmented image of the electrical device.

[0160] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing multi-scale feature extraction on the image to be processed to obtain low-level features and corresponding high-level features of the image to be processed; performing feature fusion on the low-level features and the corresponding high-level features of the image to be processed to obtain encoded features of the image to be processed.

[0161] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: extracting key semantic information from the high-level features of the image to be processed; performing feature fusion on the key semantic information and the low-level features of the image to be processed to obtain corrected features of the image to be processed; performing feature encoding on the corrected features and the key semantic information to obtain encoded features of the image to be processed.

[0162] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: extracting target space features from the corrected features; performing attention enhancement on the target space features according to the key semantic information to obtain encoded features of the image to be processed.

[0163] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing feature redrawing on the corrected features to obtain initial space features; performing max pooling on the initial space features to obtain first space features, and performing average pooling on the initial space features to obtain second space features; performing feature fusion on the first space features and the second space features to obtain target space features.

[0164] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing feature extraction on the augmented image to obtain augmented features; and performing feature extraction on the original image to obtain original features; generating a semantic perception loss according to the difference between the augmented features and the original features; adjusting the network parameters of the feature encoding network and the feature decoding network according to the semantic perception loss.

[0165] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the following steps:

[0166] Obtaining an image to be processed of an electrical device; the image to be processed includes an original image and an augmented guiding image; the image background of the augmented guiding image is at least partially different from that of the original image;

[0167] For each image to be processed, performing feature encoding processing on the image to be processed based on a feature encoding network to obtain encoded features corresponding to the image to be processed;

[0168] Based on a feature decoding network, performing feature decoding processing on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guiding image to obtain an augmented image of the electrical device.

[0169] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing multi-scale feature extraction on the image to be processed to obtain low-level features and corresponding high-level features of the image to be processed; performing feature fusion on the low-level features and the corresponding high-level features of the image to be processed to obtain encoded features of the image to be processed.

[0170] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: extracting key semantic information from the high-level features of the image to be processed; performing feature fusion on the key semantic information and the low-level features of the image to be processed to obtain corrected features of the image to be processed; performing feature encoding on the corrected features and the key semantic information to obtain encoded features of the image to be processed.

[0171] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: extracting target space features from the corrected features; enhancing the attention of the target space features according to the key semantic information to obtain encoded features of the image to be processed.

[0172] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing feature redrawing on the corrected features to obtain initial space features; performing max pooling on the initial space features to obtain first space features, and performing average pooling on the initial space features to obtain second space features; performing feature fusion on the first space features and the second space features to obtain target space features.

[0173] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing feature extraction on the augmented image to obtain augmented features; and performing feature extraction on the original image to obtain original features; generating a semantic perception loss according to the difference between the augmented features and the original features; adjusting the network parameters of the feature encoding network and the feature decoding network according to the semantic perception loss.

[0174] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0175] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0176] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. An image data augmentation method, characterized in that, The method includes: Obtaining a to-be-processed image of a power device; the to-be-processed image includes an original image and an augmented guiding image; the augmented guiding image is at least partially different from the image background of the original image; For each to-be-processed image, based on a feature encoding network, performing feature encoding processing on the to-be-processed image to obtain an encoded feature corresponding to the to-be-processed image; Based on a feature decoding network, performing feature decoding processing on the encoded feature corresponding to the original image and the encoded feature corresponding to the augmented guiding image to obtain an augmented image of the power device.

2. The method according to claim 1, wherein The performing feature encoding on the to-be-processed image to obtain an encoded feature corresponding to the to-be-processed image includes: Performing multi-scale feature extraction on the to-be-processed image to obtain low-level features and corresponding high-level features of the to-be-processed image; Fusing the low-level features and the corresponding high-level features of the to-be-processed image to obtain an encoded feature of the to-be-processed image.

3. The method according to claim 2, wherein The fusing the low-level features and the corresponding high-level features of the to-be-processed image to obtain an encoded feature of the to-be-processed image includes: Extracting key semantic information in the high-level features of the to-be-processed image; Fusing the key semantic information with the low-level features of the to-be-processed image to obtain a corrected feature of the to-be-processed image; Performing feature encoding on the corrected feature and the key semantic information to obtain an encoded feature of the to-be-processed image.

4. The method according to claim 3, wherein The performing feature encoding on the corrected feature and the key semantic information to obtain an encoded feature of the to-be-processed image includes: Extracting target spatial features in the corrected feature; Enhancing the attention of the target spatial features according to the key semantic information to obtain an encoded feature of the to-be-processed image.

5. The method according to claim 4, characterized in that, The extracting target spatial features in the corrected feature includes: Performing feature redrawing on the corrected feature to obtain initial spatial features; Performing max pooling on the initial spatial features to obtain first spatial features, and performing average pooling on the initial spatial features to obtain second spatial features; Fusing the first spatial features and the second spatial features to obtain the target spatial features.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Performing feature extraction on the augmented image to obtain augmented features; and, Performing feature extraction on the original image to obtain original features; Generating a semantic perception loss according to the difference between the augmented features and the original features; Adjusting the network parameters of the feature encoding network and the feature decoding network according to the semantic perception loss.

7. An image data augmentation device, characterized in that, The device includes: An obtaining module, configured to obtain a to-be-processed image of a power device; the to-be-processed image includes an original image and an augmented guiding image; the augmented guiding image is at least partially different from the image background of the original image; An encoding module, configured to, for each to-be-processed image, based on a feature encoding network, perform feature encoding processing on the to-be-processed image to obtain an encoded feature corresponding to the to-be-processed image; A decoding module, configured to perform feature decoding processing on the encoded features corresponding to the original image and the encoded features corresponding to the augmented guidance image based on a feature decoding network, so as to obtain an augmented image of the power equipment.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.