Power equipment image augmentation method and system and computer readable storage medium

By segmenting, encoding and feature extraction of power equipment images, combining Stable Diffusion module and U-net model, high-quality augmented power equipment images are generated, solving the problems of scarcity and uneven distribution in power equipment scenarios.

CN119942104APending Publication Date: 2025-05-06STATE GRID INFORMATION & TELECOMM BRANCH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411853196.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The lack of sample augmentation methods for power equipment scenarios in the prior art, resulting in the problems of scarcity of samples, difficulty in collecting and uneven distribution in power business scenarios.

Method used

By acquiring the image of the power device, segmenting to separate the subject object and the background image, encoding processing is performed and input into the Stable Diffusion module, a power device image containing the subject object is generated, identity and detail feature extraction is performed, and the second power device image and background image are fused to generate an augmented power device image.

Benefits of technology

The sample augmentation for power equipment scenarios has been achieved, the quality and resolution of power equipment images have been improved, and the problems of scarcity and uneven distribution have been solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942104A_ABST
    Figure CN119942104A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a power equipment image augmentation method and system and a computer readable storage medium, and belongs to the field of power equipment image processing. The augmentation method comprises the following steps: acquiring an image about power equipment; segmenting the image so as to separate a main target object in the image from a background; obtaining a segmented image, and carrying out coding processing on the segmented image; inputting the encoded image into a Stable Diffusion module, so as to obtain a complete, ordered and clear image of the power equipment; performing identity feature extraction on the electrical equipment image; performing detail feature extraction on the electrical equipment image to improve the quality and resolution of the electrical equipment image; and fusing the remaining parts of the power equipment image and the original image to obtain a complete and clear image about the power equipment. The augmentation method can realize sample augmentation for a power business scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power equipment image processing, and in particular to a method and system for augmenting a power equipment image and a computer-readable storage medium. Background Art

[0002] Image sample enhancement is to generate more data from the existing data set. Nanni et al. proposed a new data enhancement method based on Fourier transform (FT), Radon transform (RT) and discrete cosine transform (DCT), and built a new set on the data set by combining fourteen traditional data enhancement methods. In order to augment the relevant data sets for shadow detection, Li et al. proposed a shadow data enhancement network ShadowGAN based on a generative adversarial network. Given a shadow mask and a shadow-free image, it can generate labeled augmented shadow image data. In the prior art, there is no sample augmentation method specifically for power equipment scenarios due to the scarcity of power business scenarios, difficulty in sample collection, and uneven sample distribution. Therefore, a sample augmentation method for power equipment scenarios is needed. Summary of the invention

[0003] The purpose of the embodiments of the present invention is to provide a method, system and computer-readable storage medium for augmenting an image of electric power equipment. The augmentation method can implement sample augmentation for electric power business scenarios.

[0004] In order to achieve the above object, an embodiment of the present invention provides a method for augmenting an image of a power equipment, the augmenting method comprising: Acquire images of electrical equipment; Segmenting the image to obtain a main target object image and a background image in the image; Acquire a subject target object image, and perform encoding processing on the subject target object image; Inputting the encoded main target object image into a Stable Diffusion module to obtain a first power equipment image including the main target object; Extracting identity features from the first electric power equipment image, and identifying electric power equipment in the electric power equipment image according to the identity features; Extracting detail features from the electric power equipment image, and acquiring a second electric power equipment image according to the detail features; The second power equipment image and the background image are fused to obtain an augmented power equipment image.

[0005] Optionally, segmenting the image to obtain a main target object image and a background image in the image includes: Feed images of electrical equipment into the transformer model; The Transformer model divides the image into small blocks, and retains spatial information of the image divided into small blocks using position encoding; The global features of the image divided into small blocks are extracted through a multi-layer self-attention mechanism to generate a semantic segmentation mask, thereby accurately distinguishing the main object and the background; The segmentation results are optimized through post-processing and the target area is extracted to obtain the main target image.

[0006] Optionally, obtaining a subject target object image and encoding the subject target object image includes: Acquire the segmented image to obtain the main target object image; Dividing the region of the main target object image into image blocks of fixed size; Embed each image patch into a fixed-dimensional feature space through linear projection; Keep the position information of the image block in the original image; The background image is encoded into a low-dimensional global feature vector as a conditional feature.

[0007] Optionally, the encoded main target object image is input into a StableDiffusion module to obtain a first power equipment image containing the main target object, including: Acquire the linearly projected image block and conditional features; Sending the image block and conditional features to the Stable Diffusion module; Performing forward diffusion on the conditional features and the original image to obtain a noise map, thereby training the StableDiffusion module; After training, the Stable Diffusion module performs back diffusion to convert each image block into a token, adding time step embedding and position encoding; Extract global context information through multi-layer Transformer; The corresponding image is obtained by formula (1): , formula (1) in, express Step image, represents the attenuation coefficient corresponding to the reverse diffusion process, express Step image, Indicates The cumulative noise attenuation coefficient of the step, represents the predicted noise, which represents the From Remove the noise from represents the standard deviation of the noise associated with the time step, represents an additional random noise term; Estimate the noise that needs to be removed from the current image, and update the image through the back diffusion formula (1) to obtain the denoised image; Repeat the iterative process of the image using formula (1) until the set step is reached to obtain the final denoised image. According to the position information of the denoised image finally obtained, all the denoised images are fused to obtain the first electric power equipment image.

[0008] Optionally, extracting identity features from the first power equipment image, and identifying power equipment in the power equipment image according to the identity features, includes: Acquire the first electric power equipment image; Performing enhanced processing of cropping, rotating and scaling on the first electric power equipment image; The enhanced power equipment images are used as positive samples, and different power equipment images are sent to the InfoNCE optimization model as negative samples to strengthen the feature representation between positive samples and identify power equipment in power equipment images of different angles and directions.

[0009] Optionally, extracting detail features from the electric power equipment image and acquiring a second electric power equipment image according to the detail features includes: Acquiring the electric power equipment image; Converting the power equipment image from the spatial domain to the frequency domain using a two-dimensional discrete Fourier transform; After conversion, the low-frequency components in the frequency domain are moved to the center of the frequency domain; Constructing a high-pass filter, and multiplying the frequency spectrum of the power equipment image by the high-pass filter to obtain a filtered frequency spectrum; The filtered spectrum is inverse Fourier transformed to convert the frequency domain into the spatial domain to obtain the filtered power equipment image.

[0010] Optionally, extracting detail features from the electric power equipment image and acquiring a second electric power equipment image according to the detail features includes: Acquire the filtered image of the electric power equipment; Generating a binary mask according to the target area in the power equipment image; Applying an erosion operation to the binary mask to remove unnecessary details except for the electrical equipment near the target area; superimposing the eroded mask onto the filtered image of the power equipment and retaining high-frequency details of the target area; After combining the extracted high frequency details with the corrosion mask, a detail map is generated to obtain a second power device image with a higher resolution.

[0011] Optionally, fusing the second power equipment image with the background image to obtain an augmented power equipment image includes: Acquire a trained power target image, and segment the trained power target image to obtain a segmented image; The trained power target image is used as the input image and the segmented image is used as the label image to train the U-net model; After the training is completed, the second power equipment image and the background image are fed into the trained U-net model; The U-net model fine-tunes the main area in the second power equipment image and seamlessly integrates it with the background image to obtain a complete and clear image of the power equipment.

[0012] On the other hand, the present invention further provides an augmentation system for an electric power equipment image, the augmentation system comprising: An image acquisition module, used for acquiring images of electric power equipment; The background calculation module is used to execute the method for augmenting the image of the electric power equipment as described above according to the image of the electric power equipment.

[0013] On the other hand, the present invention further provides a computer-readable storage medium, on which instructions are stored, and the instructions are used to enable a machine to execute the method for augmenting an image of electric power equipment as described above.

[0014] Through the above technical scheme, the augmentation method, system and computer-readable storage medium of the power equipment image provided by the present invention can obtain the image of the power equipment, and then segment the image. After segmentation, the subject target image and the background image in the image can be obtained, so that the main target in the image can be separated from the background. Then after segmentation, the segmented image can be encoded, and the segmented image can be the main target image. After encoding, the encoded main target image can be input into the Stable Diffusion module, so that a complete, orderly and clear first power equipment image can be obtained, and the first power equipment image can contain the main target. After obtaining the first power equipment image, the identity feature extraction can be performed on the first power equipment image. After the identity feature extraction is completed, the power equipment in the power equipment image can be identified, and then the detail feature extraction can be performed on the power equipment image, and the second power equipment image can be obtained according to the detail feature, so that the quality and resolution of the power equipment image can be improved. After the extraction is completed, the second power equipment image and the segmented background image can be fused, so that a complete and clear image of the power equipment can be obtained. The augmentation method can realize sample augmentation for power business scenarios.

[0015] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings: Figure 1 is a flow chart of a method for augmenting an image of a power device according to an embodiment of the present invention; Figure 2 is a flowchart of segmenting an image according to an augmentation method of an electric power equipment image according to one embodiment of the present invention; Figure 3 is a flowchart of encoding an image according to a method for augmenting an image of a power equipment according to an embodiment of the present invention; Figure 4 is a flow chart of processing an image by a StableDiffusion module of an augmentation method for an electric power equipment image according to an embodiment of the present invention; Figure 5 is a flow chart of extracting identity features from an image according to a method for augmenting an image of electric power equipment according to an embodiment of the present invention; Figure 6is a first flow chart of extracting detail features of an image according to a method for augmenting an image of electric power equipment according to an embodiment of the present invention; Figure 7 is a second flow chart of extracting detail features from an image in a method for augmenting an image of electric power equipment according to an embodiment of the present invention; Figure 8 The present invention is a flowchart of a fusion image of a method for augmenting an image of electric power equipment according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The specific implementation of the embodiment of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiment of the present invention, and is not used to limit the embodiment of the present invention.

[0018] It should be noted that the acquisition, transmission, storage, use, and processing of data in the technical solution of this application are in compliance with the relevant provisions of national laws and regulations. In the embodiments of this application, some existing solutions in the industry such as certain software, components, and models may be mentioned, which should be considered as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use the solution.

[0019] Figure 1 is a flow chart of a method for augmenting an image of a power equipment according to an embodiment of the present invention. In the present invention, the process of image augmentation may include: In step S1 , an image of electric power equipment is acquired.

[0020] In step S2, the image is segmented to obtain a main target object image and a background image in the image.

[0021] In step S3, a main target object image is acquired and the main target object image is coded.

[0022] In step S4, the main target object image that has been coded is input into a Stable Diffusion module to obtain a first power equipment image including the main target object.

[0023] In step S5, identity features are extracted from the first electric power equipment image, and the electric power equipment in the electric power equipment image is identified according to the identity features.

[0024] In step S6, detail features are extracted from the electric power equipment image, and a second electric power equipment image is acquired according to the detail features.

[0025] In step S7, the second power equipment image and the background image are fused to obtain an augmented power equipment image.

[0026] In the present invention, by acquiring an image of an electric power device, the image can be segmented. After segmentation, the subject target image and the background image in the image can be obtained, so that the subject target in the image can be separated from the background. After segmentation, the segmented image can be encoded, and the segmented image can be the subject target image. After encoding, the encoded subject target image can be input into the Stable Diffusion module, so that a complete, orderly and clear first electric power device image can be obtained, and the first electric power device image can contain the subject target. After obtaining the first electric power device image, the identity feature extraction can be performed on the first electric power device image. After the identity feature extraction is completed, the electric power device in the electric power device image can be identified, and then the detail feature extraction can be performed on the electric power device image, and the second electric power device image can be obtained according to the detail feature, so that the quality and resolution of the electric power device image can be improved. After the extraction is completed, the second electric power device image and the segmented background image can be fused, so that a complete and clear image of the electric power device can be obtained. The augmentation method can realize sample augmentation for electric power business scenarios.

[0027] First, we collect a wide range of power equipment image data, including transmission line defect samples, substation transformers and control devices, etc. By analyzing the feature distribution of power equipment image data, we determine the optimal Transformer module configuration, including the optimal combination of multi-head self-attention mechanism, feedforward network and adaptive normalization layer, as well as the optimal depth and width of the Transformer model, so that the model can effectively capture the key features of the equipment image. Then, we use the Stable Diffusion stable diffusion model technology to design the forward diffusion process and the reverse diffusion process, process the conditional input, including the diffusion time step and class label, and convert the first layer (unitization) in the Diffusion Transformer (DiT) into Transformer tokens. A smaller unit size corresponds to a large number of Transformer tokens. Increasing the diffusion model size and the number of input tokens greatly improves the performance of DiT, so that the diffusion model can gradually restore ordered images from disordered noise. Since the basis of image-generated images is a large number of valid samples, but since the samples that need to be generated in power scenarios are mostly scarce samples (such as critical defect image samples), the traditional idea of ​​text-generated images or image-generated images is limited by valid samples. The lack of valid samples makes it difficult for the diffusion model to converge or generate more realistic images. In addition, the images of power scenes are mostly high-definition images, which require a large video memory to be generated directly through the diffusion model. However, the scarce image samples are not all valid pixels in the entire image. Most of the environmental background and normal areas can use the original image. The focus is on the generation of the main area. In order to provide a high-resolution image, the main area of ​​the scarce image to be generated is extracted from the original image, and the main area is generated by knowledge guidance. The generated main area and the remaining area are registered and fused. Secondly, the multimodal attention decoupling technology is studied. By decoupling the cross-attention mechanism, separating the cross-attention layer, extracting text features and image features, freezing the pre-trained model parameters, and designing the image encoding conversion module, the image embedding features are effectively input into the cross-attention layer, realizing the customized generation of scarce samples of multimodal images in the power industry. Combined with the powerful multi-discriminative feature extraction capability of the self-supervised model, the target in the power image is projected into the feature space; in the detail feature extraction module, the high-pass filter is used to extract the high-frequency area, the corrosion mask is used to filter out the information near the outer contour of the target object, and a series of detail images with hierarchical resolution are generated through the U-net encoder. The identity features and detail features are input into the large image and text model to achieve adaptive generation of power image samples in multiple scenarios.

[0028] In one embodiment of the present invention, Figure 2 As shown, the process of segmenting an image may include: In step S8, the image of the power equipment is fed into the Transformer model.

[0029] In step S9, the Transformer model divides the image into small blocks and uses position encoding to retain spatial information of the image divided into small blocks.

[0030] In step S10, the global features of the image divided into small blocks are extracted through a multi-layer self-attention mechanism to generate a semantic segmentation mask, thereby accurately distinguishing the main object and the background.

[0031] In step S11, the segmentation result is optimized through post-processing and the target area is extracted.

[0032] In the present invention, when segmenting an image, an image of the power equipment can be fed into a transformer Transformer model, which can divide the image into small blocks, and can retain spatial information of the image divided into small blocks by using position coding, so that after subsequent processing, the image divided into small blocks can be spliced ​​into a whole according to the position coding. The global features of the image divided into small blocks can be extracted through a multi-layer self-attention mechanism, so that a semantic segmentation mask can be generated, so that the main target and the background can be accurately distinguished, and then the segmentation result can be optimized through post-processing and the target area can be extracted to obtain the main target image. After the main target image is segmented, the background image can be left.

[0033] In one embodiment of the present invention, Figure 3 As shown, the process of encoding an image may include: In step S12, the segmented image is acquired to obtain a main target object image.

[0034] In step S13, the region of the main target object image is divided into image blocks of fixed size.

[0035] In step S14, each image block is embedded into a feature space of fixed dimension through linear projection.

[0036] In step S15, the position information of the image block in the original image is retained.

[0037] In step S16, the background image is encoded into a low-dimensional global feature vector as a conditional feature.

[0038] In the present invention, when encoding an image, a segmented image can be obtained, so that the segmentation result of the main target and the background, that is, the main target image, can be obtained, and then the area of ​​the main target image can be divided into image blocks of fixed size. After the image blocks are obtained, each image block can be embedded into a feature space of fixed dimension through linear projection. When the image blocks are projected, the position of the image blocks in the original image can be retained, so that it is convenient to fuse the image blocks into a whole according to the position information after the image blocks are processed. After the image blocks are processed, the image of the background can be encoded into a low-dimensional global feature vector, so that it can be input into the Stable Diffusion module as a conditional feature.

[0039] In one embodiment of the present invention, Figure 4 As shown, the process of image processing by the Stable Diffusion module may include: In step S17, the linearly projected image blocks and conditional features are obtained.

[0040] In step S18, the image blocks and conditional features are sent to the Stable Diffusion module.

[0041] In step S19, forward diffusion is performed on the conditional features and the original image to obtain a noise map, thereby training a Stable Diffusion module.

[0042] In step S20, after training is completed, the Stable Diffusion module performs back diffusion to convert each image block into a token, adding time step embedding and position encoding.

[0043] In step S21, global context information is extracted through a multi-layer Transformer.

[0044] In step S22, the corresponding image is obtained by formula (1): , formula (1) in, express Step image, represents the attenuation coefficient corresponding to the reverse diffusion process, express Step image, Indicates The cumulative noise attenuation coefficient of the step, represents the predicted noise, which means that in step From Remove the noise from represents the standard deviation of the noise associated with the time step, represents an additional random noise term.

[0045] In step S23, the noise to be removed from the current image is estimated, and the image is updated by the back diffusion formula (1) to obtain a denoised image.

[0046] In step S24, the image is iterated repeatedly by formula (1) until the setting step is reached to obtain the final denoised image.

[0047] In step S25, all denoised images are fused according to the position information of the denoised images finally obtained, so as to obtain a first electric power equipment image.

[0048] In the present invention, after obtaining the linearly projected image block and conditional features, the image block and the original image can be forward diffused to obtain a noise map. The Stable Diffusion module can be trained by this method. After the training is completed, the Stable Diffusion module can be reverse diffused. In the process of reverse diffusion, each image block can be converted into a token, and time step embedding and position encoding can be added. Then, the global context information can be extracted through a multi-layer Transformer. The Transformer can predict the noise. By predicting the noise, the corresponding image can be obtained according to formula (1). After estimating the noise that needs to be removed from the current image, the image can be updated by the reverse diffusion formula (1), so that the denoised image can be obtained. Through the above continuous iteration, until the iteration reaches the set step, usually the 0th step, the final denoised image can be obtained. According to the position information of the denoised image finally obtained, the image blocks can be fused, so that a clear first power equipment image can be obtained.

[0049] In one embodiment of the present invention, Figure 5 As shown, the process of extracting identity features from an image may include: In step S26 , a first electric power equipment image is acquired.

[0050] In step S27 , the first electric power equipment image is subjected to enhancement processing such as cropping, rotation and scaling.

[0051] In step S28, the enhanced power equipment image is used as a positive sample, and different power equipment images are sent to the InfoNCE optimization model as negative samples to strengthen the feature representation between positive samples and identify power equipment in power equipment images at different angles and directions.

[0052] In the present invention, after acquiring the first power equipment image, the first power equipment image can be subjected to enhancement processing such as cropping, rotation and scaling. After the power equipment image including the first power equipment image is enhanced, the enhanced power equipment image can be used as a positive sample, and different power equipment images can be used as negative samples. Multiple images in the power equipment image in the positive sample can represent the same power equipment, and different power equipment images in the negative sample can represent different power equipment. Therefore, the positive sample and the negative sample can be sent to the InfoNCE optimization model, so that the feature representation between the positive samples can be strengthened. Therefore, the power equipment in the power equipment images of different angles and directions can be identified by this method.

[0053] In one embodiment of the present invention, Figure 6 As shown, the first process of extracting detail features from an image may include: In step S29, an electric power equipment image is acquired.

[0054] In step S30, the power equipment image is converted from the spatial domain to the frequency domain using a two-dimensional discrete Fourier transform.

[0055] In step S31, after conversion, the low frequency components in the frequency domain are moved to the center of the frequency domain.

[0056] In step S32, a high-pass filter is constructed, and the frequency spectrum of the power equipment image is multiplied by the high-pass filter to obtain a filtered frequency spectrum.

[0057] In step S33, the filtered spectrum is subjected to inverse Fourier transformation, thereby converting the frequency domain into the spatial domain to obtain a filtered power equipment image.

[0058] In the present invention, after acquiring the image of the power equipment, the two-dimensional discrete Fourier transform can be used to convert the power equipment image from the space domain to the frequency domain. After the conversion, the low-frequency components in the frequency domain can be moved to the center of the frequency domain, so as to facilitate the subsequent filtering of the low-frequency components. After the movement is completed, a high-pass filter can be constructed, and the spectrum after the conversion of the power equipment image can be multiplied by the high-pass filter, so as to obtain a filtered spectrum, which can be a spectrum after filtering the low-frequency part. After the filtering is completed, the filtered spectrum can be inverse Fourier transformed, so that the frequency domain can be converted into the spatial domain, and the filtered power equipment image can be obtained.

[0059] In one embodiment of the present invention, Figure 7 As shown, the second process of extracting detail features from an image may include: In step S34, the filtered electric power equipment image is acquired.

[0060] In step S35, a binary mask is generated according to the target area in the electric power equipment image.

[0061] In step S36, an erosion operation is applied to the binary mask to remove unnecessary details around the target area except for the electrical equipment.

[0062] In step S37, the eroded mask is superimposed on the filtered power equipment image, and only the high-frequency details of the target area are retained.

[0063] In step S38, the extracted high-frequency details are combined with the corrosion mask to generate a detail map to obtain the second power equipment image with a higher resolution.

[0064] In the present invention, after obtaining the filtered power equipment image, a binary mask can be generated according to the position of the target area in the power equipment image. An erosion operation is applied to the binary mask, so that the non-essential details near the target area except for the power equipment can be eliminated. After the erosion, the eroded mask can be superimposed on the filtered power equipment image, and only the high-frequency details in the target area can be retained. After retaining the high-frequency details, the extracted high-frequency details and the erosion mask can be combined to generate a detail map, so that the key features of the edge and texture of the power equipment can be improved, thereby improving the quality and resolution of the second power equipment image.

[0065] In one embodiment of the present invention, Figure 8 As shown, the process of fusing images may include: In step S39, the trained power target image is acquired, and the trained power target image is segmented to obtain a segmented image.

[0066] In step S40, the trained power target image is used as an input image and the segmented image is used as a label image to train the U-net model.

[0067] In step S41, after the training is completed, the second power equipment image and the background image are sent to the trained U-net model.

[0068] In step S42, the U-net model fine-tunes the main area in the second power equipment image and seamlessly merges it with the background image to obtain a complete and clear image of the power equipment.

[0069] In the present invention, when performing image fusion, a trained power target image can be obtained first, and the trained power target image can be segmented, so that a segmented image can be obtained. After obtaining the segmented image, the trained power target image can be used as an input image, and the segmented image can be used as a label image to train the U-net model. After the training is completed, the second power equipment image and the background image can be sent to the trained U-net model. After the training is completed, the trained U-net model can fine-tune the main area in the second power equipment image, and can be seamlessly integrated with the background image, so that a complete image of the cleaning of the power equipment can be obtained.

[0070] On the other hand, the present invention also provides an augmentation system for an image of an electric power device, the augmentation system comprising: an image acquisition module and a background calculation module. The image acquisition module is used to acquire an image of the electric power device. The background calculation module is used to execute the augmentation method for an image of an electric power device as described above according to the image of the electric power device.

[0071] On the other hand, the present invention further provides a computer-readable storage medium, on which instructions are stored, and the instructions are used to enable a machine to execute the method for augmenting an image of electric power equipment as described above.

[0072] Through the above technical scheme, the augmentation method, system and computer-readable storage medium of the power equipment image provided by the present invention can obtain the image of the power equipment, and then segment the image. After segmentation, the subject target image and the background image in the image can be obtained, so that the main target in the image can be separated from the background. Then after segmentation, the segmented image can be encoded, and the segmented image can be the main target image. After encoding, the encoded main target image can be input into the Stable Diffusion module, so that a complete, orderly and clear first power equipment image can be obtained, and the first power equipment image can contain the main target. After obtaining the first power equipment image, the identity feature extraction can be performed on the first power equipment image. After the identity feature extraction is completed, the power equipment in the power equipment image can be identified, and then the detail feature extraction can be performed on the power equipment image, and the second power equipment image can be obtained according to the detail feature, so that the quality and resolution of the power equipment image can be improved. After the extraction is completed, the second power equipment image and the segmented background image can be fused, so that a complete and clear image of the power equipment can be obtained. The augmentation method can realize sample augmentation for power business scenarios.

[0073] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0074] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0075] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0076] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0077] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0078] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0079] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0080] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0081] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for augmenting an image of a power equipment, characterized in that: The augmentation method comprises: Acquire images of electrical equipment; Segmenting the image to obtain a main target object image and a background image in the image; Acquire a subject target object image, and perform encoding processing on the subject target object image; Inputting the encoded main target object image into a Stable Diffusion module to obtain a first power equipment image including the main target object; Extracting identity features from the first electric power equipment image, and identifying electric power equipment in the electric power equipment image according to the identity features; Extracting detail features from the electric power equipment image, and acquiring a second electric power equipment image according to the detail features; The second power equipment image and the background image are fused to obtain an augmented power equipment image.

2. The augmentation method according to claim 1, characterized in that: Segmenting the image to obtain a main target image and a background image in the image includes: Feed images of electrical equipment into the transformer model; The Transformer model divides the image into small blocks, and retains spatial information of the image divided into small blocks using position encoding; The global features of the image divided into small blocks are extracted through a multi-layer self-attention mechanism to generate a semantic segmentation mask, thereby accurately distinguishing the main object and the background; The segmentation results are optimized and the target area is extracted through post-processing to obtain the main target image.

3. The augmentation method according to claim 1, characterized in that: Obtaining a subject target image and encoding the subject target image includes: Acquire the segmented image to obtain the main target object image; Dividing the region of the main target object image into image blocks of fixed size; Embed each image patch into a fixed-dimensional feature space through linear projection; Keep the position information of the image block in the original image; The background image is encoded into a low-dimensional global feature vector as a conditional feature.

4. The augmentation method according to claim 3, characterized in that: Inputting the encoded main target object image into a Stable Diffusion module to obtain a first power equipment image containing the main target object, including: Acquire the linearly projected image block and conditional features; Sending the image block and conditional features to the Stable Diffusion module; Performing forward diffusion on the conditional features and the original image to obtain a noise map, thereby training a StableDiffusion module; After training, the Stable Diffusion module performs back diffusion to convert each image block into a token, adding time step embedding and position encoding; Extract global context information through multi-layer Transformer; The corresponding image is obtained by formula (1): , Formula (1) in, express Step image, represents the attenuation coefficient corresponding to the reverse diffusion process, express Step image, Indicates The cumulative noise attenuation coefficient of the step, represents the predicted noise, which represents the From Remove the noise from represents the standard deviation of the noise associated with the time step, represents an additional random noise term; Estimate the noise that needs to be removed from the current image, and update the image through the back diffusion formula (1) to obtain the denoised image; Repeat the iterative process of the image using formula (1) until the set step is reached to obtain the final denoised image. According to the position information of the denoised image finally obtained, all the denoised images are fused to obtain the first electric power equipment image.

5. The augmentation method according to claim 4, characterized in that: Extracting identity features from the first electric power equipment image, and identifying electric power equipment in the electric power equipment image according to the identity features, includes: Acquire the first electric power equipment image; Performing enhanced processing of cropping, rotating and scaling on the first electric power equipment image; The enhanced power equipment images are used as positive samples, and different power equipment images are sent to the InfoNCE optimization model as negative samples to strengthen the feature representation between positive samples and identify power equipment in power equipment images of different angles and directions.

6. The augmentation method according to claim 5, characterized in that: Extracting detail features from the power equipment image and acquiring a second power equipment image according to the detail features includes: Acquiring the electric power equipment image; Converting the power equipment image from the spatial domain to the frequency domain using a two-dimensional discrete Fourier transform; After conversion, the low-frequency components in the frequency domain are moved to the center of the frequency domain; Constructing a high-pass filter, and multiplying the frequency spectrum of the power equipment image by the high-pass filter to obtain a filtered frequency spectrum; The filtered spectrum is inverse Fourier transformed to convert the frequency domain into the spatial domain to obtain the filtered power equipment image.

7. The augmentation method according to claim 6, characterized in that: Extracting detail features from the power equipment image and acquiring a second power equipment image according to the detail features includes: Acquire the filtered image of the electric power equipment; Generating a binary mask according to the target area in the power equipment image; Applying an erosion operation to the binary mask to remove unnecessary details except for the electrical equipment near the target area; superimposing the eroded mask onto the filtered image of the power equipment and retaining high-frequency details of the target area; After combining the extracted high frequency details with the corrosion mask, a detail map is generated to obtain a second power device image with a higher resolution.

8. The augmentation method according to claim 7, characterized in that: The second power equipment image is fused with the background image to obtain an augmented power equipment image, including: Acquire a trained power target image, and segment the trained power target image to obtain a segmented image; The trained power target image is used as the input image and the segmented image is used as the label image to train the U-net model; After the training is completed, the second power equipment image and the background image are fed into the trained U-net model; The U-net model fine-tunes the main area in the second power equipment image and seamlessly integrates it with the background image to obtain a complete and clear image of the power equipment.

9. A system for augmenting an image of electric power equipment, characterized in that: The augmentation system comprises: An image acquisition module, used for acquiring images of electric power equipment; A background computing module is used to execute a method for augmenting an image of electric power equipment as described in any one of claims 1 to 8 according to the image of the electric power equipment.

10. A computer-readable storage medium, characterized in that: The machine-readable storage medium stores instructions for causing a machine to execute a method for augmenting an image of electric power equipment as described in any one of claims 1-8.