A data enhancement method, device, medium and equipment for industrial product images
Through improved circular generation adversarial networks, combined with U-Net network, instance normalization module and convolutional block attention module, the problem of difficulty in collecting industrial image data sets is solved, efficient data enhancement is achieved, and the generated images are closer to real images, improving diversity and effect.
Patent Information
- Application Number
- CN202411591856.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-08
AI Technical Summary
In the industrial field, it is difficult to collect image data sets of industrial objects, insufficient sample size and lack of diversity, resulting in poor effectiveness of deep learning algorithms in industrial image data augmentation applications.
Using an improved cyclic generation adversarial network, combined with the U-Net network, an instance normalization module and a convolutional block attention module, a final enhanced image that removes image stitching traces through adversarial training of the generator and discriminator is generated.
Improves the diversity and effect of industrial product image data enhancement, and the generated images are closer to real images, eliminating discord in the initial enhanced image.
Smart Images

Figure CN119515752B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data enhancement technology, and in particular to a method, device, medium and equipment for data enhancement of industrial product images. Background Art
[0002] At present, with the development of computer technology and artificial intelligence, image processing technology based on deep learning is widely used in various fields of national economy and people's livelihood. For the industrial field, accurate image analysis and processing are crucial for quality control, product inspection and equipment maintenance. Image processing technology based on deep neural network is also increasingly appearing in multiple processes such as inspection, handling and equipment related to the manufacturing industry.
[0003] Limited by the image acquisition cost and the inherent characteristics of industrial objects, the difficulty in data set acquisition, insufficient sample quantity, and lack of sample diversity are the main factors restricting the promotion of artificial intelligence technology in the industrial field. When it comes to target detection related technologies, industrial objects are also regarded as small sample objects. Therefore, the research on data enhancement of industrial images to expand the data set of industrial objects and achieve stable detection of industrial small samples is of great significance to the further promotion of artificial intelligence technology in the industrial field and promote the automation and unmanned production process.
[0004] Existing research mainly focuses on how to effectively use limited samples for accurate pattern recognition and decision-making. Many research teams and companies are exploring the use of unsupervised algorithms, transfer learning, meta-learning and other technologies to achieve image generation of target images based on original images. However, existing deep learning algorithms are usually used for image migration in general fields, while research on data enhancement of product images in the industrial field is still rare. At present, image data enhancement of industrial objects is mainly achieved through geometric transformation, color transformation and other means, and the data enhancement effect is poor. Summary of the invention
[0005] Based on this, it is necessary to provide a data enhancement method, device, medium and equipment for industrial product images to address the above technical problems.
[0006] The present invention adopts the following technical solutions:
[0007] The present invention provides a data enhancement method for industrial product images, comprising:
[0008] Obtain sample images of the target industrial product under various backgrounds with traces of image splicing after cropping and splicing product images;
[0009] Based on the cyclic generative adversarial network, the U-Net network is used as the backbone network of the generator, and an instance normalization module is added to each convolution module used for downsampling in the contraction path of the U-Net network, and a convolution block attention module is added after the contraction path to construct an improved cyclic generative adversarial network;
[0010] The sample image is input into the generator of the improved recurrent generative adversarial network to enhance the sample image to obtain an enhanced image; the enhanced image and the product image are input into the discriminator of the improved recurrent generative adversarial network to determine whether the input image is an enhanced image or a product image, and the generator is trained with the enhanced image as the goal of making the discriminator judge it as a product image, and the trained generator is used as a data enhancement model;
[0011] The initial enhanced image obtained by cropping and splicing product images is input into the data enhancement model. The background features and target industrial product features are extracted through the convolution modules in the contraction path. The extracted features are instance-normalized through the instance normalization module to obtain the style feature map of the initial enhanced image. The channel and spatial features of the style feature map are captured through the convolution block attention module to generate the final enhanced image without image splicing traces.
[0012] Optionally, obtaining sample images with image splicing traces after cropping and splicing product images of the target industrial product under multiple different backgrounds specifically includes:
[0013] Acquire product images obtained by visually capturing target industrial products under various backgrounds;
[0014] The background of the product image is cropped by various image cropping methods to obtain a partial image of the target industrial product with part of the background;
[0015] A background image different from the background carried by a partial image of a target industrial product is obtained, and the partial image of the target industrial product is embedded into the background image through a plurality of random image stitching methods to generate a sample image with image stitching traces.
[0016] Optionally, embedding the partial image of the target industrial product into the background image by using a plurality of random image stitching methods specifically includes:
[0017] The brightness of the local image of the target industrial product is adjusted by the following formula:
[0018] I'(x,y)=I(x,y)+b;
[0019] Through a variety of random image stitching methods, local images of target industrial products with different brightness are embedded into different background images.
[0020] Wherein, I(x, y) is the gray value of the pixel in the local image of the target industrial product, b is the adjustment factor, and I'(x, y) is the gray value after adjustment.
[0021] Optionally, embedding the partial image of the target industrial product into different background images by using a plurality of random image stitching methods specifically includes:
[0022] Rotate the angle of the partial image of the target industrial product by the following formula:
[0023] x'=x*cos(theta)-y*sin(theta);
[0024] y'=x*sin(theta)+y*cos(theta);
[0025] Among them, (x, y) is the position of the pixel in the local image of the target industrial product, theta is the rotation angle, and (x', y') is the pixel position of the local image of the target industrial product after rotation.
[0026] Optionally, the instance normalization module is an InstanceNorm2d module;
[0027] The step of performing instance normalization on the extracted features by the instance normalization module to obtain a feature map specifically includes:
[0028] The extracted features are instance normalized by the instance normalization module based on the following formula to obtain the style feature map of the initial enhanced image:
[0029]
[0030] Among them, z is the style feature map of the initial enhanced image obtained after instance normalization, β and γ are affine coefficients, and x is the feature matrix extracted based on the initial enhanced image. is the mean of x, ε is the standard deviation, and σ is an added term to avoid the standard deviation being zero.
[0031] Optionally, the capturing of channel and spatial features of the style feature map by a convolutional block attention module specifically includes:
[0032] The channel features of the style feature map are obtained through the global maximum pooling module and the global average pooling module in the channel attention module of the convolutional block attention module, and the channel features are sent to a shared neural network including a multi-layer perceptron and a hidden layer to generate a channel attention map;
[0033] The pooled features of the style feature map are obtained through the global maximum pooling module and the global average pooling module in the spatial attention module of the convolutional block attention module, and the pooled features are spliced to generate a spatial attention map through convolution operations;
[0034] The style feature map is channel- and spatially weighted according to the channel attention map and the spatial attention map.
[0035] Optionally, the method further comprises:
[0036] The virtual environment of the target industrial product under different background and lighting conditions is simulated by a virtual engine, and the angle, position and intensity of the target industrial product under the virtual environment are determined by a random algorithm to generate an initial enhanced image of the target industrial product under different angles, positions, intensities, lighting conditions and backgrounds; the initial enhanced image has virtual traces;
[0037] The initial enhanced image is input into a data enhancement model to generate a final enhanced image with virtual traces eliminated.
[0038] The present invention provides a data enhancement device for industrial product images, comprising:
[0039] An acquisition module, used for acquiring sample images of a target industrial product with image splicing marks after cropping and splicing product images under various backgrounds;
[0040] A construction module is used to construct an improved cyclic generative adversarial network based on a cyclic generative adversarial network, using the U-Net network as the backbone network of the generator, adding an instance normalization module to each convolution module used for downsampling in the contraction path of the U-Net network, and adding a convolution block attention module after the contraction path;
[0041] A training module is used to input the sample image into the generator of the improved recurrent generative adversarial network to enhance the sample image to obtain an enhanced image; and input the enhanced image and the product image into the discriminator of the improved recurrent generative adversarial network to determine whether the input image is an enhanced image or a product image, and train the generator with the enhanced image as the goal of making the discriminator judge it as a product image, and use the trained generator as a data enhancement model;
[0042] The enhanced image is used to input the initial enhanced image obtained by cropping and splicing the product image into the data enhancement model, extract the background features and the target industrial product features through the convolution modules in the contraction path, and perform instance normalization on the extracted features through the instance normalization module to obtain the style feature map of the initial enhanced image, and capture the channel and spatial features of the style feature map through the convolution block attention module to generate the final enhanced image with the image splicing traces removed.
[0043] The present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the data enhancement method for industrial product images is implemented.
[0044] The present invention provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned data enhancement method for industrial product images when executing the program.
[0045] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects:
[0046] The present invention first obtains an initial enhanced image by cropping and splicing product images and uses it as a sample image, then constructs an improved cyclic generative adversarial network, and uses the improved cyclic generative adversarial network to perform generative adversarial learning for data enhancement based on the sample image. On the one hand, the generator masters the feature extraction ability of the background and target industrial products in the spliced image through learning, and uses this to generate enhanced images with background style. On the other hand, the discriminator's judgment prompts the generator to generate enhanced images that are closer to the real image, so that the improved cyclic generative adversarial network can maintain the background style of the spliced image and eliminate the disharmonious parts in the initial enhanced image, so that the generated image is close to the product image obtained by visual acquisition, which fully improves the diversity of data enhancement.
[0047] Among them, the improved cyclic generative adversarial network adds an instance normalization module to the generator of the cyclic generative adversarial network. The instance normalization module can perform a separate normalization operation on each sample, so as to pay more attention to the feature distribution of each sample itself, and improve the accuracy of feature extraction of each initial enhanced image; at the same time, by adding a convolutional block attention module, the ability to capture the channel and spatial features of the initial enhanced image is improved, so that when generating the final enhanced image, the disharmonious places in the initial enhanced image can be well eliminated, thereby improving the effect of data enhancement. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0049] Figure 1 A schematic diagram of a data enhancement method for industrial product images provided by the present invention;
[0050] Figure 2 A schematic diagram of different target industrial product images generated by a virtual engine provided by the present invention;
[0051] Figure 3A schematic diagram of a strategy architecture for industrial small sample object data enhancement provided by the present invention;
[0052] Figure 4 A schematic diagram of industrial object image generation based on a generative adversarial network provided by the present invention;
[0053] Figure 5 A schematic diagram of the working principle of a CycleGAN generative adversarial network provided by the present invention;
[0054] Figure 6 A schematic diagram of an improved CycleGAN generative adversarial network provided by the present invention;
[0055] Figure 7 A schematic diagram of splicing different target industrial products generated by the present invention;
[0056] Figure 8 A schematic diagram of comparison between an image generated by a data enhancement model and an original image provided by the present invention;
[0057] Fig. 9 A schematic diagram of image comparison generated by a different method provided by the present invention;
[0058] Fig.10 A schematic diagram of a data enhancement device for industrial product images provided by the present invention;
[0059] Fig.11 A schematic diagram of a computer device for implementing a data enhancement method for industrial product images provided by the present invention. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0061] Unlike common general objects in most datasets, dedicated datasets for industrial object detection are not common. Therefore, when users try to use deep learning algorithms to complete industrial object detection, it is necessary to establish a dedicated dataset for industrial objects. Limited by the characteristics of industrial objects, image sampling costs, and working environment, the establishment of industrial object datasets has always been a difficult problem for users. Some researchers hope to achieve data enhancement during training by adding methods such as image rotation, image scaling, and contrast adjustment to the target detection network, but such methods still have limitations.
[0062] The technical solutions provided by various embodiments of the present invention are described in detail below in conjunction with the accompanying drawings.
[0063] Figure 1 The following is a flow chart of a data enhancement method for industrial product images in the present invention, which specifically includes the following steps:
[0064] S101: Obtain sample images of a target industrial product with image splicing marks after cropping and splicing product images under multiple different backgrounds.
[0065] Generally, when using deep neural network algorithms to detect or measure target objects, factors such as the distribution range of the object, the light reaction of the material surface, and background characteristics have a significant impact on the detection accuracy of the model. Establishing a uniformly distributed target object dataset through data augmentation and other data expansion methods is an important means to improve the detection accuracy of the model.
[0066] Generally, when performing data enhancement on image datasets corresponding to industrial products, the image datasets can be first enriched through traditional data enhancement methods based on real industrial product images collected in real scenes.
[0067] Based on this, in one or more embodiments of the present invention, the server of the business platform may first obtain product images obtained by visually capturing the target industrial product under various backgrounds, where the visual capture mentioned here may refer to image capture by visual capture devices such as cameras. Afterwards, the server may perform preliminary data enhancement by using various image cropping methods and image stitching methods.
[0068] Specifically, in one or more embodiments of the present invention, the server can crop the background of the product image by using a variety of image cropping methods to obtain a partial image of the target industrial product carrying part of the background. Then, a background image different from the background carried by the partial image of the target industrial product is obtained, and the partial image of the target industrial product is embedded in the background image by using a variety of random image splicing methods to generate a sample image with image splicing traces.
[0069] The server can adjust the brightness of the local image of the target industrial product by the following formula: I'(x,y) = I(x,y) + b. In the formula, I(x,y) is the gray value of the pixel in the local image of the target industrial product, b is the adjustment factor, and I'(x,y) is the adjusted gray value. Then, through a variety of random image stitching methods, the local images of the target industrial products with different brightness are embedded in different background images.
[0070] The server may also rotate the angle of the partial image of the target industrial product by the following formula: x'=x*cos(theta)-y*sin(theta), y'=x*sin(theta)+y*cos(theta).
[0071] In the formula, (x, y) is the pixel position in the partial image of the target industrial product, theta is the rotation angle, and (x', y') is the pixel position of the partial image of the target industrial product after rotation. Then, through a variety of random image stitching methods, the partial images of the target industrial product at different angles are embedded in different background images.
[0072] Table 1SplicingImageGenerationBasedonRandomFunctions
[0073]
[0074] The pseudo code of the random algorithm for generating the initial enhanced image is shown in Table 1.
[0075] In addition, in one or more embodiments of the present invention, the server may also use a virtual engine to create a virtual environment of the target industrial product under different backgrounds and lighting conditions, and determine the angle, position and intensity of the target industrial product under the virtual environment through a random algorithm to generate initial enhanced images of the target industrial product under different angles, positions, intensities, lighting conditions and backgrounds. Of course, the generated initial enhanced images usually have virtual traces, such as Figure 2 shown. Figure 2 This is a schematic diagram of different target industrial product images generated by a virtual engine in the present invention.
[0076] The server mentioned in the present invention may be a server set on a business platform, or a device such as a desktop computer, a notebook computer, etc. that can execute the solution of the present invention. For the convenience of description, the following description is only based on the server as the execution subject.
[0077] S102: Based on the cyclic generative adversarial network, the U-Net network is used as the backbone network of the generator, and an instance normalization module is added to each convolution module used for downsampling in the contraction path of the U-Net network, and a convolution block attention module is added after the contraction path to construct an improved cyclic generative adversarial network.
[0078] S103: Input the sample image into the generator of the improved recurrent generative adversarial network to enhance the sample image to obtain an enhanced image; input the enhanced image and the product image into the discriminator of the improved recurrent generative adversarial network to determine whether the input image is an enhanced image or a product image, train the generator with the goal of making the discriminator judge the enhanced image as a product image, and use the trained generator as a data enhancement model.
[0079] S104: Input the initial enhanced image obtained based on product image cropping and splicing into the data enhancement model, extract background features and target industrial product features through each convolution module in the contraction path, and perform instance normalization on the extracted features through the instance normalization module to obtain a style feature map of the initial enhanced image, and capture channel and spatial features of the style feature map through the convolution block attention module to generate a final enhanced image with image splicing traces removed.
[0080] After obtaining the initial enhanced sample image as described above, the present invention further converts it into an image consistent with the style of the camera-captured image through training based on the image generation algorithm of Cycle-Consistent Generative Adversarial Networks (CycleGAN) for further enhancement.
[0081] Figure 3 A schematic diagram of a strategy architecture for industrial small sample object data enhancement in the present invention. Figure 3 It is demonstrated that the method of the present invention can further enhance the traditional data enhancement results through a recurrent generative adversarial network, and can further enhance the data enhancement results corresponding to virtual reality technology through a recurrent generative adversarial network.
[0082] Figure 4 A schematic diagram of industrial object image generation based on a cyclic generative adversarial network in the present invention. Figure 4 The process of further enhancing traditional data augmentation results through cyclic generative adversarial networks is demonstrated in the paper.
[0083] When using a deep neural network algorithm to detect or measure a target object, the distribution range of the object, the light reaction of the material surface, and the background will all affect the accuracy of the detection model. The present invention uses a recurrent generative adversarial network to generate images of target industrial products, converting the above-mentioned initial enhanced image into a target image with the same style as the background of the target industrial product, thereby achieving data enhancement of the target industrial product.
[0084] The recurrent generative adversarial network is a deep learning architecture for unsupervised image-to-image translation, such as Figure 5 As shown, Figure 5The figure is a schematic diagram of the working principle of a CycleGAN generative adversarial network in the present invention. It consists of a generator G and a discriminator F. It uses an unsupervised learning method and unpaired image data to train the model, thereby realizing style transfer between images in different fields. By introducing cycle consistency loss, the similarity between the generated image and the original image is ensured. The present invention uses the improved CycleGAN for image data enhancement of the target object.
[0085] The traditional CycleGAN uses an encoding-decoding network as a generator. The present invention is based on a cyclic generative adversarial network and uses a U-Net network as the backbone network of the generator. The U-Net network is a U-shaped structure network, with an encoder on the left and a decoder on the right. The encoder on the left is a downsampling process that extracts features by reducing the image size and increasing the number of channels in the image. The decoder on the right is a process of restoring the image. Upsampling will gradually restore the size of the image. Here, the upsampled input feature map is not only the output of the previous step, but also contains the corresponding feature information on the left.
[0086] Furthermore, in one or more embodiments of the present invention, the generator may also use instance normalization, such as using the InstanceNorm2d module as a normalization network, and the InstanceNorm2d module normalizes a single channel of each workpiece sample separately. It calculates the mean and variance of each sample, and performs a separate normalization operation on each sample based on these. Compared with the BatchNorm2d (batch normalization) method, the InstanceNorm2d module is more effective for data enhancement of industrial objects because it pays more attention to the feature distribution of each sample itself, and InstanceNorm2d is less affected by the training batch. The style feature map of the initial enhanced image can be obtained by performing instance normalization on the extracted features based on the following formula through the instance normalization module:
[0087] Where z is the style feature map of the initial enhanced image obtained after instance normalization, β and γ are affine coefficients, and x is the feature matrix extracted based on the initial enhanced image. is the mean value of x, ε is the standard deviation, and σ is an added item to avoid the standard deviation being 0, that is, to prevent the occurrence of abnormal division by 0. The default value is 1e-5.
[0088] Furthermore, the present invention also adds a CBAM (Convolutional Block Attention Module) attention module to the generator. The Convolutional Block Attention Module (CBAM) is a lightweight attention module that can perform attention operations in the channel and spatial dimensions. It consists of a Channel Attention Module (CAM) and a Spatial Attention Module (SAM).
[0089] Specifically, the channel features of the style feature map can be obtained through the global maximum pooling module and the global average pooling module in the channel attention module of the convolutional block attention module, and the channel features are sent to a shared neural network including a multi-layer perceptron and a hidden layer to generate a channel attention map, and the pooled features of the style feature map are obtained through the global maximum pooling module and the global average pooling module in the spatial attention module of the convolutional block attention module, and the pooled features are spliced, and a spatial attention map is generated through a convolution operation, so that the style feature map is channel- and spatially weighted according to the channel attention map and the spatial attention map.
[0090] In the channel attention module, two different pooling methods, global maximum pooling (MaxPool) and global average pooling (AvgPool), are usually used to obtain channel features, and then these features are sent to a shared neural network composed of a multi-layer perceptron (MLP) and a hidden layer for calculation, and finally the channel attention map is generated through the Sigmoid activation function. In the spatial attention module, global maximum pooling and global average pooling operations are usually performed on the input feature map, and then the pooled features are spliced, and then the spatial attention map is generated through convolution operations.
[0091] Because CBAM can help the model better focus on important features and regions, and improve the performance and expression ability of the model, the present invention uses it in the discriminator to improve the learning efficiency and discrimination accuracy of the discriminator.
[0092] Figure 6 Schematic diagram of a generator in an improved recurrent generative adversarial network in the present invention. Figure 6 It can be seen that the red blocks represent the convolution blocks of each downsampled encoder. Each convolution block contains an InstanceNorm2d module. At the end of the contraction path, a convolution block attention module is added to weight the channels and space before continuing the subsequent expansion generation process to obtain the final enhanced image with a good fusion of the background and the local image of the target industrial product.
[0093] When training the improved cyclic generative adversarial network, the generator encodes, extracts features, and decodes and restores the input source domain image to convert it into the target domain image. The discriminator inputs the generated enhanced image and the visually captured product image of the target industrial product, and distinguishes the enhanced image from the product image. Among them, the generator loss includes adversarial loss and cycle consistency loss to generate realistic target domain images and maintain content consistency, while the discriminator loss is mainly adversarial loss, which aims to maximize the accuracy of distinguishing real images from generated images. Finally, the trained generator is used as the data augmentation model.
[0094] The trained data enhancement model can be used to further enhance the preliminary data enhancement results, wherein the data enhancement model can input the initial enhanced image obtained based on product image cropping and splicing, or the virtual initial enhanced image corresponding to the virtual environment of the virtual target industrial product in different background and lighting conditions through the virtual engine.
[0095] The data enhancement model can extract background features and target industrial product features through the convolution modules in the contraction path (encoder) in the generator, and perform instance normalization on the extracted features to obtain a style feature map through the instance normalization module. The style feature map is weighted by capturing channel and spatial features through the convolution block attention module, and finally the final enhanced image with image splicing traces removed is generated according to the weighted style feature map through the expansion path (decoder).
[0096] based on Figure 1 The data enhancement method for industrial product images shown in the present invention first obtains the initial enhanced image by cropping and splicing the product image and uses it as the sample image, then constructs an improved cyclic generative adversarial network, and uses the improved cyclic generative adversarial network to perform generative adversarial learning of data enhancement based on the sample image, so that the improved cyclic generative adversarial network can eliminate the discordant parts in the initial enhanced image, so that the generated image is close to the product image obtained by visual acquisition, and the diversity of data enhancement is fully improved. Among them, the instance normalization module added to the generator of the cyclic generative adversarial network can perform a separate normalization operation on each sample, so as to pay more attention to the feature distribution of each sample itself, and improve the accuracy of feature extraction of each initial enhanced image; at the same time, by adding a convolutional block attention module, the ability to capture the channel and spatial features of the initial enhanced image is improved, so that when generating the final enhanced image, the discordant parts in the initial enhanced image can be well eliminated, and the effect of data enhancement is improved.
[0097] The present invention introduces virtual reality and image generation technology based on unsupervised learning to achieve data enhancement for industrial small sample objects. The main innovations of the present invention are as follows:
[0098] 1. Combining the random algorithm with the image generation algorithm based on deep learning, a large number of target images with real-life style are generated through a small number of template images, thereby achieving data enhancement of the target industrial products in a specific background.
[0099] 2. Introduce virtual reality technology, and use the virtual engine to directly generate virtual images of target industrial products that are similar to the camera shooting style, thereby achieving data enhancement of the target industrial products.
[0100] 3. A network for target industrial product image generation is proposed. By integrating convolutional networks and self-attention modules, the image generation and style transfer of target industrial products are completed.
[0101] When applying the data enhancement method for industrial product images provided by the present invention, it is not necessary to Figure 1 The steps are executed in the order shown. The specific execution order of the steps can be determined according to needs, and the present invention does not limit this.
[0102] In addition, the present invention also provides an embodiment of a data enhancement method for applying industrial product images. Figure 7 This is a schematic diagram of splicing different target industrial products generated according to the above traditional data enhancement method in the present invention. The method used is the method in Table 1. Figure 7 It can be seen that the method in Table 1 can well generate mosaic images of different types and quantities of industrial products.
[0103] Figure 8 FIG. 1 is a schematic diagram showing a comparison between an image generated by a data enhancement model and an original image in the present invention. Figure 8 It can be seen that the generated image removes the splicing traces of the spliced image very well, which also proves the effectiveness of the method of the present invention.
[0104] Fig. 9 This is a schematic diagram of image comparison generated by different methods in the present invention, where the marked part (a) is the original image directly generated by the splicing algorithm, which is also the input image of the subsequent method; the marked part (b) is the workpiece image generated by the traditional CycleGAN; the marked part (c) is the image generated by AttentionGAN; the marked part (d) is the image generated by DCLGAN (Dual Contrastive Learning); the marked part (e) is the image generated by CycleGAN plus CBAM attention module. The marked part (f) is the image generated by the method proposed in the present invention.
[0105] Depend on Fig. 9It can be seen that the target image generated by the traditional cycleGAN does not completely remove the splicing traces in the target image, and there is still some information loss in industrial products. The AttentionGAN method hopes to improve the model's perception of key areas in the image by introducing the attention method, but for the image splicing method involved in the present invention, the image it generates has the problem of overall information loss. DCLGAN is one of the latest advances in the field of GAN network in the field of image generation. It infers the effective mapping between unpaired data through a new method based on contrastive learning and dual learning settings (using two encoders). For spliced objects, it can be seen that there is deformation of the object edges. The method proposed in the present invention is ahead of several other methods in terms of image details and generation effect, which proves the effectiveness of the improved cycleGAN proposed in the present invention.
[0106] In order to further compare the differences between the generated images and the real images, three evaluation indicators, FID (Frechet Inception Distance), SSIM (Structural Similarity Index) and PSNR (Peak Signal-to-Noise Ratio), are used to compare all generated images with the original images.
[0107] FID is an indicator used to measure the difference between images generated by generative models and real images. It is often used to measure the difference between images generated by generative models and real images. It is based on the Fréchet distance and calculates the degree of difference between the real image and the generated image by comparing the feature distribution of the real image and the generated image in the feature space. Generally speaking, the lower the FID value, the better, indicating that the difference between the generated image and the real image is smaller, and the performance of the generative model is better. On the contrary, the higher the FID value, the greater the difference between the generated image and the real image, and the worse the performance of the generative model.
[0108] Both SSIM and PSNR are indicators used to measure image quality, and are usually used to compare the similarity and quality between two images. SSIM is an indicator that measures the similarity between two images. It not only takes into account the similarity of brightness, but also the similarity of contrast and structure. The value range of SSIM is [-1,1], where 1 means that the two images are exactly the same, 0 means there is no similarity, and -1 means they are completely different. PSNR is an indicator of image quality, which measures the clarity and distortion of the image. The value range of PSNR is usually 0 to infinity, and the larger the value, the better the image quality. PSNR is calculated by the mean square error difference of the image, so it pays more attention to the difference at the pixel level of the image. Table 2 is a comparison diagram of indicators for generating images using different methods in the present invention.
[0109] Table 2 Binocular stereo vision parameters
[0110] FID SSIM PSNR CycleGAN 510.314 0.581 14.334 AttentionGAN 744.525 0.655 17.105 DCLGAN 417.177 0.682 16.224 CycleGAN+CBAM 393.034 0.571 15.465 Our method 339.459 0.833 21.094
[0111] The above is a data enhancement method for industrial product images provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding data enhancement device for industrial product images, such as Fig.10 shown.
[0112] Fig.10 A schematic diagram of a data enhancement device for industrial product images provided by the present invention includes:
[0113] An acquisition module 201 is used to acquire sample images of a target industrial product with image splicing marks after product images under various backgrounds are cropped and spliced;
[0114] A construction module 202 is used to construct an improved cyclic generative adversarial network based on a cyclic generative adversarial network, using a U-Net network as a backbone network of a generator, adding an instance normalization module to each convolution module for downsampling in a contraction path of the U-Net network, and adding a convolution block attention module after the contraction path;
[0115] The training module 203 is used to input the sample image into the generator of the improved recurrent generative adversarial network to enhance the sample image to obtain an enhanced image; and input the enhanced image and the product image into the discriminator of the improved recurrent generative adversarial network to determine whether the input image is an enhanced image or a product image, and train the generator with the enhanced image as the goal of making the discriminator judge it as a product image, and use the trained generator as a data enhancement model;
[0116] Enhanced image 204 is used to input the initial enhanced image obtained based on product image cropping and splicing into the data enhancement model, extract background features and target industrial product features through each convolution module in the contraction path, and perform instance normalization on the extracted features through the instance normalization module to obtain a style feature map of the initial enhanced image, and capture channel and spatial features of the style feature map through the convolution block attention module to generate a final enhanced image with image splicing traces removed.
[0117] The specific definition of the data enhancement device for industrial product images can be found in the definition of the data enhancement method for industrial product images above, which will not be repeated here. Each module in the above-mentioned data enhancement device for industrial product images can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0118] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 Data augmentation methods for industrial product images are provided.
[0119] The present invention also provides Fig.11 The structural diagram of the computer device shown in FIG. Fig.11 As shown in the figure, at the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Data augmentation methods for industrial product images are provided.
[0120] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0121] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.
Claims
1. A data enhancement method for industrial product images, characterized in that: include: Obtain sample images of the target industrial product under various backgrounds with traces of image splicing after cropping and splicing product images; Based on the cyclic generative adversarial network, the U-Net network is used as the backbone network of the generator, and an instance normalization module is added to each convolution module used for downsampling in the contraction path of the U-Net network, and a convolution block attention module is added after the contraction path to construct an improved cyclic generative adversarial network; The sample image is input into the generator of the improved recurrent generative adversarial network to enhance the sample image to obtain an enhanced image; the enhanced image and the product image are input into the discriminator of the improved recurrent generative adversarial network to determine whether the input image is an enhanced image or a product image, and the generator is trained with the enhanced image as the goal of making the discriminator judge it as a product image, and the trained generator is used as a data enhancement model; The initial enhanced image obtained by cropping and splicing the product image is input into the data enhancement model. The background features and target industrial product features are extracted through the convolution modules in the contraction path. The extracted features are instance-normalized by the instance normalization module to obtain the style feature map of the initial enhanced image. The channel and spatial features of the style feature map are captured by the convolution block attention module to generate the final enhanced image without the traces of image splicing. The instance normalization module is an InstanceNorm2d module; The step of performing instance normalization on the extracted features by the instance normalization module to obtain a feature map specifically includes: The extracted features are instance normalized by the instance normalization module based on the following formula to obtain the style feature map of the initial enhanced image: Among them, z is the style feature map of the initial enhanced image obtained after instance normalization, β and γ are affine coefficients, and x is the feature matrix extracted based on the initial enhanced image. is the mean value of x, ε is the standard deviation, and σ is an added term to avoid the standard deviation being 0; The channel and spatial feature capture of the style feature map through the convolution block attention module specifically includes: The channel features of the style feature map are obtained through the global maximum pooling module and the global average pooling module in the channel attention module of the convolutional block attention module, and the channel features are sent to a shared neural network including a multi-layer perceptron and a hidden layer to generate a channel attention map; The pooled features of the style feature map are obtained through the global maximum pooling module and the global average pooling module in the spatial attention module of the convolutional block attention module, and the pooled features are spliced to generate a spatial attention map through convolution operations; The style feature map is channel- and spatially weighted according to the channel attention map and the spatial attention map.
2. The data enhancement method for industrial product images according to claim 1, characterized in that: The obtaining of sample images of the target industrial product with image splicing marks after cropping and splicing product images under multiple different backgrounds specifically includes: Acquire product images obtained by visually capturing target industrial products under various backgrounds; The background of the product image is cropped by various image cropping methods to obtain a partial image of the target industrial product with part of the background; A background image different from the background carried by a partial image of a target industrial product is obtained, and the partial image of the target industrial product is embedded into the background image through a plurality of random image stitching methods to generate a sample image with image stitching traces.
3. The data enhancement method for industrial product images according to claim 2, characterized in that: The embedding of the partial image of the target industrial product into the background image by using a plurality of random image stitching methods specifically includes: The brightness of the local image of the target industrial product is adjusted by the following formula: I'(x,y)=I(x,y)+b; By using a variety of random image stitching methods, local images of target industrial products with different brightness are embedded into the background image; Wherein, I(x, y) is the gray value of the pixel in the local image of the target industrial product, b is the adjustment factor, and I'(x, y) is the gray value after adjustment.
4. The data enhancement method for industrial product images according to claim 2, characterized in that: The embedding of the partial image of the target industrial product into the background image by using a plurality of random image stitching methods specifically includes: Rotate the angle of the partial image of the target industrial product by the following formula: x'=x*cos(theta)-y*sin(theta), y'=x*sin(theta)+y*cos(theta); By using a variety of random image stitching methods, local images of target industrial products at different angles are embedded into the background image; Among them, (x, y) is the position of the pixel in the local image of the target industrial product, theta is the rotation angle, and (x', y') is the pixel position of the local image of the target industrial product after rotation.
5. The data enhancement method for industrial product images according to claim 1, characterized in that: The method further comprises: The virtual environment of the target industrial product under different background and lighting conditions is simulated by a virtual engine, and the angle, position and intensity of the target industrial product under the virtual environment are determined by a random algorithm to generate an initial enhanced image of the target industrial product under different angles, positions, intensities, lighting conditions and backgrounds; the initial enhanced image has virtual traces; The initial enhanced image is input into a data enhancement model to generate a final enhanced image with virtual traces eliminated.
6. A data enhancement device for industrial product images based on the data enhancement method for industrial product images according to any one of claims 1 to 5, characterized in that: include: An acquisition module, used for acquiring sample images of a target industrial product with image splicing marks after cropping and splicing product images under various backgrounds; A construction module is used to construct an improved cyclic generative adversarial network based on a cyclic generative adversarial network, using the U-Net network as the backbone network of the generator, adding an instance normalization module to each convolution module used for downsampling in the contraction path of the U-Net network, and adding a convolution block attention module after the contraction path; A training module, used for inputting a sample image into a generator of an improved cyclic generative adversarial network to enhance the sample image to obtain an enhanced image; The enhanced image and the product image are input into the discriminator of the improved recurrent generative adversarial network to determine whether the input image is an enhanced image or a product image, and the generator is trained with the enhanced image as the goal of making the discriminator judge it as a product image, and the trained generator is used as a data augmentation model; Enhanced image, used to input the initial enhanced image obtained by cropping and splicing the product image into the data enhancement model, extract the background features and target industrial product features through each convolution module in the contraction path, and obtain the style feature map of the initial enhanced image by instance normalization of the extracted features based on the following formula through the instance normalization module: The channel features of the style feature map are obtained through the global maximum pooling module and the global average pooling module in the channel attention module of the convolutional block attention module, and the channel features are sent to a shared neural network including a multi-layer perceptron and a hidden layer to generate a channel attention map; the pooled features of the style feature map are obtained through the global maximum pooling module and the global average pooling module in the spatial attention module of the convolutional block attention module, and the pooled features are spliced to generate a spatial attention map through a convolution operation; the style feature map is weighted in channels and spaces according to the channel attention map and the spatial attention map to generate a final enhanced image with image splicing traces removed; The instance normalization module is an InstanceNorm2d module; Among them, z is the style feature map of the initial enhanced image obtained after instance normalization, β and γ are affine coefficients, and x is the feature matrix extracted based on the initial enhanced image. is the mean of x, ε is the standard deviation, and σ is an added term to avoid the standard deviation being zero.
7. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
8. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Image enhancement method and device
CN111008940A
Image enhancement model training method and device, image enhancement method and device, equipment and medium
CN114529469A