Image Data Enhancement Method for Inner Pot of Rice Cooker Based on Mask Generative Adversarial Network
Through the combination of the mask generation adversarial network and the YOLOv5 discriminator, a high-quality rice cooker inner liner defect image data set is generated, which solves the problem of insufficient data volume and difficulty in labeling in the rice cooker inner liner defect detection, and improves the detection accuracy.
Patent Information
- Application Number
- CN202211636253.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-12-20
AI Technical Summary
In the prior art, the defect detection of rice cooker inner liner depends on manual detection efficiency and high cost. Deep learning methods require a large amount of label data while industrial defect data is scarce, resulting in poor detection accuracy. The defect samples generated by existing data enhancement methods are low in quality or uncertain in location.
A mask generation adversarial network is used to capture inner liner pictures and label defects through industrial cameras, build a real sample data set, generate a mask binary map, train a generative adversarial network and perform random transformation, and filter pseudo-samples as a discriminator to generate high-quality defect image data sets.
The detection accuracy of defect detection in the rice cooker inner liner has been improved, and the problems of insufficient data volume and difficulty in labeling have been solved. The resulting defect location and shape are rich, and the detection effect has been significantly improved.
Smart Images

Figure CN115908379B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital image processing and recognition, and particularly to a method for enhancing image data of an inner pot of an electric rice cooker based on a mask generative adversarial network. Background Art
[0002] An electric rice cooker mainly consists of a pot body and an inner pot. The inner pot is filled with rice and water, and the pot body heats the inner pot to cook the rice. As a metal product, like other metal products, during the industrial production process, due to the influence of technological processes, production equipment, and on-site environment, various defects will appear on the surface of the product. Surface defects not only affect the appearance quality and commercial value of the product itself, but also affect the performance of the product, and the safety and stability of subsequent deep processing. Therefore, surface defect detection has become a key step in industrial production. Currently, most detection tasks are completed manually, which brings high labor costs, management costs, and extremely low efficiency, and it is difficult to meet the needs of modern enterprise automated production.
[0003] To solve the problem of low efficiency of manual detection, many scholars have begun to study using deep learning methods to automatically detect defects. Although the detection accuracy of deep learning methods is high, a large amount of labeled data is required for training, while the number of industrial defect data is scarce, which will lead to poor effects of deep learning methods. To solve this problem, many scholars have begun to study expanding defect data through data augmentation, and one of the mainstream methods is to use a generative adversarial network to generate pseudo-defect data.
[0004] Generally, there are two types of defect data augmentation methods based on generative adversarial networks. The first type inputs the entire image into the generative adversarial network to generate pseudo-samples with defects. This method is simple, but only suitable for cases where the image is small and there are few defects, and the position of the defect in the image cannot be determined. The second method crops all the defect regions in all the images in the dataset into small images, then inputs the small images into the generative adversarial network to output sub-images of pseudo-defect regions, and finally randomly pastes these sub-images onto any image in the original dataset to obtain the generated pseudo-samples. The problem with this method is that the generated defects may have a large difference from the surrounding background, resulting in low-quality generated samples. Therefore, it is necessary to find a high-quality defect data augmentation method that can generate corresponding annotations. Summary of the Invention
[0005] To solve the technical problems existing in the prior art, the present invention provides a method for enhancing image data of an inner pot of an electric rice cooker based on a mask generative adversarial network, which can screen image data, generate high-quality defect images of the inner pot of an electric rice cooker. The screened pseudo-sample dataset and the real sample dataset obtain a dataset of defect images of the inner pot of an electric rice cooker after data augmentation, effectively improving the detection accuracy of inner pot defect detection of an electric rice cooker based on deep learning.
[0006] The technical solution adopted by the present invention is: an image data enhancement method for an inner pot of an electric rice cooker based on a masked generative adversarial network, including: S1. Taking pictures of the inner pot of the electric rice cooker through an industrial camera, performing defect annotation on the pictures of the inner pot of the electric rice cooker, and constructing a real sample data set of the images of the inner pot of the electric rice cooker;
[0007] S2. Generating a corresponding masked binary map for each picture of the inner pot of the electric rice cooker in the real sample data set to obtain a masked binary map data set;
[0008] S3. Constructing a generative adversarial network, initializing the model parameters of the generative adversarial network, and training the generative adversarial network in a way of gradually increasing the resolution;
[0009] S4. Performing random transformation on the masked binary map data to obtain an augmented masked binary map data set, mixing the augmented masked binary map data set with the masked binary map data set to obtain a mixed masked data set, and inputting the masked binary map data in the mixed masked data set into the trained generative adversarial network to generate a pseudo-sample data set;
[0010] S5. Training the YOLOv5 network with the real sample data set as a generation quality discriminator, using the generation quality discriminator to discriminate the quality of the pseudo-samples, and removing the pseudo-samples with low quality to obtain a filtered pseudo-sample data set;
[0011] S6. Mixing the filtered pseudo-sample data set with the real sample data set to obtain an enhanced image data set of the inner pot of the electric rice cooker.
[0012] In the preferred technical solution, the generating a corresponding masked binary map for each picture of the inner pot of the electric rice cooker in the real sample data set includes: initializing a grayscale map for the picture of the inner pot of the electric rice cooker, with the grayscale values of all pixels being 0, and setting the pixel points within the position of the defect area to 255.
[0013] Specifically, the training the generative adversarial network in a way of gradually increasing the resolution includes: training the generative adversarial network, inputting the masked binary map into the generator G to obtain a generated map g, and inputting the generated map g into the discriminator D to obtain a discrimination index; both the generator G and the discriminator D first delete some convolutional layers, and as the number of training iterations gradually increases, the number of convolutional layers increases until the resolution of the output image is the same as that of the real image.
[0014] Specifically, training the YOLOv5 network with the real sample dataset as the generation quality discriminator, using the generation quality discriminator to discriminate the quality of the pseudo-samples, and removing the pseudo-samples with low quality to obtain the filtered pseudo-sample dataset, including: inputting the generated pseudo-samples into the trained YOLOv5 network to obtain the position and category of the inner pot defects of the rice cooker. At the same time, the FID value can be calculated according to the feature vector of the YOLOv5 network, setting the FID reference value, retaining the images lower than the FID reference value, and removing the images higher than the FID reference value.
[0015] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0016] The present invention proposes an image data enhancement method for the inner pot of a rice cooker based on a masked generative adversarial network. By constructing a deep learning model based on a generative adversarial network, the original direct input of a random variable or the original image is changed to input a masked binary image, and a series of random transformations are performed on the masked binary image to make the input more abundant, so that the positions and morphologies of the generated image defects are also more abundant, and the corresponding defect positions can be directly obtained; training is carried out in a way of gradually increasing the resolution, and an additional discriminator of YOLOv5 is added to screen the generated samples, realizing the generation of high-quality inner pot defect images of a rice cooker; effectively improving the detection accuracy of the inner pot defect detection of a rice cooker based on deep learning, and effectively solving the problems of insufficient data volume and difficult data annotation in the deep learning process. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0018] Figure 1 It is a flowchart of the image data enhancement method for the inner pot of a rice cooker in an embodiment of the present invention;
[0019] Figure 2 It is a structural diagram of the generative adversarial network described in an embodiment of the present invention;
[0020] Figure 3 It is a schematic diagram of the progressive resolution training described in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] Next, the technical solution of the present invention will be further described in detail in conjunction with the accompanying drawings and embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. The implementation manners of the present invention are not limited thereto. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0022] Embodiment 1:
[0023] As Figure 1 shown, the method for enhancing the image data of the inner pot of an electric rice cooker based on a mask generative adversarial network of the present invention includes:
[0024] S1. Use an industrial camera to take pictures of the inner pot of the electric rice cooker, perform defect annotation on the inner pot image of the electric rice cooker, and construct a real sample dataset of the inner pot image of the electric rice cooker.
[0025] There is a machine on the industrial production line. The inner pot of the electric rice cooker can be conveyed and placed on its plane. There is a cantilever-mounted industrial camera above the plane. When the inner pot of the electric rice cooker is conveyed under the camera, the inner pot of the electric rice cooker is lifted to make the camera located at the center of the inner pot. At this time, the inner pot is rotated, and the camera starts to continuously take pictures of the inner pot image. After one week of rotation, the acquisition of the current electric rice cooker data is completed. Then, the current inner pot of the electric rice cooker is conveyed away, and the image acquisition of the next inner pot of the electric rice cooker is started.
[0026] Specifically, manually perform defect annotation on the inner pot pictures corresponding to the inner pots of the electric rice cookers with defects, and record the positions and categories of the defect areas in the inner pot pictures. The defect annotation format is a text file, and the specific format is (x1, y1, x2, y2, cls), where (x1, y1) is the upper left coordinate, (x2, y2) is the upper right coordinate, and cls is the category.
[0027] S2. Generate a corresponding mask binary map for each inner pot picture in the real sample dataset to obtain a mask binary map dataset;
[0028] Initialize a grayscale image for the inner pot picture, and the grayscale values of all pixels are 0. Set the pixel points within the position of the defect area to 255. In this embodiment, initialize a grayscale image for the inner pot picture, and the grayscale values of all pixels are 0. For the defect area in any format (x1, y1, x2, y2, cls), set the pixel points within the rectangular area (x1, y1, x2, y2) to 255. The finally obtained grayscale image is the mask binary map.
[0029] S3. Construct a generative adversarial network, initialize the model parameters of the generative adversarial network, and train the generative adversarial network with the mask binary map and the real picture.
[0030] Specifically, a generative adversarial network is constructed, which consists of a generator G and a discriminator D. Among them, the generator G is responsible for generating a pseudo-sample g according to the input mask, and the discriminator D is responsible for judging the difference degree between the pseudo-sample g and the real sample r.
[0031] As Figure 2 shown, the generative adversarial network includes a generator G and a discriminator D. Among them, the generator G is a UNet network, which consists of 12 convolutional layers. The first 6 convolutional layers form the first half, all of which are convolutional layers with a size of 3×3 and a stride of 2. This part will downsample the input into an output with a size of 10×10, and then input this output into the second half. The second half consists of the last six convolutional layers, and each convolutional layer is a transposed convolutional layer with a size of 3×3 and a stride of 2. This part will upsample the input into an output with a size of 640×640, and this output is the generated pseudo-sample g. There are residual connections between the corresponding convolutional layers in the first half and the second half of this network; the discriminator D is responsible for judging the difference degree between the pseudo-sample g and the real sample r, and its structure is a six-layer convolutional neural network, and each convolutional layer is a convolutional layer with a size of 3×3 and a stride of 2.
[0032] Specifically, initializing the model parameters of the generative adversarial network includes setting parameters such as the learning rate, learning rate decay coefficient, number of training iterations, batch data volume, and image size for training.
[0033] Set the initial learning rate for training to 0.001, the learning rate decay coefficient is set to decay by 10 times every 2000 iterations, the maximum number of iterations is set to 20000, the batch size is set to 128, the training image size is set to 640×640, the optimizer uses the Adma optimization algorithm, the hyperparameter beta is set to 0.5, and the kaiming initialization method is used to initialize the weight parameters of the model.
[0034] Start training the generative adversarial network. Input the binary mask image into the generator G to obtain the generated image g, and then input the generated image g into the discriminator D to obtain the discrimination index. The training goal is to make the generated image g generated by G closer and closer to the real image r, while making D have a stronger discrimination ability to distinguish the generated image from the real image. Both the generator G and the discriminator D first delete some convolutional layers, and the output image resolution is also small. As the number of training iterations gradually increases, the number of convolutional layers increases, and the output image resolution is large until the output resolution is the same as that of the real image, so that a high-quality generated image can be obtained.
[0035] The target loss function of the generative adversarial network is shown in the following formula.
[0036] V(D,G)=E x[log(D(x))]+E m [1 - log(D(G(m)))]
[0037] Among them, x is the real image, m is the mask input, G represents the generator, and D represents the discriminator.
[0038] Through backpropagation, the loss is updated during training using the stochastic gradient descent method to make the value of the objective loss function tend to be minimized as much as possible. When the change in the loss value is confined to a very small range after long-term training, it can be considered that the network has converged and the training ends.
[0039] As Figure 3 shown, when training the generative adversarial network, in the initial stage of training, the last five convolutional layers of the generator and the discriminator are deleted. At this time, the resolution size of the generator output is 10×10. Every time the training is performed 4000 times, one convolutional layer of the generator and the discriminator is restored, and the output image resolution is doubled. Then when the final training ends, the output resolution is 640×640, which is the resolution size of the input. This way of gradually increasing the convolutional layer and resolution can enable the network to generate higher-quality samples.
[0040] S4. Randomly transform the mask to obtain an augmented mask binary image dataset, mix the augmented mask binary image dataset with the mask binary image dataset to obtain a mixed mask dataset, and input the mask binary image data of the mixed mask dataset into the trained generative adversarial network to generate a pseudo-sample dataset;
[0041] After the training ends, the generator G already has good generation ability. At this time, a series of random transformations are performed on the defect area on any mask binary image, such as rotation, scaling, affine transformation, erosion, dilation, translation, etc., to generate mask binary images with random and diverse defect areas. These newly generated mask binary images are used as the augmented mask image set, and the augmented mask image set is mixed with the original mask binary images to obtain a mask dataset for subsequent generation of pseudo-samples.
[0042] Input the processed mask image into the generator G for forward inference to generate pseudo-samples. Use the mask binary image as the input of the generative adversarial network instead of a random variable.
[0043] S5. Train the YOLOv5 network with the real sample dataset as the generation quality discriminator, use the generation quality discriminator to discriminate the quality of the pseudo-samples, and remove the pseudo-samples with low quality to obtain a filtered pseudo-sample dataset.
[0044] Train a YOLOv5 network with the real images of the rice cooker inner pots in the real sample dataset. This network is used to discriminate the generation quality of the pseudo-sample images and the accurate annotation information (position and category) of the defects on them.
[0045] Specifically, the generated pseudo-samples are input into the trained YOLOv5 network to obtain the location and category of the defects on the inner liner of the rice cooker. At the same time, the FID (Frechet Inception Distance) value can be calculated based on the feature vectors of the YOLOv5 network. FID is an evaluation metric for assessing the quality of samples generated by a generative adversarial network. Generally, the Frechet distance is calculated using the feature vectors in the InceptionV3 network, so it gets this name. In this patent, the feature vectors of YOLOv5 are used to replace the feature vectors of the InceptionV3 network to calculate the FID. The higher the FID value, the greater the difference between the two. The FID reference value is set to 100. Images with FID values lower than 100 are retained, and images with FID values higher than 100 are removed, realizing the generation of high-quality images of defects on the inner liner of the rice cooker. The calculation formula of FID is as follows:
[0046]
[0047] where μ g is the mean of the feature vectors of the generated samples g, μ r is the mean of the feature vectors of the real samples r, ∑g represents the covariance matrix of the feature vectors of g, ∑r represents the covariance matrix of the feature vectors of r, and T r is the trace of the matrix.
[0048] S6. The filtered pseudo-sample dataset and the real-sample dataset are used to obtain an enhanced dataset of images of the inner liner of the rice cooker.
[0049] The filtered pseudo-samples are mixed with the original dataset to finally obtain an enhanced dataset after data augmentation. The dataset after data augmentation has more defect samples with different forms and characteristics compared to the original dataset. This dataset can be used to train a defect detection algorithm based on deep learning, enabling the defect detection algorithm to learn sufficient defect features from the dataset after data augmentation, thereby improving the defect detection effect.
[0050] In summary, the present invention proposes a method for enhancing the image data of the inner pot of an electric rice cooker based on a masked generative adversarial network. By constructing a deep learning model based on the generative adversarial network, the original input directly using random variables or the original image is changed to input a masked binary image, and a series of random transformations are performed on the masked binary image to make the input more abundant. As a result, the positions and forms of the defects in the generated images are also more abundant, and the corresponding defect positions can be directly obtained. In addition, the training is carried out by gradually increasing the resolution, and an additional discriminator of YOLOv5 is added to screen the generated samples, realizing the generation of high-quality defect images of the inner pot of the electric rice cooker; effectively improving the detection accuracy of the defect detection of the inner pot of the electric rice cooker based on deep learning; having the advantages of wide application range, intelligence and good robustness; effectively solving the problems of insufficient data volume and difficult data annotation in the deep learning process.
[0051] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An image data enhancement method for the inner pot of a rice cooker based on a masked generative adversarial network, characterized in that, Including: S1. Take pictures of the inner pot of the rice cooker through an industrial camera, label the defects of the pictures of the inner pot of the rice cooker, and construct a real sample data set of the images of the inner pot of the rice cooker; S2. Generate corresponding mask binary maps according to each picture of the inner pot of the rice cooker in the real sample data set to obtain a mask binary map data set; S3. Construct a generative adversarial network, initialize the network weights of the generative adversarial network, and train the generative adversarial network in a way of gradually increasing the resolution; The training of the generative adversarial network in the way of gradually increasing the resolution includes: training the generative adversarial network, inputting the mask binary map into the generator to obtain a generated map, and inputting the generated map into the discriminator to obtain a discrimination index; both the generator and the discriminator first delete some convolutional layers, and as the number of training iterations gradually increases, the number of convolutional layers increases until the resolution of the image output by the generator is the same as the resolution of the real image; Both the generator and the discriminator first delete some convolutional layers, and as the number of training iterations gradually increases, the number of convolutional layers increases until the resolution of the output image is the same as the resolution of the real image, including: first deleting the last five convolutional layers of the generator and the discriminator, the resolution size of the image output by the generator is 10×10, and every time 4000 training times, one convolutional layer of the generator and the discriminator is restored, and the resolution of the output image is doubled; when the training ends, the resolution of the output image is equal to the resolution of the input image, and the resolution size of the input image is 640×640; The target loss function of the generative adversarial network is: ; Among them, x is the real image, m is the mask input, G represents the generator, and D represents the discriminator; S4. Randomly transform the mask binary map data to obtain an augmented mask binary map data set, mix the augmented mask binary map data set with the mask binary map data set to obtain a mixed mask data set, and input the mask binary map data of the mixed mask data set into the trained generative adversarial network to generate a pseudo sample data set; S5. Train the YOLOv5 network with the real sample data set as a generation quality discriminator, use the generation quality discriminator to discriminate the quality of the pseudo samples, and remove the pseudo samples with low quality to obtain a filtered pseudo sample data set; The training of the YOLOv5 network with the real sample data set as a generation quality discriminator, using the generation quality discriminator to discriminate the quality of the pseudo samples, and removing the pseudo samples with low quality to obtain a filtered pseudo sample data set includes: inputting the generated pseudo samples into the trained YOLOv5 network to obtain the position, category and confidence of the defects of the inner pot of the rice cooker, the FID value can be calculated according to the feature vector of the YOLOv5 network, set the FID reference value, keep the images lower than the FID reference value, and remove the images higher than the FID reference value; S6. Mix the filtered pseudo sample data set with the real sample data set to obtain an enhanced image data set of the inner pot of the rice cooker.
2. The method for enhancing the image data of the inner pot of an electric rice cooker based on a masked generative adversarial network according to claim 1, wherein The defect labeling of the pictures of the inner pot of the rice cooker includes: defect labeling the pictures of the inner pot of the rice cooker corresponding to the inner pot of the rice cooker with defects, and recording the position and category of the defect areas in the pictures of the inner pot of the rice cooker.
3. The method for enhancing the image data of the inner pot of an electric rice cooker based on a masked generative adversarial network according to claim 2, wherein, Generating a corresponding mask binary image according to each inner pot image of the real sample dataset includes: initializing a grayscale image for the inner pot image of the rice cooker, where the grayscale values of all pixels are 0, and setting the pixel points within the position of the defect area to 255.
4. The method for enhancing the image data of the inner pot of an electric rice cooker based on a masked generative adversarial network according to claim 2, wherein, The generative adversarial network, which includes a generator and a discriminator. The generator is a 12-layer UNet network structure, and the discriminator is a six-layer convolutional neural network, with the size of each convolutional layer being 3×3 and the stride being 2.
5. The method for enhancing the image data of the inner pot of an electric rice cooker based on a masked generative adversarial network according to claim 1, characterized in that, Performing random transformation on the mask binary image data to obtain an augmented mask binary image dataset includes: generating diverse mask binary images with defect areas by performing rotation, scaling, affine transformation, erosion, dilation, and translation on the defect areas on the mask binary image.
6. The method for enhancing the image data of the inner pot of an electric rice cooker based on a masked generative adversarial network according to claim 1, characterized in that The calculation formula for the FID value is: ; Among them, is the mean of the feature vectors of the generated sample g, is the mean of the feature vectors of the real sample r, represents the covariance matrix of the feature vectors of g, represents the covariance matrix of the feature vectors of r, is the trace of the matrix.
Citation Information
Patent Citations
Pavement disease data enhancement method based on generative adversarial network
CN114022368A
Image data enhancement method and device, computer equipment and storage medium
CN114037637A